kandpal-2023-counterfactual-4-8b-params
IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s1-introduction.md
Created 2026-08-25T02:58:05+00:00
The counterfactual re-training experiment in Kandpal et al. (2023) uses a 4.8-billion-parameter language model trained with and without specific relevant documents.
Summary
The Kandpal paper's core experiment compares a roughly 4.8-billion-parameter language model trained two ways: once including certain key documents and once excluding them, to see what difference those documents make to the model's output. This sets the scale and method for everything the paper claims to show, meaning any conclusions about document influence are bounded by what a model of this size and this specific counterfactual setup can demonstrate.