kandpal-2023-corpus-spearman-087-097
IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s3-lm-accuracy-depends-on-relevant.md
Created 2026-08-25T02:58:06+00:00
Pre-training corpora (The Pile, ROOTS, C4, OpenWebText, Wikipedia) show Spearman rank correlations of 0.87–0.97 in their per-question relevant document counts.
Summary
The major pre-training datasets largely agree with each other about which questions have abundant relevant material and which are sparse, with rankings that line up 87 to 97 percent of the time. This means the choice of corpus does not radically reshape the knowledge landscape a model sees per question, so findings about coverage or gaps measured on one corpus should generalize well to the others.
Dependents
These beliefs depend on this one:
- IN corpus-redundancy-overcounts-unique-information — The high inter-correlation of pre-training corpora (Spearman 0.87–0.97) combined with marginal accuracy gains from 5× data indicates that "relevant document count" systematically overcounts unique information, making the long-tail scaling estimate an upper bound on true information scarcity rather than a lower bound on required capacity.