kandpal-2023-corpus-spearman-087-097

IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s3-lm-accuracy-depends-on-relevant.md

Created 2026-08-25T02:58:06+00:00

Pre-training corpora (The Pile, ROOTS, C4, OpenWebText, Wikipedia) show Spearman rank correlations of 0.87–0.97 in their per-question relevant document counts.

Summary

The major pre-training datasets largely agree with each other about which questions have abundant relevant material and which are sparse, with rankings that line up 87 to 97 percent of the time. This means the choice of corpus does not radically reshape the knowledge landscape a model sees per question, so findings about coverage or gaps measured on one corpus should generalize well to the others.

Dependents

These beliefs depend on this one: