corpus-redundancy-overcounts-unique-information
IN derived (depth 1)
Created 2026-08-25T03:45:56+00:00 · Reviewed 2026-08-25T04:28:09+00:00
The high inter-correlation of pre-training corpora (Spearman 0.87–0.97) combined with marginal accuracy gains from 5× data indicates that "relevant document count" systematically overcounts unique information, making the long-tail scaling estimate an upper bound on true information scarcity rather than a lower bound on required capacity.
Summary
Because the major pre-training datasets overlap so heavily in what they cover, simply counting how many documents mention a topic makes that topic look rarer than it really is. This means scaling estimates built on document counts overstate how much additional data or model capacity is actually needed, since the apparent long tail of rare topics is inflated by redundancy across corpora rather than reflecting genuine information scarcity.
Justifications
This belief has 2 justifications — it is IN if any one holds.
SL — Each antecedent independently supports the redundancy-overcounting reading: (1) high Spearman correlations mean the corpora are near-duplicates, so "relevant documents in C4" ≈ "relevant documents in The Pile" ≈ "relevant documents in Wikipedia"; (2) 5× data yielding marginal gains confirms the information is not 5×. Either alone suffices to argue the document count is a redundant proxy. ANY mode because either piece of evidence independently establishes the overcounting.
Antecedents (all must be IN):
- IN kandpal-2023-corpus-spearman-087-097 — Pre-training corpora (The Pile, ROOTS, C4, OpenWebText, Wikipedia) show Spearman rank correlations of 0.87–0.97 in their per-question relevant document counts.
SL — Each antecedent independently supports the redundancy-overcounting reading: (1) high Spearman correlations mean the corpora are near-duplicates, so "relevant documents in C4" ≈ "relevant documents in The Pile" ≈ "relevant documents in Wikipedia"; (2) 5× data yielding marginal gains confirms the information is not 5×. Either alone suffices to argue the document count is a redundant proxy. ANY mode because either piece of evidence independently establishes the overcounting.
Antecedents (all must be IN):
- IN kandpal-2023-5x-data-marginal-gains — Increasing pre-training data by 5× yields only small accuracy gains because major corpora are highly correlated (Spearman ρ ≥ 0.87) in the knowledge they cover.
Dependents
These beliefs depend on this one:
- OUT superposition-unifies-data-and-representation-redundancy — Over-complete superposition is the single structural cause of both the data-side redundancy (5× correlated corpora yield only marginal gains, Spearman 0.87–0.97 inter-correlation) and the representation-side redundancy (SAE dead features at 2–48%, reconstruction shrinkage): in both cases the available "capacity" exceeds the unique information it must encode.