kandpal-2023-5x-data-marginal-gains
IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s4-methods-to-improve-rare-fact-learning.md
Created 2026-08-25T02:58:06+00:00
Increasing pre-training data by 5× yields only small accuracy gains because major corpora are highly correlated (Spearman ρ ≥ 0.87) in the knowledge they cover.
Summary
Piling on five times as much training data barely improves accuracy because the big public datasets largely cover the same knowledge, with very little unique content. The practical takeaway is that simply buying more data is a dead end; real gains require finding genuinely diverse sources or changing how the data is used.
Dependents
These beliefs depend on this one:
- IN corpus-redundancy-overcounts-unique-information — The high inter-correlation of pre-training corpora (Spearman 0.87–0.97) combined with marginal accuracy gains from 5× data indicates that "relevant document count" systematically overcounts unique information, making the long-tail scaling estimate an upper bound on true information scarcity rather than a lower bound on required capacity.