kandpal-2023-5x-data-marginal-gains

IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s4-methods-to-improve-rare-fact-learning.md

Created 2026-08-25T02:58:06+00:00

Increasing pre-training data by 5× yields only small accuracy gains because major corpora are highly correlated (Spearman ρ ≥ 0.87) in the knowledge they cover.

Summary

Piling on five times as much training data barely improves accuracy because the big public datasets largely cover the same knowledge, with very little unique content. The practical takeaway is that simply buying more data is a dead end; real gains require finding genuinely diverse sources or changing how the data is used.

Dependents

These beliefs depend on this one: