scaling-does-not-improve-tail-knowledge
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s4-memorization-depends-on-popularity.md
Created 2026-08-25T02:58:11+00:00
The 4,000 least popular POPQA questions show only ~15–19% accuracy even for GPT-3 davinci-003, indicating scaling does not significantly improve long-tail recall at practical model sizes.
Summary
Even the largest practical language models still miss roughly 80% of the most obscure trivia questions, meaning that making models bigger does not meaningfully close the gap on rare, long-tail knowledge. For the system, this means no amount of additional scaling should be expected to fix recall failures on infact questions; the limitation is structural, not just a matter of needing a slightly larger model.