gptj-6b-vs-gpt3-longtail-scaling-insufficient
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s1-introduction.md
Created 2026-08-25T02:58:10+00:00
GPT-j 6B scores 16% versus GPT-3 davinci-003's 19% on the 4,000 least-popular POPQA questions, showing that an order-of-magnitude parameter increase yields only ~3 percentage points gain on long-tail recall.
Summary
Even with roughly ten times more parameters, the newer model only picks up about three extra correct answers on the rarest, least-frequently-asked questions. This means simply scaling up model size does not proportionally fix the problem of niche or long-tail knowledge, so systems relying on factual recall will need targeted interventions rather than just bigger models to cover the long tail.