popularity-accuracy-correlation-by-model-size

IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-evaluation-setup.md

Created 2026-08-25T02:58:11+00:00

The correlation between subject entity popularity and model accuracy is approximately 0.4 for GPT-3 (davinci-003) versus approximately 0.1 for GPT-Neo-1.3B.

Summary

Larger models show a noticeably stronger pattern of getting more accurate answers about well-known, popular topics compared to obscure ones, while smaller models show almost no such gap. In practice, this means scaling up model size widens the accuracy divide between famous and niche subjects, so a big model can still be meaningfully worse on obscure topics than it is on popular ones.