gpt3-davinci-003-popqa-longtail-19pct
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s1-introduction.md
Created 2026-08-25T02:58:10+00:00
GPT-3 davinci-003 achieves approximately 19% accuracy on the 4,000 least-popular POPQA questions.
Summary
GPT-3 (the davinci-003 version) gets only about 1 in 5 answers right when quizzed on the obscure, rarely-mentioned questions in the POPQA commonsense benchmark, meaning its knowledge drops off sharply once you move past well-known topics. This sets a baseline showing that the model is far from reliable on niche or long-tail factual questions, which matters for any system that depends on it to handle less mainstream queries.