mallen-2023-gptneo-20b-25pct
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-evaluation-setup.md
Created 2026-08-25T02:58:11+00:00
GPT-Neo 20B achieves approximately 25% accuracy on POPQA in the 15-shot closed-book setting without fine-tuning.
Summary
A 20-billion-parameter language model, given 15 example questions as a nudge and no access to outside information, gets only about a quarter of everyday common-sense questions right when it has not been specifically trained for the task. This establishes a modest baseline that highlights a large gap in the model's raw knowledge, showing that retrieval, fine-tuning, or larger architectures would need to do significant additional work to close it.