mallen-2023-gpt3-zero-shot-35pct

IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s3-evaluation-setup.md

Created 2026-08-25T02:58:11+00:00

Zero-shot GPT-3 achieves approximately 35% accuracy on POPQA without fine-tuning.

Summary

GPT-3 can answer roughly one in three deliberately tricky common-sense questions about obscure topics correctly even without any task-specific training, which sets a baseline for how much useful knowledge a large pre-trained model already carries before any further work is invested in improving it. This matters because it tells the system what to expect as a starting point when evaluating whether additional fine-tuning or context provides a meaningful gain.