mallen-2023-2-7b-retriever-beats-davinci-longtail
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s2-related-work.md
Created 2026-08-25T02:58:10+00:00
A 2.7B GPT-Neo model augmented with a dense retriever outperforms GPT-3 davinci-003 on the 4,000 least-popular POPQA questions.
Summary
A small 2.7-billion-parameter model paired with a search tool outperforms the much larger GPT-3 davinci on the rarest, most obscure trivia questions. The practical takeaway is that for long-tail knowledge, bolting on a retrieval component to a modest model can beat simply scaling up parameters, which matters when deciding whether to invest in bigger models or better search infrastructure.