liu-2023-closed-book-below-mid-context
IN premise — summaries/2026/08/24/liu-2023-lost-in-middle-s1-introduction.md
Created 2026-08-25T02:58:08+00:00
When the relevant document is placed mid-context among 20 documents, GPT-3.5-Turbo scores below its closed-book (no-document) accuracy of approximately 56.1%, indicating distractors actively harm performance.
Summary
Cramming a correct answer document into a long list of 19 other documents actually makes GPT-3.5-Turbo perform worse than it would with no documents at all, meaning the irrelevant context is doing real harm rather than just adding noise. This matters for any system that tries to improve accuracy by stuffing more reference material into the prompt, because the extra material can backfire and drag the model below its baseline ability.