flan-ul2-position-spread-2048
IN premise — summaries/2026/08/24/liu-2023-lost-in-middle-s4-why-are-language-models-not-robust.md
Created 2026-08-25T02:58:09+00:00
Flan-UL2 (2048-token encoder / 512-token decoder) shows only 1.9% absolute difference between best- and worst-case positions within its trained 2048-token context window, but exhibits U-shaped degradation when evaluated beyond 2048 tokens.
Summary
Within its designed range, Flan-UL2 handles information roughly equally well no matter where it appears in the input, with less than 2% performance variation across positions. However, once you push it past the 2048-token window it was trained on, performance drops off in a predictable U-shape, meaning the model is reliable in its comfort zone but degrades noticeably when stretched beyond it.