xu-2024-linear-order-universal-model-failure

IN premise — summaries/2026/08/24/xu-2024-hallucination-innate-sA-appendix.md

Created 2026-08-24T17:11:30+00:00

All tested models in Xu et al. 2024 (Llama 2/3 70B, GPT-3.5, GPT-4, GPT-4-turbo) fail linear-order reasoning tasks ω(m) at m=1000₂, exhibiting specific errors including inability to apply transitivity and inconsistent answers for 'x$y' versus 'y$x'.

Summary

Every major language model tested in Xu et al.'s 2024 study breaks down on basic ordering logic once the sequence reaches just eight items, making errors like forgetting that if A comes before B and B before C, then A must come before C, or contradicting themselves when the same relationship is simply restated in reverse. This sets a concrete floor for where current models fail at structured reasoning, meaning any system relying on them for ordered comparisons needs to account for these gaps at surprisingly small scales.