reasoning-models-are-genuine-cognitive-discontinuity
OUT derived (depth 2)
Created 2026-06-21T10:00:59+00:00
Reasoning-specialized models demonstrate a genuine cognitive discontinuity — with o1 scoring 83% vs GPT-4o's 13% on IMO problems, consistent with the broader observation that emergent abilities appear discontinuously at scale thresholds — unless the apparent discontinuity is an artifact of metric choice rather than a real capability transition.
Justifications
SL — reasoning model breakthroughs are genuine only if emergent abilities are not metric artifacts
Antecedents (all must be IN):
- IN reasoning-models-represent-distinct-capability-tier — Reasoning-specialized models — OpenAI o1 scoring 83% vs GPT-4o's 13% on IMO qualifying problems, DeepSeek R1 matching proprietary models at lower cost — represent a distinct capability tier above standard LLMs, achievable through both proprietary and open-weight approaches.
- IN emergent-abilities-discontinuous-scale — Emergent abilities in LLMs appear discontinuously at certain scale thresholds, not linearly (Wei et al., 2022, arXiv:2206.07682)
Unless (any of these IN defeats this justification):
- IN emergent-abilities-metric-artifact-debate — The appearance of emergent abilities in LLMs depends on metric choice: accuracy metrics show step-function discontinuities while log-probability metrics show smooth scaling curves (Schaeffer et al.).