reasoning-advances-reflect-genuine-discontinuity

OUT derived (depth 2)

Created 2026-06-21T10:12:39+00:00

Both training-time reasoning specialization (o1 scoring 6x better than GPT-4o on IMO problems, R1 matching proprietary models at open-weight cost) and inference-time structured reasoning evolution (CoT → self-consistency → tree-of-thoughts) reflect genuine cognitive discontinuities rather than smooth scaling artifacts.

Justifications

SL — If emergent abilities are metric artifacts rather than real discontinuities, both training-time and inference-time reasoning advances may be overstated

Antecedents (all must be IN):

  • IN reasoning-models-represent-distinct-capability-tier — Reasoning-specialized models — OpenAI o1 scoring 83% vs GPT-4o's 13% on IMO qualifying problems, DeepSeek R1 matching proprietary models at lower cost — represent a distinct capability tier above standard LLMs, achievable through both proprietary and open-weight approaches.
  • IN structured-reasoning-prompting-evolved-from-linear-to-branching — Prompting for reasoning evolved from linear chain-of-thought (single path) to self-consistency (multiple paths, majority vote) to tree-of-thoughts (branching with backtracking), progressively adding search structure.

Unless (any of these IN defeats this justification):

  • IN emergent-abilities-metric-artifact-debate — The appearance of emergent abilities in LLMs depends on metric choice: accuracy metrics show step-function discontinuities while log-probability metrics show smooth scaling curves (Schaeffer et al.).