emergent-abilities-metric-artifact-debate
IN premise — entries/2026/06/21/wiki-Large_language_model-chunk-2.md
Created 2026-06-21T09:50:09+00:00
The appearance of emergent abilities in LLMs depends on metric choice: accuracy metrics show step-function discontinuities while log-probability metrics show smooth scaling curves (Schaeffer et al.).
Summary
The sudden "emergence" of new abilities in large language models may not be a real phase change in the model at all, but rather an illusion created by choosing a step-like measurement scale. This matters because if the apparent breakthroughs dissolve under a smoother metric, then the story of LLMs crossing a magical threshold at a certain size loses its footing, and planning around predicted emergence events becomes unreliable.
Dependents
These beliefs depend on this one:
- OUT cot-threshold-validates-emergent-discontinuity — Chain-of-thought prompting's empirically measured threshold of ~62B parameters is a specific documented instance of emergent abilities' discontinuous appearance at scale, validating that reasoning itself is an emergent property rather than a gradually improving one.
- OUT llm-evaluation-provides-stable-capability-assessment — LLM evaluation through diverse benchmarking frameworks (MMLU, HLE, HELM, LMArena) provides stable, convergent capability assessments that reliably rank models and track progress.
- OUT llm-learning-involves-genuine-phase-transitions — LLM capability acquisition involves genuine discontinuous phase transitions — both grokking (sudden generalization after memorization within training) and emergent abilities (capabilities appearing at scale thresholds) — rather than smooth, predictable improvement curves.
- OUT reasoning-advances-reflect-genuine-discontinuity — Both training-time reasoning specialization (o1 scoring 6x better than GPT-4o on IMO problems, R1 matching proprietary models at open-weight cost) and inference-time structured reasoning evolution (CoT → self-consistency → tree-of-thoughts) reflect genuine cognitive discontinuities rather than smooth scaling artifacts.
- OUT reasoning-models-are-genuine-cognitive-discontinuity — Reasoning-specialized models demonstrate a genuine cognitive discontinuity — with o1 scoring 83% vs GPT-4o's 13% on IMO problems, consistent with the broader observation that emergent abilities appear discontinuously at scale thresholds — unless the apparent discontinuity is an artifact of metric choice rather than a real capability transition.