emergent-abilities-discontinuous-scale
IN premise — entries/2026/06/21/wiki-Generative_pre-trained_transformer.md
Created 2026-06-21T09:50:09+00:00
Emergent abilities in LLMs appear discontinuously at certain scale thresholds, not linearly (Wei et al., 2022, arXiv:2206.07682)
Summary
Large language models don't gain new capabilities gradually as they grow; instead, certain abilities snap into existence all at once when the model crosses a size threshold. This makes it genuinely hard to predict when a model will suddenly start doing something it previously couldn't, because performance doesn't scale in a smooth, linear way.
Dependents
These beliefs depend on this one:
- OUT cot-threshold-validates-emergent-discontinuity — Chain-of-thought prompting's empirically measured threshold of ~62B parameters is a specific documented instance of emergent abilities' discontinuous appearance at scale, validating that reasoning itself is an emergent property rather than a gradually improving one.
- OUT llm-learning-involves-genuine-phase-transitions — LLM capability acquisition involves genuine discontinuous phase transitions — both grokking (sudden generalization after memorization within training) and emergent abilities (capabilities appearing at scale thresholds) — rather than smooth, predictable improvement curves.
- OUT reasoning-models-are-genuine-cognitive-discontinuity — Reasoning-specialized models demonstrate a genuine cognitive discontinuity — with o1 scoring 83% vs GPT-4o's 13% on IMO problems, consistent with the broader observation that emergent abilities appear discontinuously at scale thresholds — unless the apparent discontinuity is an artifact of metric choice rather than a real capability transition.