emergent-abilities-threshold-not-linear
IN premise — entries/2026/06/21/wiki-Generative_pre-trained_transformer-chunk-1.md
Created 2026-06-21T09:50:09+00:00
Emergent abilities in LLMs (multi-step reasoning, in-context learning) appear only above certain scale thresholds and do not scale linearly with model size.
Summary
Capabilities like multi-step reasoning or in-context learning don't grow gradually as models get bigger; they switch on at specific size thresholds, more like a light turning on than a dimmer sliding up. This means you can't reliably predict when a model will gain a new ability just by scaling it up, and it means smaller models are missing entire capabilities rather than having just weaker versions of them.
Dependents
These beliefs depend on this one:
- IN scaling-predictions-fail-at-both-capability-and-resource-levels — Scaling predictions systematically fail at both capability and resource allocation levels: emergent abilities appear discontinuously at unpredictable thresholds rather than following smooth power-law trends, and Chinchilla-optimal compute allocation is systematically violated by successful models (Llama 3 8B trained at 75x the prescribed data-to-parameter ratio with continued improvement) — the field's quantitative scaling framework provides trend guidance but not operational prediction.