scaling-predictions-fail-at-both-capability-and-resource-levels
IN derived (depth 1)
Created 2026-06-21T13:22:51+00:00 · Reviewed 2026-06-21T14:41:08+00:00
Scaling predictions systematically fail at both capability and resource allocation levels: emergent abilities appear discontinuously at unpredictable thresholds rather than following smooth power-law trends, and Chinchilla-optimal compute allocation is systematically violated by successful models (Llama 3 8B trained at 75x the prescribed data-to-parameter ratio with continued improvement) — the field's quantitative scaling framework provides trend guidance but not operational prediction.
Summary
The formulas the AI field uses to predict model size, training data, and compute spend are more like a rough compass than a GPS — new capabilities pop up suddenly at scale points nobody predicted, and successful models routinely use far more training data than the math says is optimal while still improving. In practice, the field's quantitative scaling laws can suggest a general direction but cannot reliably tell you where performance will plateau, when a new skill will emerge, or exactly how much compute to budget, making both capability forecasting and resource planning far more uncertain than the framework implies.
Justifications
SL — Scaling laws fail predictively at both the capability emergence level and the resource allocation level
Antecedents (all must be IN):
- IN emergent-abilities-threshold-not-linear — Emergent abilities in LLMs (multi-step reasoning, in-context learning) appear only above certain scale thresholds and do not scale linearly with model size.
- IN llama3-8b-trained-15t-tokens-chinchilla-suboptimal — Llama 3 8B was trained on 15T tokens — 75x more than the Chinchilla-optimal 200B tokens — and performance continued to scale log-linearly, challenging Chinchilla scaling assumptions