chinchilla-balance-prescribes-optimal-resource-allocation

OUT derived (depth 1)

Created 2026-06-21T10:16:20+00:00

Chinchilla's prescription to scale parameters and data in equal proportion provides the optimal training resource allocation strategy.

Justifications

SL — Chinchilla balance is a theoretical optimum, but Llama 3's 75x overtraining with continued improvement shows inference-optimal allocation diverges from compute-optimal

Antecedents (all must be IN):

  • IN chinchilla-compute-optimal-training — Hoffmann et al. (2022, 'Chinchilla', arXiv:2203.15556) showed prior models were undertrained relative to dataset size and that optimal training requires scaling data proportionally with parameters
  • IN chinchilla-scaling-law-constants — The Chinchilla scaling law is L = A/N^alpha + B/D^beta + L0 with alpha=0.34, beta=0.28, L0=1.69, and training cost C = 6·N·D FLOPs.

Unless (any of these IN defeats this justification):