llama3-8b-trained-15t-tokens-chinchilla-suboptimal

IN premiseentries/2026/06/21/wiki-LLaMA-chunk-1.md

Created 2026-06-21T09:50:09+00:00

Llama 3 8B was trained on 15T tokens — 75x more than the Chinchilla-optimal 200B tokens — and performance continued to scale log-linearly, challenging Chinchilla scaling assumptions

Summary

Meta trained an 8-billion-parameter model on roughly 75 times more data than the standard Chinchilla guideline would call "optimal," yet the model kept improving instead of plateauing. This matters because it suggests the usual scaling rules of thumb may be too conservative — teams can spend a lot of extra compute overtraining a smaller model and still see real gains, which shifts the cost calculus when deciding between building a bigger model or just training a smaller one longer.

Dependents

These beliefs depend on this one: