llama-training-token-volumes-per-generation
IN premise — summaries/2026/08/24/wiki-LLaMA-chunk-1.md
Created 2026-08-24T17:11:14+00:00
Llama training data volumes were approximately: Llama 1 â 1.4T tokens, Llama 2 = 2T tokens, Llama 3 = 15T tokens, Llama 3.1 = 15T tokens (July 2024), and Llama 4 = 22Tâ40T tokens.
Summary
Between Llama 1 and Llama 4, the amount of training text Meta feeds into each generation has grown roughly 28-fold, from about 1.4 trillion to up to 40 trillion tokens. This matters because it shows data volume is one of Meta's primary levers for capability gains, and each jump carries enormous compute, curation, and licensing costs that shape how quickly and how often new models can ship.