llama-training-token-volumes-per-generation

IN premisesummaries/2026/08/24/wiki-LLaMA-chunk-1.md

Created 2026-08-24T17:11:14+00:00

Llama training data volumes were approximately: Llama 1 ≈ 1.4T tokens, Llama 2 = 2T tokens, Llama 3 = 15T tokens, Llama 3.1 = 15T tokens (July 2024), and Llama 4 = 22T–40T tokens.

Summary

Between Llama 1 and Llama 4, the amount of training text Meta feeds into each generation has grown roughly 28-fold, from about 1.4 trillion to up to 40 trillion tokens. This matters because it shows data volume is one of Meta's primary levers for capability gains, and each jump carries enormous compute, curation, and licensing costs that shape how quickly and how often new models can ship.