original-transformer-100m-parameters
IN premise — entries/2026/06/21/wiki-Transformer_deep_learning_architecture.md
Created 2026-06-21T09:50:10+00:00
The original Transformer model from 'Attention Is All You Need' had approximately 100 million parameters.
Summary
The original Transformer architecture that launched the modern language-model era was a compact design with roughly 100 million tunable numbers, which serves as the baseline against which all later scaling is measured. Knowing this anchor point lets the system reason about how many times larger today's frontier models are and what that implies for training cost and capability.