sae-dictionary-size-gpt2-mistral
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-s1-create-a-graph-g-out-of-the-dictionary-elements-by-adding-di.md
Created 2026-08-25T02:58:01+00:00
GPT-2-small layer 7 has approximately 25,000 SAE dictionary elements (Bloom, 2024); Mistral 7B layer 8 has 216,000 SAE dictionary elements.
Summary
These numbers show how many distinct, sparse features a decoder can identify in a given layer of a model, and they reveal a big scaling gap: the much larger Mistral model exposes roughly nine times as many identifiable features at a comparable intermediate depth. This matters because it tells interpretability researchers that the feature space of modern models is far richer than early models like GPT-2, so methods and tools built around small dictionaries won't generalize without major rework.