sae-dead-feature-proportions-by-size

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-3.md

Created 2026-08-25T02:58:37+00:00

Dead feature proportions (zero activation across 10⁷ tokens) are approximately 2% for the 1M SAE, 35% for the 4M SAE, and 65% for the 34M SAE.

Summary

As SAEs get bigger, a rapidly growing share of their features never fire at all, jumping from 2 percent in the smallest model to two-thirds in the largest. This means the effective vocabulary of interpretable concepts in a large SAE is far smaller than its nominal feature count suggests, and most of the added capacity is wasted on features that contribute nothing.

Dependents

These beliefs depend on this one: