feature-splitting-sae-width

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity.md

Created 2026-08-25T02:57:56+00:00

As SAE width increases, a single coarse feature splits into multiple finer-grained interpretable sub-features (e.g., one base64 feature becomes three more specific base64 sub-features), indicating the smaller SAE was under-resolved

Summary

Widening a sparse autoencoder tends to break apart vague, lumped-together features into several sharper, more specific ones, meaning a smaller model was blurring distinct concepts into one blob. In practice, any single "feature" you see in a narrow SAE should be treated as potentially hiding multiple separate behaviors, so interpretability conclusions drawn from small models are under-resolved by default.