anthropic-sae-expansion-factor-range
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity.md
Created 2026-08-25T02:57:56+00:00
SAEs in the Bricken et al. paper are trained with expansion factors from 1x (512 features) to 256x (131,072 features), with the A/1 run focusing on 8x (4,096 features)
Summary
The Bricken et al. study breaks a model's internal activations down into sparse, interpretable features at a wide range of granularities, from just 512 features up to over 131,000. The primary analysis (the A/1 run) locks in on a 4,096-feature level, which serves as the concrete reference point for what the model's representational structure looks like at that specific resolution.