mistral-sae-optimizer-dead-features
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references-chunk-2.md
Created 2026-08-25T02:58:03+00:00
Mistral 7B SAEs use AdamW with weight decay 10⁻³, learning rate 2×10⁻⁴, and linear warmup; after 5× dead-feature resampling, approximately 1000 dead features remain out of 65,536.
Summary
Even with a well-tuned training setup and five rounds of explicitly resampling dead features, the Mistral 7B sparse autoencoder still carries roughly 1.5 percent of its feature slots as effectively unused, meaning interpreters should expect a small but persistent gap in coverage rather than a fully dense dictionary. This sets a practical floor on how much of the model's internal representation the SAE can meaningfully account for under these hyperparameters.