mistral-sae-l1-2-sparse-penalty

IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references-chunk-2.md

Created 2026-08-25T02:58:02+00:00

Mistral 7B SAEs use an L_{1/2} (half-norm) sparsity penalty with λ = 0.012, which produces sparser codes than L1.

Summary

The Mistral 7B sparse autoencoder is trained with an especially aggressive sparsity pressure (the L1/2 penalty) so that, for any given input, only a very small handful of internal features light up — far fewer than you would get with a standard L1 penalty. This is a deliberate design trade: it makes each active feature easier to interpret as a single clear concept, but it means the system compresses more information into fewer active dimensions, so downstream reasoning has to work with a narrower, more selective set of signals.