nanda-decoder-weight-sparsity-distribution
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity-chunk-8.md
Created 2026-08-25T02:57:56+00:00
Nanda's SAE replication found that 4% of decoder weights are well explained by a single neuron, 4% by 2-10 neurons, and 92% are dense across many neurons in the neuron basis
Summary
When people try to express SAE feature directions back in terms of the network's original neurons, the overwhelming majority (92%) smear across many neurons rather than lining up with one or a few. This means most "features" the sparse autoencoder discovers are not simple single-neuron concepts but distributed, multi-neuron patterns, which makes them harder to interpret as clean, localized ideas about what the network is doing.