nanda-decoder-weight-sparsity-distribution

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-8.md

Created 2026-08-25T02:57:56+00:00

Nanda's SAE replication found that 4% of decoder weights are well explained by a single neuron, 4% by 2-10 neurons, and 92% are dense across many neurons in the neuron basis

Summary

When people try to express SAE feature directions back in terms of the network's original neurons, the overwhelming majority (92%) smear across many neurons rather than lining up with one or a few. This means most "features" the sparse autoencoder discovers are not simple single-neuron concepts but distributed, multi-neuron patterns, which makes them harder to interpret as clean, localized ideas about what the network is doing.