saey-paper-rejects-sparse-architecture-approach

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-1.md

Created 2026-08-25T02:57:54+00:00

The Bricken et al. 2023 paper explicitly rejects the 'sparse architecture' approach (encouraging activation sparsity at training time) as insufficient to eliminate polysemanticity.

Summary

The Bricken et al. 2023 paper found that simply training neural networks to fire fewer neurons at once is not enough to prevent a single neuron from carrying multiple unrelated meanings. This matters because it rules out a relatively cheap fix for polysemanticity, meaning the system should not treat sparsity-based regularization as a solution to the core interpretability problem.