sae-shrinkage-problem
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-10.md
Created 2026-08-25T02:58:36+00:00
The SAE shrinkage problem refers to sparse autoencoders under-reconstructing the original activations, with mitigations including finetuning approaches (Wright & Sharkey) and gating activation functions (Rajamanoharan et al.).
Summary
Sparse autoencoders, a key tool for interpreting what AI models are doing internally, tend to produce reconstructions that are systematically weaker than the original signals they are trying to capture. This matters because any analysis built on those reconstructions risks missing or underestimating real features of the model's behavior, and while fixes like finetuning or gating have been proposed, the underlying loss is a known limitation of the interpretability pipeline itself.
Dependents
These beliefs depend on this one:
- OUT superposition-unifies-data-and-representation-redundancy — Over-complete superposition is the single structural cause of both the data-side redundancy (5× correlated corpora yield only marginal gains, Spearman 0.87–0.97 inter-correlation) and the representation-side redundancy (SAE dead features at 2–48%, reconstruction shrinkage): in both cases the available "capacity" exceeds the unique information it must encode.