sae-shrinkage-problem

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-10.md

Created 2026-08-25T02:58:36+00:00

The SAE shrinkage problem refers to sparse autoencoders under-reconstructing the original activations, with mitigations including finetuning approaches (Wright & Sharkey) and gating activation functions (Rajamanoharan et al.).

Summary

Sparse autoencoders, a key tool for interpreting what AI models are doing internally, tend to produce reconstructions that are systematically weaker than the original signals they are trying to capture. This matters because any analysis built on those reconstructions risks missing or underestimating real features of the model's behavior, and while fixes like finetuning or gating have been proposed, the underlying loss is a known limitation of the interpretability pipeline itself.

Dependents

These beliefs depend on this one: