engels-2024-sae-m10-sparse-subset
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references.md
Created 2026-08-25T02:58:03+00:00
When a 10-feature SAE is trained on a 2D unit circle, dictionary elements spread around the circle and only a sparse subset activates per input (Engels et al., Appendix D).
Summary
Even on a simple 2D circle, a small 10-feature SAE naturally spreads its learned directions around the shape and lights up only a few per point, confirming that sparsity and local coverage emerge from the training objective itself rather than from data complexity. This gives a minimal geometric intuition for why SAE features in larger models tend to be interpretable: each one is tuned to a narrow region of input space rather than being a global blend.