engels-2024-sae-m2-circle-both-features-fire
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references.md
Created 2026-08-25T02:58:03+00:00
When a 2-feature SAE is trained to reconstruct a 2D unit circle, both dictionary features must activate simultaneously on every input, violating sparsity (Engels et al., Appendix D).
Summary
A simple two-neuron sparse autoencoder cannot capture the geometry of a circle without both neurons firing at the same time on every input, which breaks the core assumption that most features stay silent. This means SAEs have a structural blind spot for circular or rotational patterns, so interpreters relying on sparsity will miss coordinated two-feature activity that actually carries the signal.