three-discrete-feature-learning-regimes
IN premise — summaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-5.md
Created 2026-08-25T02:58:00+00:00
Neural network feature learning exhibits three discrete regimes (not learned, superposition, dedicated orthogonal dimension) with sharp discontinuous transitions between them.
Summary
Features in a neural network don't learn on a smooth gradient; they snap between three distinct states — a feature is either absent, crammed into shared space with other features, or given its own clean dedicated representation. This matters because it means feature learning behaves more like flipping a switch than dialing a knob: small changes in training conditions can trigger abrupt reorganizations of how information is stored, rather than gradual improvements.