three-discrete-feature-learning-regimes

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-5.md

Created 2026-08-25T02:58:00+00:00

Neural network feature learning exhibits three discrete regimes (not learned, superposition, dedicated orthogonal dimension) with sharp discontinuous transitions between them.

Summary

Features in a neural network don't learn on a smooth gradient; they snap between three distinct states — a feature is either absent, crammed into shared space with other features, or given its own clean dedicated representation. This matters because it means feature learning behaves more like flipping a switch than dialing a knob: small changes in training conditions can trigger abrupt reorganizations of how information is stored, rather than gradual improvements.