feature-dimensionality-discrete-jumps-training

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-7.md

Created 2026-08-25T02:58:00+00:00

Feature dimensionalities in toy superposition models exhibit discrete jumps ('energy level jumps') during training rather than smooth transitions, despite continuous gradient-based optimization.

Summary

In small models designed to study how neural networks pack multiple concepts into the same parameters, the number of representable concepts doesn't grow gradually as training proceeds. Instead, it snaps to new levels in sudden steps, which is surprising because the training process itself is smooth and continuous. This implies that feature learning is inherently lumpy, not a slow linear ramp, and any system relying on these models should expect abrupt capability changes rather than predictable intermediate states.