l1-regularization-reduces-superposition

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-11.md

Created 2026-08-25T02:57:59+00:00

Adding L1 regularization (λ||h||₁) to the loss function in the superposition toy model kills features below an importance threshold, particularly non-basis-aligned ones, thereby reducing superposition.

Summary

In the toy superposition model, applying L1 regularization acts like a sparsifying filter that zeros out small or poorly-aligned features, forcing the model to keep only its most important, cleanly separated ones. This matters because it gives the system a simple, well-understood knob to control how much information is packed into shared dimensions, trading representational capacity for clearer and more interpretable structure.