eh2022-two-model-disentangled-superposed
IN premise — summaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-10.md
Created 2026-08-25T02:57:58+00:00
The two-model decomposition for studying superposition assigns input and output layers to a disentangled (privileged-basis) representation and the hidden layer to a lower-dimensional superposed representation
Summary
This is a controlled experimental setup for studying how neural networks pack more features than they have neurons: the input and output sides are given a clean, one-feature-per-slot representation, while the hidden layer is forced to represent those same features in a smaller, overlapping space. The "so what" is that it gives researchers a way to isolate and measure exactly how superposition works in the middle of a network, rather than trying to observe it in a fully uncontrolled model.