non-privileged-basis-equivalence-lemma
IN premise — summaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-3.md
Created 2026-08-25T02:57:59+00:00
In a non-privileged basis (e.g., word embedding space), applying a random invertible linear matrix M to the embedding space and M⁻¹ to all downstream weights produces a functionally identical model with a different basis, demonstrating no direction is inherently special.
Summary
Any specific direction you might spot in a word-embedding space, like an axis that appears to encode gender or sentiment, is not a fixed, discoverable feature of the model. You can mathematically re-rotate the entire space and adjust the downstream weights to compensate, producing an identical model, which means no single axis is inherently privileged or independently interpretable.