feature-manifolds-higher-dimensional
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity-chunk-7.md
Created 2026-08-25T02:57:55+00:00
Features in transformer networks may not be constrained to one-dimensional directions but may occupy higher-dimensional subspaces called feature manifolds, where a single semantic concept varies across a patch of activation space
Summary
A concept like "dog" in a transformer may not correspond to a single on/off axis but to a small neighborhood of activation patterns, meaning the same idea can light up slightly different combinations of units depending on context. This matters because any technique that treats each feature as a simple single knob will miss the internal variation of a concept, making interpretation, editing, and alignment harder than they would be under the simpler one-direction picture.