superposition-features-exceed-dimensionality

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-6.md

Created 2026-08-25T02:57:55+00:00

Superposition hypothesis posits that neural networks encode more features than the dimensionality of their activation space, whereas disentanglement seeks features equal to or fewer than the dimensionality

Summary

Neural networks likely pack far more distinct concepts into their internal representations than the raw number of neurons or dimensions would suggest, meaning each "channel" carries a blend of multiple ideas rather than one clean signal. This matters because it implies the network's inner logic is more compressed and entangled than a simple feature-by-feature reading would reveal, making interpretability harder and forcing any system that reasons about these models to account for hidden, overlapping structure.