feature-clumping-downstream-actions
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity-chunk-7.md
Created 2026-08-25T02:57:55+00:00
The most central explanation for feature clumping in activation space is similar downstream actions (output-effect similarity), not merely correlated input activations
Summary
When neurons or directions in a network's activation space tend to cluster together, the main reason is that they push the network toward similar outputs or behaviors, not simply because they happen to light up on similar inputs. This matters because it means the best way to understand what a feature is doing is to trace what it causes downstream, rather than just cataloging which inputs trigger it.