sae-raw-activation-not-causal-importance
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-6.md
Created 2026-08-25T02:58:38+00:00
Ranking features by raw activation magnitude does not reliably identify causally important features; a feature that fires on the token 'alone' is distinct from one encoding the concept of wanting solitude.
Summary
Just because a feature lights up strongly on a particular word doesn't mean it's the feature actually driving the model's reasoning. This matters because it warns that shortcutting interpretability by ranking features on how loudly they fire can mislead you — the feature encoding the word "alone" and the one capturing the deeper idea of desiring solitude are different things, and confusing them means you'd be tracking surface statistics rather than the actual causal structure behind model behavior.
Dependents
These beliefs depend on this one:
- OUT editable-semantic-space — The residual stream, equipped with its covariance-geometric structure, constitutes a well-defined editable semantic space in which knowledge can be read (SAE feature activations, Park polytope coordinates) and written (ROME rank-one value-projection updates) as addressable, independently manipulable units.
- OUT feature-level-editing-reliability — SAE-identified features can serve as interpretable, causally-grounded targets for knowledge editing—specifying edits in semantic feature space rather than raw weight matrices—because the residual stream is the universal substrate and knowledge is locally stored, contingent on feature activations being causally meaningful rather than mere statistical correlates.
- OUT sae-feature-causal-reliability — SAE-identified features serve as causally meaningful interpretability primitives, with ablation producing predictable logit-space effects that correlate strongly with downstream behavior.
- OUT sae-guided-feature-space-editing — SAE-identified features provide a semantically-interpretable coordinate system for specifying knowledge edits—enabling edits to be expressed as feature-space operations (e.g., "suppress feature 34M/31164353 and amplify its neighborhood") that the covariance geometry guarantees map to valid rank-one weight-space operations—thereby bridging the interpretability and editing literatures through the shared second-moment structure.