sae-raw-activation-not-causal-importance

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-6.md

Created 2026-08-25T02:58:38+00:00

Ranking features by raw activation magnitude does not reliably identify causally important features; a feature that fires on the token 'alone' is distinct from one encoding the concept of wanting solitude.

Summary

Just because a feature lights up strongly on a particular word doesn't mean it's the feature actually driving the model's reasoning. This matters because it warns that shortcutting interpretability by ranking features on how loudly they fire can mislead you — the feature encoding the word "alone" and the one capturing the deeper idea of desiring solitude are different things, and confusing them means you'd be tracking surface statistics rather than the actual causal structure behind model behavior.

Dependents

These beliefs depend on this one: