sae-feature-causal-reliability

OUT derived (depth 1)

Created 2026-08-25T03:00:40+00:00

SAE-identified features serve as causally meaningful interpretability primitives, with ablation producing predictable logit-space effects that correlate strongly with downstream behavior.

Justifications

SL — Human raters confirm SAE features are substantially more interpretable than neurons, and the 0.8 ablation-logit correlation validates causal efficacy. However, the claim that "features that fire more are more important" is retracted when the raw-activation caveat is IN, as it shows activation magnitude is a poor proxy for causal contribution—requiring attribution-based (not magnitude-based) reasoning.

Antecedents (all must be IN):

  • IN sae-features-more-interpretable-than-neurons — Autointerpretability analysis (Bills et al. method applied by Cunningham) showed SAE-identified features are substantially more interpretable than individual transformer neurons
  • IN sae-attribution-ablation-0-8-correlation — In the worked example (John says 'I want to be alone' → John feels ___), ablating every active feature yielded a 0.8 correlation with attribution scores, validating attribution as a reasonable proxy.

Unless (any of these IN defeats this justification):

  • IN sae-raw-activation-not-causal-importance — Ranking features by raw activation magnitude does not reliably identify causally important features; a feature that fires on the token 'alone' is distinct from one encoding the concept of wanting solitude.