feature-level-editing-reliability
OUT derived (depth 2)
Created 2026-08-25T03:02:15+00:00
SAE-identified features can serve as interpretable, causally-grounded targets for knowledge editing—specifying edits in semantic feature space rather than raw weight matrices—because the residual stream is the universal substrate and knowledge is locally stored, contingent on feature activations being causally meaningful rather than mere statistical correlates.
Justifications
SL — Universal substrate + local storage → features on that substrate are valid editing targets. BUT if raw activation magnitude does not reliably track causal importance (the known SAE finding), then selecting which feature to edit by its activation is unsound. Gate is currently OUT, correctly flagging that feature-level editing lacks a validated causal selection criterion.
Antecedents (all must be IN):
- OUT residual-stream-universal-substrate — SAE (middle-layer residual stream), ROME (mid-layer MLP value projection), and Park (final-layer unembedding) all identify the residual stream at different depths as the primary locus of interpretable geometric structure.
- IN local-storage-distributed-acquisition — Factual knowledge is acquired through distributed corpus exposure (Kandpal's log-linear document-count dependence) but stored in a locally addressable MLP slot (ROME's single-layer FFN edit), revealing a two-phase knowledge pipeline.
Unless (any of these IN defeats this justification):
- IN sae-raw-activation-not-causal-importance — Ranking features by raw activation magnitude does not reliably identify causally important features; a feature that fires on the token 'alone' is distinct from one encoding the concept of wanting solitude.