editable-semantic-space
OUT derived (depth 4)
Created 2026-08-25T03:05:14+00:00
The residual stream, equipped with its covariance-geometric structure, constitutes a well-defined editable semantic space in which knowledge can be read (SAE feature activations, Park polytope coordinates) and written (ROME rank-one value-projection updates) as addressable, independently manipulable units.
Justifications
SL — The addressability bound (depth-3) and the superposition→whitening necessity (depth-2) jointly establish a well-defined coordinate system; the raw-activation-not-causal caveat, if IN, breaks the "read" half by showing activation magnitude does not track causal function.
Antecedents (all must be IN):
- OUT geometric-editing-addressability-bound — The covariance-geometry framework defines a precise and minimal addressable space for knowledge editing (rank-one updates to a single MLP value projection), but the combination of superposition and distributed corpus acquisition structurally bounds this to single-fact local corrections—edits cannot create novel multi-hop associations because the target knowledge was never locally consolidated in the first place.
- OUT superposition-necessitates-covariance-whitening — Over-complete superposition in the shared residual-stream substrate is the precise structural condition that necessitates covariance/whitening (second-moment projection) as the canonical tool for isolating individual features and performing targeted rank-one edits; without superposition, raw Euclidean geometry would suffice and the entire covariance-geometry framework would be unnecessary.
Unless (any of these IN defeats this justification):
- IN sae-raw-activation-not-causal-importance — Ranking features by raw activation magnitude does not reliably identify causally important features; a feature that fires on the token 'alone' is distinct from one encoding the concept of wanting solitude.