editing-reliability-under-superposition
OUT derived (depth 4)
Created 2026-08-25T03:07:18+00:00
Rank-one knowledge editing is reliable as a single-fact correction mechanism in superposed models because covariance whitening provides sufficient feature separation to isolate the target association from the background superposition, making the edit direction well-defined and locally confined.
Justifications
SL — Both antecedents are required (superposition→covariance separation AND covariance→bounded editing) to establish reliability. The gate on fluency-loss acknowledges that even when the geometry is correct, the edit can still introduce collateral degradation—separating this quality concern from the scope concerns that gate other editing conclusions.
Antecedents (all must be IN):
- OUT superposition-necessitates-covariance-whitening — Over-complete superposition in the shared residual-stream substrate is the precise structural condition that necessitates covariance/whitening (second-moment projection) as the canonical tool for isolating individual features and performing targeted rank-one edits; without superposition, raw Euclidean geometry would suffice and the entire covariance-geometry framework would be unnecessary.
- OUT geometric-editing-addressability-bound — The covariance-geometry framework defines a precise and minimal addressable space for knowledge editing (rank-one updates to a single MLP value projection), but the combination of superposition and distributed corpus acquisition structurally bounds this to single-fact local corrections—edits cannot create novel multi-hop associations because the target knowledge was never locally consolidated in the first place.
Unless (any of these IN defeats this justification):
- IN rome-entropy-metrics-miss-fluency-loss — Human raters detected subtle fluency losses in ROME's generated output that an entropy-based automatic metric failed to capture, indicating a gap between statistical and perceptual quality assessment.