superposition-necessitates-covariance-whitening
OUT derived (depth 2)
Created 2026-08-25T03:03:22+00:00 · Reviewed 2026-08-25T04:02:18+00:00
Over-complete superposition in the shared residual-stream substrate is the precise structural condition that necessitates covariance/whitening (second-moment projection) as the canonical tool for isolating individual features and performing targeted rank-one edits; without superposition, raw Euclidean geometry would suffice and the entire covariance-geometry framework would be unnecessary.
Justifications
SL — Superposition creates the interference problem, covariance geometry is the specific algebraic solution to that problem, and the residual stream is the shared substrate where both co-occur; retracting any one breaks the logical chain
Antecedents (all must be IN):
- OUT superposition-as-compositional-basis — Superposition is the fundamental compositional mechanism in LLMs: the 10–200× over-complete expansion (SAE), the key-value memory structure (ROME's W_fc/W_proj), and the direct-sum space decomposition (Park's polytope+orthogonality) are three independent geometric consequences of the same over-completeness.
- IN covariance-geometry-as-canonical-tool — Independent lines of work (ROME's key-space projection and Park's unembedding whitening) converge on using empirical second-moment matrices to define the "correct" inner product for reasoning about transformer representations.
- OUT residual-stream-universal-substrate — SAE (middle-layer residual stream), ROME (mid-layer MLP value projection), and Park (final-layer unembedding) all identify the residual stream at different depths as the primary locus of interpretable geometric structure.
Unless (any of these IN defeats this justification):
- IN superposition-necessitates-covariance-whitening-v2 — Over-complete superposition in the residual stream (identified as a locus of interpretable structure across multiple depths) is a primary structural condition that motivates the use of empirical second-moment matrices as a natural inner product for reasoning about individual features and performing targeted interventions; independent lines of work (ROME's key-space projection, Park's unembedding whitening) converge on this covariance-geometry approach as a practical framework for feature manipulation in over-complete representation spaces.
Dependents
These beliefs depend on this one:
- OUT edit-complexity-is-geometrically-necessary — The O(D²) cost of a rank-one knowledge edit is a fundamental lower bound imposed by the geometry of superposition, not an implementation artifact: any edit that preserves the covariance-geometric structure of the residual stream must operate in the whitened D-dimensional subspace, incurring at least O(D²) parameter modification
- OUT editable-semantic-space — The residual stream, equipped with its covariance-geometric structure, constitutes a well-defined editable semantic space in which knowledge can be read (SAE feature activations, Park polytope coordinates) and written (ROME rank-one value-projection updates) as addressable, independently manipulable units.
- OUT editing-reliability-under-superposition — Rank-one knowledge editing is reliable as a single-fact correction mechanism in superposed models because covariance whitening provides sufficient feature separation to isolate the target association from the background superposition, making the edit direction well-defined and locally confined.
- OUT rome-edit-as-partial-whitening — A rank-one ROME weight update is operationally equivalent to a local, single-direction whitening of the residual stream: the C⁻¹k* projection in the update formula performs precisely the covariance-normalization that superposition necessitates, but confined to one key direction—making each edit a "partial whitening" that corrects one superposed feature without disturbing the orthogonal complement.
- OUT sae-shrinkage-as-finite-resolution-limit — The SAE shrinkage problem (under-reconstruction with finite expansion ratios) is the operational signature of finite resolution in the superposition→whitening framework: any finite dictionary size leaves irreducible reconstruction loss because the superposed structure is fundamentally over-complete, and the power-law decrease in loss with compute is the scaling signature of approaching (but never reaching) full resolution.
- OUT superposition-covariance-editability-triangle — Superposition, covariance whitening, and rank-one editability form a closed logical triangle in which each property necessitates the others: over-complete superposition requires covariance separation for feature addressability, covariance separation defines the geometric space in which rank-one updates are well-defined, and the boundedness of rank-one editing confirms the addressable space is finite.