local-mlp-editing-principle

OUT derived (depth 3)

Created 2026-08-25T03:03:22+00:00

The MLP-as-key-value-memory structure provides a principled, minimal-intervention editing mechanism for locally-stored factual (entity-relation-object) associations, with the edit's efficacy and specificity guaranteed by the rank-one update's geometric isolation in the covariance-projected key space.

Justifications

SL — Local storage establishes the target exists in MLP; covariance geometry establishes the edit mechanism is well-defined and isolated; the gate on ROME's scope limitation (factual associations only, not logical/spatial/numerical) correctly bounds the claim to its valid domain

Antecedents (all must be IN):

  • IN local-storage-distributed-acquisition — Factual knowledge is acquired through distributed corpus exposure (Kandpal's log-linear document-count dependence) but stored in a locally addressable MLP slot (ROME's single-layer FFN edit), revealing a two-phase knowledge pipeline.
  • OUT covariance-geometry-unifies-analysis-and-editing — The mathematically principled framework for both interpreting (SAE feature extraction, Park polytope analysis) and modifying (ROME rank-one edits) LLM representations is second-moment covariance geometry applied to the residual stream, since C = KKᵀ whitening defines the canonical coordinate system in which all three operations become linear algebra on the same substrate.

Unless (any of these IN defeats this justification):

  • IN rome-scope-limitation — ROME and Causal Tracing address factual (entity-relation-object) associations only; logical, spatial, and numerical knowledge are explicitly out of scope.