feature-neighborhood-as-geometric-theorem-instantiation

OUT derived (depth 4)

Created 2026-08-25T03:10:25+00:00 · Reviewed 2026-08-25T04:02:18+00:00

SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.

Justifications

SL — The empirical SAE neighborhoods (what is observed) combined with the covariance-geometric framework (what explains them) together establish that interpretable features are theorem-level predictions, not ad hoc discoveries.

Antecedents (all must be IN):

  • OUT feature-hierarchy-empirical-validation — The hierarchical feature structure predicted geometrically by Park's orthogonality theorem (parent⊥child−parent) is empirically instantiated in SAE decoder-space neighborhoods: abstract/general features (transit infrastructure) subsume concrete/specific features (Golden Gate Bridge, Alcatraz), and this neighborhood hierarchy mirrors the WordNet synset hierarchy that Park validates across Gemma-2B and LLaMA-3-8B.
  • OUT covariance-geometry-as-operational-semantic-space — The covariance/whitening geometry (second-moment matrices) is the operational definition of semantic coordinate space in LLMs: it simultaneously parameterises feature interpretation (SAE decoder space, Park polytopes), similarity evaluation (cosine→Spearman pipeline), and knowledge modification (ROME rank-one updates), and this structure converges across model families.

Unless (any of these IN defeats this justification):

  • IN feature-neighborhood-as-geometric-theorem-instantiation-v2 — SAE decoder-space feature neighborhoods (e.g., concrete features such as the Golden Gate Bridge grouped under more abstract categories like transit infrastructure) provide an empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity mirrors subordination relations of the kind validated by Park's orthogonality theorem across Gemma-2B and LLaMA-3-8B, linking interpretability with the geometric account of semantic space.

Dependents

These beliefs depend on this one: