feature-neighborhood-as-geometric-theorem-instantiation
OUT derived (depth 4)
Created 2026-08-25T03:10:25+00:00 · Reviewed 2026-08-25T04:02:18+00:00
SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.
Justifications
SL — The empirical SAE neighborhoods (what is observed) combined with the covariance-geometric framework (what explains them) together establish that interpretable features are theorem-level predictions, not ad hoc discoveries.
Antecedents (all must be IN):
- OUT feature-hierarchy-empirical-validation — The hierarchical feature structure predicted geometrically by Park's orthogonality theorem (parent⊥child−parent) is empirically instantiated in SAE decoder-space neighborhoods: abstract/general features (transit infrastructure) subsume concrete/specific features (Golden Gate Bridge, Alcatraz), and this neighborhood hierarchy mirrors the WordNet synset hierarchy that Park validates across Gemma-2B and LLaMA-3-8B.
- OUT covariance-geometry-as-operational-semantic-space — The covariance/whitening geometry (second-moment matrices) is the operational definition of semantic coordinate space in LLMs: it simultaneously parameterises feature interpretation (SAE decoder space, Park polytopes), similarity evaluation (cosine→Spearman pipeline), and knowledge modification (ROME rank-one updates), and this structure converges across model families.
Unless (any of these IN defeats this justification):
- IN feature-neighborhood-as-geometric-theorem-instantiation-v2 — SAE decoder-space feature neighborhoods (e.g., concrete features such as the Golden Gate Bridge grouped under more abstract categories like transit infrastructure) provide an empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity mirrors subordination relations of the kind validated by Park's orthogonality theorem across Gemma-2B and LLaMA-3-8B, linking interpretability with the geometric account of semantic space.
Dependents
These beliefs depend on this one:
- OUT geometry-as-universal-semantic-currency — The covariance geometry is the single operational definition of "meaning" in LLMs, simultaneously determining what can be measured (evaluation via cosine/Spearman), what can be modified (rank-one editing via C⁻¹k*), what converges across architectures (SAE/Park universality), and what is hierarchically structured (feature neighborhoods as theorem instantiations)—making it a model-independent semantic currency rather than an architecture-specific artifact.
- OUT sae-guided-feature-space-editing — SAE-identified features provide a semantically-interpretable coordinate system for specifying knowledge edits—enabling edits to be expressed as feature-space operations (e.g., "suppress feature 34M/31164353 and amplify its neighborhood") that the covariance geometry guarantees map to valid rank-one weight-space operations—thereby bridging the interpretability and editing literatures through the shared second-moment structure.
- OUT sae-neighborhood-as-polytope-navigation — SAE feature neighborhoods (e.g., Golden Gate Bridge → Alcatraz → San Francisco → California) are the operational navigation algorithm for the categorical polytope geometry: each SAE feature is a polytope vertex, the neighborhood structure is the polytope edge adjacency, and cross-model universality confirms this polytope is a shared semantic object rather than a model-specific artifact.
- OUT sae-resolution-of-space-decomposition — SAE's expansion-ratio scaling (broad→specific features) is the operational resolution mechanism for the algebraic space decomposition: coarse SAEs resolve the polytope hulls (categorical structure), fine SAEs resolve individual vertices (entity-level features), and the neighborhood adjacency graph is the polytope edge structure visible at whichever resolution is chosen.