feature-neighborhood-as-geometric-theorem-instantiation-v2
IN premise
Created 2026-08-25T04:06:47+00:00
SAE decoder-space feature neighborhoods (e.g., concrete features such as the Golden Gate Bridge grouped under more abstract categories like transit infrastructure) provide an empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity mirrors subordination relations of the kind validated by Park's orthogonality theorem across Gemma-2B and LLaMA-3-8B, linking interpretability with the geometric account of semantic space.
Summary
When you look at which interpretable features sit near each other in the model's decoded feature space, the spatial layout isn't arbitrary: specific concepts cluster near their broader categories the same way the model's abstract geometry predicts. This means the names we attach to features are not just a human-friendly overlay; they track the same mathematical structure that governs how the model organizes meaning, so interpretability and the model's internal geometry are two views of one underlying organization rather than two separate stories.
Dependents
These beliefs depend on this one:
- OUT feature-neighborhood-as-geometric-theorem-instantiation — SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.