sae-neighborhood-as-polytope-navigation-v2
IN premise
Created 2026-08-25T04:08:39+00:00
SAE feature neighborhoods (e.g., Golden Gate Bridge → San Francisco → California) provide a concrete interpretable instantiation of the categorical polytope geometry identified by Park et al.: decoder-space proximity among SAE features reflects the same subordination relations that Park et al. formalize as polytope vertex adjacency, and cross-model feature similarity suggests this geometric structure reflects shared data properties rather than being a model-specific architectural artifact.
Summary
When SAE features form chains like a bridge, its city, and its state, the spatial arrangement of those features in decoder space mirrors the same hierarchical part-of structure that Park et al. define mathematically as polytope vertex adjacency. Because this geometry appears across different models, it likely captures real structure in the data rather than an artifact of any single architecture, which makes feature-to-feature relationships a reliable basis for interpretation.
Dependents
These beliefs depend on this one:
- OUT sae-neighborhood-as-polytope-navigation — SAE feature neighborhoods (e.g., Golden Gate Bridge → Alcatraz → San Francisco → California) are the operational navigation algorithm for the categorical polytope geometry: each SAE feature is a polytope vertex, the neighborhood structure is the polytope edge adjacency, and cross-model universality confirms this polytope is a shared semantic object rather than a model-specific artifact.