sae-neighborhood-as-polytope-navigation
OUT derived (depth 5)
Created 2026-08-25T03:48:00+00:00 · Reviewed 2026-08-25T04:02:18+00:00
SAE feature neighborhoods (e.g., Golden Gate Bridge → Alcatraz → San Francisco → California) are the operational navigation algorithm for the categorical polytope geometry: each SAE feature is a polytope vertex, the neighborhood structure is the polytope edge adjacency, and cross-model universality confirms this polytope is a shared semantic object rather than a model-specific artifact.
Justifications
SL — The polytope geometry (base) provides the theorem; the feature neighborhood (d4) provides the empirical instantiation; cross-model universality (base) confirms it is not model-specific. Together they show SAE neighborhoods *are* the polytope navigation, making interpretability and geometry the same operation.
Antecedents (all must be IN):
- OUT feature-neighborhood-as-geometric-theorem-instantiation — SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.
- IN sae-universality-across-models — SAEs applied to different transformer models produce mostly similar features—more similar to each other than to their own model's neurons—suggesting features reflect data structure rather than architecture
- IN park-2025-iclr-categorical-polytope-geometry — Park et al. (ICLR 2025) prove that categorical concepts in LLM representation spaces are geometrically represented as polytopes (convex hulls of vertex vectors), with 'natural' concepts forming (k−1)-simplices.
Unless (any of these IN defeats this justification):
- IN sae-neighborhood-as-polytope-navigation-v2 — SAE feature neighborhoods (e.g., Golden Gate Bridge → San Francisco → California) provide a concrete interpretable instantiation of the categorical polytope geometry identified by Park et al.: decoder-space proximity among SAE features reflects the same subordination relations that Park et al. formalize as polytope vertex adjacency, and cross-model feature similarity suggests this geometric structure reflects shared data properties rather than being a model-specific architectural artifact.
Dependents
These beliefs depend on this one:
- OUT geometric-closed-loop-eval-edit-navigate — The MTEB evaluation coordinate system, ROME's rank-one editing, and SAE feature navigation form a single closed geometric loop: evaluation identifies the whitened directions to read, ROME modifies one whitened direction to write, and SAE neighborhood traversal navigates between whitened directions—each operation is a different linear functional on the same covariance-geometric space, unified by the Riesz map.
- OUT sae-neighborhood-as-directsum-navigation — The SAE feature neighborhood (Golden Gate Bridge → Alcatraz → San Francisco → California) is not merely a navigation path within a single categorical polytope but the operational traversal algorithm for the full direct-sum decomposition of the semantic space: it simultaneously navigates the categorical subspace (discrete concepts) and the hierarchical orthogonal subspace (WordNet parent-child structure), validated across Gemma-2B and LLaMA-3-8B.