sae-neighborhood-as-polytope-navigation

OUT derived (depth 5)

Created 2026-08-25T03:48:00+00:00 · Reviewed 2026-08-25T04:02:18+00:00

SAE feature neighborhoods (e.g., Golden Gate Bridge → Alcatraz → San Francisco → California) are the operational navigation algorithm for the categorical polytope geometry: each SAE feature is a polytope vertex, the neighborhood structure is the polytope edge adjacency, and cross-model universality confirms this polytope is a shared semantic object rather than a model-specific artifact.

Justifications

SL — The polytope geometry (base) provides the theorem; the feature neighborhood (d4) provides the empirical instantiation; cross-model universality (base) confirms it is not model-specific. Together they show SAE neighborhoods *are* the polytope navigation, making interpretability and geometry the same operation.

Antecedents (all must be IN):

  • OUT feature-neighborhood-as-geometric-theorem-instantiation — SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.
  • IN sae-universality-across-models — SAEs applied to different transformer models produce mostly similar features—more similar to each other than to their own model's neurons—suggesting features reflect data structure rather than architecture
  • IN park-2025-iclr-categorical-polytope-geometry — Park et al. (ICLR 2025) prove that categorical concepts in LLM representation spaces are geometrically represented as polytopes (convex hulls of vertex vectors), with 'natural' concepts forming (k−1)-simplices.

Unless (any of these IN defeats this justification):

  • IN sae-neighborhood-as-polytope-navigation-v2 — SAE feature neighborhoods (e.g., Golden Gate Bridge → San Francisco → California) provide a concrete interpretable instantiation of the categorical polytope geometry identified by Park et al.: decoder-space proximity among SAE features reflects the same subordination relations that Park et al. formalize as polytope vertex adjacency, and cross-model feature similarity suggests this geometric structure reflects shared data properties rather than being a model-specific architectural artifact.

Dependents

These beliefs depend on this one: