sae-resolution-of-space-decomposition

OUT derived (depth 6)

Created 2026-08-25T03:47:59+00:00

SAE's expansion-ratio scaling (broad→specific features) is the operational resolution mechanism for the algebraic space decomposition: coarse SAEs resolve the polytope hulls (categorical structure), fine SAEs resolve individual vertices (entity-level features), and the neighborhood adjacency graph is the polytope edge structure visible at whichever resolution is chosen.

Justifications

SL — The space decomposition (d5) gives the algebraic structure; SAE granularity (d2) gives the operational lens; feature neighborhoods (d4) give the visible geometry. Together they show SAE is not merely "an interpretability tool" but the *resolution operator* for the full algebraic decomposition.

Antecedents (all must be IN):

  • OUT sae-granularity-as-superposition-resolution — SAE's empirical granularity scaling (broad categorical features at small run sizes → specific entity features at large run sizes) is the operational resolution of superposition: the 10–200× expansion ratio creates discrete "zoom levels" at which the same underlying covariance geometry manifests as features of different semantic specificity.
  • OUT space-decomposition-under-superposition — The full LLM semantic space admits a clean algebraic decomposition into categorical polytope subspaces (discrete concepts) and hierarchical orthogonality subspaces (graded taxonomic structure) within the covariance-geometric framework, providing a complete account of how discrete and graded meaning coexist in a single over-complete vector space
  • OUT feature-neighborhood-as-geometric-theorem-instantiation — SAE feature neighborhood structure (e.g., Golden Gate Bridge → San Francisco → California) is the concrete empirical instantiation of the covariance-geometric semantic space at the interpretable level: decoder-space proximity reflects the same subordination relations predicted by Park's orthogonality theorem, unifying interpretability with geometric theory.