superposition-geometry-explains-universality
OUT derived (depth 2)
Created 2026-08-25T03:05:14+00:00 · Reviewed 2026-08-25T04:02:18+00:00
The cross-model universality of feature geometry (SAE features more similar across architectures than within, Park orthogonality validated on both Gemma and LLaMA) is a consequence of superposition: the over-complete compositional basis is determined by the shared semantic grammar of language, making geometric structure an architectural invariant rather than a model-specific artifact.
Justifications
SL — The compositional mechanism (superposition as basis, depth-1) provides the *why*; the observed convergence (depth-1) provides the *that*; together they establish universality as a structural necessity of language composition rather than a training coincidence.
Antecedents (all must be IN):
- OUT superposition-as-compositional-basis — Superposition is the fundamental compositional mechanism in LLMs: the 10–200× over-complete expansion (SAE), the key-value memory structure (ROME's W_fc/W_proj), and the direct-sum space decomposition (Park's polytope+orthogonality) are three independent geometric consequences of the same over-completeness.
- IN multi-model-geometric-convergence — Both the polytope/orthogonality geometry (Park, validated on Gemma-2B and LLaMA-3-8B) and sparse feature structure (SAE, universal across architectures) converge on the finding that transformer representation spaces carry model-independent geometric invariants.
Dependents
These beliefs depend on this one:
- OUT geometric-convergence-as-mathematical-attractor — The cross-model universality of feature geometry combined with its ontological status as a model-independent semantic object implies LLMs are converging to a shared mathematical attractor: the covariance/whitening geometry is the unique fixed point that any differentiable language model must instantiate, not an architectural artifact.
- OUT geometry-ontological-status — The covariance/whitening geometry is not merely a convenient analytical tool but possesses ontological status as a genuine model-independent semantic structure, because three independent lines converge: it is the operational metric for editing and interpretation (depth-3), it is universal across architectures (depth-2), and it converges with externally-validated human-judgment metrics (depth-3).