convergence-is-necessary-not-contingent

OUT derived (depth 7)

Created 2026-08-25T03:13:05+00:00 · Reviewed 2026-08-25T03:24:16+00:00

Cross-model geometric convergence (SAE feature similarity, Park orthogonality) is a logical necessity of superposition in a shared residual stream rather than a contingent empirical coincidence: any system that encodes d concepts in an over-complete m > d basis within a common substrate MUST produce the same covariance structure

Justifications

SL — superposition-as-single-root-cause provides the sufficient condition (superposition is the unique root); sae-universality-across-models and multi-model-geometric-convergence provide convergent empirical evidence from two independent analysis methods (sparse features and polytope geometry). The ALL chain: root cause + two independent empirical confirmations = necessity theorem.

Antecedents (all must be IN):

  • OUT superposition-as-single-root-cause — Superposition is the unique root cause from which the full read/write/editing logical structure follows: over-complete representation necessitates covariance whitening (making geometry well-defined), which in turn explains both the broad-read/narrow-write asymmetry (projection vs. rank-one injection) and the closed triangle linking all three properties—there is no independent second principle needed.
  • IN sae-universality-across-models — SAEs applied to different transformer models produce mostly similar features—more similar to each other than to their own model's neurons—suggesting features reflect data structure rather than architecture
  • IN multi-model-geometric-convergence — Both the polytope/orthogonality geometry (Park, validated on Gemma-2B and LLaMA-3-8B) and sparse feature structure (SAE, universal across architectures) converge on the finding that transformer representation spaces carry model-independent geometric invariants.

Unless (any of these IN defeats this justification):

  • IN convergence-is-necessary-not-contingent-v2 — Cross-model geometric convergence (SAE feature similarity, Park orthogonality) is strongly predicted by the superposition framework: over-complete representation (m > d) in a shared substrate motivates a common covariance structure, providing a unified explanation for why geometric invariants appear across architectures. Empirical convergence (validated on Gemma-2B, LLaMA-3-8B, with SAE features 'mostly similar' across models) is consistent with this theoretical account and argues against a pure contingent coincidence, though the evidence establishes a well-motivated expectation rather than a proven logical necessity for all systems that encode concepts in an over-complete basis.

Dependents

These beliefs depend on this one: