sae-shrinkage-as-finite-resolution-limit

OUT derived (depth 3)

Created 2026-08-25T03:48:00+00:00 · Reviewed 2026-08-25T04:02:18+00:00

The SAE shrinkage problem (under-reconstruction with finite expansion ratios) is the operational signature of finite resolution in the superposition→whitening framework: any finite dictionary size leaves irreducible reconstruction loss because the superposed structure is fundamentally over-complete, and the power-law decrease in loss with compute is the scaling signature of approaching (but never reaching) full resolution.

Justifications

SL — The superposition→whitening necessitation (d2) provides the theoretical reason; the granularity scaling (d2) provides the operational mechanism; the power-law scaling (base) provides the empirical signature. Together they reframe shrinkage from an engineering bug to a geometric necessity of finite resolution.

Antecedents (all must be IN):

  • OUT superposition-necessitates-covariance-whitening — Over-complete superposition in the shared residual-stream substrate is the precise structural condition that necessitates covariance/whitening (second-moment projection) as the canonical tool for isolating individual features and performing targeted rank-one edits; without superposition, raw Euclidean geometry would suffice and the entire covariance-geometry framework would be unnecessary.
  • OUT sae-granularity-as-superposition-resolution — SAE's empirical granularity scaling (broad categorical features at small run sizes → specific entity features at large run sizes) is the operational resolution of superposition: the 10–200× expansion ratio creates discrete "zoom levels" at which the same underlying covariance geometry manifests as features of different semantic specificity.
  • IN sae-scaling-law-power-law-compute — SAE loss decreases approximately as a power law with respect to compute (features × training steps), and training runs for exactly one epoch so steps map linearly to data volume.

Dependents

These beliefs depend on this one: