sae-shrinkage-as-finite-resolution-limit
OUT derived (depth 3)
Created 2026-08-25T03:48:00+00:00 · Reviewed 2026-08-25T04:02:18+00:00
The SAE shrinkage problem (under-reconstruction with finite expansion ratios) is the operational signature of finite resolution in the superposition→whitening framework: any finite dictionary size leaves irreducible reconstruction loss because the superposed structure is fundamentally over-complete, and the power-law decrease in loss with compute is the scaling signature of approaching (but never reaching) full resolution.
Justifications
SL — The superposition→whitening necessitation (d2) provides the theoretical reason; the granularity scaling (d2) provides the operational mechanism; the power-law scaling (base) provides the empirical signature. Together they reframe shrinkage from an engineering bug to a geometric necessity of finite resolution.
Antecedents (all must be IN):
- OUT superposition-necessitates-covariance-whitening — Over-complete superposition in the shared residual-stream substrate is the precise structural condition that necessitates covariance/whitening (second-moment projection) as the canonical tool for isolating individual features and performing targeted rank-one edits; without superposition, raw Euclidean geometry would suffice and the entire covariance-geometry framework would be unnecessary.
- OUT sae-granularity-as-superposition-resolution — SAE's empirical granularity scaling (broad categorical features at small run sizes → specific entity features at large run sizes) is the operational resolution of superposition: the 10–200× expansion ratio creates discrete "zoom levels" at which the same underlying covariance geometry manifests as features of different semantic specificity.
- IN sae-scaling-law-power-law-compute — SAE loss decreases approximately as a power law with respect to compute (features × training steps), and training runs for exactly one epoch so steps map linearly to data volume.
Dependents
These beliefs depend on this one:
- OUT mteb-task-specificity-as-shrinkage-readout — MTEB's "no dominant model" result is the readout-side signature of the same superposition that causes SAE shrinkage on the write side: the over-complete representation that prevents any finite SAE expansion ratio from fully decomposing the space also prevents any single embedding model from simultaneously optimizing all task-specific readout directions.