mteb-task-specificity-as-shrinkage-readout

OUT derived (depth 6)

Created 2026-08-25T03:50:14+00:00 · Reviewed 2026-08-25T04:02:18+00:00

MTEB's "no dominant model" result is the readout-side signature of the same superposition that causes SAE shrinkage on the write side: the over-complete representation that prevents any finite SAE expansion ratio from fully decomposing the space also prevents any single embedding model from simultaneously optimizing all task-specific readout directions.

Justifications

SL — Both phenomena are "finite resolution" artifacts of superposition: SAE under-reconstructs at finite m/d (shrinkage), and embedding models under-optimize at finite capacity (task-specificity). The key emergent claim is that these are *the same geometric constraint* viewed from the decoder side (write) and the readout side (read), not independent limitations.

Antecedents (all must be IN):

  • OUT mteb-no-dominant-model-as-geometric-signature — MTEB's observation that no single embedding model dominates all eight tasks is the expected geometric signature of a shared canonical semantic space probed by task-specific linear readouts, and the same second-moment structure that generates this task-specificity pattern is precisely what makes rank-one knowledge editing possible.
  • OUT sae-shrinkage-as-finite-resolution-limit — The SAE shrinkage problem (under-reconstruction with finite expansion ratios) is the operational signature of finite resolution in the superposition→whitening framework: any finite dictionary size leaves irreducible reconstruction loss because the superposed structure is fundamentally over-complete, and the power-law decrease in loss with compute is the scaling signature of approaching (but never reaching) full resolution.
  • OUT task-specificity-emerges-from-readout — Task-specificity in embedding quality is a readout phenomenon: the internal feature geometry is largely model-independent (convergent across architectures), while MTEB's no-dominant-model result arises because each task's unembedding/projection head selects a different subspace of the same shared geometric structure.