mteb-aggregate-as-readout-bias

OUT derived (depth 3)

Created 2026-08-25T03:10:25+00:00

MTEB's dataset-count-weighted aggregate score measures readout calibration breadth across task types rather than internal representation quality, because task-specificity is a readout-level phenomenon operating on top of a shared geometric substrate.

Justifications

SL — The structural bias (what is measured) combined with the readout-level nature of task specificity (why it is biased) together yield the insight that the aggregate is a calibration-breadth metric, not a geometry-quality metric.

Antecedents (all must be IN):

  • IN mteb-aggregate-structural-bias — MTEB's aggregate "best model" ranking is structurally biased toward task-breadth coverage rather than semantic quality, because (a) the score is dataset-count-weighted (15 retrieval datasets dominate), (b) summarization is underrepresented by a single dataset, and (c) retrieval/clustering are English-only while other tasks are multilingual.
  • OUT task-specificity-emerges-from-readout — Task-specificity in embedding quality is a readout phenomenon: the internal feature geometry is largely model-independent (convergent across architectures), while MTEB's no-dominant-model result arises because each task's unembedding/projection head selects a different subspace of the same shared geometric structure.