mteb-simultaneously-measures-read-and-edit-addressability
OUT derived (depth 8)
Created 2026-08-25T03:45:56+00:00 ยท Reviewed 2026-08-25T04:02:18+00:00
MTEB's "no dominant model" observation is simultaneously a statement about readout diversity AND edit-addressability diversity: different embedding models optimize different directions in the same shared editable geometric space, making the benchmark a dual measure of read-quality and edit-coverage.
Justifications
SL — The no-dominant-model result (readout diversity) is reinterpreted through the lens of evaluation-geometry=editing-geometry: if the evaluation pipeline IS the editing coordinate system, then model differentiation on MTEB reflects different "addressable subspaces" for editing, not merely different readout calibrations. Both antecedents are load-bearing: without the geometric unification, MTEB differences are just calibration; without MTEB's breadth, the editing-coordinate claim lacks empirical grounding.
Antecedents (all must be IN):
- OUT mteb-no-dominant-model-as-geometric-signature โ MTEB's observation that no single embedding model dominates all eight tasks is the expected geometric signature of a shared canonical semantic space probed by task-specific linear readouts, and the same second-moment structure that generates this task-specificity pattern is precisely what makes rank-one knowledge editing possible.
- OUT evaluation-geometry-is-editing-coordinate-system โ The optimal evaluation metric for an embedding model (cosine/Spearman pipeline) is simultaneously the optimal coordinate system for specifying knowledge edits, because both are readouts of the same universal covariance geometry
Dependents
These beliefs depend on this one:
- OUT mteb-task-diversity-as-geometric-anisotropy โ MTEB's "no dominant model" result across 8 task types is the expected geometric signature of anisotropic readout in a shared semantic space: each task type (STS, retrieval, clustering, classification) probes a different projection direction, and task-specificity is the *predicted* outcome of geometric anisotropy, not a benchmark failure or model deficiency.