mteb-four-identified-limitations
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-2.md
Created 2026-08-25T02:58:23+00:00
MTEB has 4 identified limitations: no very-long-document datasets, task imbalance in average score, retrieval and clustering are English-only, and no multimodal benchmarks.
Summary
The standard benchmark used to rank text-embedding models has known gaps: it can't measure performance on very long documents, the overall score is skewed toward certain task types, retrieval and clustering are tested in English only, and multimodal inputs are entirely unrepresented. This matters because a model's ranking on that benchmark may not reflect how well it handles the kinds of content your system actually needs to embed and retrieve.