mteb-retrieval-clustering-english-only
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s1-long-document-datasets-mteb-covers-mul.md
Created 2026-08-25T02:58:18+00:00
MTEB's retrieval and clustering tasks are English-only; multilingual coverage in MTEB exists only for classification, STS, and bitext mining tasks.
Summary
If you are evaluating an embedding model for retrieval or clustering work in any language other than English, MTEB scores will not reflect that capability, because those two task types in the benchmark only cover English. Multilingual performance in MTEB is only visible through classification, semantic similarity, and bitext-mining tasks, so a model that ranks well overall could still be weak at, say, French or Japanese document search and the benchmark would not surface that.
Dependents
These beliefs depend on this one:
- IN mteb-aggregate-structural-bias — MTEB's aggregate "best model" ranking is structurally biased toward task-breadth coverage rather than semantic quality, because (a) the score is dataset-count-weighted (15 retrieval datasets dominate), (b) summarization is underrepresented by a single dataset, and (c) retrieval/clustering are English-only while other tasks are multilingual.
- OUT mteb-triple-structural-bias — MTEB's aggregate score is triply biased—by dataset count weighting, by language restriction (retrieval/clustering English-only), and by task underrepresentation (summarization = 1 dataset)—making it a measure of readout-breadth coverage rather than of semantic quality