mteb-average-score-dataset-count-weighted
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s2-task-imbalance-tasks-in-mteb-have-a-differ.md
Created 2026-08-25T02:58:19+00:00
MTEB's overall average score is dataset-count-weighted rather than task-weighted, biasing results toward tasks with more datasets (retrieval, classification, clustering).
Summary
MTEB's headline score treats every individual dataset as equal, so the task categories that happen to have more datasets (retrieval, classification, clustering) pull the average in their direction more than they should. This means the overall number overrepresents those task types, making it misleading to treat the single score as a balanced picture of a model's abilities across all task categories.
Dependents
These beliefs depend on this one:
- IN mteb-aggregate-structural-bias — MTEB's aggregate "best model" ranking is structurally biased toward task-breadth coverage rather than semantic quality, because (a) the score is dataset-count-weighted (15 retrieval datasets dominate), (b) summarization is underrepresented by a single dataset, and (c) retrieval/clustering are English-only while other tasks are multilingual.
- OUT mteb-triple-structural-bias — MTEB's aggregate score is triply biased—by dataset count weighting, by language restriction (retrieval/clustering English-only), and by task underrepresentation (summarization = 1 dataset)—making it a measure of readout-breadth coverage rather than of semantic quality