mteb-summarization-single-dataset
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s1-long-document-datasets-mteb-covers-mul.md
Created 2026-08-25T02:58:18+00:00
MTEB includes only one summarization dataset (SummEval, evaluated on CNN/DailyMail), making summarization the most imbalanced task in the benchmark.
Summary
Because MTEB tests summarization on just one dataset while every other task category draws on many, a model's summarization score reflects performance on a single narrow set of news articles rather than a broad picture of its ability to summarize. This means the benchmark's summarization number is the least reliable and most brittle of all its task scores, so it should carry less weight when comparing models.
Dependents
These beliefs depend on this one:
- IN mteb-aggregate-structural-bias — MTEB's aggregate "best model" ranking is structurally biased toward task-breadth coverage rather than semantic quality, because (a) the score is dataset-count-weighted (15 retrieval datasets dominate), (b) summarization is underrepresented by a single dataset, and (c) retrieval/clustering are English-only while other tasks are multilingual.
- OUT mteb-triple-structural-bias — MTEB's aggregate score is triply biased—by dataset count weighting, by language restriction (retrieval/clustering English-only), and by task underrepresentation (summarization = 1 dataset)—making it a measure of readout-breadth coverage rather than of semantic quality