mteb-summarization-single-dataset

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s1-long-document-datasets-mteb-covers-mul.md

Created 2026-08-25T02:58:18+00:00

MTEB includes only one summarization dataset (SummEval, evaluated on CNN/DailyMail), making summarization the most imbalanced task in the benchmark.

Summary

Because MTEB tests summarization on just one dataset while every other task category draws on many, a model's summarization score reflects performance on a single narrow set of news articles rather than a broad picture of its ability to summarize. This means the benchmark's summarization number is the least reliable and most brittle of all its task scores, so it should carry less weight when comparing models.

Dependents

These beliefs depend on this one: