mteb-total-datasets-count

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s8-tasks.md

Created 2026-08-25T02:58:20+00:00

MTEB contains 58 constituent datasets, of which 56 are English subsets used in the main results table

Summary

MTEB is built from 58 individual benchmark datasets, and the main published results table reports on the 56 English-language subsets among them. This count anchors every downstream claim about what MTEB covers and how representative its headline scores are.