mteb-massive-51-languages
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md
Created 2026-08-25T02:58:22+00:00
The Massive Intent/Scenario dataset in MTEB covers 51 typologically diverse languages.
Summary
This means the benchmark can evaluate whether a text-embedding model actually works across wildly different language structures and families, not just a handful of closely related ones. It sets a high bar for claiming multilingual performance, since 51 typologically diverse languages represent a genuinely broad cross-section of human linguistic variety.