mteb-massive-51-languages

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md

Created 2026-08-25T02:58:22+00:00

The Massive Intent/Scenario dataset in MTEB covers 51 typologically diverse languages.

Summary

This means the benchmark can evaluate whether a text-embedding model actually works across wildly different language structures and families, not just a handful of closely related ones. It sets a high bar for claiming multilingual performance, since 51 typologically diverse languages represent a genuinely broad cross-section of human linguistic variety.