mteb-tatoeba-112-languages

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md

Created 2026-08-25T02:58:22+00:00

Tatoeba Bitext Mining in MTEB covers 112 languages.

Summary

The Tatoeba Bitext Mining task in the MTEB benchmark evaluates text embedding models across 112 different languages, meaning results from this task reflect broad multilingual performance rather than just a few major languages. Any conclusions drawn from MTEB scores should be understood as spanning this wide linguistic range.