stsbenchmark-monolingual-english

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s4-additional-modalities-text-embeddings-are.md

Created 2026-08-25T02:58:22+00:00

STSBenchmark in MTEB is a monolingual English Semantic Textual Similarity dataset

Summary

STSBenchmark tests whether a model can judge how close two English sentences are in meaning, and it only uses English. Any ranking or comparison built on this dataset is therefore limited to English-only, similarity-judgment performance and tells you nothing about a model's multilingual ability or other text tasks.