mteb-prior-benchmarks-narrow-scope
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s1-introduction.md
Created 2026-08-25T02:58:18+00:00
SemEval/SentEval cover only STS and classifier-based evaluation, USEB focuses on reranking, and BEIR covers zero-shot retrieval; none provided a unified cross-task embedding evaluation prior to MTEB.
Summary
Before MTEB, the main benchmarks each tested embeddings on just one narrow slice of work, so there was no single yardstick for comparing how well a model performs across different tasks like similarity, classification, and retrieval. This matters because it means earlier claims about which embedding model was "best" were based on incomplete, non-comparable evidence, and any system that relies on those older rankings is working with a fragmented picture of real capability.