mteb-zero-shot-evaluation-datasets
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references.md
Created 2026-08-25T02:58:23+00:00
Several MTEB datasets have no train split (Train=0), including ArxivClustering, STS12–STS17, and most retrieval benchmarks, making them zero-shot evaluations.
Summary
Because these benchmark datasets contain only test examples and no training data, any model score reported on them reflects pure out-of-the-box generalization from pre-training alone, not task-specific learning. This matters because it makes the scores a cleaner measure of a model's inherent ability, but it also means you can't use them to compare fine-tuning strategies or to check whether a model actually learned from examples on that task.