mteb-msmarco-v2-largest-retrieval
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md
Created 2026-08-25T02:58:23+00:00
MSMARCOv2 is the largest retrieval dataset in MTEB with approximately 138 million train and test samples.
Summary
MSMARCOv2 dwarfs every other retrieval task in the MTEB benchmark, carrying roughly 138 million samples between training and testing. That scale means evaluating or tuning a retrieval model on it is computationally expensive, and its results reflect performance on a volume of data no other MTEB retrieval task can match.