msmarco-dev-split
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s8-tasks.md
Created 2026-08-25T02:58:21+00:00
MS-MARCO is the one retrieval dataset in MTEB evaluated on its dev split rather than test split, following BEIR convention (Thakur et al., 2021)
Summary
When you compare retrieval results across the MTEB benchmark, MS-MARCO is the odd one out: it scores against its dev split while every other retrieval dataset scores against its test split, because that's how the original BEIR paper set it up. This matters because dev splits are more likely to have leaked into model training, so MS-MARCO scores may look artificially better than the other retrieval benchmarks and are not directly comparable to them.