mteb-arxiv-clustering-three-splits
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md
Created 2026-08-25T02:58:22+00:00
arXiv clustering in MTEB uses three split strategies: main category (coarse), secondary category within main (fine-grained), and secondary category across all (multi-scale).
Summary
When MTEB clusters arXiv papers, it doesn't just group them one way. It tests three different cuts of the same data: broad top-level topics, fine-grained subtopics nested under each broad topic, and the same subtopics pulled together across all broad topics. This matters because it means a model is being evaluated on its ability to capture structure at multiple granularities, not just a single flat categorization, which is a harder and more realistic test of how well it understands relationships between documents.