mteb-average-score-dataset-count-weighted

IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s2-task-imbalance-tasks-in-mteb-have-a-differ.md

Created 2026-08-25T02:58:19+00:00

MTEB's overall average score is dataset-count-weighted rather than task-weighted, biasing results toward tasks with more datasets (retrieval, classification, clustering).

Summary

MTEB's headline score treats every individual dataset as equal, so the task categories that happen to have more datasets (retrieval, classification, clustering) pull the average in their direction more than they should. This means the overall number overrepresents those task types, making it misleading to treat the single score as a balanced picture of a model's abilities across all task categories.

Dependents

These beliefs depend on this one: