mteb-triple-structural-bias
OUT derived (depth 1)
Created 2026-08-25T03:13:05+00:00
MTEB's aggregate score is triply biased—by dataset count weighting, by language restriction (retrieval/clustering English-only), and by task underrepresentation (summarization = 1 dataset)—making it a measure of readout-breadth coverage rather than of semantic quality
Justifications
This belief has 3 justifications — it is IN if any one holds.
SL — Each base belief independently identifies a distinct axis of structural bias in MTEB's aggregate. Together they form a complete taxonomy of the bias (task-weighting, language, task-coverage). ANY mode because any single axis is sufficient to invalidate the aggregate as a pure quality signal.
Antecedents (all must be IN):
- IN mteb-average-score-dataset-count-weighted — MTEB's overall average score is dataset-count-weighted rather than task-weighted, biasing results toward tasks with more datasets (retrieval, classification, clustering).
SL — Each base belief independently identifies a distinct axis of structural bias in MTEB's aggregate. Together they form a complete taxonomy of the bias (task-weighting, language, task-coverage). ANY mode because any single axis is sufficient to invalidate the aggregate as a pure quality signal.
Antecedents (all must be IN):
- IN mteb-retrieval-clustering-english-only — MTEB's retrieval and clustering tasks are English-only; multilingual coverage in MTEB exists only for classification, STS, and bitext mining tasks.
SL — Each base belief independently identifies a distinct axis of structural bias in MTEB's aggregate. Together they form a complete taxonomy of the bias (task-weighting, language, task-coverage). ANY mode because any single axis is sufficient to invalidate the aggregate as a pure quality signal.
Antecedents (all must be IN):
- IN mteb-summarization-single-dataset — MTEB includes only one summarization dataset (SummEval, evaluated on CNN/DailyMail), making summarization the most imbalanced task in the benchmark.