mteb-sts-evaluation-metric
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-sR-references-chunk-1.md
Created 2026-08-25T02:58:23+00:00
STS evaluation in MTEB correlates model-predicted similarity scores with human-annotated ratings using Spearman or Pearson correlation.
Summary
The STS task in the MTEB benchmark grades a model by checking whether its predicted similarity scores for pairs of texts track the way humans actually rated those same pairs, using a rank-correlation statistic like Spearman or Pearson. This means the system defines "good performance" as agreement with human intuition about how similar two passages are, rather than as matching some fixed numeric target.