mteb-bioss-100-sentence-pairs
IN premise — summaries/2026/08/24/muennighoff-2022-mteb-s1-long-document-datasets-mteb-covers-mul.md
Created 2026-08-25T02:58:18+00:00
The BIOSS biomedical STS dataset in MTEB contains 100 sentence pairs.
Summary
This is a direct observation confirming that the BIOSS biomedical similarity benchmark in the MTEB collection is built from exactly 100 paired sentences. The size matters because it caps how much signal any evaluation on this subset can carry — with only 100 pairs, results are sensitive to individual examples and should be read with caution rather than treated as a broad generalization.