sbert-wiki-sec-triplet-data-size
IN premise — summaries/2026/08/24/reimers-2019-sentence-bert-s5-evaluation-senteval.md
Created 2026-08-25T02:58:29+00:00
Wikipedia section triplet training uses ~1.8M training triplets with 222,957 held-out test triplets from distinct articles, evaluated by rank-order accuracy.
Summary
This is a factual record of the dataset used in a Wikipedia-based embedding training run: roughly 1.8 million triplets for training and about 223,000 held-out triplets from completely different articles, scored by how well the model ranks the correct match against alternatives. It matters because any performance claim about that model should be read against this specific scale, this clean train/test split, and this ranking-style metric rather than a simpler binary hit rate.