vec2vec-datasets-nq-2m-tweettopic-8192-mimic-8192
IN premise — summaries/2026/08/24/jha-2025-vec2vec-s2-problem-formulation-unsupervised-embedding-translation.md
Created 2026-08-24T17:10:57+00:00
vec2vec is evaluated on Natural Questions (2M train / 65,536 eval, 1M 64-token sequences per source), TweetTopic (8,192 records, 19 topics), and MIMIC-III pseudo-re-identified (8,192 records, 2,673 MedCAT labels), with training models spanning four size categories, five transformer backbones, and two output dimensionalities.
Summary
This records the specific experimental conditions under which vec2vec was tested: three datasets of very different sizes and domains (large-scale QA, small topic classification, and medical records), combined with a wide range of model sizes, architectures, and output dimensions. It matters because any performance claims about vec2vec in this system are only valid within these exact boundaries, and anyone comparing results against other work needs to know precisely what was measured and how.