vec2vec-data-efficiency-50k-near-1m
IN premise — summaries/2026/08/24/jha-2025-vec2vec-s7-ablations.md
Created 2026-08-24T17:10:58+00:00
50K training embeddings yield cos 0.74, within 0.01 of the 1M-embedding result (cos 0.75), on the gte→gtr NQ 8192-record evaluation
Summary
You can map one embedding space to another using just 50,000 training pairs and land almost as close to the target (cosine 0.74) as you would with a full million pairs (cosine 0.75), on a standard natural-questions evaluation. In practical terms, this means the mapping saturates fast: spending twenty times more data buys you only a negligible bump in quality, so the system can rely on small, cheap training sets for this kind of cross-space translation.