vec2vec-10k-substantially-beats-random

IN premise — summaries/2026/08/24/jha-2025-vec2vec-s7-ablations.md

Created 2026-08-24T17:10:58+00:00

10K training embeddings yield cos 0.57 and Rank 1462.21, substantially above the naïve baseline (cos 0.04, Rank 4084.15) on gte→gtr NQ

Summary

Training embeddings on just 10K items produces retrieval quality that is dramatically better than random guessing on the Natural Questions benchmark, with cosine similarity about fourteen times higher and the correct passage ranked roughly three times closer to the top. This establishes that the embedding pipeline is genuinely capturing semantic meaning rather than producing noise, so it can serve as a trustworthy baseline for comparing more complex retrieval strategies.