bijective-relaxation-decreases-mutual-knn-alignment

IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references-chunk-2.md

Created 2026-08-24T17:11:00+00:00

Relaxing the one-to-one image-text correspondence assumption consistently decreases mutual kNN alignment: on non-synthetic WIT data, 7.1% of captions map to >1 image and 24.6% of images map to >1 caption; on synthetic CycleReward data (11 captions/image or 12 images/caption), mutual kNN still decreases even with a generous 'same source' matching criterion.

Summary

Allowing multiple captions to share one image, or one caption to describe multiple images, measurably degrades how well image and text representations align with each other, and this holds whether you look at real-world data or deliberately constructed many-to-many pairs. In practice, this means the assumption that every image has exactly one "correct" caption is load-bearing for vision-language model quality; benchmarks and training pipelines that quietly enforce that one-to-one pairing may be hiding a real tension in how these models encode cross-modal similarity.