many-to-many-correspondence-reduces-measured-alignment

IN premise — summaries/2026-08-24/koepke-2026-back-into-cave-s3-experimental-setup-chunk-1.md

Created 2026-08-24T17:11:00+00:00

Densifying one modality in the CycleReward dataset (11 captions per image for I2T, 12 images per caption for T2I) consistently decreases mutual kNN scores for both k=1 and k=10, demonstrating that relaxing the bijective one-to-one pairing assumption further reduces measured cross-modal alignment even when retrieved neighbors remain semantically valid.

Summary

Adding more captions to each image and more images to each caption in the CycleReward dataset makes the measured text-to-image alignment score drop, even though the matched pairs are still semantically reasonable. This matters because it shows the alignment metric is sensitive to how densely the two modalities are paired, so a system evaluating cross-modal quality cannot treat a many-to-many mapping as equivalent to a one-to-one one without expecting the scores to shift downward.