cross-modal-knn-alignment-degrades-with-gallery-scale

OUT premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references.md

Created 2026-08-24T17:11:01+00:00

Cross-modal mutual kNN alignment between vision and language models degrades as the evaluation gallery scales from 1,024 to 1,000,000 samples (WIT dataset), while within-modality mutual kNN alignment remains stable across the same range, as shown by Koepke et al. (2026).