recaptioning-does-not-prevent-cross-modal-alignment-decline

IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references.md

Created 2026-08-24T17:11:01+00:00

Recaptioning WIT-1M using gemini-3-flash-preview to produce ~500-word descriptions raises absolute mutual kNN alignment scores but does not prevent the cross-modal alignment decline as gallery scale increases, ruling out caption quality as the primary driver.

Summary

Even when the image descriptions are swapped out for much longer, richer captions, the retrieval system still degrades as the candidate pool grows. This tells us the bottleneck is not poor text quality but something deeper in how the model represents and compares items at scale, so pouring effort into better captions will not fix the core problem.