cross-modal-alignment-finding-generalizes-to-text-video-audio
IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-s3-experimental-setup.md
Created 2026-08-24T17:11:01+00:00
The cross-modal mutual kNN alignment drop with gallery scale generalizes beyond text-image to text-video (PVD-100k, VideoMAE-v2) and text-audio (LAION-Audio-100k, Dasheng audio encoder) pairs, where the LLM-performance alignment trend is weak (video) or flat (audio), and the finding is verified across DINOv2/Pixio/CLIP vision encoders and OpenLlama/Gemma/Mistral/LLaMA language models.
Summary
The finding that mutual alignment quality degrades as the gallery of items grows is not unique to text-image pairs; it also shows up when text is paired with video or audio, across a wide range of vision encoders and language models. What matters is that the degradation tracks with the non-text side of the pairing, while LLM-side performance stays relatively flat, meaning the scaling problem lives in how visual and auditory embeddings organize rather than in the language model itself.