prh-findings-generalize-across-three-modality-pairs
IN premise — summaries/2026-08-24/koepke-2026-back-into-cave-s3-experimental-setup-chunk-1.md
Created 2026-08-24T17:11:00+00:00
The degradation of cross-modal alignment with gallery size at fixed small k, and the weak/flat LLM-capability-to-alignment trend, are observed consistently for text–image (DINOv2 + OpenLlama), text–video (VideoMAE-v2 on PVD-100k), and text–audio (Dasheng on LAION-Audio-100k), indicating the pattern is a structural property of independent training rather than a text-image-specific artifact.
Summary
The fact that alignment quality falls as the candidate pool grows, and that a stronger language model doesn't meaningfully help, shows up the same way across images, video, and audio. That consistency means the problem is baked into how these cross-modal systems are trained independently, so the fix has to target the training architecture itself rather than just swapping in a different modality or a bigger model.