coarse-grained-cross-modal-convergence-stable-at-k-n100
IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references.md
Created 2026-08-24T17:11:01+00:00
At k = n/100 (coarse, semantic-category level), cross-modal mutual kNN alignment remains stable as gallery density increases, but at fixed small k (k=1, k=10) alignment drops—indicating vision and language models share broad semantic categories but do not achieve fine-grained representational convergence.
Summary
Vision and language models line up well when you compare them at a broad, category level (like "animal" or "vehicle"), and that agreement holds even as the dataset grows, but the alignment crumbles when you demand fine-grained, instance-level matching. The practical implication is that any system relying on cross-modal matching should operate at a coarse semantic level rather than expecting pixel-or-token-level correspondence between the two representations.