cross-modal-alignment-drop-is-model-agnostic-across-3b-to-65b

IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references.md

Created 2026-08-24T17:11:01+00:00

The cross-modal mutual kNN alignment degradation pattern holds across DINOv2-base, DINOv2-giant, Pixio-ViT-B/16, and CLIP-B/16 (vision axis) paired with OpenLlama-3b, OpenLlama-13b, OpenLlama-65b, Gemma-7B, and Mistral-7B (language axis), with embedding dimensions ranging from 768 to 8192.

Summary

The pattern of misalignment between vision and language model embeddings is not a quirk of any particular architecture or scale; it shows up the same way whether you pair a small 3B language model with a compact vision encoder or a 65B language model with a large one. This matters because it means the alignment gap is a structural problem across the model ecosystem, not something you can solve by simply swapping in a different or bigger model family.