ipc1-alignment-exceeds-both-models-correct-class-rate

IN premise — summaries/2026-08-24/koepke-2026-back-into-cave-s3-experimental-setup-chunk-1.md

Created 2026-08-24T17:11:00+00:00

On ImageNet at 1 image per class (ipc=1), strict cross-modal alignment (23.1%) exceeds the rate at which both DINOv2 and OpenLlama retrieve correct-class neighbors (11.7%), meaning the models frequently agree on the same incorrect item (e.g., both matching a bookstore query to a library image).

Summary

When two very different models (one for images, one for text) both pick the same neighbor, that agreement looks like strong confirmation, but this observation shows they are usually wrong together — they share the same confusion rather than the same accuracy. For the system, this means "both models agree" cannot be treated as a reliable correctness signal; correlated mistakes will inflate confidence in the wrong answer.