prh-cross-modal-alignment-predicts-downstream-performance
IN premise — summaries/2026/08/24/huh-2024-prh-s2-representations-are-converging-chunk-1.md
Created 2026-08-24T17:10:56+00:00
Higher alignment to vision model DINOv2 correlates linearly with Hellaswag (commonsense reasoning) scores in LLMs and shows an emergence-like trend for GSM8K (math) scores, linking representational convergence to downstream task performance.
Summary
When a language model's internal representations grow more similar to those of a vision model, its commonsense reasoning improves in a steady, predictable way, while its math ability seems to jump in sudden steps rather than climb smoothly. In practical terms, how well a model "sees" the world visually is a useful early signal for how well it will handle everyday and mathematical reasoning tasks.