stronger-llm-alignment-trend-refuted-with-negative-r2

IN premise — summaries/2026-08-24/koepke-2026-back-into-cave-s3-experimental-setup-chunk-2.md

Created 2026-08-24T17:11:00+00:00

The previously reported trend that stronger LLMs exhibit higher alignment with DINOv2 vision features does not continue for the most recent model generations, with R² values of −1.41 to −0.58 on broader reasoning benchmarks (ARC, GSM8K, MMLU, LogiQA2), indicating the extrapolated scaling line is worse than a flat prediction.

Summary

The earlier observation that larger LLMs tracked DINOv2 vision features more closely was expected to keep holding as models got stronger, but the newest model generations actually fall below a flat-line prediction on major reasoning benchmarks, meaning the scaling relationship has broken down rather than just plateaued. This matters because any downstream reasoning in the system that assumed "bigger model, better vision-feature alignment" as a continuing trend can no longer rely on that extrapolation.