alignment-performance-trend-holds-for-language-benchmarks-fails-for-reasoning

IN premise — summaries/2026/08/24/koepke-2026-back-into-cave-sR-references-chunk-2.md

Created 2026-08-24T17:11:00+00:00

The alignment–language-performance linear trend from Huh et al. (2024) holds for HellaSwag (R²_new = 0.297) and Wikitext (R²_new = 0.489) but fails for ARC (R²_new = −0.575), GSM8K (R²_new = −1.753), MMLU (R²_new = −0.662), and LogiQA2 (R²_new = −1.414) across 36 recent LLMs, where negative R² means the regression line predicts worse than the mean.

Summary

The alignment-to-performance relationship that Huh et al. described works for basic language comprehension tasks but completely breaks down for reasoning, math, and logic benchmarks, where the linear model predicts worse than simply guessing the average. In practice, this means you cannot use a single rule of thumb to estimate how much capability a model trades off for alignment — the cost looks very different depending on whether the task involves understanding language or actually reasoning through a problem.