base-models-show-stronger-inversion-than-instruction-tuned

OUT premise — summaries/2026/08/24/convergence-without-understanding-2026-s4-discussion.md

Created 2026-08-24T17:10:52+00:00

The difficulty inversion is stronger in non-instruction-tuned base models (ΔCKA = +0.117, monotonic 0.962 → 0.827 across bins) than in instruction-tuned models (ΔCKA = +0.067), ruling out RLHF/RLAIF alignment training as the source of the phenomenon.