convergence-effects-stronger-in-base-models-than-instruction-tuned
IN premise — summaries/2026/08/24/convergence-without-understanding-2026-s6-limitations.md
Created 2026-08-24T17:10:52+00:00
Difficulty inversion, generation gap, and epiphenomenal correctness effects are stronger in base (pre-trained) models than in instruction-tuned variants, implying fine-tuning partially mitigates but does not eliminate the dissociation.
Summary
Even after a model is fine-tuned to follow instructions, it still shows telltale signs of shallow pattern-matching rather than genuine reasoning, such as getting harder questions right while missing easier ones or producing correct answers for the wrong reasons. This matters because instruction-tuning only softens those gaps without closing them, so a well-formatted or compliant output should not be treated as evidence that the model actually understands the underlying task.