correctness-subspace-ablation-flips-3-to-7-percent-of-answers

IN premise — summaries/2026/08/24/convergence-without-understanding-2026-sR-references.md

Created 2026-08-24T17:10:53+00:00

Projecting out the correctness subspace (spanned by probe-activating directions at the peak probe layer) flips model answers at rates of 3–7% across 8 small models (Qwen, SmolLM2, Gemma, LLaMA, Phi-3.5).

Summary

Removing a small set of internal directions that track whether a model's answer will be correct causes 3 to 7 percent of its outputs to change, and this holds consistently across eight different model families. The takeaway is that the model's sense of "am I right here" is not just a passive byproduct but a causally active signal that nudges the final answer, and this mechanism appears to be a general feature of how these small models work rather than something specific to one architecture.