epiphenomenal-correctness-probe-vs-ablation

IN premise — summaries/2026/08/24/convergence-without-understanding-2026-s3-results.md

Created 2026-08-24T17:10:52+00:00

Cross-model transfer probe accuracy for correctness prediction is 66% (exceeding 55% permutation and 62.9% majority-class baselines), yet full-subspace causal ablation produces only a 1.5% flip rate (5.5% under relaxed protocol), demonstrating shared correctness information is encoded but not causally deployed.

Summary

A model's internal representations contain enough signal about whether an answer is correct that an outside probe can predict correctness at 66% accuracy, but when you surgically remove that signal, the model's outputs barely change. In other words, the model "knows" when it is wrong in a passive, readable sense, but it never actually consults that knowledge when producing its answers, so we cannot rely on these internal correctness signals as a self-monitoring or safety mechanism.