epiphenomenal-correctness-transfer-vs-causal

IN premise — summaries/2026/08/24/convergence-without-understanding-2026-s0-abstract.md

Created 2026-08-24T17:10:52+00:00

Shared information across LLMs is linearly decodable (66% transfer accuracy) but exerts minimal causal influence on predictions (1.5%–5.5% flip rate under ablation), demonstrating epiphenomenal correctness

Summary

When you look inside different large language models, their internal representations overlap a lot, but that overlap is mostly decorative: the models don't actually rely on the shared information to produce their answers. This means that apparent similarity between models is a training side-effect rather than evidence of a common underlying mechanism, so interpretability findings based on representation overlap can overstate how much we truly understand what drives model behavior.