rome-gpt2xl-clean-27pct-corrupted-8.47pct
IN premise — summaries/2026/08/24/meng-2022-rome-sR-references-chunk-1.md
Created 2026-08-25T02:58:16+00:00
In GPT-2 XL Causal Tracing, clean p(o_c) is 27.0%, drops to 8.47% after 3σ_t corruption, and best single hidden-state restoration recovers to 19.5% (layer 15, last subject token).
Summary
When the internal representations of GPT-2 XL are heavily disturbed, the model's ability to produce the correct word falls sharply (from about 27% down to under 9%), but surgically repairing just one spot — the last subject token at layer 15 — brings performance back to roughly two-thirds of its original level. This shows that a small, identifiable slice of the model's internal computation carries the bulk of what drives the correct answer, and that information is distributed across many layers rather than living in one place, since even the best single repair cannot fully restore original performance.