rome-mlp-run-10-recovery-23.6pct

IN premise — summaries/2026/08/24/meng-2022-rome-sR-references-chunk-1.md

Created 2026-08-25T02:58:16+00:00

Restoring runs of 10 consecutive MLP lookup values at the early causal site achieves up to 23.6% average max score recovery, while attention at the late site achieves up to 19.4% in GPT-2 XL.

Summary

In GPT-2 XL, selectively restoring ten consecutive MLP lookup values in an early layer recovers more of the model's output score than restoring attention in a late layer, suggesting that the early feedforward stages carry more recoverable signal per unit of intervention. This means targeted, sparse fixes to early MLP lookups are a more efficient path to partial recovery than adjusting late attention, which matters for any strategy that tries to repair or compress model behavior with minimal changes.