rome-corruption-noise-protocol

IN premise — summaries/2026/08/24/meng-2022-rome-sR-references.md

Created 2026-08-25T02:58:17+00:00

Causal Tracing corruption adds Gaussian noise ε ~ N(0; 3σ_t) to subject token embeddings (ν = 3 times the observed standard deviation), reducing correct-object probability from ~27% to ~8.47%.

Summary

The system recorded a direct measurement: when the experiment intentionally jiggles a model's internal representation of a sentence's subject by about three times its natural variation, the model's chance of picking the right object drops from roughly 27 percent to under 9 percent. This matters because it confirms the corruption method is strong enough to meaningfully disrupt the model's reasoning, giving the team a reliable knob to turn when they want to trace which parts of the model actually drive a given output.