ripple-edits-ice-outperforms-rome
IN premise — summaries/2026/08/24/cohen-2023-ripple-effects-s5-experiments.md
Created 2026-08-25T02:57:57+00:00
ICE outperforms ROME by more than 10 points on GPT-NeoX and more than 29 points on LLaMA in average accuracy across RECENT, RANDOM, and POPULAR subsets.
Summary
When editing the internal knowledge of large language models, the ICE method produces substantially more accurate results than ROME, with margins of 10 to nearly 30 points depending on the base model, and this advantage is consistent whether the edits target recently learned, randomly sampled, or popular concepts. In practice, this means the system should prefer ICE over ROME when it needs to surgically update a model's stored facts, since the gap is large enough to matter in real deployment.