c2023-rome-vs-ice-failure-modes
IN premise — summaries/2026/08/24/cohen-2023-ripple-effects-s6-conclusion-and-discussion.md
Created 2026-08-25T02:57:57+00:00
ROME tends to produce more incorrect/noisy changes after editing, while ICE tends to cause the model to generate abstention responses (e.g., 'unknown', 'a mystery')
Summary
These two editing techniques break in fundamentally different ways: ROME injects wrong or garbled content into the model's output, while ICE makes the model give up and say things like "I don't know." That distinction matters because downstream systems need to handle the two failure modes differently — one requires filtering out corrupted text, the other requires recovering a usable answer from an empty refusal.