c2023-ripple-edits-scope-1-to-2-hop

IN premise — summaries/2026/08/24/cohen-2023-ripple-effects-s6-conclusion-and-discussion.md

Created 2026-08-25T02:57:57+00:00

RIPPLE EDITS benchmark covers only the close neighbourhood of an edit (1–2 hops) and does not test paraphrase robustness, subject specificity, or distantly-related fact retention

Summary

The RIPPLE EDITS benchmark only checks whether a model correctly updates facts in the immediate neighborhood of a change, leaving out harder real-world scenarios like reworded questions, making sure an edit doesn't bleed onto a different subject, or preserving facts that are further away in the knowledge graph. In practice, a model can score well on this test and still fail at the kinds of targeted, robust knowledge updates you would actually need in deployment.