rome-human-ranker-rome-over-ftl-consistency

IN premise — summaries/2026/08/24/meng-2022-rome-sR-references-chunk-3.md

Created 2026-08-25T02:58:17+00:00

In the ROME human evaluation, all three sampled raters ranked ROME > FT+L > original GPT in factual consistency with the injected counterfactual.

Summary

In the ROME human evaluation, every rater independently agreed on the same ordering, putting ROME first for staying factually consistent with the newly injected knowledge, then FT+L, then the unedited model last. Because the ranking is unanimous rather than contested, the system can rely on it as a firm, uncontested piece of evidence that ROME is the strongest method when preserving factual consistency after an edit is the priority.