rome-human-eval-ratios
IN premise — summaries/2026/08/24/meng-2022-rome-s4-related-work.md
Created 2026-08-25T02:58:15+00:00
In human evaluation, ROME is rated 1.8× more likely to be consistent with the inserted fact than FT+L, but 1.3× less likely to be more fluent than FT+L.
Summary
When people judge the two editing methods, they find ROME more likely to actually reflect the new fact you asked for, but its output reads less naturally than FT+L. This is a real tradeoff: picking ROME buys you factual accuracy at the cost of slightly awkward phrasing, so the right choice depends on whether you care more about the model getting the fact right or sounding smooth.