rome-human-eval-150-raters-per-axis

IN premise — summaries/2026/08/24/meng-2022-rome-sR-references-chunk-3.md

Created 2026-08-25T02:58:17+00:00

The ROME human evaluation study used n=150 raters for factual consistency and n=150 for fluency, with 10 independent test cases per participant, using shuffled AI ordering and forced rankings.

Summary

This records how the ROME model-editing paper actually ran its human evaluation: it hired 150 separate raters for each quality dimension (accuracy of edited facts and naturalness of the text), gave each person only 10 examples to judge, and used randomized order plus forced ranking to reduce bias. It matters because every claim the paper makes about ROME being "better" or "competitive" rests on whether this particular sample size, case count, and ranking setup was strong enough to support those conclusions.