rome-entropy-metrics-miss-fluency-loss
IN premise — summaries/2026/08/24/meng-2022-rome-sR-references-chunk-3.md
Created 2026-08-25T02:58:17+00:00
Human raters detected subtle fluency losses in ROME's generated output that an entropy-based automatic metric failed to capture, indicating a gap between statistical and perceptual quality assessment.
Summary
Raters caught small stumbles in ROME's rephrased text that a standard entropy score would have waved through as fine, meaning the automatic quality check has a blind spot. In practice, this means relying solely on that metric could let degraded edits pass review unnoticed, so the pipeline needs either a more sensitive measure or a human-in-the-loop step to catch what the numbers miss.
Dependents
These beliefs depend on this one:
- OUT editing-reliability-under-superposition — Rank-one knowledge editing is reliable as a single-fact correction mechanism in superposed models because covariance whitening provides sufficient feature separation to isolate the target association from the background superposition, making the edit direction well-defined and locally confined.