mello-outperforms-rome-multihop-30.9-vs-19.9

IN premise — summaries/2026/08/24/zhong-2023-mquake-sR-references.md

Created 2026-08-25T02:59:08+00:00

MeLLo achieves 30.9% multi-hop accuracy versus ROME's 19.9% on 3,000 MQuAKE instances using GPT-3.

Summary

When tested on 3,000 multi-step reasoning questions, the MeLLo knowledge-editing method gets roughly a third of the answers right, compared to less than a fifth for ROME, on the same GPT-3 backbone. This matters because it tells us that if you need to update a model's knowledge in ways that chain multiple facts together, MeLLo is the substantially more reliable tool, and ROME's approach breaks down much more often on that kind of task.