mend-cot-mquake-t-anomaly

IN premise — summaries/2026/08/24/zhong-2023-mquake-s4-mq-uake-challenges-model-editors.md

Created 2026-08-25T02:59:06+00:00

MEND achieves 38.2% multi-hop accuracy with CoT prompting on MQuAKE-T, substantially higher than ~4-12% for ROME and MEMIT, attributed to relation-specific editing effectiveness

Summary

When you edit a language model's knowledge, MEND leaves its ability to chain multiple reasoning steps together much more intact than ROME or MEMIT, scoring roughly three to nine times higher on multi-hop questions when guided step by step. This matters because it suggests MEND's approach of updating specific relation slots rather than overwriting broader weight patterns causes less collateral damage to the model's downstream reasoning.