memit-gptj-mquake-cf-multi-hop-drop

IN premise — summaries/2026/08/24/zhong-2023-mquake-s4-mq-uake-challenges-model-editors.md

Created 2026-08-25T02:59:04+00:00

MEMIT-edited GPT-J drops from 40.5% to 7.0% multi-hop accuracy on MQuAKE-CF

Summary

Editing a GPT-J model's internal knowledge with MEMIT causes its ability to chain together multiple facts to collapse by more than 80%, falling from roughly 40% to 7% accuracy on a multi-hop reasoning benchmark. This matters because it shows that even targeted memory edits can badly damage the model's broader reasoning skills, so any system relying on MEMIT-style interventions needs to weigh that tradeoff carefully.