ft-catastrophic-mquake-cf-multipath

IN premise — summaries/2026/08/24/zhong-2023-mquake-sR-references-chunk-1.md

Created 2026-08-25T02:59:07+00:00

Fine-tuning achieves only 1.4% multi-hop accuracy on MQuAKE-CF versus 39.5% for the unedited base model, while ROME achieves 37.3% on 2-hop but drops to 10.0% on 3-hop and 7.7% on 4-hop.

Summary

Editing a model's knowledge in a targeted way (whether through fine-tuning or the ROME method) severely breaks its ability to reason across multiple connected facts. Fine-tuning drops multi-hop accuracy to near zero compared to the untouched model, and ROME's performance collapses as the reasoning chain gets longer, meaning these methods are unreliable whenever the system needs to follow a chain of inferences beyond a single step.