memit-10k-edit-scores-cOUNTERFACT

IN premise — summaries/2026/08/24/meng-2022-memit-s3-p-reliminaries-l-anguage-modeling-and-memory-editing-chunk-2.md

Created 2026-08-25T02:58:13+00:00

At 10,000 edits on GPT-J (COUNTERFACT), MEMIT achieves a harmonic-mean Score of ≈85.8, compared to ROME ≈50.3, MEND ≈23.1, and FT-W ≈67.6 (with generation failure).

Summary

When you ask a language model to change 10,000 facts at once, MEMIT keeps the model producing good, coherent answers, while the other editing methods either degrade heavily or stop working altogether. This matters because it shows MEMIT is the only method that can handle the scale of edits a real deployment would actually need, without the model collapsing into broken output.