rome-value-optimization-objective
IN premise — summaries/2026/08/24/meng-2022-rome-s3-interventions-on-weights-for-understanding-factual-associati.md
Created 2026-08-25T02:58:15+00:00
ROME optimizes the value v* by minimizing a two-term objective: maximizing P[o*|prompt] (efficacy) plus a KL-divergence penalty on a 'subject is a ___' prompt to control essence drift.
Summary
When ROME edits a fact in a model, it balances two goals: making sure the new answer actually gets produced, and making sure the model's broader understanding of the subject stays intact. This two-part scoring prevents a simple "fix one answer" from accidentally scrambling other related knowledge about the same entity.