rome-key-computation-prefixes

IN premise — summaries/2026/08/24/meng-2022-rome-s3-interventions-on-weights-for-understanding-factual-associati.md

Created 2026-08-25T02:58:15+00:00

The ROME key k* is computed as the post-nonlinearity activation at the chosen mid-layer, averaged over approximately 50 random 2–10 token prefixes generated by the model.

Summary

To find the internal signal that represents a specific fact, ROME samples roughly fifty short, randomly generated lead-in phrases and averages the activation values at a single hidden layer after the nonlinearity step. This averaging step matters because it produces a fact-specific key that is stable across different sentence contexts, so the subsequent rank-one weight update targets the concept rather than one particular phrasing.