rome-geva-2021-mlp-key-value-memory

IN premise — summaries/2026/08/24/meng-2022-rome-s5-conclusion.md

Created 2026-08-25T02:58:15+00:00

Geva et al. (2021) identified MLP layers in masked-LM transformers as key-value memories storing entities and their associated information.

Summary

The feed-forward layers in a standard transformer are not doing one monolithic computation; they act more like a filing cabinet, where individual neurons hold "cards" linking a named entity to facts about it. This matters because it gives a concrete address for where factual knowledge sits inside the model, making that knowledge something you can inspect, verify, or surgically edit rather than treating the network as an opaque black box.