memit-mlp-not-attention-mediation

IN premise — summaries/2026/08/24/meng-2022-memit-sR-references-chunk-1.md

Created 2026-08-25T02:58:13+00:00

Path-dependent ablation experiments confirm that mid-layer MLPs (not Attention modules) at the last subject token are the causal mediators of factual recall in transformers.

Summary

Experiments that surgically disable parts of a transformer model show that factual recall is carried out by the feedforward blocks in the middle layers, not by the attention mechanism that links words together. This pinpoints exactly where in the model's architecture facts are computed, which tells us where to look if we want to understand, edit, or control what a model "knows."