memit-ftw-hyperparameters-gptj

IN premise — summaries/2026/08/24/meng-2022-memit-sR-references.md

Created 2026-08-25T02:58:14+00:00

FT-W baseline on GPT-J targets layer 21 with weight decay 5×10⁻⁴, max 25 steps, learning rate 5×10⁻⁴, early-stopping at loss ≤ 10⁻², completing 10,000 edits in ~0.48 hours.

Summary

This records the specific recipe used to fine-tune a single layer (layer 21) of the GPT-J model for a baseline edit task, including its exact hyperparameters and the observation that it finishes 10,000 edits in under half an hour. It matters because it sets the performance and speed benchmark that any improved editing method must beat, and its narrowness (one layer, 25 steps) shows how minimal a tuning intervention can be for this task.