memit-covariance-lambda-gptj
IN premise — summaries/2026/08/24/meng-2022-memit-s3-p-reliminaries-l-anguage-modeling-and-memory-editing-chunk-2.md
Created 2026-08-25T02:58:13+00:00
The covariance adjustment factor λ is set to 15,000 for GPT-J and 20,000 for GPT-NeoX-20B, with an optimal range near 10⁴.
Summary
The system has tuned a single scaling knob for covariance corrections: larger models (GPT-NeoX-20B) get a slightly higher setting than smaller ones (GPT-J), and both land in the ten-thousands range. This matters because any future model or reconfiguration should fall near that 10,000 zone; drifting far outside it signals a likely miscalibration that will degrade the quality of the covariance adjustments the system relies on.