grokking-memorize-then-generalize

IN premiseentries/2026/06/21/wiki-Large_language_model-chunk-2.md

Created 2026-06-21T09:50:09+00:00

Grokking is the phenomenon where a model first memorizes training data (overfitting), then suddenly learns the underlying algorithm and generalizes, discovered via mechanistic interpretability of modular arithmetic models.

Summary

A model can first just memorize its training examples, then suddenly extract the underlying rule and start generalizing to problems it has never seen. This matters because overfitting is not always a dead end; it can be a necessary stepping stone, and stopping training at the point of best memorization would miss the real understanding that comes next.

Dependents

These beliefs depend on this one: