memorization-superposition-phase-transition

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-6.md

Created 2026-08-25T02:57:55+00:00

Small datasets are memorized in superposed (entangled) form rather than generalized as separate features, with a sharp phase transition separating memorization from generalization regimes (Henighan et al.)

Summary

When a model is trained on a small dataset, it doesn't learn clean, separable patterns so much as it stores the examples tangled together in a way that's hard to disentangle, and there's a sharp threshold where the behavior flips from that rote storage to genuine generalization. This matters because it means overfitting and memorization aren't a gradual spectrum but a cliff: a model is either memorizing in a fragile, entangled way or it is actually learning, and small shifts in dataset size relative to model capacity can push it across that line.