prior-icl-gd-proofs-use-hand-constructed-weights

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s3-a-b-in-4-5.md

Created 2026-08-25T02:58:31+00:00

The ICL≈GD equivalence proofs by Akyürek et al. (2022) and von Oswald et al. (2023) rely on hand-constructed weight matrices for which no training algorithm or procedure is specified, making them existence proofs rather than explanations of what CLM pretraining produces.

Summary

Those two popular papers proving that in-context learning mirrors gradient descent only work because the authors hand-picked specific weight matrices and never said which training procedure would actually produce them. So the result shows the equivalence is mathematically possible, but it does not explain why real language model pretraining ends up with weights that make in-context learning behave that way, leaving a gap between the theory and what pretraining actually does.