ahn-2024-preconditioned-gd
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-sR-references.md
Created 2026-08-25T02:58:33+00:00
Ahn et al. (2024, NeurIPS) refined the gradient descent claim to preconditioned gradient descent, addressing the gap between vanilla GD and what transformers actually compute.
Summary
Ahn et al. sharpened an earlier gradient-descent argument by specifying that the analysis should apply to preconditioned gradient descent, the kind of optimizer transformers actually use in practice rather than the idealized vanilla version. This matters because it closes the gap between what theory predicts and what real training loops compute, making the theoretical guarantees actually relevant to how transformers learn.