icl-gd-hypothesis-1-universal

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s1-introduction.md

Created 2026-08-25T02:58:31+00:00

Hypothesis 1 in Shen et al. claims that for any Transformer weights from self-supervised pretraining and any well-defined task, ICL is algorithmically equivalent to GD (whole-model or sub-model updates).

Summary

Shen et al. make a sweeping claim that the way a pretrained Transformer "learns" from examples in its prompt is not a new trick at all but is computationally identical to standard gradient descent, whether you adjust the whole model or just a slice of it. If this holds universally across every pretrained Transformer and every task, it collapses two seemingly different learning mechanisms into one, meaning in-context learning adds no extra capacity beyond what weight updates already provide and any apparent prompt-based "learning" is secretly just gradient descent in disguise.