icl-gd-hypothesis-2-existence
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s1-introduction.md
Created 2026-08-25T02:58:31+00:00
Hypothesis 2 in Shen et al. claims that for a given task, there exist Transformer weights (possibly hand-constructed) such that non-emergent ICL (dICL) is equivalent to GD; prior works (Akyürek et al. 2022; von Oswald et al. 2023) target this weaker claim.
Summary
Shen et al.'s Hypothesis 2 is an existence claim: for any given task, there is at least one set of Transformer weights (possibly hand-built) where in-context learning mathematically coincides with gradient descent. The practical upshot is that the prior work it references was aiming at this weaker "does some weight exist" target, which leaves open the harder question of whether Transformers that actually emerge from standard training exhibit this equivalence.