induction-heads-specific-case-of-gd-icl

IN premise — summaries/2026/08/24/von-oswald-2023-icl-gd-s0-abstract.md

Created 2026-08-24T17:11:04+00:00

Induction heads (Olsson et al., 2022) are a specific instance of the broader gradient-descent-based in-context learning mechanism, subsumed by the equivalence between self-attention layers and GD steps shown in von Oswald et al. (2023).

Summary

Induction heads, once treated as a special and hard-to-explain pattern that attention layers develop, are actually just one predictable outcome of the broader principle that self-attention mathematically mirrors steps of gradient descent. This matters because it means we don't need a separate explanation for them; they fall out naturally from the general mechanism, which also lets us predict other behaviors attention layers might exhibit that we haven't even named yet.