dai2023-kendall-icl-ft-019-021-vs-random-000

IN premise — summaries/2026/08/24/dai-2023-icl-gradient-descent-s5-momentum-based-attention-inspired.md

Created 2026-08-24T17:10:54+00:00

Kendall rank correlation between ICL and finetuning attention to training tokens is 0.193 (GPT 1.3B) and 0.214 (GPT 2.7B), while Kendall correlation between ICL and random attention is 0.000 for both model sizes.

Summary

When a language model does in-context learning, the way it pays attention to the example tokens it was given is not random noise, but it only loosely resembles the attention patterns that come from explicitly finetuning on those same tokens. This matters because it tells us ICL is doing something real and structured, yet it is operating in a meaningfully different attention space than standard finetuning, so the two learning mechanisms can't be treated as interchangeable.