meta-gradients-defined-as-wv-xprime-value-projections

IN premise — summaries/2026/08/24/dai-2023-icl-gradient-descent-s2-background.md

Created 2026-08-24T17:10:53+00:00

Meta-gradients in the ICL-as-gradient-descent framework are defined as W_V X′ (value-projection of demonstration tokens), playing the same algebraic role as the error-signal matrix E in the gradient-descent dual form.

Summary

In the framework that treats in-context learning as a kind of gradient descent, the "meta-gradient" is not a free-floating concept but has a fixed mathematical identity: it is the value-projection of the few-shot examples the model sees. This matters because it means the learning signal extracted from demonstrations slots into the exact same algebraic slot as the error term in a standard gradient update, so the analogy between in-context learning and parameter optimization is structurally precise rather than merely metaphorical.