cgd-simulated-on-single-random-middle-layer

OUT premise — summaries/2026/08/24/shen-2023-icl-not-gd-s2-icl-demonstrations.md

Created 2026-08-25T02:58:32+00:00

In Shen et al. (2023), the continual gradient descent (cGD) baseline is simulated by optimizing on a single randomly chosen middle layer of LLaMA rather than the full model, as a deliberate simplification.