cgd-simulated-on-single-random-middle-layer
OUT premise — summaries/2026/08/24/shen-2023-icl-not-gd-s2-icl-demonstrations.md
Created 2026-08-25T02:58:32+00:00
In Shen et al. (2023), the continual gradient descent (cGD) baseline is simulated by optimizing on a single randomly chosen middle layer of LLaMA rather than the full model, as a deliberate simplification.