icl-gd-gap-persists-across-scales
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s8-icl-demonstrations.md
Created 2026-08-25T02:58:33+00:00
The performance gap between ICL and GD/cGD does not significantly close as model size increases from 1.5B to 7B parameters or as demonstration count increases from 1 to 8 on RTE.
Summary
Getting a bigger model or adding more example prompts doesn't meaningfully narrow the edge that gradient-based training has over in-context learning on textual entailment. This means the limitation is structural rather than a resource problem, so strategies that lean on scale or volume of demonstrations won't close the gap and the system should treat that shortfall as a persistent design constraint.