model-capacity-15b-7b-does-not-close-icl-gd-gap
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s2-icl-demonstrations.md
Created 2026-08-25T02:58:32+00:00
On the AGNews dataset with N=8 demonstrations, increasing model capacity from 1.5B (GPT2-XL) to 7B (LLaMA) parameters does not significantly reduce the performance gap between ICL and explicit gradient descent.
Summary
Simply using a larger model does not make in-context learning catch up to the performance you'd get by actually training the model on the task, even with more than four times the parameters. This matters because it undercuts the assumption that scaling up model size is a shortcut to matching fine-tuning, and suggests the gap between prompting and gradient-based learning is a structural limitation rather than a capacity problem.