xie-2021-scaling-improves-icl-constant-loss
IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s2-in-context-learning-setting.md
Created 2026-08-25T02:58:55+00:00
In GINC experiments, in-context learning accuracy improves with model scale even when pretraining loss is unchanged, suggesting larger models better approximate the true posterior p(θ|prompt) rather than merely fitting the training distribution.
Summary
Larger models get better at in-context learning even when they are not fitting their training data any more tightly, meaning the gain comes from genuinely improving their ability to infer the rule behind a prompt rather than from memorization. This matters because it suggests that simply scaling up model size unlocks better reasoning about unseen patterns, which is a qualitatively different capability than recognizing patterns the model has already seen.