icl-scale-independent-of-pretraining-loss
IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s4-simulations.md
Created 2026-08-25T02:58:56+00:00
In Xie et al. (2021) GINC experiments, 12-layer (85M) and 16-layer (115M) Transformers achieve identical pretraining validation loss (~1.33 at vocabulary size 50) yet differ in in-context learning accuracy, demonstrating ICL quality is not a monotonic function of perplexity alone.
Summary
Two transformer models that reach the same score on their training data can still perform very differently at the practical task of learning from a few examples given at test time, so pretraining loss alone cannot predict how well a model will actually use in-context demonstrations. This means the system needs its own separate ICL evaluation metric rather than relying on perplexity as a proxy, since two models tied on loss may be meaningfully apart in real-world few-shot performance.