gptj-checkpoints-no-icl-degradation

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s6-billion-parameter-autoregressive-language.md

Created 2026-08-25T02:58:32+00:00

GPT-J checkpoints from 310k to 380k pretraining steps show no significant difference in ICL performance on AGNews, SST-2, CB, and RTE with 8 demonstrations, despite measurable parameter changes.

Summary

During a specific stretch of GPT-J's pretraining (roughly 310k to 380k steps), the model's ability to learn new tasks from just a few examples had already plateaued, even though its internal weights were still shifting. This means you don't need to wait for the final checkpoint to get reliable in-context learning performance in that window, and it suggests that the ICL capability settles in before full training is complete, which can save compute and evaluation effort.