icl-gd-capacity-sweep-models

IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s1-icl-demonstrations.md

Created 2026-08-25T02:58:31+00:00

The capacity invariance experiment (Section F) uses four models: GPT2-XL (1.5B), GPT-NEO (2.7B), GPT-J (6B), and LLaMA (7B) on AGNews with N=8 demonstrations.

Summary

This records the specific setup of the capacity invariance test: four models of increasing size (1.5B through 7B parameters) are each asked to classify news articles using just 8 example prompts. The point of spanning that range is to check whether the observed effect holds regardless of model size, so the result isn't an artifact of one particular architecture or scale.