kandpal-2023-gpt-neo-range

IN premise — summaries/2026/08/24/kandpal-2023-long-tail-knowledge-s6-conclusion-and-future-work.md

Created 2026-08-25T02:58:07+00:00

Kandpal et al. (ICML 2023) tested GPT-Neo variants ranging from 125M to 20B parameters for long-tail knowledge retention.

Summary

The Kandpal team ran a controlled experiment across a wide range of model sizes, from a modest 125-million-parameter network all the way up to 20 billion parameters, specifically to measure how well each size retains rare, uncommon facts rather than just popular knowledge. This gives the system a concrete, size-graded data point for understanding where long-tail knowledge starts to break down as models grow.