fv-models-and-scope-iclr-2024
IN premise — summaries/2026/08/24/todd-2023-function-vectors-s0-abstract.md
Created 2026-08-25T02:58:40+00:00
Todd et al. (ICLR 2024) demonstrate FVs across GPT-J (6B, 28L, 16 heads/layer), GPT-NeoX (20B, 44L, 64 heads), and Llama-2 7B/13B/70B, testing 40+ diverse ICL tasks.
Summary
Feature vectors that support in-context learning have been confirmed across at least three different model families (GPT-J, GPT-NeoX, Llama-2) ranging from 6B to 70B parameters, and tested on more than 40 distinct task types. This means the phenomenon is a robust, general property of how large language models handle in-context learning, not a quirk of one particular architecture or training setup.