icl-gd-llama7b-probed
IN premise — summaries/2026/08/24/shen-2023-icl-not-gd-s0-abstract.md
Created 2026-08-25T02:58:30+00:00
The Shen et al. ICL vs. GD paper uses LLaMa-7B (pre-trained on natural data) as the primary model for empirical probing.
Summary
This is a factual anchor: the empirical results in the Shen et al. paper are grounded specifically in LLaMa-7B, meaning any generalizations or comparisons built on top of their findings are tied to that model's size, training data, and architecture. It matters because if later reasoning depends on those results, the conclusions inherit the limitations and quirks of that particular model rather than applying universally across LLMs.