hendel-2023-layer-l-relative-position
IN premise — summaries/2026/08/24/hendel-2023-icl-task-vectors-s2-a-hypothesis-class-view-of-icl.md
Created 2026-08-25T02:58:03+00:00
Across all tested models in Hendel et al. (2023), the optimal boundary layer L between task-encoding (A) and task-application (f) peaks at a similar relative (fractional) position regardless of total model depth or parameter count.
Summary
In Hendel et al.'s 2023 experiments, the layer where a network transitions from understanding a task to executing it falls at roughly the same fraction of the total depth, whether the model is small or large. This scale invariance means findings about where and how computation is organized can be generalized across model sizes without re-deriving them for every new architecture.