dai-2023-empirical-scope-limits

IN premise — summaries/2026/08/24/dai-2023-icl-gradient-descent-sR-references.md

Created 2026-08-24T17:10:54+00:00

Dai et al. 2023 empirical scope is restricted to Transformer architectures, models of 2.7B parameters or fewer, and classification tasks only; results do not automatically extend to LSTMs, generation tasks, or larger models.

Summary

The Dai et al. 2023 findings only cover small Transformer models doing classification, so they should not be cited as evidence for how LSTMs, generative tasks, or models larger than 2.7B parameters behave. Any reasoning in the system that leans on this work for a broader architecture, task type, or scale is overreaching.