xie-2021-icl-both-architectures-ginc
IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s0-abstract.md
Created 2026-08-25T02:58:55+00:00
On the synthetic GINC dataset, both Transformers and LSTMs exhibit in-context learning, demonstrating the phenomenon is not specific to the attention mechanism.
Summary
In-context learning isn't a trick unique to attention-based models. The fact that older LSTM architectures can also pick up patterns from a prompt on this synthetic task means the phenomenon is more general, likely rooted in how sequence models learn rather than in any special property of the attention mechanism.