xie-2021-icl-both-architectures-ginc

IN premise — summaries/2026/08/24/xie-2021-icl-bayesian-s0-abstract.md

Created 2026-08-25T02:58:55+00:00

On the synthetic GINC dataset, both Transformers and LSTMs exhibit in-context learning, demonstrating the phenomenon is not specific to the attention mechanism.

Summary

In-context learning isn't a trick unique to attention-based models. The fact that older LSTM architectures can also pick up patterns from a prompt on this synthetic task means the phenomenon is more general, likely rooted in how sequence models learn rather than in any special property of the attention mechanism.