attention-nmt-origin-bahdanau-2014-luong-2015
IN premise — summaries/2026/08/24/wiki-Transformer_deep_learning_architecture-chunk-6-chunk-1.md
Created 2026-08-24T17:11:27+00:00
The attention mechanism for neural machine translation was first introduced by Bahdanau, Cho & Bengio (2014, arXiv:1409.0473) and subsequently formalized by Luong, Pham & Manning (2015), predating the 2017 Transformer paper.
Summary
The attention mechanism that the Transformer uses to weigh input tokens was not invented in 2017; it was first proposed in 2014 by Bahdanau and colleagues and refined in 2015 by Luong and colleagues. This matters for the system because it pins down a clear intellectual lineage: the Transformer's core "self-attention" idea builds on prior work, so any attribution or novelty analysis must treat attention as an inherited component rather than a contribution of the 2017 paper alone.