decoder-only-removes-cross-attention

IN premiseentries/2026/06/21/wiki-Transformer_deep_learning_architecture-chunk-4.md

Created 2026-06-21T09:55:55+00:00

In decoder-only models (e.g., GPT), the cross-attention sublayer is removed entirely, leaving only 2 sublayers: masked self-attention and FFN.