transformer-2017-quadratic-context

IN premiseentries/2026/06/21/wiki-Neural_network_28machine_learning29-chunk-2.md

Created 2026-06-21T09:55:51+00:00

The Transformer architecture (2017, 'Attention Is All You Need') uses self-attention with quadratic computation cost in context window size and became the basis for GPT, Gemini, Grok, DeepSeek, and Qwen

Dependents

These beliefs depend on this one: