gpt2-context-window-1k-tokens

IN premiseentries/2026/06/21/wiki-Large_language_model.md

Created 2026-06-21T09:50:09+00:00

GPT-2 had 12 attention heads and a 1,024-token context window

Summary

GPT-2 could only keep about 750 words in its working memory at once and processed each passage through 12 parallel "lenses" that tracked different kinds of relationships in the text. This sets a hard ceiling on what the model could reason about in a single pass, since anything outside that short window was simply invisible to it.

Dependents

These beliefs depend on this one: