transformer-early-mlp-extremely-polysemantic

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-11.md

Created 2026-08-25T02:57:59+00:00

The first MLP layers in Transformer models are predicted by the superposition theory and confirmed empirically to be extremely polysemantic.

Summary

Early in a Transformer's processing, individual neurons do not each track one clean concept — they simultaneously encode a jumble of many different meanings at once. This matters because it means you cannot look at a single early-layer unit and read off a single interpretable idea; the model relies on these densely mixed representations as raw building blocks, and later layers are what resolve them into more discrete, usable concepts.