inceptionv1-early-layers-monosemantic
IN premise — summaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-11.md
Created 2026-08-25T02:57:59+00:00
In InceptionV1, early-layer neurons are monosemantic (aligned with a privileged basis, no superposition), while later-layer neurons become polysemantic as features become sparser.
Summary
Early layers in InceptionV1 behave like specialized detectors where each neuron captures one distinct visual feature, while deeper layers gradually blend multiple concepts into a single neuron. This matters because it means the network builds understanding by progressively mixing clean, single-purpose signals into richer but less transparent representations, so any analysis targeting individual neurons will only get a one-to-one mapping in the first few layers.