eh2022-hidden-nonlinearity-essential-for-superposition
IN premise — summaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-10.md
Created 2026-08-25T02:57:58+00:00
The hidden-layer non-linearity is essential for computation in superposition; without it, the superposed hidden layer cannot implement the non-linear mapping
Summary
The non-linear activation in a hidden layer is what lets a superposed representation actually perform useful computation; without it, the shared neurons can only do simple linear mixing and the whole mechanism for packing multiple features into the same space collapses. In practical terms, any system that relies on superposition must treat that non-linearity as load-bearing, not optional, or the approach stops working.