solu-activation-neuron-basis-alignment
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity-chunk-6.md
Created 2026-08-25T02:57:55+00:00
The SoLU activation function (Elhage et al.) was designed to replace ReLU to push features toward neuron-basis alignment, but may make other neurons less interpretable
Summary
SoLU was introduced as a drop-in replacement for ReLU to make individual neurons fire for cleaner, single concepts, but the trade-off is that it can muddle the meaning of other neurons in the network. In practice, swapping in this activation function does not give you a uniformly more interpretable model; it shifts clarity from one set of neurons to another, so any interpretability gains are local rather than global.