sae-activation-normalization-rule

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-3.md

Created 2026-08-25T02:58:37+00:00

Activations are scalar-normalized so their average squared L2 norm equals the residual stream dimension D before SAE training.

Summary

Before training a sparse autoencoder, the model's internal signals are rescaled so their average magnitude lands on a fixed, predictable baseline tied to the model's width. This matters because it keeps the SAE from wasting its capacity learning to compensate for arbitrary input sizes, letting it focus on discovering which features actually light up.