sae-expansion-ratio-10-to-200x

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity.md

Created 2026-08-25T02:58:39+00:00

Sparse Autoencoders decompose dense d-dimensional activations into a typically 10× to 200× larger set of mostly-sparse, near-orthogonal latent features.

Summary

Instead of working with a single compressed signal, the system unpacks neural network activity into roughly 10 to 200 times as many mostly-silent, independent features. This means every downstream calculation operates on a much larger but far more separable representation, trading raw density for clarity at the cost of significant expansion in dimensionality.

Dependents

These beliefs depend on this one: