superposition-capacity-linear-in-dimension

IN premisesummaries/2026/08/24/elhage-2022-toy-models-superposition-chunk-12.md

Created 2026-08-25T02:57:59+00:00

In the Elhage et al. 2022 toy model, the maximum number of recoverable features is linear (not exponential) in the embedding dimension m, with the proportionality constant depending on sparsity density S: m = Ω(-n*(1-S) log(1-S)).

Summary

In the Elhage et al. toy model, the number of features you can reliably extract from a representation grows only in direct proportion to its size, not exponentially — so doubling the embedding dimension at most doubles the recoverable feature count. This means there is a hard linear ceiling on how many distinct features a given representation width can carry, and sparsity only shifts the constant, not the scaling law itself.