saey-target-architecture-one-layer-512-mlp

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-1.md

Created 2026-08-25T02:57:54+00:00

The Bricken et al. 2023 monosemanticity paper uses a one-layer transformer with a 512-neuron MLP layer as the base model for SAE training.

Summary

The Bricken et al. 2023 monosemanticity work was demonstrated on a deliberately tiny, single-layer model with only 512 hidden neurons, not on a full-scale language model. This matters because it means the paper's results prove the SAE/monosemanticity approach works in principle on a simple toy architecture, but do not by themselves establish how the method scales to the large, multi-layer models actually used in practice.