anthropic-sae-512-neuron-mlp-architecture

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity.md

Created 2026-08-25T02:57:56+00:00

The Bricken et al. (2023) paper uses a one-layer transformer with a 512-neuron MLP using ReLU activation, trained on 8 billion tokens

Summary

This establishes the experimental setup for the Bricken et al. interpretability work: a deliberately small, single-layer transformer with a modest 512-feature feed-forward block, trained on a large but not enormous corpus. Any downstream findings about how the model processes information are bounded by this simple architecture, so conclusions drawn from it may not transfer directly to deeper or larger production models.