saey-activations-collected-from-8b-datapoints

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-1.md

Created 2026-08-25T02:57:54+00:00

SAE training activations were collected from 8 billion data points of the one-layer transformer's MLP layer.

Summary

This is a straightforward record of where the training data came from: 8 billion activation snapshots were pulled from the feedforward (MLP) layer of a single-layer transformer to train the sparse autoencoder. It establishes the scale and provenance of the data the SAE learned from, so any feature or interpretation built on top of it ultimately traces back to that specific layer and that volume of examples.