sae-universality-across-models

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity.md

Created 2026-08-25T02:57:56+00:00

SAEs applied to different transformer models produce mostly similar features—more similar to each other than to their own model's neurons—suggesting features reflect data structure rather than architecture

Summary

The features that sparse autoencoders extract from different transformer models look nearly the same across models, and are actually closer to each other than to the original neurons they were decoded from. This matters because it means the features are capturing something real about the data itself, not quirks of a particular model's wiring, so a single feature dictionary is likely to be meaningful no matter which model you plug it into.

Dependents

These beliefs depend on this one: