saey-feature-splitting-base64-three-subfeatures

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-1.md

Created 2026-08-25T02:57:54+00:00

As SAE width increases, a single 'base64' feature in a small dictionary decomposes into three more specific sub-features in a larger dictionary.

Summary

When you give a sparse autoencoder more capacity, a single broad feature that was loosely tagging "base64-encoded text" actually splits into three narrower, more specific signals that were being forced into one bucket before. This matters because it tells us that coarse-grained features in smaller models are often mixtures of finer-grained concepts, so any interpretation built on a narrow dictionary should be treated as a rough sketch rather than a final decomposition.