text-sae-generalizes-to-image-activations

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-9.md

Created 2026-08-25T02:58:39+00:00

SAEs trained purely on text activations generalize zero-shot to image activations (off-distribution), as reported by Anthropic in the scaling monosemanticity paper.

Summary

Anthropic found that a feature-detection tool trained only on text data works immediately on image data, without any image-specific training, suggesting the underlying structure of concepts is shared across modalities rather than being modality-specific. This matters because it implies interpretability tools built for one type of data may transfer broadly, reducing the cost of understanding what a model is doing across text and images.