sae-cross-modal-text-trained-image-activation

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md

Created 2026-08-25T02:58:38+00:00

SAE features trained on text-only data also activate on relevant image inputs, indicating shared latent cross-modal representations in Claude 3 Sonnet.

Summary

Features learned purely from text in Claude 3 Sonnet still light up when the model processes matching images, suggesting the model builds a shared internal concept space that doesn't care whether information arrived as words or pixels. This matters because it means interpretability tools trained on text can still be used to probe and explain the model's visual reasoning, rather than requiring a completely separate toolkit for each input type.

Dependents

These beliefs depend on this one: