sae-cross-modal-text-trained-image-activation
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md
Created 2026-08-25T02:58:38+00:00
SAE features trained on text-only data also activate on relevant image inputs, indicating shared latent cross-modal representations in Claude 3 Sonnet.
Summary
Features learned purely from text in Claude 3 Sonnet still light up when the model processes matching images, suggesting the model builds a shared internal concept space that doesn't care whether information arrived as words or pixels. This matters because it means interpretability tools trained on text can still be used to probe and explain the model's visual reasoning, rather than requiring a completely separate toolkit for each input type.
Dependents
These beliefs depend on this one:
- OUT sae-functional-abstraction-extends-geometric-scope — SAE features activating on functional analogies (transit feature on wormholes) and cross-modal inputs (text-trained features firing on images) demonstrate the geometric space encodes intensional and relational structure beyond Park's extensional categorical polytopes, broadening the geometric framework's explanatory scope to include non-lexical, compositional semantics.