safety-feature-cross-modal-activation
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-8.md
Created 2026-08-25T02:58:39+00:00
Safety-relevant features such as 'unsafe code' and 'backdoor' activate on both text prompts and image inputs (e.g., hidden cameras, keylogger ads), indicating shared representational structure across modalities.
Summary
The model recognizes dangerous concepts like backdoors and malicious code through a single shared internal representation, whether those concepts arrive as written text or as images. This means you can't lock down safety by filtering one input type and ignoring the other, since the underlying "danger detector" fires on both, and an attacker could route a harmful request through whichever modality has looser guardrails.