safety-feature-steering-causal-not-correlational

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-8.md

Created 2026-08-25T02:58:39+00:00

Steering experiments (positive and negative activation scaling) demonstrate that safety-relevant SAE features causally influence model generated output, not merely correlate with unsafe content.

Summary

When researchers deliberately turn specific safety-related circuits in a language model up or down, the model's output actually changes in predictable ways. This means those circuits are not just patterns that happen to sit near unsafe text; they are the actual machinery producing the behavior, which gives the system a real lever it can pull to intervene.