steering-clamp-effective-range
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-10.md
Created 2026-08-25T02:58:36+00:00
In the scaling monosemanticity paper, effective feature steering clamping values range from −10 to +10 times the maximum observed activity, while values of ±100× cause degenerate output.
Summary
When you try to manually steer a specific feature in a language model, there is a practical window of roughly ten times its normal peak activity where the intervention produces clean, controlled changes to the output. Push that much further, to around a hundred times, and the model's output collapses into incoherent garbage, so any useful steering work must stay well within that safe band.