sae-vs-fewshot-steering-comparison
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-10.md
Created 2026-08-25T02:58:36+00:00
In the scaling monosemanticity paper's tests, SAE-based features outperformed few-shot probe steering vectors in 5 of 7 test cases (secrecy, sycophancy, code errors, self-improving AI, methamphetamine), while both were effective for gender bias and agreement features.
Summary
When comparing two ways to nudge an AI model's behavior, the more sophisticated technique of steering along learned monosemantic feature directions beat the simpler few-shot probe method in most tested scenarios, including harder cases like secrecy, sycophancy, and self-improving AI. This implies that for broad and subtle behavioral control, investing in SAE-based feature extraction is worth it, though for well-understood traits like gender bias the cheaper probe approach is just as good.