sae-attribution-ablation-0-8-correlation

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-6.md

Created 2026-08-25T02:58:38+00:00

In the worked example (John says 'I want to be alone' → John feels ___), ablating every active feature yielded a 0.8 correlation with attribution scores, validating attribution as a reasonable proxy.

Summary

In a simple test case (filling in what John feels after saying "I want to be alone"), turning off each feature one at a time produced results that matched the system's attribution scores 80 percent of the time, which is strong agreement. This means the system can trust its attribution scores as a reliable shortcut for identifying which features actually drive an output, without having to run a full ablation every time.

Dependents

These beliefs depend on this one: