sae-interpretability-rubric-0-3-scale
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md
Created 2026-08-25T02:58:37+00:00
The automated interpretability rubric uses a 0–3 scale where 0 means feature is irrelevant to context and 3 means the feature cleanly identifies the activating text.
Summary
This is the scoring yardstick the system uses to judge whether an interpretability feature actually tells you something useful about what a model is computing on a given input. It matters because every downstream assessment of feature quality, reliability, or usefulness in the SAE interpretability pipeline is anchored to what counts as a meaningful hit on this four-point scale.