sae-interpretability-rubric-0-3-scale

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md

Created 2026-08-25T02:58:37+00:00

The automated interpretability rubric uses a 0–3 scale where 0 means feature is irrelevant to context and 3 means the feature cleanly identifies the activating text.

Summary

This is the scoring yardstick the system uses to judge whether an interpretability feature actually tells you something useful about what a model is computing on a given input. It matters because every downstream assessment of feature quality, reliability, or usefulness in the SAE interpretability pipeline is anchored to what counts as a meaningful hit on this four-point scale.