sae-claude-opus-as-interpretability-judge

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md

Created 2026-08-25T02:58:38+00:00

Claude 3 Opus is used as the interpretability judge to rate SAE feature activations, distinct from Claude 3 Sonnet which is the model being analyzed, to reduce self-reference bias.

Summary

The system uses one model (Opus) to evaluate and interpret features extracted from a different model (Sonnet), so the judge isn't essentially grading its own internal representations. This separation keeps the interpretability ratings from being skewed by a model's natural tendency to rationalize or favor its own patterns.