saey-scoring-rubric-four-dimensions

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-4.md

Created 2026-08-25T02:57:55+00:00

The human scoring rubric has four dimensions: confidence in the explanation, consistency of activations with the explanation, consistency of logit output weights with the explanation, and specificity of the explanation.

Summary

This defines the four specific lenses through which a human will judge whether an explanation of model behavior is good: whether the explainer sounds sure of themselves, whether the raw activations actually line up with what they claim, whether the output-weight patterns support the story, and whether the explanation avoids vague hand-waving. It matters because every downstream scoring decision in the system inherits these four criteria as the fixed yardstick, so any explanation that scores well must satisfy all four at once rather than just one.