saey-five-validation-criteria-monosemantic-feature
IN premise — summaries/2026/08/24/bricken-2023-monosemanticity-chunk-3.md
Created 2026-08-25T02:57:54+00:00
The five validation criteria for a monosemantic feature are: high specificity, high sensitivity, causal downstream effect, non-correspondence to any single neuron, and replication in an independently trained model.
Summary
This sets a five-point checklist for confirming that a pattern found inside a neural network is a genuine, single concept rather than a statistical fluke. It matters because without all five checks — the pattern fires only for the right inputs, reliably catches them when present, actually changes the model's output when poked, isn't just one quirky neuron, and shows up in a separately trained model — anyone could mistake noise for real interpretability and build wrong conclusions on top of it.