saey-five-validation-criteria-monosemantic-feature

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-3.md

Created 2026-08-25T02:57:54+00:00

The five validation criteria for a monosemantic feature are: high specificity, high sensitivity, causal downstream effect, non-correspondence to any single neuron, and replication in an independently trained model.

Summary

This sets a five-point checklist for confirming that a pattern found inside a neural network is a genuine, single concept rather than a statistical fluke. It matters because without all five checks — the pattern fires only for the right inputs, reliably catches them when present, actually changes the model's output when poked, isn't just one quirky neuron, and shows up in a separately trained model — anyone could mistake noise for real interpretability and build wrong conclusions on top of it.