saey-412-intervals-scored-162-features-neurons

IN premisesummaries/2026/08/24/bricken-2023-monosemanticity-chunk-4.md

Created 2026-08-25T02:57:55+00:00

The human evaluation in the global interpretability analysis scored 412 feature activation intervals across 162 features and neurons.

Summary

A human evaluator manually cataloged 412 specific activation windows spread across 162 features and neurons during a whole-model interpretability review, giving the system a concrete, human-verified map of what the network is actually doing internally. This anchors the interpretability record to a specific, person-checked level of detail rather than an automated approximation, and sets the scope for what downstream reasoning can reasonably assume has been understood.