sae-feature-search-four-methods
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-7.md
Created 2026-08-25T02:58:39+00:00
Anthropic's SAE feature search employs four distinct methods: single-prompt activation, multi-prompt positive/negative filtering, geometric cosine-similarity nearest-neighbor search, and logit-difference attribution.
Summary
Anthropic's pipeline for surfacing interpretable features in their sparse autoencoder relies on a fixed toolkit of four specific search techniques, from single-input activation to geometric similarity matching. This bounds what the system can actually discover: if a feature isn't caught by one of those four methods, it simply does not enter the interpretability picture, so the search is only as broad as that small menu of approaches allows.