multi-prompt-filtering-logic
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-7.md
Created 2026-08-25T02:58:39+00:00
Multi-prompt filtering retains only features that are active on ALL positive prompts AND inactive on ALL negative prompts, reducing noise from syntax- or content-nonspecific activations.
Summary
When multiple prompts are used together as a filter, only the signals that light up across every single positive example and stay quiet across every single negative one get kept, while everything else is discarded. This matters because it strips out generic, context-independent noise — the kind of activation that fires for any prompt regardless of what's actually being asked — leaving behind only the features that genuinely distinguish "right" from "wrong."