three-superposition-types

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-9.md

Created 2026-08-25T02:58:39+00:00

The Templeton 2024 paper identifies three distinct types of superposition that complicate mechanistic understanding: activation superposition (addressed by SAEs), attention superposition (features packed across attention heads), and weight superposition (interference in learned weights).

Summary

This comes from the Templeton 2024 interpretability paper and matters because "understanding a neural network" isn't one problem — features get mixed up in at least three separate places (inside individual neurons, across attention heads, and within the learned weight matrices), and each requires its own diagnostic tool. For the system, any claim about what a model is doing should specify which of these three layers it's talking about, or the conclusion is likely incomplete.