mistral-sae-layers-8-16-24
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references-chunk-2.md
Created 2026-08-25T02:58:02+00:00
Mistral 7B SAEs are trained on layers 8, 16, and 24 of the 32-layer model, targeting different abstraction levels.
Summary
The system observes how Mistral 7B processes information by looking at three specific points along its 32-layer pipeline — an early stage, a middle stage, and a late stage — rather than just one cross-section. This gives it a multi-resolution view of what the model is doing, so it can track how raw patterns get shaped into higher-level concepts as data flows through.