gpt2-small-sae-25k-features-layer7
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references-chunk-2.md
Created 2026-08-25T02:58:03+00:00
GPT-2-small clustering uses approximately 25k SAE features from layer 7 (sourced from Bloom 2024) with spectral clustering at n_clusters=1000.
Summary
This sets the scope for an interpretability analysis of GPT-2-small: researchers are grouping roughly 25,000 learned features pulled from the model's seventh layer into 1,000 natural clusters to see how the model organizes its internal representations. Every downstream finding about what the model "knows" or "thinks" at that layer is bounded by this particular feature set and clustering method, so conclusions drawn from it only apply to this specific decomposition of the model's internal states.