engels-2025-sae-clustering-extension

IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sA-acknowledgments.md

Created 2026-08-25T02:58:02+00:00

The paper extends SAE dictionary learning from Bricken et al. (2023) and Cunningham et al. (2023) by adding a clustering step over dictionary elements to recover multi-dimensional irreducible subspaces.

Summary

This paper builds on earlier work that learned individual features from neural network internals by adding a grouping step that collects related features into clusters, each representing a higher-level concept that cannot be broken down further. It matters for interpretability because it reveals the hierarchical structure inside a network, showing that some "features" are actually parts of a larger, irreducible idea rather than independent atomic concepts.