engels-2024-sae-train-hyperparams
IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references.md
Created 2026-08-25T02:58:03+00:00
The SAE toy experiments in Engels et al. (Appendix D) use Adam optimizer, learning rate 10⁻³, sparsity penalty λ = 0.1, 20,000 training steps, and 1000-step warmup.
Summary
This records the exact training recipe behind the small-scale sparse autoencoder demos in the Engels et al. appendix, so that anyone trying to reproduce or compare against those results knows precisely what optimizer, learning rate, sparsity setting, and step counts were used. Without this pinned down, it would be easy to misattribute differences in behavior to the SAE architecture itself rather than to an unknown training configuration.