sae-training-objective-mse-plus-l1
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-9.md
Created 2026-08-25T02:58:39+00:00
The SAE training objective minimizes reconstruction error (mean squared error) plus an L1 sparsity penalty on activations, with no principled 'ground truth' for the trade-off between the two terms.
Summary
When training a Sparse Autoencoder, you balance two competing goals — faithfully reconstructing the input and keeping most features switched off — but there is no objective rule for how much of each to prioritize. That means the number and granularity of features the SAE discovers depend largely on a manually chosen knob, so any interpretation of those features carries an extra layer of uncertainty that isn't grounded in the data itself.