tanh-l1-penalty-reduces-interpretability

IN premisesummaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-9.md

Created 2026-08-25T02:58:39+00:00

Replacing the standard L1 sparsity penalty with a tanh L1 penalty improved proxy metrics but made SAE features less interpretable.

Summary

Switching to a smoothed (tanh) version of the sparsity penalty made the SAE's quantitative scores look better, but the resulting features became harder for a human to identify and describe. This matters because the whole point of training an SAE is to get features you can actually understand, so a method that inflates proxy metrics while eroding interpretability is moving in the wrong direction.