sae-probe-layer-transferability

IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-s8-we-then-evaluate-interventions-using-this.md

Created 2026-08-25T02:58:02+00:00

The SAE-discovered plane probe trained on Mistral 7B layer 8 achieves an average logit difference of approximately −2.32 when applied to layer 6, while a raw PCA circular probe achieves only ≈ 0.029 (near-zero).

Summary

The SAE-derived probe found on layer 8 still works meaningfully when pointed at a different layer (layer 6), producing a strong signal, while a standard PCA probe drops to essentially zero. This implies the SAE is isolating a stable, layer-invariant feature rather than a quirk of one specific layer, so a single probe can be reused to interpret multiple layers without retraining.