mistral-sae-1b-tokens-datasets

IN premise — summaries/2026/08/24/engels-2024-not-all-features-linear-sR-references-chunk-2.md

Created 2026-08-25T02:58:03+00:00

Mistral 7B SAEs are trained on more than 1 billion tokens drawn from the Pile and Alpaca datasets.

Summary

The interpretability layer built on top of Mistral 7B was trained on over a billion text examples pulled from two specific corpora, the Pile and Alpaca, so the features it extracts will reflect whatever kinds of content those datasets emphasize. This matters because it sets a boundary on what the system can interpret: concepts or patterns that are underrepresented in those two sources may be invisible to the autoencoder, and any downstream reasoning about model behavior inherits that blind spot.