residual-stream-universal-substrate-v2

IN premise

Created 2026-08-25T04:31:46+00:00

SAE targets the middle-layer residual stream, ROME edits the mid-layer FFN value projection (W_V), and Park's whitening applies only to the final-layer unembedding matrix. Each approach selects a specific depth as its intervention or representation target, but they do not jointly identify the residual stream as the primary locus of interpretable geometric structure; in particular, Park explicitly leaves internal-layer geometry as an open problem.

Summary

Each of the major interpretability tools we track — sparse autoencoders, ROME-style edits, and whitening — picks its own single layer to work with, and none of them step back to treat the residual stream as the place where the model's meaningful structure actually lives. This matters because our current toolkit ends up treating model internals as a collection of independent layer-by-layer hacks rather than a unified geometric picture, and at least one of these approaches openly admits that understanding the internal layers is still an unsolved problem.

Dependents

These beliefs depend on this one: