park-2025-whitening-final-layer-only
IN premise — summaries/2026/08/24/park-2024-categorical-hierarchical-concepts-s2-using-this-result-we-show-that-semantic-hierarchy-between-co-chunk-2.md
Created 2026-08-25T02:58:24+00:00
The whitening transformation used to define the canonical representation space applies only to the final layer's unembedding matrix, leaving internal-layer geometry as an open problem.
Summary
The mathematical normalization that gives the system a clean, standardized coordinate space for analyzing representations currently only works for the very last layer that produces output. Everything happening in the intermediate layers still lacks a comparable standardized geometry, so any analysis of how representations evolve through a model is built on an incomplete foundation.
Dependents
These beliefs depend on this one:
- OUT residual-stream-universal-substrate — SAE (middle-layer residual stream), ROME (mid-layer MLP value projection), and Park (final-layer unembedding) all identify the residual stream at different depths as the primary locus of interpretable geometric structure.