park-2024-70-percent-training-tokens-experiment
IN premise — summaries/2026/08/24/park-2024-categorical-hierarchical-concepts-sR-references.md
Created 2026-08-25T02:58:26+00:00
In the 70% training-tokens experiment, shuffled unembeddings lose orthogonality when set inclusion is violated, while original unembeddings retain orthogonality, evidencing genuine semantic encoding.
Summary
In the 70% training experiment, shuffling the model's output-layer weights causes their distinct semantic directions to blur together, while the original arrangement keeps them cleanly separated, showing the geometry is doing real work rather than being accidental. This matters because it confirms the model has built a genuine internal map of meaning, not just a scrambled lookup table, which means downstream reasoning and interpretability work can rely on that structure being stable and meaningful.