park-2025-empirical-controls-shuffle-random
IN premise — summaries/2026/08/24/park-2024-categorical-hierarchical-concepts-s2-using-this-result-we-show-that-semantic-hierarchy-between-co.md
Created 2026-08-25T02:58:25+00:00
Three empirical controls rule out artifacts: (1) shuffling unembeddings destroys the structure, (2) random parent selection yields non-zero cosine (ruling out high-dim geometry), (3) independent 70% train splits break set-inclusion orthogonality only in shuffled embeddings.
Summary
Three sanity checks confirm the observed embedding structure is genuine rather than a methodological artifact: shuffling destroys the pattern, random assignments still produce meaningful similarity (so it is not just a quirk of high-dimensional geometry), and the orthogonality relationship only holds with the true structure, not with shuffled or independently split data. This matters because it means the system can trust that the relationships it reads out of the embeddings reflect real learned content, not noise or data leakage.
Dependents
These beliefs depend on this one:
- IN shuffling-contradiction-resolution — The apparent contradiction between the 2025 empirical control (shuffling unembeddings *destroys* the full orthogonality structure) and the set-inclusion finding (shuffled unembeddings *reproduce* the child-parent⊥parent sub-orthogonality) is resolved by scope: the *full* four-condition Theorem 8 structure is genuine and shuffle-destroyed, while a *specific* sub-condition (consecutive-level parent⊥child-parent) is a set-inclusion artifact—meaning the geometric framework captures the genuine full structure, not the spurious sub-structure.