park-2025-empirical-controls-shuffle-random

IN premise — summaries/2026/08/24/park-2024-categorical-hierarchical-concepts-s2-using-this-result-we-show-that-semantic-hierarchy-between-co.md

Created 2026-08-25T02:58:25+00:00

Three empirical controls rule out artifacts: (1) shuffling unembeddings destroys the structure, (2) random parent selection yields non-zero cosine (ruling out high-dim geometry), (3) independent 70% train splits break set-inclusion orthogonality only in shuffled embeddings.

Summary

Three sanity checks confirm the observed embedding structure is genuine rather than a methodological artifact: shuffling destroys the pattern, random assignments still produce meaningful similarity (so it is not just a quirk of high-dimensional geometry), and the orthogonality relationship only holds with the true structure, not with shuffled or independently split data. This matters because it means the system can trust that the relationships it reads out of the embeddings reflect real learned content, not noise or data leakage.

Dependents

These beliefs depend on this one: