feature-hierarchy-empirical-validation

OUT derived (depth 1)

Created 2026-08-25T03:08:54+00:00 · Reviewed 2026-08-25T04:02:18+00:00

The hierarchical feature structure predicted geometrically by Park's orthogonality theorem (parent⊥child−parent) is empirically instantiated in SAE decoder-space neighborhoods: abstract/general features (transit infrastructure) subsume concrete/specific features (Golden Gate Bridge, Alcatraz), and this neighborhood hierarchy mirrors the WordNet synset hierarchy that Park validates across Gemma-2B and LLaMA-3-8B.

Justifications

This belief has 3 justifications — it is IN if any one holds.

SL — Each antecedent independently provides evidence of hierarchical feature structure (SAE concrete→abstract neighborhood, SAE generalization to wormholes, Park's geometric theorem); the conclusion is convergent evidence from both SAE and Park perspectives.

Antecedents (all must be IN):

  • IN sae-golden-gate-bridge-neighborhood-structure — Features near the Golden Gate Bridge feature (34M/31164353) include San Francisco locations (Alcatraz, Presidio), California landmarks (Lake Tahoe, Yosemite, Solano County), and further-out tourist attractions (Médoc wine region, Isle of Skye).

Unless (any of these IN defeats this justification):

  • IN feature-hierarchy-empirical-validation-v2 — The SAE decoder-space neighborhood structure exhibits a general-to-specific gradient—broader California landmarks (Lake Tahoe, Yosemite) appear near the Golden Gate Bridge feature, and the transit-infrastructure feature activates across diverse modalities (trains, ferries, wormholes), suggesting it encodes a higher-level transport abstraction rather than a single literal entity. This hierarchical, abstraction-to-concreteness pattern is consistent with the geometric property Park characterizes in LLaMA-3's representation space, where the (child − parent) vector for WordNet taxonomic relations is approximately orthogonal to the parent vector. The two observations arise in different representational spaces (SAE decoder neighborhoods vs. continuous LLM embeddings) and are best read as parallel evidence for a shared hierarchical organization rather than a direct instantiation of one by the other.
SL — Each antecedent independently provides evidence of hierarchical feature structure (SAE concrete→abstract neighborhood, SAE generalization to wormholes, Park's geometric theorem); the conclusion is convergent evidence from both SAE and Park perspectives.

Antecedents (all must be IN):

  • IN sae-transit-feature-abstract-generalization — The transit infrastructure feature (1M/3) activates on wormholes alongside physical transit entities like trains, ferries, and tunnels, suggesting it captures a higher-level transit/transport concept rather than literal entities.

Unless (any of these IN defeats this justification):

  • IN feature-hierarchy-empirical-validation-v2 — The SAE decoder-space neighborhood structure exhibits a general-to-specific gradient—broader California landmarks (Lake Tahoe, Yosemite) appear near the Golden Gate Bridge feature, and the transit-infrastructure feature activates across diverse modalities (trains, ferries, wormholes), suggesting it encodes a higher-level transport abstraction rather than a single literal entity. This hierarchical, abstraction-to-concreteness pattern is consistent with the geometric property Park characterizes in LLaMA-3's representation space, where the (child − parent) vector for WordNet taxonomic relations is approximately orthogonal to the parent vector. The two observations arise in different representational spaces (SAE decoder neighborhoods vs. continuous LLM embeddings) and are best read as parallel evidence for a shared hierarchical organization rather than a direct instantiation of one by the other.
SL — Each antecedent independently provides evidence of hierarchical feature structure (SAE concrete→abstract neighborhood, SAE generalization to wormholes, Park's geometric theorem); the conclusion is convergent evidence from both SAE and Park perspectives.

Antecedents (all must be IN):

  • IN park-2024-llama3-hierarchy-orthogonality — In LLaMA-3's representation space, the vector (child − parent) is approximately orthogonal to the parent vector for WordNet taxonomic relations, with cosine similarity approaching zero.

Unless (any of these IN defeats this justification):

  • IN feature-hierarchy-empirical-validation-v2 — The SAE decoder-space neighborhood structure exhibits a general-to-specific gradient—broader California landmarks (Lake Tahoe, Yosemite) appear near the Golden Gate Bridge feature, and the transit-infrastructure feature activates across diverse modalities (trains, ferries, wormholes), suggesting it encodes a higher-level transport abstraction rather than a single literal entity. This hierarchical, abstraction-to-concreteness pattern is consistent with the geometric property Park characterizes in LLaMA-3's representation space, where the (child − parent) vector for WordNet taxonomic relations is approximately orthogonal to the parent vector. The two observations arise in different representational spaces (SAE decoder neighborhoods vs. continuous LLM embeddings) and are best read as parallel evidence for a shared hierarchical organization rather than a direct instantiation of one by the other.

Dependents

These beliefs depend on this one: