feature-hierarchy-empirical-validation-v2

IN premise

Created 2026-08-25T04:06:11+00:00

The SAE decoder-space neighborhood structure exhibits a general-to-specific gradient—broader California landmarks (Lake Tahoe, Yosemite) appear near the Golden Gate Bridge feature, and the transit-infrastructure feature activates across diverse modalities (trains, ferries, wormholes), suggesting it encodes a higher-level transport abstraction rather than a single literal entity. This hierarchical, abstraction-to-concreteness pattern is consistent with the geometric property Park characterizes in LLaMA-3's representation space, where the (child − parent) vector for WordNet taxonomic relations is approximately orthogonal to the parent vector. The two observations arise in different representational spaces (SAE decoder neighborhoods vs. continuous LLM embeddings) and are best read as parallel evidence for a shared hierarchical organization rather than a direct instantiation of one by the other.

Summary

The model's internal features are arranged on a spectrum from broad categories (like "transportation" or "California landmarks") down to specific instances (like "Golden Gate Bridge" or "Yosemite"), and this same general-to-specific geometry shows up independently in two different measurement spaces. This matters because it suggests the model genuinely organizes knowledge hierarchically rather than just clustering similar-sounding concepts, giving the system a principled way to navigate from abstract to concrete.

Dependents

These beliefs depend on this one: