sae-transit-feature-abstract-generalization
IN premise — summaries/2026/08/24/templeton-2024-scaling-monosemanticity-chunk-4.md
Created 2026-08-25T02:58:38+00:00
The transit infrastructure feature (1M/3) activates on wormholes alongside physical transit entities like trains, ferries, and tunnels, suggesting it captures a higher-level transit/transport concept rather than literal entities.
Summary
A feature in the sparse autoencoder that we'd expect to track literal transit hardware (trains, ferries, tunnels) also lights up when the text is about wormholes, which are fictional or abstract transport mechanisms. This tells us the feature has latched onto the underlying idea of "a way to move between places" rather than just the specific physical objects, which is stronger evidence that these features capture functional concepts instead of surface-level labels.
Dependents
These beliefs depend on this one:
- OUT feature-hierarchy-empirical-validation — The hierarchical feature structure predicted geometrically by Park's orthogonality theorem (parent⊥child−parent) is empirically instantiated in SAE decoder-space neighborhoods: abstract/general features (transit infrastructure) subsume concrete/specific features (Golden Gate Bridge, Alcatraz), and this neighborhood hierarchy mirrors the WordNet synset hierarchy that Park validates across Gemma-2B and LLaMA-3-8B.
- OUT sae-functional-abstraction-extends-geometric-scope — SAE features activating on functional analogies (transit feature on wormholes) and cross-modal inputs (text-trained features firing on images) demonstrate the geometric space encodes intensional and relational structure beyond Park's extensional categorical polytopes, broadening the geometric framework's explanatory scope to include non-lexical, compositional semantics.