sae-functional-abstraction-extends-geometric-scope-v2
IN premise
Created 2026-08-25T04:33:13+00:00
SAE features activating on functional analogies (a transit feature firing on wormholes alongside physical transit entities) and on cross-modal inputs (text-trained features responding to image inputs) suggest that LLM representation spaces capture abstract, generalized concepts and shared cross-modal structure alongside the categorical polytope geometry described by Park et al., pointing toward a representational picture that extends beyond strict category membership to include non-literal generalization.
Summary
The system has noticed that individual pattern-detectors inside LLMs light up not just for literal categories but for functional equivalents across domains, like a movement-related pattern firing on both buses and wormholes, and even for inputs in a different modality than the one they were trained on. This means the internal representation space is organized around abstract, shared principles rather than strict category boxes, so any interpretability or alignment work built on top of it must account for this non-literal, cross-modal generalization instead of assuming each feature maps to one narrow concept.
Dependents
These beliefs depend on this one:
- OUT sae-functional-abstraction-extends-geometric-scope — SAE features activating on functional analogies (transit feature on wormholes) and cross-modal inputs (text-trained features firing on images) demonstrate the geometric space encodes intensional and relational structure beyond Park's extensional categorical polytopes, broadening the geometric framework's explanatory scope to include non-lexical, compositional semantics.