park-2023-27-concepts-thing-part-not-encoded
IN premise — summaries/2026/08/24/park-2023-linear-representation-s4-arthur-was-a-legendary.md
Created 2026-08-24T17:11:03+00:00
Park et al. (2023) test 27 concepts in LLaMA-2's unembedding space Γ; most are linearly encoded, but 'thing⇒part' is a named counterexample that is absent from the unembedding space, demonstrating not all concepts have linear representations.
Summary
A 2023 study of LLaMA-2 showed that while most conceptual relationships (like synonymy or antonymy) can be read off the model's output geometry as simple linear directions, the basic "a part belongs to a whole" relationship does not appear there at all. This matters because it means any tool or audit that assumes it can find every concept by inspecting the model's unembedding space will silently miss important knowledge, and the model's internal representations are more tangled than a flat map of its vocabulary would suggest.
Dependents
These beliefs depend on this one:
- IN linearity-boundary-of-geometric-framework — The geometric framework's explanatory boundary is precisely the boundary of linearity: ROME cannot edit non-factual (logical/spatial/numerical) associations because they are not rank-one key-value pairs, and Park's framework cannot encode the 'thing⇒part' relation because it is a non-linear constraint in the unembedding space—both failures share the root cause of the linear-subspace assumption.