llama2-unembedding-not-universal-concept-encoder
IN premise — summaries/2026/08/24/park-2023-linear-representation-s11-he-wore-a-crown-signifying-he-was-the.md
Created 2026-08-24T17:11:02+00:00
Not all concepts are linearly encoded in LLaMA-2's unembedding space (Γ); 'thing ⇒ part' is explicitly cited as a counterexample where the unembedding representation fails to capture the concept, establishing a boundary condition for the Linear Representation Hypothesis.
Summary
You can't assume every semantic relationship in LLaMA-2 shows up as a simple, readable direction in its final output layer; the part-whole relationship (e.g., "a tree has branches") is a known case where that clean linear readout breaks down. This sets a hard limit on how far the "concepts are just directions in the output space" shortcut can be pushed, meaning anyone interpreting or steering the model needs to check whether a given relationship actually lives linearly before relying on that assumption.