lrh-empirical-limit-thing-part-not-in-unembedding
IN premise — summaries/2026/08/24/park-2023-linear-representation-s7-he-sat-on-the-throne-the.md
Created 2026-08-24T17:11:04+00:00
Of 27 linguistic concepts tested in LLaMA-2's unembedding space (spanning morphology, semantics, and cross-lingual mapping), the majority are linearly encoded, but some (e.g., thing⇒part) fail to show a clear signal, indicating the Linear Representation Hypothesis has empirical boundaries.
Summary
When researchers poked at 27 different linguistic features inside LLaMA-2's internal math, most lined up neatly as simple directions you could point to and extract, but a few (like the thing-to-part relationship) refused to show up cleanly. This matters because it warns against assuming every concept a model "knows" sits in a tidy, addressable slot; some knowledge is blended in ways that simple linear extraction just won't find, so tools built on that assumption will hit blind spots.