lrh-three-interpretations-subspace-measurement-intervention
IN premise — summaries/2026/08/24/park-2023-linear-representation-s1-introduction.md
Created 2026-08-24T17:11:01+00:00
The Linear Representation Hypothesis has three distinct formal interpretations: subspace (a concept is a 1-dimensional direction, e.g., γ("queen") − γ("king")), measurement (read out by a linear probe, logit-linear in the representation), and intervention (changed by adding a steering vector), with the paper proving unembedding-subspace maps to measurement and embedding-subspace maps to intervention.
Summary
These three ways people describe how a concept is encoded in a neural network — as a geometric direction, as something you can read out with a linear test, and as something you can nudge by adding a vector — are not independent ideas. The paper shows the geometric direction directly tells you the measurement recipe, and the embedding direction directly tells you the intervention recipe, so knowing one unlocks the others and unifies the vocabulary.