causal-separability-not-universal-among-concepts
IN premise — summaries/2026/08/24/park-2023-linear-representation-s2-the-linear-representation-hypothesis.md
Created 2026-08-24T17:11:02+00:00
In Park et al. (2023), concepts are causally separable only if they can be varied freely: English⇒French and Male⇒Female are separable, but English⇒French and English⇒Russian are not (they share a source language), and the concept 'thing⇒part' is explicitly not encoded in the unembedding space Γ at all.
Summary
Not every pair of concept dimensions in a model can be pulled independently; some are locked together because they share a common cause, like two target languages both depending on the same source, and some concept relationships simply do not exist in the model's output space at all. This means any system that reasons about or manipulates concepts must account for these structural couplings and gaps rather than assuming the concept space is a free, fully independent grid where every axis can be set on its own.