intervention-linearly-shifts-target-logit

IN premise — summaries/2026/08/24/park-2023-linear-representation-s5-discussion-and-related-work.md

Created 2026-08-24T17:11:03+00:00

Adding α·λ̄_W to the model's representation (λ_C,α(x) = λ(x) + α·λ̄_W) with α ∈ [0, 0.4] linearly increases the target concept's logit while leaving causally separable concept logits unchanged; at α=0.4 in the 'Long live the' / male⇒female experiment, 'king' fell entirely out of the top-5 predictions.

Summary

Nudging the model's internal state along a specific concept direction shifts its prediction for that concept in a predictable, proportional way without disturbing scores for unrelated concepts. This suggests the model stores concepts in relatively isolated, linear directions, meaning you can steer or edit its behavior along one axis with minimal collateral effect on everything else.