iti-truthfulqa-attention-head-probing

IN premise — summaries/2026/08/24/xu-2024-knowledge-conflicts-survey-s4-intra-memory-conflict.md

Created 2026-08-25T02:59:03+00:00

ITI (Li et al., 2023c) identifies sparse attention heads with high TruthfulQA linear-probing accuracy, then shifts activations along the truth-correlated direction at each autoregressive step.

Summary

A small handful of attention heads in a language model carries most of the signal that distinguishes truthful from untruthful answers, and by gently nudging the model's internal state toward that signal at every step of token generation, you can steer it to produce more factually grounded output without retraining. In practice, this means truthfulness is not spread diffusely across the network but lives in a few identifiable components that can be read from and adjusted.