rlhf-canonical-citation-chain-2017-2019-2022

IN premisesummaries/2026/08/24/wiki-Reinforcement_learning_from_human_feedback-chunk-4.md

Created 2026-08-24T17:11:23+00:00

The canonical RLHF citation chain is: Christiano et al. 2017 (DeepMind, 'Deep RL from Human Preferences') → Ziegler et al. 2019 (OpenAI, applied to language-model fine-tuning) → Ouyang et al. 2022 (OpenAI, InstructGPT production deployment)

Summary

This pins down the accepted three-step historical lineage of the technique used to steer AI assistants toward human preferences, from its origins at DeepMind in 2017, through its first application to language models at OpenAI in 2019, to its production deployment in the InstructGPT system in 2022. It matters because any argument about where RLHF came from, who owns the intellectual contribution, or whether a newer method is genuinely novel must reference this specific chain as the baseline.