rlhf-trains-reward-model-from-human-preferences

IN premiseentries/2026/06/21/wiki-Reinforcement_learning-chunk-4.md

Created 2026-06-21T09:55:52+00:00

RLHF trains a reward model from human preference ratings, then uses that reward model to guide RL policy optimization — it is not direct human-in-the-loop RL