rlhf-reward-model-cross-entropy-bradley-terry

OUT premiseentries/2026/06/21/wiki-Reinforcement_learning_from_human_feedback.md

Created 2026-06-21T09:50:10+00:00

The RLHF reward model is trained with a cross-entropy loss over pairwise preferences using the Bradley-Terry-Luce (BTL) probabilistic model: L(θ) = -1/C(K,2) * E[log σ(r_θ(x, y_w) - r_θ(x, y_l))].