rlhf-reward-model-cross-entropy-bradley-terry
OUT premise — entries/2026/06/21/wiki-Reinforcement_learning_from_human_feedback.md
Created 2026-06-21T09:50:10+00:00
The RLHF reward model is trained with a cross-entropy loss over pairwise preferences using the Bradley-Terry-Luce (BTL) probabilistic model: L(θ) = -1/C(K,2) * E[log σ(r_θ(x, y_w) - r_θ(x, y_l))].