rlhf-hard-to-specify-easy-to-judge

IN premiseentries/2026/06/21/wiki-Reinforcement_learning_from_human_feedback-chunk-1.md

Created 2026-06-21T09:50:10+00:00

RLHF is motivated by tasks that are hard to specify but easy to judge — humans can compare outputs more easily than they can write explicit reward functions

Summary

For many tasks, writing a precise rule for what "good" looks like is nearly impossible, but glancing at two examples and picking the better one is effortless. This gap is the core reason RLHF exists: it sidesteps the need for explicit specifications and instead trains on the comparisons humans can actually make, which means the whole approach is only as strong as the quality of those pairwise judgments.

Dependents

These beliefs depend on this one: