daa-beta-controls-deviation-from-reference

IN premiseentries/2026/06/21/wiki-Reinforcement_learning_from_human_feedback-chunk-3.md

Created 2026-06-21T09:50:10+00:00

The β parameter in all Direct Alignment Algorithms (DPO, IPO, KTO) serves as KL regularization strength controlling deviation from the reference SFT policy — higher β keeps the policy closer to the reference

Summary

The beta parameter in alignment methods like DPO, IPO, and KTO acts as a safety dial: it sets how far the model's behavior is allowed to drift from its original, pre-alignment version. Cranking beta higher keeps the model closer to what it already knew, trading alignment aggressiveness for stability, so it's the key knob for balancing how much personality the alignment process is allowed to reshape.

Dependents

These beliefs depend on this one: