Part of Language AI Handbook
Derive the DPO loss function from first principles. Explains how the optimal RLHF policy leads to reward reparameterization and direct preference optimization.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
DPO Derivation
In the previous chapter, we introduced Direct Preference Optimization as a simpler alternative to RLHF that eliminates the need for an explicit reward model. We saw that DPO directly optimizes a language model using preference data, but we didn't examine why this works or how the DPO loss function is derived. This chapter fills that gap by tracing every algebraic step, from the original reinforcement learning objective all the way to the final training loss.
The DPO derivation is a key result in alignment research. It shows that the optimal policy for the RLHF objective has a closed-form solution, and this solution can be rearranged to express the implicit reward in terms of the policy itself. When we substitute this reparameterized reward into the Bradley-Terry preference model, we obtain the DPO loss function directly, with no reinforcement learning required. This chain of reasoning is surprisingly elegant, and understanding it will fundamentally change how you think about what a language model is doing when it trains on human preference data.
Understanding this derivation reveals why DPO works, its underlying assumptions, and how to interpret the model's learning process. By the end of this chapter, you'll see that DPO is essentially solving a classification problem where the model learns to assign higher probability to preferred responses. Each gradient step simultaneously rewards the model for raising the probability of good outputs and penalizes it for raising the probability of bad ones, all without ever explicitly parameterizing a reward function.
Think of the whole derivation as a three-act story. Act one asks: if we had the ideal reward function, what would the optimal policy look like? Act two reverses that question and asks: if we have the optimal policy, what reward function does it imply? Act three substitutes that reversed reward into the preference model and discovers that the messy partition function term vanishes, leaving a clean, tractable loss. Each act is mathematically precise, and each act builds directly on the previous one. When you finish reading, the chain from RLHF to DPO will feel inevitable rather than mysterious.
One reason this derivation matters beyond academic interest is that it reveals the hidden assumptions baked into DPO. The loss function is not a heuristic: it follows from specific, clearly stated premises about the form of the RLHF objective, the Boltzmann structure of optimal policies, and the Bradley-Terry preference model. Knowing those premises tells you exactly when DPO is appropriate, what might break it, and how to modify it for different settings. We will examine those assumptions carefully in the section on limitations.
Direct Preference Optimization was introduced by Rafailov et al. in 2023 in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model." The central insight, that the RLHF objective admits a closed-form optimal policy whose reward can be reparameterized in terms of policy ratios, appeared almost simultaneously with work on "SLiC-HF" (Zhao et al., 2023) and related ideas in reward-free offline RL. What made DPO distinctive was its clean derivation that tied together three well-known components (the KL-regularized RL objective, the Boltzmann optimal policy, and the Bradley-Terry model) into a single differentiable loss. Within months, DPO became one of the most widely used alignment techniques in practice, appearing in the training recipes of Llama 2, Mistral Instruct, and dozens of other models. The paper has become one of the most cited alignment papers of the 2020s, and its derivation has inspired a family of variants including IPO, KTO, and ORPO, which we cover in the next chapter.
The RLHF Objective
Let's begin with the objective that RLHF aims to optimize. As we discussed in the chapters on PPO for Language Models and KL Divergence Penalty, the goal is to find a policy that maximizes expected reward while staying close to a reference policy (typically the supervised fine-tuned model).
Before diving into the mathematics, it helps to understand the intuition behind this objective. We want our language model to generate responses that humans prefer, which is captured by the reward function. However, if we optimize the reward too aggressively, the model might find unexpected shortcuts or produce degenerate outputs that technically achieve high reward but don't represent helpful behavior. The reference policy is an anchor, representing the model's pre-trained knowledge and natural language capabilities. By penalizing deviations from this anchor, we encourage the model to improve its outputs while maintaining coherent, fluent generation.
Think of the reference policy as the starting point on a map, and the reward function as a compass pointing toward a better destination. The KL penalty acts like a tether that keeps you from wandering too far afield in a single step. You want to move in the direction of higher reward, but you don't want to teleport so far that you lose the path entirely. In language modeling, that "losing the path" corresponds to mode collapse or reward hacking: outputs that score well on a proxy reward metric but are nonsensical or harmful. The KL constraint is what prevents those degenerate solutions.
The constrained optimization problem is:
where:
- : the policy being optimized
- : the expectation operator
- : a prompt sampled from the data distribution
- : a response sampled from the policy
- : the learned reward function
- : a hyperparameter controlling the strength of the KL constraint
- : the Kullback-Leibler divergence
- : the reference policy
To work with this objective more directly, we need to expand the KL divergence term. Recall that the KL divergence measures how different one probability distribution is from another, expressed as the expected log ratio of the two distributions. Writing out the KL divergence explicitly, this becomes:
where:
- : the expectation operator
- : the data distribution
- : the reward function
- : the KL penalty coefficient
- : the natural logarithm
- : probability of response given prompt under the policy
- : probability of response given prompt under the reference model
This reformulation reveals the objective's structure more clearly. The log ratio measures how much the policy has shifted away from the reference for a particular response. When this ratio is positive (the policy assigns more probability than the reference), the KL term subtracts from the objective, penalizing the deviation. When the ratio is negative (the policy assigns less probability), the term adds to the objective, which might seem like a reward, but since we're sampling from , we're unlikely to generate responses where assigns low probability.
The key insight is that this scalar objective, combining a reward signal and a distributional constraint, can be solved analytically. That analytical solution is the foundation of everything that follows. The reward-KL tradeoff creates a specific mathematical structure that admits an exact solution in closed form, rather than acting as an arbitrary heuristic regularizer.
This objective captures a basic tension in alignment: we want the model to produce high-reward outputs (as judged by human preferences), but we don't want it to deviate too far from its pre-trained behavior. Without the KL penalty, the model might find degenerate solutions that exploit flaws in the reward model, a phenomenon we explored in the chapter on Reward Hacking.
Why the KL Penalty Has This Form
Before proceeding to the derivation, it is worth spending a moment on why the KL divergence is the natural choice of regularizer here. Many other distance measures between distributions exist, such as total variation distance, Renyi divergences, or the Wasserstein distance. Each would produce a different optimization problem and a different optimal policy. The KL divergence is particularly convenient because it has exactly the right algebraic structure to yield a clean closed-form solution. Its definition as means that when we write the full objective as an expectation under , everything becomes a single expectation of a per-token scalar quantity, which is tractable.
The KL divergence also has information-theoretic grounding. It measures the number of extra "bits" required to encode samples from using a code optimized for . Keeping this quantity small means the policy's outputs remain within the statistical neighborhood of the reference, which in practice means they retain the grammatical fluency, factual grounding, and stylistic coherence of the original supervised fine-tuned model.
Deriving the Optimal Policy
The key insight behind DPO is that this optimization problem has a closed-form solution. Unlike many optimization problems in machine learning that require iterative gradient descent, this particular formulation admits an analytical answer. For a fixed prompt , we're optimizing over the distribution for all possible responses . This is a constrained optimization problem over probability distributions: we must ensure our solution is a valid probability distribution that sums to one and assigns non-negative probability to every possible response.
The special structure of the RLHF objective allows for this closed-form solution. The expectation under combined with the KL divergence creates what's known as a variational problem over distributions. Such problems often have elegant solutions when the constraints are simple probability simplex constraints.
Let's work through the derivation step by step. For a single prompt , the objective is:
where:
- : sum over all possible responses in the vocabulary
- : the entire probability distribution over responses for prompt
- : probability of response given prompt
- : the reward function
- : the KL penalty coefficient
- : the natural logarithm
- : the reference policy probability
We need to maximize this subject to the constraint that is a valid probability distribution: and for all .
To solve this constrained optimization problem, we use the method of Lagrange multipliers, a classical technique from calculus. The idea is to incorporate the constraint directly into the objective by introducing a new variable (the Lagrange multiplier) that penalizes violations of the constraint. At the optimal solution, the gradient of the objective with respect to the decision variables must be proportional to the gradient of the constraint.
The Lagrangian is:
where:
- : the Lagrangian function
- : sum over all responses
- : probability of response
- : the reward function
- : the KL penalty coefficient
- : the natural logarithm
- : the reference policy probability
- : the Lagrange multiplier for the constraint that probabilities sum to 1
Notice that we've expanded the log ratio term from the original objective, separating it into and . This separation makes taking derivatives more straightforward. The term enforces our normalization constraint: if the probabilities don't sum to one, this term will be nonzero, and the Lagrange multiplier will adjust to push us toward a valid distribution.
Taking the derivative with respect to for a specific :
where:
- : the partial derivative of the Lagrangian with respect to the probability of response
- : the reward function
- : the KL penalty coefficient
- : the natural logarithm
- : the policy probability
- : the reference policy probability
- : the Lagrange multiplier
The term comes from differentiating , which gives . This is a standard result from calculus: when we differentiate with respect to , we apply the product rule, obtaining . Setting this derivative equal to zero gives us the first-order optimality conditions.
Solving for :
where:
- : the log-probability of response
- : the reward function
- : the KL penalty coefficient
- : the reference policy probability
- : constant terms related to the normalization constraint
This expression for the log-probability has an illuminating structure. The log-probability under our optimal policy equals the log-probability under the reference, adjusted by the scaled reward , plus some constants that don't depend on . Higher reward responses get higher log-probabilities, with controlling how much the reward matters relative to staying close to the reference.
Exponentiating both sides:
where:
- : the policy probability
- : the reference policy probability
- : the exponential function
- : the reward-scaling term
- : the reward function
- : the KL penalty coefficient
- : the normalization term constant across all
The term is just a normalizing constant that ensures the distribution sums to 1. This term doesn't depend on ; it only depends on through the constraint that all probabilities must sum to one. We can absorb it into a partition function :
where:
- : the optimal policy distribution
- : the partition function (normalizing constant) for prompt
- : the reference policy distribution
- : the reward function
- : the temperature parameter
This is the optimal policy for the RLHF objective. It has a clear interpretation: the optimal policy takes the reference distribution and reweights each response by the exponentiated reward, then normalizes. The exponential function ensures all probabilities are positive, naturally satisfying the constraint. Responses with higher reward get exponentially more probability mass, with controlling how aggressively we reweight.
Consider what happens at extreme values of beta to see how this reweighting works. When is very large, the reward term becomes small, and for all responses. In this limit, the optimal policy stays very close to the reference distribution, barely adjusting for reward at all. Conversely, when is very small, the exponential amplifies small reward differences into massive probability shifts, putting all probability mass on the highest-reward response. The choice of therefore controls the exploration-exploitation trade-off: larger values favor staying close to the reference (exploration), while smaller values favor chasing high reward (exploitation).




The optimal policy has the form of a Boltzmann (or Gibbs) distribution from statistical mechanics, where plays the role of negative energy and plays the role of temperature. Lower temperature (smaller ) concentrates probability on the highest-reward responses.
The Partition Function Problem
Before moving on, we should acknowledge a serious computational obstacle lurking in the optimal policy formula. The partition function is defined as:
where the sum ranges over all possible responses . For a language model, "all possible responses" means all sequences of tokens up to some maximum length. The number of such sequences is astronomical: a modest vocabulary of 32,000 tokens with a maximum length of 200 tokens gives possible responses, a number larger than the estimated number of atoms in the observable universe. Computing directly is completely intractable.
This intractability is precisely why naive approaches to optimizing the RLHF objective are so expensive. Algorithms like PPO work around it by using sample-based estimates and importance weighting, but those approaches have their own difficulties, including high variance and sensitivity to hyperparameters. The DPO derivation sidesteps this problem entirely, and the key step is coming up next in the reparameterization.
Reparameterizing the Reward
The next step uses the relationship between the optimal policy and the reward function in reverse. Given an optimal policy, we can solve for the reward function it implies.
This reverse direction might seem like a mathematical curiosity, but it turns out to be the key insight that makes DPO possible. If we can express the reward purely in terms of policies (without needing to train a separate reward model), then we can substitute this expression into the preference model and optimize directly. The reward becomes implicit in the policy rather than explicit in a separate neural network.
Think of it this way: the optimal policy and the reward function are two sides of the same coin. Knowing one completely determines the other, up to an additive constant per prompt. The optimal policy is just a "softmax" over rewards (scaled by ) relative to the reference. Inverting that relationship is like going from probabilities back to log-probabilities: algebraically straightforward once you know the connection exists.
Starting from:
where:
- : the optimal policy distribution
- : the partition function
- : the reference policy distribution
- : the reward function
- : the temperature parameter
We can rearrange to isolate the reward. First, take the log of both sides:
where:
- : log-probability under the optimal policy
- : log-probability under the reference policy
- : the reward function
- : the KL penalty coefficient
- : log of the partition function (depends only on )
Taking the logarithm turns our multiplicative relationship into an additive one. This makes algebraic manipulation simpler. The logarithm of a product becomes a sum of logarithms, and the logarithm of the exponential simply returns its argument.
Now, we rearrange terms to solve for :
where:
- : the implicit reward
- : the KL penalty coefficient
- : the optimal policy
- : the reference policy
- : the partition function
- : the scaled log-ratio term
This is the reward reparameterization. It tells us that for any optimal policy , we can express the implicit reward purely in terms of log probability ratios, plus a prompt-dependent constant .
The log ratio has a natural interpretation: it measures how much the optimal policy has increased or decreased the probability of response compared to the reference. Responses that the optimal policy strongly prefers will have large positive log ratios, while responses it disfavors will have large negative log ratios. This log ratio, scaled by , gives us the implicit reward up to an additive constant.
The partition function doesn't depend on ; it only depends on the prompt . This will turn out to be important because when we compute preference probabilities, this term cancels out.
Why This Reparameterization Is Significant
The reward reparameterization eliminates the need for an explicit reward model. In standard RLHF, you train a separate neural network on human preference data, and then you use that trained reward model to provide training signal to the language model policy. That two-step pipeline doubles the computational cost, introduces an additional source of approximation error, and creates the reward hacking problem when the policy finds outputs that score well on but do not match human preferences.
The reparameterization shows that you never needed a separate reward model. The reward is already encoded in the ratio between two language models: the policy being trained and the reference model. Every time you update the policy's weights, the implicit reward function changes accordingly. The reward model and the policy are the same object, just viewed from different angles. This is the deep reason why DPO is simpler than RLHF: it reduces the engineering work and the number of conceptual entities you have to reason about.
The DPO Loss Function
Now we can derive the DPO loss by substituting our reward reparameterization into the Bradley-Terry preference model. Recall from the chapter on the Bradley-Terry Model that the probability of preferring response over response is:
where:
- : probability that response is preferred over
- : the sigmoid function mapping values to
- : reward for the winning response
- : reward for the losing response
The Bradley-Terry model is elegant in its simplicity: the probability of preferring one response over another depends only on the difference in their rewards, passed through a sigmoid function. This means responses with much higher reward are strongly preferred, while responses with similar rewards have preference probabilities close to 0.5.
Substituting our reparameterized reward:
where:
- : the sigmoid function
- : the temperature parameter
- : the optimal policy
- : the reference policy
- : the partition function terms which appear with opposite signs
Notice that the terms cancel! This cancellation is not a coincidence; it's a direct consequence of using reward differences in the Bradley-Terry model. The partition function contributes equally to both rewards, so when we subtract them, it disappears. This cancellation is needed for the practical success of DPO because computing would require summing over all possible responses, which is computationally intractable for language models with exponentially large output spaces.
The key insight is that this cancellation is guaranteed by the structure of the Bradley-Terry model. Because preferences are defined as comparisons between pairs of responses (rather than absolute scores), any term that depends only on the prompt and not on the response will cancel in the subtraction. The partition function is exactly such a term, depending on through the sum over all possible responses but not on any particular response or . This is why the DPO derivation only works cleanly with pairwise comparison preference models and would need modification for other preference structures.
This leaves:
where:
- : the sigmoid function
- : temperature parameter
- : the optimal policy
- : the reference policy
We can simplify this using logarithm properties:
where:
- : the sigmoid function
- : the KL penalty coefficient
- : the log-likelihood ratio of the optimal policy vs reference
This expression is remarkably clean. The probability of the correct preference depends entirely on the difference between two log ratios: how much the optimal policy prefers the winning response (relative to the reference) versus how much it prefers the losing response. When this difference is large and positive, the sigmoid outputs a value close to 1, showing strong confidence in the correct preference.
This is the probability that the optimal policy assigns to the correct preference. To train a policy to match this optimal policy, we maximize the log-likelihood of the observed preferences:
where:
- : the DPO loss function
- : the policy network being trained (parameterized by )
- : the reference policy
- : the dataset of preference pairs
- : expectation over the dataset
- : the KL penalty coefficient
- : the sigmoid function
The negative sign appears because we're minimizing a loss rather than maximizing likelihood.
To make the notation more compact, define the log-ratio for a response :
where:
- : the implicit reward assigned by the current model
- : the KL penalty coefficient scaling the reward
- : the policy probability
- : the reference policy probability
- : the log-ratio of the policy probability to the reference probability
This is sometimes called the "implicit reward" because it's what the reward would be if were optimal. The terminology is fitting: we never explicitly compute or learn a reward function, yet the policy implicitly defines one through its log-probability ratios with the reference model.
The DPO loss becomes:
where:
- : the DPO loss function
- : the implicit reward function
- : the margin between implicit rewards for winning and losing responses
- : the sigmoid function
- : expectation over the dataset
This is remarkably simple. The loss encourages the model to increase its log probability (relative to the reference) on preferred responses and decrease it on dispreferred responses.
DPO as Classification
The DPO loss has a revealing interpretation: it's essentially binary classification with a specific parameterization. Consider a standard binary cross-entropy loss for classifying which of two responses is preferred:
where:
- : Binary Cross-Entropy loss
- : a scoring function showing preference strength
- : the sigmoid function
- : expectation over the dataset
In this context, is some function that should output a positive value when is truly preferred. In DPO, this function is:
where:
- : the logit fed into the sigmoid classifier
- : the KL penalty coefficient
- : the policy network
- : the reference policy
The model is learning to classify preference pairs by adjusting its output probabilities. When the loss is low, the model assigns higher relative probability to the preferred response.
This classification view explains why DPO is so stable compared to RLHF. Instead of learning a separate reward model and then using policy gradients with their high variance, DPO directly optimizes the likelihood of preference labels. The gradients flow directly through the model's log probabilities, just like in standard language model training.
The gradient of the DPO loss with respect to has a particularly intuitive form. For a single example:
where:
- : gradient with respect to model parameters
- : the DPO loss
- : the sigmoid function
- : the implicit reward function
- : the weighting term derived from the sigmoid derivative (probability of the wrong preference)
- : the geometric gradient direction
- : the KL penalty coefficient
This gradient expression reveals the geometry of DPO optimization. The term points in the direction that increases the probability of the winning response while decreasing the probability of the losing response. This is the natural direction to push given our objective.
The term is the probability that the model currently assigns to the wrong preference. This acts as an implicit weighting:
- When the model is confident and correct (low wrong-preference probability), the gradient is small
- When the model is confident but wrong, the gradient is large
- When the model is uncertain, the gradient is moderate
This automatic weighting helps the model focus on examples it's getting wrong, similar to how hard negative mining works in contrastive learning.

Connecting the Gradient to Language Model Training
It is instructive to compare the DPO gradient to what happens during ordinary language model pre-training. In standard cross-entropy training on tokens, the gradient for a single token at position increases the log-probability of the correct next token and the magnitude is governed by the current probability of that token: if the model already assigns high probability, the gradient is small. DPO follows the same pattern but at the sequence level: it increases log-probability for the entire chosen sequence and decreases it for the rejected sequence , weighted by how wrong the model currently is.
This parallel makes DPO feel like a natural extension of language modeling rather than a fundamentally different training paradigm. You can implement DPO in virtually any training framework that supports computing sequence log-probabilities, because the loss is just a scalar function of those log-probabilities. The practical implementation details, including how to handle variable-length sequences, padding masks, and batched reference model inference, are covered in the next chapter on DPO Implementation.
Worked Example
Let's trace through the derivation with concrete numbers to build intuition. Suppose we have a prompt = "Write a haiku about autumn" and two responses:
- (preferred): "Crimson leaves descend / Dancing on October wind / Earth prepares to sleep"
- (dispreferred): "Leaves fall down in fall / The weather gets cold outside / I like pumpkin spice"
The preferred response demonstrates poetic craft through concrete imagery ("crimson," "dancing"), sensory detail, and thematic resonance with the natural world. The dispreferred response is grammatically correct but prosaic, repeating "fall" clumsily and ending on an unrelated tangent about pumpkin spice. A human evaluator would clearly prefer the first response, and a well-trained reward model should assign it a higher score.
Assume our reference model assigns:
- (log probability of generating the preferred response)
- (log probability of generating the dispreferred response)
The reference model assigns higher probability to the dispreferred response because it's simpler and uses more common words. The word "fall" is far more common in training data than "crimson" or "descend," and the dispreferred response is shorter and uses everyday vocabulary throughout. This is exactly the distribution shift DPO is designed to fix: the reference model, trained to predict natural text, has no built-in preference for poetic quality over simplicity.
Now suppose our current policy assigns:
With , the implicit rewards are:
The difference is:
The loss contribution from this example:
The model is doing okay on this example (probability of correct preference is about 60%), but there's room to improve. The gradient will push the model to further increase relative to the reference and decrease relative to the reference.
Let's also trace through what a "good" and a "bad" model state look like for this same example. Suppose after more training the policy reaches:
- (probability has increased from reference)
- (probability has decreased from reference)
Now the implicit rewards are:
The reward margin is now , and the preference probability is . The loss drops to . The model has learned to "see" the quality difference between these two responses, even though it never directly observed a reward signal during this training phase.



Code Implementation
Let's implement the DPO loss function and verify our derivation with code.
import torch
import torch.nn.functional as F
def dpo_loss(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
ref_chosen_logps: torch.Tensor,
ref_rejected_logps: torch.Tensor,
beta: float = 0.1,
) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
"""
Compute the DPO loss for a batch of preference pairs.
Args:
policy_chosen_logps: Log probs of chosen responses under policy [batch_size]
policy_rejected_logps: Log probs of rejected responses under policy [batch_size]
ref_chosen_logps: Log probs of chosen responses under reference [batch_size]
ref_rejected_logps: Log probs of rejected responses under reference [batch_size]
beta: Temperature parameter controlling deviation from reference
Returns:
loss: Scalar DPO loss
chosen_rewards: Implicit rewards for chosen responses [batch_size]
rejected_rewards: Implicit rewards for rejected responses [batch_size]
"""
# Compute implicit rewards (log-ratios scaled by beta)
chosen_rewards = beta * (policy_chosen_logps - ref_chosen_logps)
rejected_rewards = beta * (policy_rejected_logps - ref_rejected_logps)
# DPO loss: negative log-sigmoid of reward difference
logits = chosen_rewards - rejected_rewards
loss = -F.logsigmoid(logits).mean()
return loss, chosen_rewards, rejected_rewardsNow let's verify with our worked example:
# Values from our worked example
policy_chosen_logps = torch.tensor([-42.1])
policy_rejected_logps = torch.tensor([-39.5])
ref_chosen_logps = torch.tensor([-45.2])
ref_rejected_logps = torch.tensor([-38.7])
loss, chosen_rewards, rejected_rewards = dpo_loss(
policy_chosen_logps,
policy_rejected_logps,
ref_chosen_logps,
ref_rejected_logps,
beta=0.1,
)Chosen implicit reward: 0.310 Rejected implicit reward: -0.080 Reward difference: 0.390 Preference probability: 0.596 DPO loss: 0.517
The implicit rewards align with our derivation: the chosen response has a positive reward (0.310), while the rejected response has a negative reward (-0.080). The positive difference (0.390) results in a preference probability of 0.596. This indicates the model correctly prefers the chosen response, though the probability is only moderately high.
Let's also visualize how the loss and preference probability change as the model learns:


The loss approaches zero as the reward difference increases, meaning the model strongly prefers chosen responses. Conversely, when the model prefers rejected responses (negative reward difference), the loss grows large.
Let's also implement a function to compute the gradients and see how the implicit weighting works:
def dpo_loss_with_grad_weights(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
ref_chosen_logps: torch.Tensor,
ref_rejected_logps: torch.Tensor,
beta: float = 0.1,
) -> tuple[torch.Tensor, torch.Tensor]:
"""
Compute DPO loss and the implicit gradient weights.
The gradient weight is sigma(-reward_diff), which is the probability
the model assigns to the WRONG preference.
"""
chosen_rewards = beta * (policy_chosen_logps - ref_chosen_logps)
rejected_rewards = beta * (policy_rejected_logps - ref_rejected_logps)
logits = chosen_rewards - rejected_rewards
loss = -F.logsigmoid(logits).mean()
# Gradient weight: probability of wrong preference
grad_weight = torch.sigmoid(-logits)
return loss, grad_weight# Test with different scenarios
scenarios = [
("Model strongly correct", -20.0, -83.5, -45.2, -38.7),
("Model weakly correct", -42.1, -50.6, -45.2, -38.7),
("Model uncertain", -45.0, -38.5, -45.2, -38.7),
("Model weakly wrong", -47.0, -25.5, -45.2, -38.7),
("Model strongly wrong", -75.5, -19.0, -45.2, -38.7),
]
results = []
for name, pc, pr, rc, rr in scenarios:
loss, grad_weight = dpo_loss_with_grad_weights(
torch.tensor([pc]),
torch.tensor([pr]),
torch.tensor([rc]),
torch.tensor([rr]),
beta=0.1,
)
results.append((name, loss.item(), grad_weight.item()))Scenario Loss Grad Weight -------------------------------------------------- Model strongly correct 0.001 0.001 Model weakly correct 0.201 0.182 Model uncertain 0.693 0.500 Model weakly wrong 1.701 0.818 Model strongly wrong 5.007 0.993
Notice how the gradient weight automatically adjusts based on how wrong the model is. In the "Model strongly correct" case, the gradient weight is near zero (0.000), meaning there's little to learn. Conversely, when the model is "Strongly wrong", the gradient weight approaches one (0.993), pushing hard to correct the mistake.

Key Parameters
The key parameter for DPO is:
- beta: The temperature parameter (often denoted as ) that scales the log-ratio of the policy and reference probabilities. It controls the strength of the KL divergence penalty, with larger values keeping the policy closer to the reference model.
Choosing in practice requires balancing two failure modes. If is too small, the DPO loss will push the policy aggressively in the direction of human preferences without adequate constraint from the reference model. This can cause the policy to drift far from the reference distribution and start generating fluent-but-wrong outputs, such as responses that contain the right topics but with incorrect factual content, or responses that are overly verbose because longer responses were slightly preferred in the training data. If is too large, the policy can barely move away from the reference, which means it learns slowly and may not improve much even after thousands of training steps. Empirically, values in the range 0.01 to 0.5 are common, with 0.1 being a popular default starting point.


Assumptions and Validity
The DPO derivation makes several assumptions worth examining carefully, because each one represents a potential point of failure in real-world applications.
The first assumption is that the Bradley-Terry model correctly captures human preferences. Specifically, DPO assumes that the probability of preferring response over can be expressed as a sigmoid of a scalar reward difference. Real human preferences are far messier. They are noisy (the same annotator might rate the same pair differently on different days), inconsistent across annotators (different people have different values), and often non-transitive (a human might prefer over , prefer over , but also prefer over ). The Bradley-Terry model cannot represent this complexity. DPO inherits this limitation from the underlying preference model: if your preference data is very noisy or inconsistent, the loss will receive contradictory training signals and may not converge to a meaningful policy.
The second assumption is that the optimal policy has the Boltzmann form derived earlier. This is guaranteed only for the specific objective we started with, namely expected reward minus a KL divergence from the reference. Different alignment objectives, such as objectives using different divergence measures or adding additional safety constraints, would yield different optimal policies and potentially different direct alignment algorithms. The DPO loss is not a general-purpose alignment loss; it is the specific loss that corresponds to the KL-regularized RLHF objective with the Bradley-Terry preference model. When researchers have proposed modifications, such as using an -divergence instead of KL, they obtain modified versions of DPO with different properties. The next chapter discusses several of these variants.
The third assumption is that model capacity is sufficient. DPO optimizes toward the optimal policy but doesn't guarantee reaching it. The policy is constrained by the model's architecture and capacity. A small model may not be able to represent the true optimal policy, in which case DPO finds the best approximation within the model class. This is not unique to DPO, since all finite-parameter models face this limitation, but the theoretical guarantees assume that the policy class is expressive enough to represent the optimal Boltzmann distribution.
A fourth, more subtle assumption concerns the relationship between the data distribution and the behavior of the optimal policy. The derivation assumes we optimize over the true expectation under the training distribution . In practice, we minimize an empirical loss over a finite dataset. If the preference dataset is small or unrepresentative of the prompts the model will encounter at deployment, DPO may overfit to the dataset distribution. This manifests as mode collapse on training prompts or failure to generalize to held-out prompt types. This issue is partly mitigated by the KL constraint (via the reference policy), but a badly distributed training set will still cause problems.
Finally, DPO requires that we can compute exactly for the same responses we're evaluating under . This is straightforward when both are the same model architecture, but becomes complex if the reference model is unavailable or uses different tokenization. The next chapter on DPO Implementation will address these practical considerations in detail.
Summary
This chapter derived the DPO loss function from first principles, showing how it emerges naturally from the RLHF objective.
The key steps in the derivation were:
- RLHF objective: Maximize expected reward with a KL penalty to stay close to the reference policy
- Optimal policy: Using Lagrange multipliers, we found that the optimal policy is a Boltzmann distribution that reweights the reference by exponentiated rewards
- Reward reparameterization: Inverting this relationship, we expressed the implicit reward as a log-ratio between the optimal and reference policies, plus a partition function
- Bradley-Terry substitution: Plugging the reparameterized rewards into the preference model, the partition functions cancel, leaving a loss in terms of log-ratios only
- Classification view: The resulting DPO loss is binary cross-entropy, learning to classify which response is preferred based on relative log-probabilities
The gradient of the DPO loss has an automatic weighting scheme: examples where the model is confidently wrong receive larger gradients, while examples where it's already correct receive smaller gradients. This makes DPO training stable and efficient.
The derivation also reveals DPO's key assumptions: that human preferences follow the Bradley-Terry model, that the KL-regularized RLHF objective is the right starting point, and that the policy class is expressive enough to represent the optimal Boltzmann distribution. When those assumptions hold, DPO provides a principled, efficient alternative to RLHF. When they break down, the variants we explore in the following chapter offer modified objectives that relax different aspects of those assumptions.
In the next chapter, we'll implement DPO end-to-end, covering practical details like computing sequence log-probabilities, handling padding, and integrating with Hugging Face's training infrastructure.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the DPO derivation.
DPO Derivation
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!