DPO Derivation: From RLHF Objective to Direct Optimization

Michael BrenndoerferJanuary 1, 202645 min read

Part of Language AI Handbook

Derive the DPO loss function from first principles. Explains how the optimal RLHF policy leads to reward reparameterization and direct preference optimization.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

DPO Derivation

In the previous chapter, we introduced Direct Preference Optimization as a simpler alternative to RLHF that eliminates the need for an explicit reward model. We saw that DPO directly optimizes a language model using preference data, but we didn't examine why this works or how the DPO loss function is derived. This chapter fills that gap by tracing every algebraic step, from the original reinforcement learning objective all the way to the final training loss.

The DPO derivation is a key result in alignment research. It shows that the optimal policy for the RLHF objective has a closed-form solution, and this solution can be rearranged to express the implicit reward in terms of the policy itself. When we substitute this reparameterized reward into the Bradley-Terry preference model, we obtain the DPO loss function directly, with no reinforcement learning required. This chain of reasoning is surprisingly elegant, and understanding it will fundamentally change how you think about what a language model is doing when it trains on human preference data.

Understanding this derivation reveals why DPO works, its underlying assumptions, and how to interpret the model's learning process. By the end of this chapter, you'll see that DPO is essentially solving a classification problem where the model learns to assign higher probability to preferred responses. Each gradient step simultaneously rewards the model for raising the probability of good outputs and penalizes it for raising the probability of bad ones, all without ever explicitly parameterizing a reward function.

Think of the whole derivation as a three-act story. Act one asks: if we had the ideal reward function, what would the optimal policy look like? Act two reverses that question and asks: if we have the optimal policy, what reward function does it imply? Act three substitutes that reversed reward into the preference model and discovers that the messy partition function term vanishes, leaving a clean, tractable loss. Each act is mathematically precise, and each act builds directly on the previous one. When you finish reading, the chain from RLHF to DPO will feel inevitable rather than mysterious.

One reason this derivation matters beyond academic interest is that it reveals the hidden assumptions baked into DPO. The loss function is not a heuristic: it follows from specific, clearly stated premises about the form of the RLHF objective, the Boltzmann structure of optimal policies, and the Bradley-Terry preference model. Knowing those premises tells you exactly when DPO is appropriate, what might break it, and how to modify it for different settings. We will examine those assumptions carefully in the section on limitations.

Historical Context

Direct Preference Optimization was introduced by Rafailov et al. in 2023 in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model." The central insight, that the RLHF objective admits a closed-form optimal policy whose reward can be reparameterized in terms of policy ratios, appeared almost simultaneously with work on "SLiC-HF" (Zhao et al., 2023) and related ideas in reward-free offline RL. What made DPO distinctive was its clean derivation that tied together three well-known components (the KL-regularized RL objective, the Boltzmann optimal policy, and the Bradley-Terry model) into a single differentiable loss. Within months, DPO became one of the most widely used alignment techniques in practice, appearing in the training recipes of Llama 2, Mistral Instruct, and dozens of other models. The paper has become one of the most cited alignment papers of the 2020s, and its derivation has inspired a family of variants including IPO, KTO, and ORPO, which we cover in the next chapter.

The RLHF Objective

Let's begin with the objective that RLHF aims to optimize. As we discussed in the chapters on PPO for Language Models and KL Divergence Penalty, the goal is to find a policy π\pi that maximizes expected reward while staying close to a reference policy πref\pi_{\text{ref}} (typically the supervised fine-tuned model).

Before diving into the mathematics, it helps to understand the intuition behind this objective. We want our language model to generate responses that humans prefer, which is captured by the reward function. However, if we optimize the reward too aggressively, the model might find unexpected shortcuts or produce degenerate outputs that technically achieve high reward but don't represent helpful behavior. The reference policy is an anchor, representing the model's pre-trained knowledge and natural language capabilities. By penalizing deviations from this anchor, we encourage the model to improve its outputs while maintaining coherent, fluent generation.

Think of the reference policy as the starting point on a map, and the reward function as a compass pointing toward a better destination. The KL penalty acts like a tether that keeps you from wandering too far afield in a single step. You want to move in the direction of higher reward, but you don't want to teleport so far that you lose the path entirely. In language modeling, that "losing the path" corresponds to mode collapse or reward hacking: outputs that score well on a proxy reward metric but are nonsensical or harmful. The KL constraint is what prevents those degenerate solutions.

The constrained optimization problem is:

max⁡πEx∼D,y∼π(y∣x)[r(x,y)]−β⋅DKL(π(y∣x)∥πref(y∣x))\max_{\pi} \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi(y|x)} \left[ r(x, y) \right] - \beta \cdot D_{\text{KL}}\left( \pi(y|x) \| \pi_{\text{ref}}(y|x) \right)

where:

  • π\pi: the policy being optimized
  • E\mathbb{E}: the expectation operator
  • xx: a prompt sampled from the data distribution D\mathcal{D}
  • yy: a response sampled from the policy
  • r(x,y)r(x, y): the learned reward function
  • β\beta: a hyperparameter controlling the strength of the KL constraint
  • DKLD_{\text{KL}}: the Kullback-Leibler divergence
  • πref\pi_{\text{ref}}: the reference policy

To work with this objective more directly, we need to expand the KL divergence term. Recall that the KL divergence measures how different one probability distribution is from another, expressed as the expected log ratio of the two distributions. Writing out the KL divergence explicitly, this becomes:

max⁡πEx∼D,y∼π(y∣x)[r(x,y)−βlog⁡π(y∣x)πref(y∣x)]\max_{\pi} \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi(y|x)} \left[ r(x, y) - \beta \log \frac{\pi(y|x)}{\pi_{\text{ref}}(y|x)} \right]

where:

  • E\mathbb{E}: the expectation operator
  • D\mathcal{D}: the data distribution
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • log⁡\log: the natural logarithm
  • π(y∣x)\pi(y|x): probability of response yy given prompt xx under the policy
  • πref(y∣x)\pi_{\text{ref}}(y|x): probability of response yy given prompt xx under the reference model

This reformulation reveals the objective's structure more clearly. The log ratio log⁡π(y∣x)πref(y∣x)\log \frac{\pi(y|x)}{\pi_{\text{ref}}(y|x)} measures how much the policy has shifted away from the reference for a particular response. When this ratio is positive (the policy assigns more probability than the reference), the KL term subtracts from the objective, penalizing the deviation. When the ratio is negative (the policy assigns less probability), the term adds to the objective, which might seem like a reward, but since we're sampling from π\pi, we're unlikely to generate responses where π\pi assigns low probability.

The key insight is that this scalar objective, combining a reward signal and a distributional constraint, can be solved analytically. That analytical solution is the foundation of everything that follows. The reward-KL tradeoff creates a specific mathematical structure that admits an exact solution in closed form, rather than acting as an arbitrary heuristic regularizer.

This objective captures a basic tension in alignment: we want the model to produce high-reward outputs (as judged by human preferences), but we don't want it to deviate too far from its pre-trained behavior. Without the KL penalty, the model might find degenerate solutions that exploit flaws in the reward model, a phenomenon we explored in the chapter on Reward Hacking.

Why the KL Penalty Has This Form

Before proceeding to the derivation, it is worth spending a moment on why the KL divergence is the natural choice of regularizer here. Many other distance measures between distributions exist, such as total variation distance, Renyi divergences, or the Wasserstein distance. Each would produce a different optimization problem and a different optimal policy. The KL divergence DKL(π∥πref)D_{\text{KL}}(\pi \| \pi_{\text{ref}}) is particularly convenient because it has exactly the right algebraic structure to yield a clean closed-form solution. Its definition as Eπ[log⁡(π/πref)]\mathbb{E}_\pi[\log(\pi / \pi_{\text{ref}})] means that when we write the full objective as an expectation under π\pi, everything becomes a single expectation of a per-token scalar quantity, which is tractable.

The KL divergence also has information-theoretic grounding. It measures the number of extra "bits" required to encode samples from π\pi using a code optimized for πref\pi_{\text{ref}}. Keeping this quantity small means the policy's outputs remain within the statistical neighborhood of the reference, which in practice means they retain the grammatical fluency, factual grounding, and stylistic coherence of the original supervised fine-tuned model.

Deriving the Optimal Policy

The key insight behind DPO is that this optimization problem has a closed-form solution. Unlike many optimization problems in machine learning that require iterative gradient descent, this particular formulation admits an analytical answer. For a fixed prompt xx, we're optimizing over the distribution π(y∣x)\pi(y|x) for all possible responses yy. This is a constrained optimization problem over probability distributions: we must ensure our solution is a valid probability distribution that sums to one and assigns non-negative probability to every possible response.

The special structure of the RLHF objective allows for this closed-form solution. The expectation under π\pi combined with the KL divergence creates what's known as a variational problem over distributions. Such problems often have elegant solutions when the constraints are simple probability simplex constraints.

Let's work through the derivation step by step. For a single prompt xx, the objective is:

max⁡π(⋅∣x)∑yπ(y∣x)[r(x,y)−βlog⁡π(y∣x)πref(y∣x)]\max_{\pi(\cdot|x)} \sum_{y} \pi(y|x) \left[ r(x, y) - \beta \log \frac{\pi(y|x)}{\pi_{\text{ref}}(y|x)} \right]

where:

  • ∑y\sum_{y}: sum over all possible responses in the vocabulary
  • π(⋅∣x)\pi(\cdot|x): the entire probability distribution over responses for prompt xx
  • π(y∣x)\pi(y|x): probability of response yy given prompt xx
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • log⁡\log: the natural logarithm
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability

We need to maximize this subject to the constraint that π(⋅∣x)\pi(\cdot|x) is a valid probability distribution: ∑yπ(y∣x)=1\sum_y \pi(y|x) = 1 and π(y∣x)≥0\pi(y|x) \geq 0 for all yy.

To solve this constrained optimization problem, we use the method of Lagrange multipliers, a classical technique from calculus. The idea is to incorporate the constraint directly into the objective by introducing a new variable (the Lagrange multiplier) that penalizes violations of the constraint. At the optimal solution, the gradient of the objective with respect to the decision variables must be proportional to the gradient of the constraint.

The Lagrangian is:

L=∑yπ(y∣x)[r(x,y)−βlog⁡π(y∣x)+βlog⁡πref(y∣x)]+λ(1−∑yπ(y∣x))\mathcal{L} = \sum_{y} \pi(y|x) \left[ r(x, y) - \beta \log \pi(y|x) + \beta \log \pi_{\text{ref}}(y|x) \right] + \lambda \left( 1 - \sum_y \pi(y|x) \right)

where:

  • L\mathcal{L}: the Lagrangian function
  • ∑y\sum_{y}: sum over all responses
  • π(y∣x)\pi(y|x): probability of response yy
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • log⁡\log: the natural logarithm
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability
  • λ\lambda: the Lagrange multiplier for the constraint that probabilities sum to 1

Notice that we've expanded the log ratio term from the original objective, separating it into −βlog⁡π(y∣x)-\beta \log \pi(y|x) and +βlog⁡πref(y∣x)+\beta \log \pi_{\text{ref}}(y|x). This separation makes taking derivatives more straightforward. The term λ(1−∑yπ(y∣x))\lambda(1 - \sum_y \pi(y|x)) enforces our normalization constraint: if the probabilities don't sum to one, this term will be nonzero, and the Lagrange multiplier λ\lambda will adjust to push us toward a valid distribution.

Taking the derivative with respect to π(y∣x)\pi(y|x) for a specific yy:

∂L∂π(y∣x)=r(x,y)−βlog⁡π(y∣x)−β+βlog⁡πref(y∣x)−λ=0\frac{\partial \mathcal{L}}{\partial \pi(y|x)} = r(x, y) - \beta \log \pi(y|x) - \beta + \beta \log \pi_{\text{ref}}(y|x) - \lambda = 0

where:

  • ∂L∂π(y∣x)\frac{\partial \mathcal{L}}{\partial \pi(y|x)}: the partial derivative of the Lagrangian with respect to the probability of response yy
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • log⁡\log: the natural logarithm
  • π(y∣x)\pi(y|x): the policy probability
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability
  • λ\lambda: the Lagrange multiplier

The −β-\beta term comes from differentiating πlog⁡π\pi \log \pi, which gives log⁡π+1\log \pi + 1. This is a standard result from calculus: when we differentiate πlog⁡π\pi \log \pi with respect to π\pi, we apply the product rule, obtaining log⁡π+π⋅(1/π)=log⁡π+1\log \pi + \pi \cdot (1/\pi) = \log \pi + 1. Setting this derivative equal to zero gives us the first-order optimality conditions.

Solving for log⁡π(y∣x)\log \pi(y|x):

βlog⁡π(y∣x)=r(x,y)+βlog⁡πref(y∣x)−β−λlog⁡π(y∣x)=1βr(x,y)+log⁡πref(y∣x)−1−λβ\begin{aligned} \beta \log \pi(y|x) &= r(x, y) + \beta \log \pi_{\text{ref}}(y|x) - \beta - \lambda \\ \log \pi(y|x) &= \frac{1}{\beta} r(x, y) + \log \pi_{\text{ref}}(y|x) - 1 - \frac{\lambda}{\beta} \end{aligned}

where:

  • log⁡π(y∣x)\log \pi(y|x): the log-probability of response yy
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability
  • λ/β\lambda/\beta: constant terms related to the normalization constraint

This expression for the log-probability has an illuminating structure. The log-probability under our optimal policy equals the log-probability under the reference, adjusted by the scaled reward r(x,y)/βr(x,y)/\beta, plus some constants that don't depend on yy. Higher reward responses get higher log-probabilities, with β\beta controlling how much the reward matters relative to staying close to the reference.

Exponentiating both sides:

π(y∣x)=πref(y∣x)⋅exp⁡(r(x,y)β)⋅exp⁡(−1−λβ)\pi(y|x) = \pi_{\text{ref}}(y|x) \cdot \exp\left( \frac{r(x, y)}{\beta} \right) \cdot \exp\left( -1 - \frac{\lambda}{\beta} \right)

where:

  • π(y∣x)\pi(y|x): the policy probability
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability
  • exp⁡(⋅)\exp(\cdot): the exponential function
  • exp⁡(r(x,y)/β)\exp(r(x, y)/\beta): the reward-scaling term
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • exp⁡(−1−λ/β)\exp(-1 - \lambda/\beta): the normalization term constant across all yy

The term exp⁡(−1−λ/β)\exp(-1 - \lambda/\beta) is just a normalizing constant that ensures the distribution sums to 1. This term doesn't depend on yy; it only depends on xx through the constraint that all probabilities must sum to one. We can absorb it into a partition function Z(x)Z(x):

π∗(y∣x)=1Z(x)πref(y∣x)exp⁡(r(x,y)β)\pi^*(y|x) = \frac{1}{Z(x)} \pi_{\text{ref}}(y|x) \exp\left( \frac{r(x, y)}{\beta} \right)

where:

  • π∗(y∣x)\pi^*(y|x): the optimal policy distribution
  • Z(x)=∑y′πref(y′∣x)exp⁡(r(x,y′)β)Z(x) = \sum_{y'} \pi_{\text{ref}}(y'|x) \exp\left( \frac{r(x, y')}{\beta} \right): the partition function (normalizing constant) for prompt xx
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy distribution
  • r(x,y)r(x, y): the reward function
  • β\beta: the temperature parameter

This is the optimal policy for the RLHF objective. It has a clear interpretation: the optimal policy takes the reference distribution and reweights each response by the exponentiated reward, then normalizes. The exponential function ensures all probabilities are positive, naturally satisfying the π(y∣x)≥0\pi(y|x) \ge 0 constraint. Responses with higher reward get exponentially more probability mass, with β\beta controlling how aggressively we reweight.

Consider what happens at extreme values of beta to see how this reweighting works. When β\beta is very large, the reward term r(x,y)/βr(x,y)/\beta becomes small, and exp⁡(r(x,y)/β)≈1\exp(r(x,y)/\beta) \approx 1 for all responses. In this limit, the optimal policy stays very close to the reference distribution, barely adjusting for reward at all. Conversely, when β\beta is very small, the exponential amplifies small reward differences into massive probability shifts, putting all probability mass on the highest-reward response. The choice of β\beta therefore controls the exploration-exploitation trade-off: larger values favor staying close to the reference (exploration), while smaller values favor chasing high reward (exploitation).

Out[3]:
Visualization
Bar chart for beta=0.2 showing reference and optimal policy distributions, with optimal distribution heavily concentrated on the highest-reward response D.
Reference vs optimal policy distribution for $\beta = 0.2$. Low temperature concentrates probability heavily on the highest-reward response (D).
Bar chart for beta=0.5 showing reference and optimal policy distributions with moderate concentration on the highest-reward response D.
Reference vs optimal policy distribution for $\beta = 0.5$. Moderate-low temperature still noticeably shifts mass toward high-reward responses.
Bar chart for beta=1.0 showing reference and optimal policy distributions with balanced reweighting toward higher-reward responses.
Reference vs optimal policy distribution for $\beta = 1.0$. Moderate temperature produces a balanced reweighting of the reference distribution.
Bar chart for beta=2.0 showing reference and optimal policy distributions nearly identical, preserving the reference distribution.
Reference vs optimal policy distribution for $\beta = 2.0$. High temperature preserves the shape of the reference distribution closely.
Boltzmann Distribution

The optimal policy has the form of a Boltzmann (or Gibbs) distribution from statistical mechanics, where r(x,y)r(x,y) plays the role of negative energy and β\beta plays the role of temperature. Lower temperature (smaller β\beta) concentrates probability on the highest-reward responses.

The Partition Function Problem

Before moving on, we should acknowledge a serious computational obstacle lurking in the optimal policy formula. The partition function Z(x)Z(x) is defined as:

Z(x)=∑y′πref(y′∣x)exp⁡(r(x,y′)β)Z(x) = \sum_{y'} \pi_{\text{ref}}(y'|x) \exp\left( \frac{r(x, y')}{\beta} \right)

where the sum ranges over all possible responses y′y'. For a language model, "all possible responses" means all sequences of tokens up to some maximum length. The number of such sequences is astronomical: a modest vocabulary of 32,000 tokens with a maximum length of 200 tokens gives 3200020032000^{200} possible responses, a number larger than the estimated number of atoms in the observable universe. Computing Z(x)Z(x) directly is completely intractable.

This intractability is precisely why naive approaches to optimizing the RLHF objective are so expensive. Algorithms like PPO work around it by using sample-based estimates and importance weighting, but those approaches have their own difficulties, including high variance and sensitivity to hyperparameters. The DPO derivation sidesteps this problem entirely, and the key step is coming up next in the reparameterization.

Reparameterizing the Reward

The next step uses the relationship between the optimal policy and the reward function in reverse. Given an optimal policy, we can solve for the reward function it implies.

This reverse direction might seem like a mathematical curiosity, but it turns out to be the key insight that makes DPO possible. If we can express the reward purely in terms of policies (without needing to train a separate reward model), then we can substitute this expression into the preference model and optimize directly. The reward becomes implicit in the policy rather than explicit in a separate neural network.

Think of it this way: the optimal policy and the reward function are two sides of the same coin. Knowing one completely determines the other, up to an additive constant per prompt. The optimal policy is just a "softmax" over rewards (scaled by 1/β1/\beta) relative to the reference. Inverting that relationship is like going from probabilities back to log-probabilities: algebraically straightforward once you know the connection exists.

Starting from:

π∗(y∣x)=1Z(x)πref(y∣x)exp⁡(r(x,y)β)\pi^*(y|x) = \frac{1}{Z(x)} \pi_{\text{ref}}(y|x) \exp\left( \frac{r(x, y)}{\beta} \right)

where:

  • π∗(y∣x)\pi^*(y|x): the optimal policy distribution
  • Z(x)Z(x): the partition function
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy distribution
  • r(x,y)r(x, y): the reward function
  • β\beta: the temperature parameter

We can rearrange to isolate the reward. First, take the log of both sides:

log⁡π∗(y∣x)=log⁡πref(y∣x)+r(x,y)β−log⁡Z(x)\log \pi^*(y|x) = \log \pi_{\text{ref}}(y|x) + \frac{r(x, y)}{\beta} - \log Z(x)

where:

  • log⁡π∗\log \pi^*: log-probability under the optimal policy
  • log⁡πref(y∣x)\log \pi_{\text{ref}}(y|x): log-probability under the reference policy
  • r(x,y)r(x, y): the reward function
  • β\beta: the KL penalty coefficient
  • log⁡Z(x)\log Z(x): log of the partition function (depends only on xx)

Taking the logarithm turns our multiplicative relationship into an additive one. This makes algebraic manipulation simpler. The logarithm of a product becomes a sum of logarithms, and the logarithm of the exponential simply returns its argument.

Now, we rearrange terms to solve for r(x,y)r(x, y):

r(x,y)β=log⁡π∗(y∣x)−log⁡πref(y∣x)+log⁡Z(x)(isolate reward term)r(x,y)β=log⁡π∗(y∣x)πref(y∣x)+log⁡Z(x)(combine logs)r(x,y)=βlog⁡π∗(y∣x)πref(y∣x)+βlog⁡Z(x)(multiply by β)\begin{aligned} \frac{r(x, y)}{\beta} &= \log \pi^*(y|x) - \log \pi_{\text{ref}}(y|x) + \log Z(x) && \text{(isolate reward term)} \\ \frac{r(x, y)}{\beta} &= \log \frac{\pi^*(y|x)}{\pi_{\text{ref}}(y|x)} + \log Z(x) && \text{(combine logs)} \\ r(x, y) &= \beta \log \frac{\pi^*(y|x)}{\pi_{\text{ref}}(y|x)} + \beta \log Z(x) && \text{(multiply by } \beta \text{)} \end{aligned}

where:

  • r(x,y)r(x, y): the implicit reward
  • β\beta: the KL penalty coefficient
  • π∗(y∣x)\pi^*(y|x): the optimal policy
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy
  • Z(x)Z(x): the partition function
  • βlog⁡(⋅)\beta \log (\cdot): the scaled log-ratio term

This is the reward reparameterization. It tells us that for any optimal policy π∗\pi^*, we can express the implicit reward purely in terms of log probability ratios, plus a prompt-dependent constant βlog⁡Z(x)\beta \log Z(x).

The log ratio log⁡π∗(y∣x)πref(y∣x)\log \frac{\pi^*(y|x)}{\pi_{\text{ref}}(y|x)} has a natural interpretation: it measures how much the optimal policy has increased or decreased the probability of response yy compared to the reference. Responses that the optimal policy strongly prefers will have large positive log ratios, while responses it disfavors will have large negative log ratios. This log ratio, scaled by β\beta, gives us the implicit reward up to an additive constant.

The partition function Z(x)Z(x) doesn't depend on yy; it only depends on the prompt xx. This will turn out to be important because when we compute preference probabilities, this term cancels out.

Why This Reparameterization Is Significant

The reward reparameterization eliminates the need for an explicit reward model. In standard RLHF, you train a separate neural network rϕ(x,y)r_\phi(x, y) on human preference data, and then you use that trained reward model to provide training signal to the language model policy. That two-step pipeline doubles the computational cost, introduces an additional source of approximation error, and creates the reward hacking problem when the policy finds outputs that score well on rϕr_\phi but do not match human preferences.

The reparameterization shows that you never needed a separate reward model. The reward is already encoded in the ratio between two language models: the policy being trained and the reference model. Every time you update the policy's weights, the implicit reward function changes accordingly. The reward model and the policy are the same object, just viewed from different angles. This is the deep reason why DPO is simpler than RLHF: it reduces the engineering work and the number of conceptual entities you have to reason about.

The DPO Loss Function

Now we can derive the DPO loss by substituting our reward reparameterization into the Bradley-Terry preference model. Recall from the chapter on the Bradley-Terry Model that the probability of preferring response ywy_w over response yly_l is:

p(yw≻yl∣x)=σ(r(x,yw)−r(x,yl))p(y_w \succ y_l | x) = \sigma\left( r(x, y_w) - r(x, y_l) \right)

where:

  • p(yw≻yl∣x)p(y_w \succ y_l | x): probability that response ywy_w is preferred over yly_l
  • σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}: the sigmoid function mapping values to (0,1)(0, 1)
  • r(x,yw)r(x, y_w): reward for the winning response
  • r(x,yl)r(x, y_l): reward for the losing response

The Bradley-Terry model is elegant in its simplicity: the probability of preferring one response over another depends only on the difference in their rewards, passed through a sigmoid function. This means responses with much higher reward are strongly preferred, while responses with similar rewards have preference probabilities close to 0.5.

Substituting our reparameterized reward:

p(yw≻yl∣x)=σ([βlog⁡π∗(yw∣x)πref(yw∣x)+βlog⁡Z(x)]−[βlog⁡π∗(yl∣x)πref(yl∣x)+βlog⁡Z(x)])p(y_w \succ y_l | x) = \sigma\left( \left[ \beta \log \frac{\pi^*(y_w|x)}{\pi_{\text{ref}}(y_w|x)} + \beta \log Z(x) \right] - \left[ \beta \log \frac{\pi^*(y_l|x)}{\pi_{\text{ref}}(y_l|x)} + \beta \log Z(x) \right] \right)

where:

  • σ\sigma: the sigmoid function
  • β\beta: the temperature parameter
  • π∗\pi^*: the optimal policy
  • πref\pi_{\text{ref}}: the reference policy
  • Z(x)Z(x): the partition function terms which appear with opposite signs

Notice that the βlog⁡Z(x)\beta \log Z(x) terms cancel! This cancellation is not a coincidence; it's a direct consequence of using reward differences in the Bradley-Terry model. The partition function contributes equally to both rewards, so when we subtract them, it disappears. This cancellation is needed for the practical success of DPO because computing Z(x)Z(x) would require summing over all possible responses, which is computationally intractable for language models with exponentially large output spaces.

The key insight is that this cancellation is guaranteed by the structure of the Bradley-Terry model. Because preferences are defined as comparisons between pairs of responses (rather than absolute scores), any term that depends only on the prompt xx and not on the response yy will cancel in the subtraction. The partition function is exactly such a term, depending on xx through the sum over all possible responses but not on any particular response ywy_w or yly_l. This is why the DPO derivation only works cleanly with pairwise comparison preference models and would need modification for other preference structures.

This leaves:

p(yw≻yl∣x)=σ(βlog⁡π∗(yw∣x)πref(yw∣x)−βlog⁡π∗(yl∣x)πref(yl∣x))p(y_w \succ y_l | x) = \sigma\left( \beta \log \frac{\pi^*(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \beta \log \frac{\pi^*(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right)

where:

We can simplify this using logarithm properties:

p(yw≻yl∣x)=σ(β[log⁡π∗(yw∣x)πref(yw∣x)−log⁡π∗(yl∣x)πref(yl∣x)])p(y_w \succ y_l | x) = \sigma\left( \beta \left[ \log \frac{\pi^*(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi^*(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right] \right)

where:

  • σ\sigma: the sigmoid function
  • β\beta: the KL penalty coefficient
  • log⁡π∗(y∣x)πref(y∣x)\log \frac{\pi^*(y|x)}{\pi_{\text{ref}}(y|x)}: the log-likelihood ratio of the optimal policy vs reference

This expression is remarkably clean. The probability of the correct preference depends entirely on the difference between two log ratios: how much the optimal policy prefers the winning response (relative to the reference) versus how much it prefers the losing response. When this difference is large and positive, the sigmoid outputs a value close to 1, showing strong confidence in the correct preference.

This is the probability that the optimal policy assigns to the correct preference. To train a policy πθ\pi_\theta to match this optimal policy, we maximize the log-likelihood of the observed preferences:

LDPO(πθ;πref)=−E(x,yw,yl)∼D[log⁡σ(β[log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)])]\mathcal{L}_{\text{DPO}}(\pi_\theta; \pi_{\text{ref}}) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma\left( \beta \left[ \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right] \right) \right]

where:

  • LDPO\mathcal{L}_{\text{DPO}}: the DPO loss function
  • πθ\pi_\theta: the policy network being trained (parameterized by θ\theta)
  • πref\pi_{\text{ref}}: the reference policy
  • D\mathcal{D}: the dataset of preference pairs (x,yw,yl)(x, y_w, y_l)
  • E\mathbb{E}: expectation over the dataset
  • β\beta: the KL penalty coefficient
  • σ\sigma: the sigmoid function

The negative sign appears because we're minimizing a loss rather than maximizing likelihood.

To make the notation more compact, define the log-ratio for a response yy:

r^θ(x,y)=βlog⁡πθ(y∣x)πref(y∣x)\hat{r}_\theta(x, y) = \beta \log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)}

where:

  • r^θ(x,y)\hat{r}_\theta(x, y): the implicit reward assigned by the current model πθ\pi_\theta
  • β\beta: the KL penalty coefficient scaling the reward
  • πθ(y∣x)\pi_\theta(y|x): the policy probability
  • πref(y∣x)\pi_{\text{ref}}(y|x): the reference policy probability
  • log⁡πθ(y∣x)πref(y∣x)\log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)}: the log-ratio of the policy probability to the reference probability

This is sometimes called the "implicit reward" because it's what the reward would be if πθ\pi_\theta were optimal. The terminology is fitting: we never explicitly compute or learn a reward function, yet the policy implicitly defines one through its log-probability ratios with the reference model.

The DPO loss becomes:

LDPO=−E(x,yw,yl)∼D[log⁡σ(r^θ(x,yw)−r^θ(x,yl))]\mathcal{L}_{\text{DPO}} = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma\left( \hat{r}_\theta(x, y_w) - \hat{r}_\theta(x, y_l) \right) \right]

where:

  • LDPO\mathcal{L}_{\text{DPO}}: the DPO loss function
  • r^θ(x,y)\hat{r}_\theta(x, y): the implicit reward function
  • r^θ(x,yw)−r^θ(x,yl)\hat{r}_\theta(x, y_w) - \hat{r}_\theta(x, y_l): the margin between implicit rewards for winning and losing responses
  • σ\sigma: the sigmoid function
  • E\mathbb{E}: expectation over the dataset

This is remarkably simple. The loss encourages the model to increase its log probability (relative to the reference) on preferred responses and decrease it on dispreferred responses.

DPO as Classification

The DPO loss has a revealing interpretation: it's essentially binary classification with a specific parameterization. Consider a standard binary cross-entropy loss for classifying which of two responses is preferred:

LBCE=−E[log⁡σ(f(x,yw,yl))]\mathcal{L}_{\text{BCE}} = -\mathbb{E} \left[ \log \sigma(f(x, y_w, y_l)) \right]

where:

  • LBCE\mathcal{L}_{\text{BCE}}: Binary Cross-Entropy loss
  • f(x,yw,yl)f(x, y_w, y_l): a scoring function showing preference strength
  • σ\sigma: the sigmoid function
  • E\mathbb{E}: expectation over the dataset

In this context, ff is some function that should output a positive value when ywy_w is truly preferred. In DPO, this function is:

f(x,yw,yl)=β[log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)]f(x, y_w, y_l) = \beta \left[ \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right]

where:

  • f(x,yw,yl)f(x, y_w, y_l): the logit fed into the sigmoid classifier
  • β\beta: the KL penalty coefficient
  • πθ\pi_\theta: the policy network
  • πref\pi_{\text{ref}}: the reference policy

The model is learning to classify preference pairs by adjusting its output probabilities. When the loss is low, the model assigns higher relative probability to the preferred response.

This classification view explains why DPO is so stable compared to RLHF. Instead of learning a separate reward model and then using policy gradients with their high variance, DPO directly optimizes the likelihood of preference labels. The gradients flow directly through the model's log probabilities, just like in standard language model training.

The gradient of the DPO loss with respect to θ\theta has a particularly intuitive form. For a single example:

∇θLDPO=−βσ(−r^θ(x,yw)+r^θ(x,yl))[∇θlog⁡πθ(yw∣x)−∇θlog⁡πθ(yl∣x)]\nabla_\theta \mathcal{L}_{\text{DPO}} = -\beta \sigma(-\hat{r}_\theta(x, y_w) + \hat{r}_\theta(x, y_l)) \left[ \nabla_\theta \log \pi_\theta(y_w|x) - \nabla_\theta \log \pi_\theta(y_l|x) \right]

where:

  • ∇θ\nabla_\theta: gradient with respect to model parameters θ\theta
  • LDPO\mathcal{L}_{\text{DPO}}: the DPO loss
  • σ\sigma: the sigmoid function
  • r^θ\hat{r}_\theta: the implicit reward function
  • σ(… )\sigma(\dots): the weighting term derived from the sigmoid derivative (probability of the wrong preference)
  • ∇θlog⁡…\nabla_\theta \log \dots: the geometric gradient direction
  • β\beta: the KL penalty coefficient

This gradient expression reveals the geometry of DPO optimization. The term ∇θlog⁡πθ(yw∣x)−∇θlog⁡πθ(yl∣x)\nabla_\theta \log \pi_\theta(y_w|x) - \nabla_\theta \log \pi_\theta(y_l|x) points in the direction that increases the probability of the winning response while decreasing the probability of the losing response. This is the natural direction to push given our objective.

The term σ(−r^θ(x,yw)+r^θ(x,yl))\sigma(-\hat{r}_\theta(x, y_w) + \hat{r}_\theta(x, y_l)) is the probability that the model currently assigns to the wrong preference. This acts as an implicit weighting:

  • When the model is confident and correct (low wrong-preference probability), the gradient is small
  • When the model is confident but wrong, the gradient is large
  • When the model is uncertain, the gradient is moderate

This automatic weighting helps the model focus on examples it's getting wrong, similar to how hard negative mining works in contrastive learning.

Out[4]:
Visualization
Line plot showing gradient weight decreasing sigmoidally from 1 to 0 as reward difference increases from -4 to 4.
The DPO gradient weighting function (probability of the wrong preference) across implicit reward differences. Large weights occur when the model incorrectly favors the rejected response (negative difference), triggering strong updates. Small weights occur when the model correctly favors the chosen response (positive difference), stabilizing training.

Connecting the Gradient to Language Model Training

It is instructive to compare the DPO gradient to what happens during ordinary language model pre-training. In standard cross-entropy training on tokens, the gradient for a single token at position tt increases the log-probability of the correct next token wtw_t and the magnitude is governed by the current probability of that token: if the model already assigns high probability, the gradient is small. DPO follows the same pattern but at the sequence level: it increases log-probability for the entire chosen sequence ywy_w and decreases it for the rejected sequence yly_l, weighted by how wrong the model currently is.

This parallel makes DPO feel like a natural extension of language modeling rather than a fundamentally different training paradigm. You can implement DPO in virtually any training framework that supports computing sequence log-probabilities, because the loss is just a scalar function of those log-probabilities. The practical implementation details, including how to handle variable-length sequences, padding masks, and batched reference model inference, are covered in the next chapter on DPO Implementation.

Worked Example

Let's trace through the derivation with concrete numbers to build intuition. Suppose we have a prompt xx = "Write a haiku about autumn" and two responses:

  • ywy_w (preferred): "Crimson leaves descend / Dancing on October wind / Earth prepares to sleep"
  • yly_l (dispreferred): "Leaves fall down in fall / The weather gets cold outside / I like pumpkin spice"

The preferred response demonstrates poetic craft through concrete imagery ("crimson," "dancing"), sensory detail, and thematic resonance with the natural world. The dispreferred response is grammatically correct but prosaic, repeating "fall" clumsily and ending on an unrelated tangent about pumpkin spice. A human evaluator would clearly prefer the first response, and a well-trained reward model should assign it a higher score.

Assume our reference model πref\pi_{\text{ref}} assigns:

  • log⁡πref(yw∣x)=−45.2\log \pi_{\text{ref}}(y_w|x) = -45.2 (log probability of generating the preferred response)
  • log⁡πref(yl∣x)=−38.7\log \pi_{\text{ref}}(y_l|x) = -38.7 (log probability of generating the dispreferred response)

The reference model assigns higher probability to the dispreferred response because it's simpler and uses more common words. The word "fall" is far more common in training data than "crimson" or "descend," and the dispreferred response is shorter and uses everyday vocabulary throughout. This is exactly the distribution shift DPO is designed to fix: the reference model, trained to predict natural text, has no built-in preference for poetic quality over simplicity.

Now suppose our current policy πθ\pi_\theta assigns:

  • log⁡πθ(yw∣x)=−42.1\log \pi_\theta(y_w|x) = -42.1
  • log⁡πθ(yl∣x)=−39.5\log \pi_\theta(y_l|x) = -39.5

With β=0.1\beta = 0.1, the implicit rewards are:

r^θ(x,yw)=0.1×(−42.1−(−45.2))=0.1×3.1=0.31\begin{aligned} \hat{r}_\theta(x, y_w) &= 0.1 \times (-42.1 - (-45.2)) \\ &= 0.1 \times 3.1 \\ &= 0.31 \end{aligned} r^θ(x,yl)=0.1×(−39.5−(−38.7))=0.1×(−0.8)=−0.08\begin{aligned} \hat{r}_\theta(x, y_l) &= 0.1 \times (-39.5 - (-38.7)) \\ &= 0.1 \times (-0.8) \\ &= -0.08 \end{aligned}

The difference is:

r^θ(x,yw)−r^θ(x,yl)=0.31−(−0.08)=0.39\begin{aligned} \hat{r}_\theta(x, y_w) - \hat{r}_\theta(x, y_l) &= 0.31 - (-0.08) \\ &= 0.39 \end{aligned}

The loss contribution from this example:

−log⁡σ(0.39)=−log⁡(0.596)=0.517\begin{aligned} -\log \sigma(0.39) &= -\log(0.596) \\ &= 0.517 \end{aligned}

The model is doing okay on this example (probability of correct preference is about 60%), but there's room to improve. The gradient will push the model to further increase πθ(yw∣x)\pi_\theta(y_w|x) relative to the reference and decrease πθ(yl∣x)\pi_\theta(y_l|x) relative to the reference.

Let's also trace through what a "good" and a "bad" model state look like for this same example. Suppose after more training the policy reaches:

  • log⁡πθ(yw∣x)=−38.0\log \pi_\theta(y_w|x) = -38.0 (probability has increased from reference)
  • log⁡πθ(yl∣x)=−41.5\log \pi_\theta(y_l|x) = -41.5 (probability has decreased from reference)

Now the implicit rewards are:

r^θ(x,yw)=0.1×(−38.0−(−45.2))=0.1×7.2=0.72r^θ(x,yl)=0.1×(−41.5−(−38.7))=0.1×(−2.8)=−0.28\begin{aligned} \hat{r}_\theta(x, y_w) &= 0.1 \times (-38.0 - (-45.2)) = 0.1 \times 7.2 = 0.72 \\ \hat{r}_\theta(x, y_l) &= 0.1 \times (-41.5 - (-38.7)) = 0.1 \times (-2.8) = -0.28 \end{aligned}

The reward margin is now 0.72−(−0.28)=1.000.72 - (-0.28) = 1.00, and the preference probability is σ(1.00)≈0.73\sigma(1.00) \approx 0.73. The loss drops to −log⁡(0.73)≈0.31-\log(0.73) \approx 0.31. The model has learned to "see" the quality difference between these two responses, even though it never directly observed a reward signal during this training phase.

Out[5]:
Visualization
Grouped bar chart showing log probabilities for preferred and rejected responses under both the reference model and the current policy.
Log probabilities of the preferred and rejected responses under the reference and policy models. The policy shifts mass toward the preferred haiku compared to the reference.
Bar chart showing log-ratios of policy to reference, with a positive bar for the preferred response and a negative bar for the rejected response.
Log-ratio of policy to reference probabilities. A positive ratio for the preferred response and negative for the rejected indicates the model is moving in the right direction.
Bar chart showing implicit rewards for preferred and rejected responses, with an annotation showing a positive margin of 0.39.
Implicit rewards scaled by $\beta = 0.1$. The positive margin of 0.39 shows the model correctly assigns a higher implicit reward to the preferred response.

Code Implementation

Let's implement the DPO loss function and verify our derivation with code.

In[6]:
Code
import torch
import torch.nn.functional as F


def dpo_loss(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    ref_chosen_logps: torch.Tensor,
    ref_rejected_logps: torch.Tensor,
    beta: float = 0.1,
) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
    """
    Compute the DPO loss for a batch of preference pairs.

    Args:
        policy_chosen_logps: Log probs of chosen responses under policy [batch_size]
        policy_rejected_logps: Log probs of rejected responses under policy [batch_size]
        ref_chosen_logps: Log probs of chosen responses under reference [batch_size]
        ref_rejected_logps: Log probs of rejected responses under reference [batch_size]
        beta: Temperature parameter controlling deviation from reference

    Returns:
        loss: Scalar DPO loss
        chosen_rewards: Implicit rewards for chosen responses [batch_size]
        rejected_rewards: Implicit rewards for rejected responses [batch_size]
    """
    # Compute implicit rewards (log-ratios scaled by beta)
    chosen_rewards = beta * (policy_chosen_logps - ref_chosen_logps)
    rejected_rewards = beta * (policy_rejected_logps - ref_rejected_logps)

    # DPO loss: negative log-sigmoid of reward difference
    logits = chosen_rewards - rejected_rewards
    loss = -F.logsigmoid(logits).mean()

    return loss, chosen_rewards, rejected_rewards

Now let's verify with our worked example:

In[7]:
Code
# Values from our worked example
policy_chosen_logps = torch.tensor([-42.1])
policy_rejected_logps = torch.tensor([-39.5])
ref_chosen_logps = torch.tensor([-45.2])
ref_rejected_logps = torch.tensor([-38.7])

loss, chosen_rewards, rejected_rewards = dpo_loss(
    policy_chosen_logps,
    policy_rejected_logps,
    ref_chosen_logps,
    ref_rejected_logps,
    beta=0.1,
)
Out[8]:
Console
Chosen implicit reward: 0.310
Rejected implicit reward: -0.080
Reward difference: 0.390
Preference probability: 0.596
DPO loss: 0.517

The implicit rewards align with our derivation: the chosen response has a positive reward (0.310), while the rejected response has a negative reward (-0.080). The positive difference (0.390) results in a preference probability of 0.596. This indicates the model correctly prefers the chosen response, though the probability is only moderately high.

Let's also visualize how the loss and preference probability change as the model learns:

Out[9]:
Visualization
Line chart showing sigmoid preference probability rising from near 0 to near 1 as the implicit reward difference increases from -4 to 4.
Sigmoid preference probability as a function of the implicit reward difference. The probability increases from near 0 to near 1 as the model learns to prefer chosen responses.
Line chart showing DPO loss decreasing from high values to near zero as the implicit reward difference increases, with a steep drop around zero.
DPO loss as a function of the implicit reward difference. The loss decreases toward zero as the reward margin grows, rewarding confident and correct preferences.

The loss approaches zero as the reward difference increases, meaning the model strongly prefers chosen responses. Conversely, when the model prefers rejected responses (negative reward difference), the loss grows large.

Let's also implement a function to compute the gradients and see how the implicit weighting works:

In[10]:
Code
def dpo_loss_with_grad_weights(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    ref_chosen_logps: torch.Tensor,
    ref_rejected_logps: torch.Tensor,
    beta: float = 0.1,
) -> tuple[torch.Tensor, torch.Tensor]:
    """
    Compute DPO loss and the implicit gradient weights.

    The gradient weight is sigma(-reward_diff), which is the probability
    the model assigns to the WRONG preference.
    """
    chosen_rewards = beta * (policy_chosen_logps - ref_chosen_logps)
    rejected_rewards = beta * (policy_rejected_logps - ref_rejected_logps)

    logits = chosen_rewards - rejected_rewards
    loss = -F.logsigmoid(logits).mean()

    # Gradient weight: probability of wrong preference
    grad_weight = torch.sigmoid(-logits)

    return loss, grad_weight
In[11]:
Code
# Test with different scenarios
scenarios = [
    ("Model strongly correct", -20.0, -83.5, -45.2, -38.7),
    ("Model weakly correct", -42.1, -50.6, -45.2, -38.7),
    ("Model uncertain", -45.0, -38.5, -45.2, -38.7),
    ("Model weakly wrong", -47.0, -25.5, -45.2, -38.7),
    ("Model strongly wrong", -75.5, -19.0, -45.2, -38.7),
]

results = []
for name, pc, pr, rc, rr in scenarios:
    loss, grad_weight = dpo_loss_with_grad_weights(
        torch.tensor([pc]),
        torch.tensor([pr]),
        torch.tensor([rc]),
        torch.tensor([rr]),
        beta=0.1,
    )
    results.append((name, loss.item(), grad_weight.item()))
Out[12]:
Console
Scenario                    Loss    Grad Weight
--------------------------------------------------
Model strongly correct      0.001      0.001
Model weakly correct        0.201      0.182
Model uncertain             0.693      0.500
Model weakly wrong          1.701      0.818
Model strongly wrong        5.007      0.993

Notice how the gradient weight automatically adjusts based on how wrong the model is. In the "Model strongly correct" case, the gradient weight is near zero (0.000), meaning there's little to learn. Conversely, when the model is "Strongly wrong", the gradient weight approaches one (0.993), pushing hard to correct the mistake.

Out[13]:
Visualization
Grouped bar chart comparing loss and gradient weight across five scenarios from strongly correct to strongly wrong.
Comparison of DPO loss and gradient weights across five model performance scenarios. The loss and gradient weights are minimal when the model is 'Strongly correct'. Both metrics increase as the model's predictions worsen, with 'Strongly wrong' predictions triggering the largest gradient updates to correct the behavior.

Key Parameters

The key parameter for DPO is:

  • beta: The temperature parameter (often denoted as β\beta) that scales the log-ratio of the policy and reference probabilities. It controls the strength of the KL divergence penalty, with larger values keeping the policy closer to the reference model.

Choosing β\beta in practice requires balancing two failure modes. If β\beta is too small, the DPO loss will push the policy aggressively in the direction of human preferences without adequate constraint from the reference model. This can cause the policy to drift far from the reference distribution and start generating fluent-but-wrong outputs, such as responses that contain the right topics but with incorrect factual content, or responses that are overly verbose because longer responses were slightly preferred in the training data. If β\beta is too large, the policy can barely move away from the reference, which means it learns slowly and may not improve much even after thousands of training steps. Empirically, values in the range 0.01 to 0.5 are common, with 0.1 being a popular default starting point.

Out[14]:
Visualization
Line chart showing implicit reward margin growing linearly as beta increases from 0.05 to 1.0, with an annotation showing larger margin equals stronger signal.
Implicit reward margin as a function of beta for fixed log-probability ratios. Increasing beta linearly scales the margin, amplifying the reward signal.
Line chart showing preference probability rising toward 1.0 as beta increases, with an annotation showing higher confidence with larger beta.
Preference probability as a function of beta. Larger beta drives the probability closer to 1.0, which reflects higher confidence in the learned preference.

Assumptions and Validity

The DPO derivation makes several assumptions worth examining carefully, because each one represents a potential point of failure in real-world applications.

The first assumption is that the Bradley-Terry model correctly captures human preferences. Specifically, DPO assumes that the probability of preferring response ywy_w over yly_l can be expressed as a sigmoid of a scalar reward difference. Real human preferences are far messier. They are noisy (the same annotator might rate the same pair differently on different days), inconsistent across annotators (different people have different values), and often non-transitive (a human might prefer y1y_1 over y2y_2, prefer y2y_2 over y3y_3, but also prefer y3y_3 over y1y_1). The Bradley-Terry model cannot represent this complexity. DPO inherits this limitation from the underlying preference model: if your preference data is very noisy or inconsistent, the loss will receive contradictory training signals and may not converge to a meaningful policy.

The second assumption is that the optimal policy has the Boltzmann form derived earlier. This is guaranteed only for the specific objective we started with, namely expected reward minus a KL divergence from the reference. Different alignment objectives, such as objectives using different divergence measures or adding additional safety constraints, would yield different optimal policies and potentially different direct alignment algorithms. The DPO loss is not a general-purpose alignment loss; it is the specific loss that corresponds to the KL-regularized RLHF objective with the Bradley-Terry preference model. When researchers have proposed modifications, such as using an ff-divergence instead of KL, they obtain modified versions of DPO with different properties. The next chapter discusses several of these variants.

The third assumption is that model capacity is sufficient. DPO optimizes toward the optimal policy but doesn't guarantee reaching it. The policy πθ\pi_\theta is constrained by the model's architecture and capacity. A small model may not be able to represent the true optimal policy, in which case DPO finds the best approximation within the model class. This is not unique to DPO, since all finite-parameter models face this limitation, but the theoretical guarantees assume that the policy class is expressive enough to represent the optimal Boltzmann distribution.

A fourth, more subtle assumption concerns the relationship between the data distribution and the behavior of the optimal policy. The derivation assumes we optimize over the true expectation under the training distribution D\mathcal{D}. In practice, we minimize an empirical loss over a finite dataset. If the preference dataset is small or unrepresentative of the prompts the model will encounter at deployment, DPO may overfit to the dataset distribution. This manifests as mode collapse on training prompts or failure to generalize to held-out prompt types. This issue is partly mitigated by the KL constraint (via the reference policy), but a badly distributed training set will still cause problems.

Finally, DPO requires that we can compute πref(y∣x)\pi_{\text{ref}}(y|x) exactly for the same responses we're evaluating under πθ\pi_\theta. This is straightforward when both are the same model architecture, but becomes complex if the reference model is unavailable or uses different tokenization. The next chapter on DPO Implementation will address these practical considerations in detail.

Summary

This chapter derived the DPO loss function from first principles, showing how it emerges naturally from the RLHF objective.

The key steps in the derivation were:

  1. RLHF objective: Maximize expected reward with a KL penalty to stay close to the reference policy
  2. Optimal policy: Using Lagrange multipliers, we found that the optimal policy is a Boltzmann distribution that reweights the reference by exponentiated rewards
  3. Reward reparameterization: Inverting this relationship, we expressed the implicit reward as a log-ratio between the optimal and reference policies, plus a partition function
  4. Bradley-Terry substitution: Plugging the reparameterized rewards into the preference model, the partition functions cancel, leaving a loss in terms of log-ratios only
  5. Classification view: The resulting DPO loss is binary cross-entropy, learning to classify which response is preferred based on relative log-probabilities

The gradient of the DPO loss has an automatic weighting scheme: examples where the model is confidently wrong receive larger gradients, while examples where it's already correct receive smaller gradients. This makes DPO training stable and efficient.

The derivation also reveals DPO's key assumptions: that human preferences follow the Bradley-Terry model, that the KL-regularized RLHF objective is the right starting point, and that the policy class is expressive enough to represent the optimal Boltzmann distribution. When those assumptions hold, DPO provides a principled, efficient alternative to RLHF. When they break down, the variants we explore in the following chapter offer modified objectives that relax different aspects of those assumptions.

In the next chapter, we'll implement DPO end-to-end, covering practical details like computing sequence log-probabilities, handling padding, and integrating with Hugging Face's training infrastructure.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the DPO derivation.

DPO Derivation

Question 1 of 80 of 8 completed
What is the primary purpose of the KL divergence term in the RLHF objective?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026dpoderivation, author = {Michael Brenndoerfer}, title = {DPO Derivation: From RLHF Objective to Direct Optimization}, year = {2026}, url = {https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). DPO Derivation: From RLHF Objective to Direct Optimization. Retrieved from https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization
MLAAcademic
Michael Brenndoerfer. "DPO Derivation: From RLHF Objective to Direct Optimization." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization>.
CHICAGOAcademic
Michael Brenndoerfer. "DPO Derivation: From RLHF Objective to Direct Optimization." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization.
HARVARDAcademic
Michael Brenndoerfer (2026) 'DPO Derivation: From RLHF Objective to Direct Optimization'. Available at: https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). DPO Derivation: From RLHF Objective to Direct Optimization. https://mbrenndoerfer.com/writing/dpo-derivation-rlhf-optimal-policy-reward-reparameterization

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.