DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment

Michael BrenndoerferJanuary 3, 202665 min read

Part of Language AI Handbook

DPO variants include IPO, KTO, ORPO, and cDPO. Compare their objectives, data requirements, computational costs, and suitable alignment tasks.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

DPO Variants: IPO, KTO, ORPO, and cDPO

In the previous chapters, we derived Direct Preference Optimization from first principles and implemented it as a simpler alternative to RLHF. DPO was a breakthrough: it showed that you could align language models with human preferences using a simple classification-style loss, without ever explicitly training a reward model or running a reinforcement learning loop. The key derivation revealed that the optimal policy under the KL-constrained reward maximization objective could be expressed directly in terms of the policy's own log-probabilities, collapsing what had been a two-stage pipeline into a single supervised objective. That elegance made DPO enormously attractive.

Yet elegance often conceals hidden fragility. As practitioners began applying DPO at scale, they encountered a cluster of practical problems that the theoretical derivation had glossed over. The training loss continued to decrease even when models were creating clearly overconfident predictions. Systems trained on crowd-sourced preference data often behaved erratically, apparently because some annotators had labeled responses incorrectly. Teams working with limited GPU memory struggled to keep two copies of a large model in VRAM simultaneously. And researchers collecting feedback from deployed products found that most of their data came in the form of thumbs-up and thumbs-down ratings on individual responses, not as head-to-head comparisons that DPO's paired format requires. Each of these problems was real, each caused real engineering pain, and each motivated a distinct line of research.

This chapter explores four of the most influential responses to those problems. Identity Preference Optimization (IPO) addresses the unbounded reward growth problem by replacing DPO's monotone loss with a regression target, giving optimization a well-defined stopping point. Kahneman-Tversky Optimization (KTO) discards the paired data requirement entirely, drawing on behavioral economics to define a learning objective that works with simple binary labels. Odds Ratio Preference Optimization (ORPO) eliminates the reference model, combining supervised fine-tuning and preference learning into a single forward pass that requires only one copy of the model. Conservative DPO (cDPO) tackles annotation noise through label smoothing, acknowledging that human judgments are probabilistic rather than deterministic.

Think of these four variants as a family of tools, each sharpened for a different job. Understanding when to reach for each one requires understanding the problem each one solves and the mechanism it uses to solve it. We will work through the mathematics of all four variants carefully, compare their gradient behavior, implement them in code, and walk through a concrete numerical example that makes the differences tangible. By the end of this chapter, you will have the conceptual toolkit to select the right alignment method for your data characteristics and computational constraints, rather than defaulting to the original DPO when a better-matched alternative exists.

Historical Context

The DPO variants covered in this chapter appeared in rapid succession between 2023 and 2024. This shows how quickly the alignment field moved once DPO demonstrated that RL-free preference learning was viable. IPO (Azar et al., 2024) identified a theoretical gap in DPO's convergence guarantees. KTO (Ethayarajh et al., 2024) drew inspiration from the Nobel Prize-winning work of Kahneman and Tversky on human decision-making under uncertainty. This produces a method that mirrors how behavioral economists model value perception. ORPO (Hong et al., 2024) emerged partly from the practical constraints of researchers at institutions without access to large multi-GPU clusters. cDPO (Mitchell et al., 2023) applied a decades-old classification technique, label smoothing, to the preference learning setting. The fact that all four addressed different limitations of the same base algorithm within roughly twelve months illustrates both how fertile the space was and how quickly the community can respond when a tractable framework becomes available.

Identity Preference Optimization (IPO)

IPO emerged from a careful analysis of DPO's theoretical properties. Azar et al. (2024) identified a subtle but important issue: as DPO training progresses, the implicit reward gap between chosen and rejected responses can grow unboundedly, causing the policy to assign extreme probabilities that do not reflect the actual strength of human preferences. This section explains exactly why that happens, what consequences it produces, and how IPO's squared-error reformulation resolves the problem.

Think of DPO's loss as a one-sided ramp: no matter how strongly the model already prefers the chosen response, the ramp keeps tilting in the same direction, always pushing the model to make the chosen response even more probable. IPO replaces the ramp with a bowl: there is a specific level of preference the model should reach, and any deviation from that level, whether too weak or too strong, incurs a penalty. The bowl has a well-defined minimum, and once the model reaches it, training stops pushing.

The Overfitting Problem in DPO

To understand why IPO was developed, we must first examine a basic tension within DPO's optimization dynamics. Recall from our DPO derivation that the loss function is:

LDPO=−E(x,yw,yl)[log⁡σ(β(log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)))]\mathcal{L}_{\text{DPO}} = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \beta \left( \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} \right) \right) \right]

where:

  • LDPO\mathcal{L}_{\text{DPO}}: the DPO loss function
  • E(x,yw,yl)\mathbb{E}_{(x, y_w, y_l)}: the expectation over the dataset of preference pairs
  • πθ(y∣x)\pi_\theta(y|x): the probability of response yy given prompt xx under the policy being trained
  • πref(y∣x)\pi_{\text{ref}}(y|x): the probability under the frozen reference model
  • yw,yly_w, y_l: the chosen (winner) and rejected (loser) responses
  • β\beta: the temperature parameter controlling the strength of the KL divergence penalty
  • σ\sigma: the sigmoid function

To appreciate the overfitting problem, consider what this loss function is asking the model to do. The expression inside the sigmoid is the difference between two log-ratios: how much more likely the policy makes the chosen response compared to the reference, minus the same quantity for the rejected response. The negative log-sigmoid of this quantity becomes the loss. When training minimizes this loss, it maximizes what is inside the log-sigmoid.

The sigmoid function σ\sigma approaches 1 as its argument grows large. To minimize this loss, the model is incentivized to make the log-ratio difference as large as possible. In the limit, this pushes the model toward assigning probability approaching 1 to chosen responses and probability approaching 0 to rejected ones. There is no natural stopping point in this formulation: the gradient always points toward making the gap larger, even when the gap is already enormous.

This behavior is problematic for several reasons. First, human preferences are not absolute. A response labeled "chosen" is not infinitely better than the "rejected" alternative; it is merely preferred in that comparison. Two responses might differ only slightly in quality, yet the labeling treats one as definitively better. When the model learns to assign extreme probabilities based on these labels, it develops a false sense of certainty that does not match the underlying reality of human judgment. Second, training instability compounds as log probabilities approach −∞-\infty for rejected responses, causing gradients to become numerically unreliable. Third, and most subtly, the model may generalize poorly: it in effect memorizes that specific responses are "infinitely good" or "infinitely bad" rather than learning the generalizable patterns that distinguish better responses from worse ones.

Out[3]:
Visualization
Line plot of the DPO negative log-sigmoid loss versus the log-ratio difference, with a red shaded region showing the overfitting risk zone at high values.
DPO's negative log-sigmoid loss as a function of the log-ratio difference. The loss decreases monotonically without bound, creating an ever-present incentive to increase the preference gap. The red shaded region marks where extreme values signal overfitting risk.

IPO's Squared Error Objective

IPO addresses the overfitting problem by reformulating preference learning as a regression problem with a specific target. The reason is that instead of letting the reward gap to grow without bound, we should specify how large we want that gap to be and then train the model to reach exactly that target. This turns an unbounded optimization problem into a bounded one with a clear convergence criterion.

The main point is: bounded optimization requires a target. DPO has no target for the log-ratio difference; it simply wants that difference to be as large as possible. IPO sets the target at 12β\frac{1}{2\beta} and penalizes any deviation from it, whether the model has not learned a strong enough preference or has learned an excessively strong one. This single design choice fundamentally changes the optimization landscape.

Instead of pushing the reward gap to infinity, IPO targets a fixed margin:

LIPO=E(x,yw,yl)[(log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)−12β)2]\mathcal{L}_{\text{IPO}} = \mathbb{E}_{(x, y_w, y_l)} \left[ \left( \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)} - \frac{1}{2\beta} \right)^2 \right]

where:

  • LIPO\mathcal{L}_{\text{IPO}}: the IPO loss function
  • E(x,yw,yl)\mathbb{E}_{(x, y_w, y_l)}: the expectation over the dataset of preference pairs
  • 12β\frac{1}{2\beta}: the target margin derived from the regularization strength β\beta
  • πθ,πref\pi_\theta, \pi_{\text{ref}}: the policy and reference models

The key innovation is the squared loss with a specific target. Rather than rewarding the model for making the gap as large as possible, IPO rewards the model for making the gap equal to a specific value. Any deviation from this target, whether too small or too large, incurs a penalty. This is the essence of regression: we have a target value and we minimize the squared distance to that target.

To analyze these dynamics more precisely, let's define the log-ratio difference as:

hθ(x,yw,yl)=log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)h_\theta(x, y_w, y_l) = \log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}

where:

  • hθ(x,yw,yl)h_\theta(x, y_w, y_l): the difference in log-probability ratios between chosen and rejected responses
  • log⁡πθ(y∣x)πref(y∣x)\log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)}: the "implicit reward" for a specific response yy

This quantity hθh_\theta captures the essence of what preference optimization reaches. It measures how much the policy has learned to prefer the chosen response over the rejected one, relative to what the reference model would predict. A positive value means the policy has shifted probability mass toward the chosen response; a larger positive value means a stronger learned preference.

With this notation, the IPO loss becomes elegantly simple:

LIPO=E[(hθ−12β)2]\mathcal{L}_{\text{IPO}} = \mathbb{E} \left[ \left( h_\theta - \frac{1}{2\beta} \right)^2 \right]

where:

  • LIPO\mathcal{L}_{\text{IPO}}: the IPO loss function
  • E\mathbb{E}: the expectation over the dataset
  • hθh_\theta: the log-ratio difference computed by the model
  • 12β\frac{1}{2\beta}: the specific target value that the difference should converge to

This formulation makes the regression nature of IPO crystal clear. We are asking the model to make hθh_\theta equal to 12β\frac{1}{2\beta} for every preference pair in the dataset. The squared loss penalizes deviations in either direction: if the model has not learned a strong enough preference (hθ<12βh_\theta < \frac{1}{2\beta}), the loss is positive and gradients push toward a larger gap; if the model has learned too strong a preference (hθ>12βh_\theta > \frac{1}{2\beta}), the loss is also positive and gradients push toward a smaller gap.

The gradient with respect to hθh_\theta reveals this self-correcting behavior:

∂LIPO∂hθ=2(hθ−12β)\frac{\partial \mathcal{L}_{\text{IPO}}}{\partial h_\theta} = 2 \left( h_\theta - \frac{1}{2\beta} \right)

where:

  • ∂LIPO∂hθ\frac{\partial \mathcal{L}_{\text{IPO}}}{\partial h_\theta}: the gradient of the loss with respect to the log-ratio difference
  • hθ−12βh_\theta - \frac{1}{2\beta}: the error term (distance from target) that drives the update

This gradient is zero precisely when hθ=12βh_\theta = \frac{1}{2\beta}, creating a stable equilibrium. Once the log-ratio difference reaches the target margin, there is no further pressure to increase it. The equilibrium is stable and attractive: regardless of where training starts, the gradients always point toward the target, and the strength of the gradient is proportional to the distance from the target.

Interpreting the Target Margin

The target 12β\frac{1}{2\beta} has a principled interpretation that connects IPO back to the theoretical foundations of preference learning. The β\beta parameter controls the strength of the KL constraint. A larger β\beta means weaker regularization, letting larger deviations from the reference policy. With IPO's target:

  • When β\beta is large (weak KL penalty), the target margin 12β\frac{1}{2\beta} is small.
  • When β\beta is small (strong KL penalty), the target margin is large.

This might seem counterintuitive until you consider a budget analogy. If you have a strict overall spending limit, you can afford to splurge on individual items because you are being careful everywhere else. Conversely, if your overall limit is loose, you must be more conservative on each purchase to avoid overshooting. Stronger KL constraints (small β\beta) permit larger per-example margins because the overall deviation is more tightly controlled.

The target margin also has an interpretation in terms of the Bradley-Terry model that underlies preference learning. In that model, the probability that response ywy_w is preferred to yly_l depends on the difference in their rewards. The target 12β\frac{1}{2\beta} corresponds to a specific preference probability representing a moderate preference rather than absolute certainty. This aligns with the reality that human annotations express preferences, not certainties. An annotator who chooses response A over response B is not asserting that A is infinitely better; they are expressing a judgment that may be uncertain, context-dependent, and occasionally reversed.

Out[4]:
Visualization
Line plot showing IPO target margin as a function of beta, with red dots marking common beta values at 0.1, 0.2, 0.3, and 0.5 and their corresponding target margins.
IPO target margin as a function of the beta parameter. As beta increases (weaker KL penalty), the target margin decreases. Red dots mark commonly used beta values and their corresponding target margins.

Gradient Comparison

Different loss functions produce different gradient behaviors, and those differences explain why methods perform differently in practice. The gradient determines how the model updates its parameters at each step, so differences in gradient behavior translate directly into differences in training dynamics. Understanding these gradients is the key to predicting how a method will behave when you deploy it on your data.

Think of the gradient as the engine that drives training. DPO's engine runs at full throttle when the model is uncertain and idles as the model becomes confident, but it never turns off. IPO's engine runs proportionally to how far the model is from the target and cuts off completely once the target is reached.

The DPO gradient magnitude is:

∣∂LDPO∂hθ∣=βσ(−βhθ)\left| \frac{\partial \mathcal{L}_{\text{DPO}}}{\partial h_\theta} \right| = \beta \sigma(-\beta h_\theta)

where:

  • β\beta: the temperature parameter
  • σ(−βhθ)\sigma(-\beta h_\theta): the sigmoid of the negative scaled log-ratio, acting as a weighting factor

This sigmoid term means DPO's gradient is largest when hθh_\theta is near zero (uncertain predictions) and vanishes as hθ→∞h_\theta \to \infty (confident predictions). The intuition is that DPO pushes hardest when the model is uncertain about which response is preferred and pushes more gently when the model is already confident. While vanishing gradients prevent infinite growth in principle, they also mean learning slows sharply once the model becomes confident. This creates a problematic dynamic: the model can still drift toward extreme probabilities because the gradients, though small, remain consistently positive, never reversing direction.

The IPO gradient magnitude is:

∣∂LIPO∂hθ∣=2∣hθ−12β∣\left| \frac{\partial \mathcal{L}_{\text{IPO}}}{\partial h_\theta} \right| = 2 \left| h_\theta - \frac{1}{2\beta} \right|

where:

  • ∣hθ−12β∣| h_\theta - \frac{1}{2\beta} |: the absolute distance from the target margin

IPO's gradient magnitude is proportional to distance from the target. If the model overshoots the target (hθ>12βh_\theta > \frac{1}{2\beta}), the gradient reverses direction and pushes back. This self-correcting behavior is entirely absent in DPO. The gradient grows stronger as the model strays further from the target, which means that even large deviations are corrected rather than allowed to persist indefinitely.

Kahneman-Tversky Optimization (KTO)

While DPO and IPO improve upon RLHF's computational complexity, they share a basic data requirement: paired preferences. Each training example must contain a prompt together with both a chosen and rejected response. This pairing constraint creates practical challenges that are easy to underestimate if you have only worked with carefully curated benchmark datasets.

Think of the difference between a controlled taste test and a product review. In a taste test, each participant explicitly compares two options side by side and says which they prefer. Product reviews, by contrast, arrive as standalone ratings: a customer gives three stars or five stars to a single item, with no explicit comparison to an alternative. Most of the feedback you collect from deployed language models looks much more like product reviews than taste tests, and forcing it into paired-comparison format requires either discarding most of your data or constructing artificial comparisons that may not reflect real user preferences.

KTO, introduced by Ethayarajh et al. (2024), eliminates the paired data requirement by designing an objective that works directly with unpaired binary feedback. The result is a method that can absorb the full breadth of available user signal without the preprocessing overhead of converting it into an unnatural format.

The Unpaired Feedback Problem

Real-world human feedback often comes in unpaired form, and this mismatch between how feedback is collected and how preference optimization algorithms expect data creates significant friction in practical alignment pipelines. Consider the forms that user feedback naturally takes:

  • Binary ratings: users click thumbs up or thumbs down on individual responses, with no stated alternative.
  • Flagging systems: users report problematic outputs without giving a corrected version.
  • Implicit signals: engagement metrics, session lengths, and follow-up question rates indicate whether a response was helpful, but never identify a counterfactual.

Converting this abundant unpaired feedback into DPO's paired format requires either discarding data or artificially constructing pairs. Discarding wastes useful signal. Constructing pairs introduces artifacts: if you pair a thumbs-up response to a randomly selected thumbs-down response, the pair may not reflect a real comparison because the two responses came from different conversations with different prompts and different contexts.

Consider a scenario where you have 10,000 thumbs-up ratings and 5,000 thumbs-down ratings, but these ratings come from different conversations. To use DPO, you would need to either match these into pairs somehow, losing most of your data, or generate new responses to create artificial comparisons. KTO avoids this entirely by treating each labeled response as an independent training example.

Inspiration from Prospect Theory

KTO draws inspiration from Kahneman and Tversky's prospect theory, which describes how humans make decisions under uncertainty. This connection to behavioral economics is more than a naming convention: it gives principled guidance for how to weight different types of feedback. Two key insights from behavioral economics directly inform KTO's design.

Reference dependence is the first insight. People evaluate outcomes relative to a reference point, not in absolute terms. A gain of $100 feels different depending on whether you expected $0 or $200. For alignment, this suggests that the quality of a model response should be measured relative to some baseline expectation, not in absolute terms. KTO implements this by measuring implicit rewards relative to the expected KL divergence across the training distribution, creating a dynamic baseline that adapts as training progresses.

Loss aversion is the second insight. Losses loom larger than equivalent gains. Losing $100 feels psychologically worse than gaining $100 feels good, and behavioral experiments consistently estimate that losses are weighted roughly 2x more heavily than gains in human value assessments. For alignment, this suggests that suppressing bad outputs may be more important than promoting good ones. You might forgive a bland response, but you remember a harmful or embarrassingly incorrect one. KTO incorporates this asymmetry by letting different loss weights for desirable and undesirable examples.

KTO incorporates both principles into its loss function, treating "desirable" and "undesirable" responses asymmetrically in a way that mirrors how humans experience quality differences.

Out[5]:
Visualization
Line plot of the prospect theory value function compared to a symmetric reference, with shaded red and green regions showing the loss and gain domains respectively.
Prospect theory's value function compared to a symmetric baseline. Losses (negative x-axis) produce steeper decreases in perceived value than equivalent gains produce increases, which shows the empirical phenomenon of loss aversion. KTO applies this asymmetry to alignment training.

The KTO Loss Function

The construction of KTO's loss function proceeds in several stages, each motivated by the behavioral economics principles described above. We begin by defining the implicit reward, which is the raw signal that KTO turns into a learning objective.

For a response yy to prompt xx with binary label z∈{desirable,undesirable}z \in \{\text{desirable}, \text{undesirable}\}, KTO defines:

rθ(x,y)=log⁡πθ(y∣x)πref(y∣x)r_\theta(x, y) = \log \frac{\pi_\theta(y|x)}{\pi_{\text{ref}}(y|x)}

where:

  • rθ(x,y)r_\theta(x, y): the implicit reward assigned to response yy given prompt xx
  • πθ(y∣x)\pi_\theta(y|x): the probability of the response under the current policy
  • πref(y∣x)\pi_{\text{ref}}(y|x): the probability of the response under the reference model

This is the same implicit reward used in DPO. It measures how much more likely the current policy makes this response compared to the reference. A positive implicit reward means the policy has learned to favor this response; a negative implicit reward means the policy has learned to disfavor it. The important property is that this quantity is well-defined for a single response, without needing a comparison partner.

KTO then defines a reference point that implements the reference dependence principle from prospect theory:

z0=Ex′∼D[KL(πθ(⋅∣x′)∥πref(⋅∣x′))]z_0 = \mathbb{E}_{x' \sim \mathcal{D}} \left[ \text{KL}(\pi_\theta(\cdot|x') \| \pi_{\text{ref}}(\cdot|x')) \right]

where:

  • z0z_0: the reference point (baseline) for evaluation
  • πθ,πref\pi_\theta, \pi_{\text{ref}}: the policy and reference models
  • D\mathcal{D}: the training dataset distribution
  • Ex′∼D\mathbb{E}_{x' \sim \mathcal{D}}: the expectation over prompts x′x' sampled from D\mathcal{D}
  • KL\text{KL}: the Kullback-Leibler divergence measuring the drift of the policy from the reference

The reference point z0z_0 is the expected KL divergence between policy and reference across the training distribution. In practice, this is estimated from a running average during training. The reference point defines what counts as "above average" versus "below average" performance. Rather than using an arbitrary fixed threshold, KTO adapts the reference point to the current state of training, creating a dynamic baseline that evolves as the model improves.

With the implicit reward and reference point defined, KTO constructs a value function that differs based on whether the response is desirable:

v(x,y)={σ(β⋅rθ(x,y)−z0)if z=desirableσ(z0−β⋅rθ(x,y))if z=undesirablev(x, y) = \begin{cases} \sigma(\beta \cdot r_\theta(x, y) - z_0) & \text{if } z = \text{desirable} \\ \sigma(z_0 - \beta \cdot r_\theta(x, y)) & \text{if } z = \text{undesirable} \end{cases}

where:

  • v(x,y)v(x, y): the value assigned to the response, bounded between 0 and 1
  • σ\sigma: the sigmoid function
  • zz: the label showing if the response is desirable or undesirable
  • β⋅rθ(x,y)−z0\beta \cdot r_\theta(x, y) - z_0: the shifted implicit reward used for desirable examples

This asymmetric definition is the mathematical implementation of reference dependence. For desirable responses, we ask whether the implicit reward exceeds the reference point. For undesirable responses, we ask whether the implicit reward falls below the reference point. The sigmoid function squashes these comparisons into the range (0, 1), creating a smooth value measure that is easy to differentiate.

The loss weights desirable and undesirable examples differently, implementing the loss aversion principle:

LKTO=E(x,y,z)[λz⋅(1−v(x,y))]\mathcal{L}_{\text{KTO}} = \mathbb{E}_{(x, y, z)} \left[ \lambda_z \cdot (1 - v(x, y)) \right]

where:

  • LKTO\mathcal{L}_{\text{KTO}}: the KTO loss function
  • E(x,y,z)\mathbb{E}_{(x, y, z)}: the expectation over the dataset of labeled examples
  • λz\lambda_z: the weighting factor specific to the label type
  • v(x,y)v(x, y): the value computed by the value function

The weighting factors are λdesirable=λd\lambda_{\text{desirable}} = \lambda_d and λundesirable=λu\lambda_{\text{undesirable}} = \lambda_u, typically with λu>λd\lambda_u > \lambda_d to implement loss aversion. By setting λu>λd\lambda_u > \lambda_d, we tell the model that failing to suppress a bad response is worse than failing to promote a good response. This asymmetry shows the empirical finding that people are more bothered by failures than impressed by equivalent successes.

Understanding the Value Function

The asymmetric value function encodes prospect theory's core insights, and understanding its behavior illuminates why KTO works. The function turns implicit rewards into values differently depending on the label. This creates distinct learning signals for positive and negative feedback.

For desirable responses, the sigmoid argument is β⋅rθ−z0\beta \cdot r_\theta - z_0. The model is rewarded with a low loss when the implicit reward exceeds the reference point. Intuitively, the model should push desirable responses to have higher implicit rewards than average. When β⋅rθ>z0\beta \cdot r_\theta > z_0, the sigmoid output is greater than 0.5, meaning the value is high and the loss (1−v)(1 - v) is low.

For undesirable responses, the sigmoid argument is z0−β⋅rθz_0 - \beta \cdot r_\theta. The model is rewarded when the implicit reward falls below the reference point. When β⋅rθ<z0\beta \cdot r_\theta < z_0, the sigmoid output is greater than 0.5, meaning the value is high and the loss is low. The model learns to associate undesirable responses with below-average implicit rewards.

The reference point z0z_0 is the dividing line between "gains" and "losses" in the prospect theory sense. This relative framing means KTO does not need paired comparisons; it learns to push good responses up and bad responses down relative to an adaptive baseline. The elegance of this approach is that it converts an inherently comparative problem into a classification problem, which is exactly the form that unpaired feedback gives.

Out[6]:
Visualization
Line plot showing sigmoid-shaped value functions for desirable and undesirable responses, with a vertical dashed line marking the reference point that separates the two response classes.
KTO value functions for desirable and undesirable responses. The reference point at the vertical dashed line divides the implicit reward space: desirable responses gain value above it, while undesirable responses gain value below it.
Line plot showing KTO loss curves for desirable and undesirable responses as functions of implicit reward, illustrating the opposing slopes around the reference point.
KTO loss functions for desirable and undesirable responses. Each curve is minimized in the region where the model correctly orders the implicit reward relative to the reference point.

Practical Implementation Details

KTO requires tracking the reference point z0z_0 during training. This running average must be estimated from the current batch of data and updated incrementally as training proceeds.

In[7]:
Code
import torch


def kto_loss(
    policy_logps: torch.Tensor,  # log π_θ(y|x) for batch
    reference_logps: torch.Tensor,  # log π_ref(y|x) for batch
    is_desirable: torch.Tensor,  # binary mask: 1 for desirable, 0 for undesirable
    kl_reference: float,  # z_0: running average KL
    beta: float = 0.1,
    lambda_d: float = 1.0,  # weight for desirable
    lambda_u: float = 1.0,  # weight for undesirable (often > lambda_d)
) -> torch.Tensor:
    """
    Compute KTO loss for a batch of (potentially unpaired) examples.
    """
    # Implicit reward: log ratio
    implicit_reward = policy_logps - reference_logps

    # Value function differs by desirability
    desirable_mask = is_desirable.bool()

    values = torch.zeros_like(policy_logps)
    values[desirable_mask] = torch.sigmoid(
        beta * implicit_reward[desirable_mask] - kl_reference
    )
    values[~desirable_mask] = torch.sigmoid(
        kl_reference - beta * implicit_reward[~desirable_mask]
    )

    # Weighted loss
    weights = torch.where(desirable_mask, lambda_d, lambda_u)
    loss = weights * (1 - values)

    return loss.mean()

The running average KL is typically computed as:

In[8]:
Code
def update_kl_reference(
    kl_reference: float, batch_kl: float, momentum: float = 0.99
):
    """Update running average of KL divergence."""
    return momentum * kl_reference + (1 - momentum) * batch_kl

The momentum parameter controls how quickly the reference point adapts to changes in the policy. A high momentum (close to 1) means the reference point changes slowly. This gives a stable baseline. A low momentum means it tracks the current policy more aggressively, which can be destabilizing if the policy is changing rapidly at the start of training.

KTO's Advantages

KTO offers several practical benefits for real-world alignment pipelines:

  • Data efficiency: Uses all available binary feedback without discarding unpaired examples.
  • Natural data format: Matches how most user feedback is collected in deployed products.
  • Principled asymmetry: The loss-aversion weighting shows empirical findings about human judgment rather than an arbitrary design choice.
  • Stable training: The reference point gives a grounding mechanism similar to IPO's target, preventing unbounded optimization.

Odds Ratio Preference Optimization (ORPO)

ORPO takes a more radical departure from the DPO framework by eliminating the reference model entirely. Introduced by Hong et al. (2024), ORPO combines supervised fine-tuning with preference optimization into a single training objective, compressing what had been a two-stage pipeline into a single pass.

Think of ORPO as designing a building that gives its own structural support rather than leaning against an adjacent building. DPO leans against the reference model: it measures every change relative to a frozen copy of the initial policy. ORPO's odds formulation gives internal stability by coupling the SFT loss and the odds ratio loss in a way that creates natural constraints on how far the policy can drift.

Motivation: The Reference Model Burden

All previous methods, including RLHF, DPO, IPO, and KTO, require maintaining a reference policy πref\pi_{\text{ref}}. This creates practical complications that matter most for researchers and engineers working with large models under tight resource constraints.

Memory overhead is the most immediate problem. Two copies of the model must reside in GPU memory simultaneously, or the reference model must be frequently loaded and unloaded from disk. For a 70-billion-parameter model, the reference copy alone requires roughly 140 GB of VRAM in BF16 precision, before accounting for gradients, optimizer states, and activations. This pushes many alignment experiments beyond the reach of single-node setups.

Two-stage training adds pipeline complexity. Models typically undergo supervised fine-tuning first to create a sensible reference policy, then preference training relative to that reference. Every bug, hyperparameter choice, and data quality issue must be diagnosed across two separate training runs rather than one.

Computational cost compounds the memory pressure. Every forward pass during preference training requires evaluating both policies: the current trainable policy and the frozen reference. This doubles the compute cost of every gradient step relative to supervised fine-tuning alone.

ORPO asks whether we can eliminate the reference model while still preventing the policy from drifting arbitrarily far from creating sensible outputs. The answer turns out to be yes, provided we design the objective carefully.

The Odds Ratio Approach

Instead of comparing log probabilities to a reference, ORPO uses the odds ratio between chosen and rejected responses within the current policy itself. This shift is significant. Instead of measuring change from a fixed reference, ORPO measures the relative likelihood of responses within the current policy. The odds of generating response yy given prompt xx are:

oddsθ(y∣x)=πθ(y∣x)1−πθ(y∣x)\text{odds}_\theta(y|x) = \frac{\pi_\theta(y|x)}{1 - \pi_\theta(y|x)}

where:

  • oddsθ(y∣x)\text{odds}_\theta(y|x): the odds of generating response yy under policy θ\theta
  • πθ(y∣x)\pi_\theta(y|x): the probability of the response

The odds representation has a natural interpretation: it tells us how likely the response is compared to everything else. If the odds are 2:1, the response is twice as likely as all alternatives combined. The odds formulation is particularly useful because ratios of odds have clean mathematical properties that are difficult to reach with plain probability ratios.

For a language model, the probability of a specific sequence becomes numerically insignificant as length increases, making direct odds calculation unstable. Consider that even a moderately long sequence might have a probability of 10−5010^{-50} or smaller, which would make both the numerator and denominator of the odds calculation problematically small. ORPO instead defines the sequence-level odds as the geometric mean of the token-level odds:

oddsθ(y∣x)≈exp⁡(1∣y∣∑t=1∣y∣log⁡πθ(yt∣x,y<t)1−πθ(yt∣x,y<t))\text{odds}_\theta(y|x) \approx \exp\left( \frac{1}{|y|} \sum_{t=1}^{|y|} \log \frac{\pi_\theta(y_t|x, y_{<t})}{1 - \pi_\theta(y_t|x, y_{<t})} \right)

where:

  • xx: the input prompt
  • ∣y∣|y|: the length of the response in tokens
  • yty_t: the token at step tt
  • y<ty_{<t}: the sequence of tokens preceding step tt (the context)
  • πθ(yt∣x,y<t)\pi_\theta(y_t|x, y_{<t}): the probability of the next token given the prompt xx and previous tokens y<ty_{<t}
  • exp⁡(… )\exp(\dots): the exponentiation to convert average log-odds back to the odds scale

This formulation works at the token level, where probabilities are large enough to be numerically stable, and then aggregates these token-level odds into a sequence-level measure. The averaging by sequence length ensures that longer sequences are not automatically penalized, since we take a geometric mean rather than a product.

The ratio of odds between chosen (ywy_w) and rejected (yly_l) responses becomes:

ORθ(yw,yl)=oddsθ(yw∣x)oddsθ(yl∣x)\text{OR}_\theta(y_w, y_l) = \frac{\text{odds}_\theta(y_w|x)}{\text{odds}_\theta(y_l|x)}

where:

  • ORθ\text{OR}_\theta: the odds ratio between the chosen and rejected responses
  • yw,yly_w, y_l: the chosen and rejected responses

This odds ratio captures the relative preference of the current policy for the chosen response over the rejected one. An odds ratio greater than 1 means the policy favors the chosen response; we want to train the policy to increase this ratio.

The ORPO Loss Function

ORPO combines two components that work together to give both language modeling signal and preference signal. This combination is what allows ORPO to eliminate the separate SFT stage and the reference model.

The first component is the supervised fine-tuning loss on the chosen response:

LSFT=−E(x,yw)[log⁡πθ(yw∣x)]\mathcal{L}_{\text{SFT}} = -\mathbb{E}_{(x, y_w)} \left[ \log \pi_\theta(y_w|x) \right]

where:

  • LSFT\mathcal{L}_{\text{SFT}}: the supervised fine-tuning loss component
  • E(x,yw)\mathbb{E}_{(x, y_w)}: the expectation over prompts and chosen responses
  • πθ(yw∣x)\pi_\theta(y_w|x): the likelihood of the chosen response

This component serves two purposes: it teaches the model to generate fluent, coherent text (the standard language modeling objective), and it gives an anchor that prevents the model from drifting too far from creating sensible outputs. By training on the chosen responses, the model learns what good outputs look like and maintains a probability distribution that assigns reasonable mass to those outputs.

The second component is the odds ratio loss that increases the relative odds of chosen over rejected:

LOR=−E(x,yw,yl)[log⁡σ(log⁡ORθ(yw,yl))]\mathcal{L}_{\text{OR}} = -\mathbb{E}_{(x, y_w, y_l)} \left[ \log \sigma \left( \log \text{OR}_\theta(y_w, y_l) \right) \right]

where:

  • LOR\mathcal{L}_{\text{OR}}: the odds ratio loss component
  • E(x,yw,yl)\mathbb{E}_{(x, y_w, y_l)}: the expectation over the dataset of preference pairs
  • yw,yly_w, y_l: the chosen and rejected responses
  • log⁡ORθ(yw,yl)\log \text{OR}_\theta(y_w, y_l): the log odds ratio (log of the ratio of odds)
  • σ\sigma: the sigmoid function

This component implements the preference learning objective. By maximizing the log-sigmoid of the log odds ratio, we push the model to make the chosen response relatively more likely than the rejected one. The structure is similar to DPO's loss, but operates on odds ratios within a single policy rather than log-probability ratios between two policies.

The combined ORPO objective is:

LORPO=LSFT+λ⋅LOR\mathcal{L}_{\text{ORPO}} = \mathcal{L}_{\text{SFT}} + \lambda \cdot \mathcal{L}_{\text{OR}}

where:

  • LORPO\mathcal{L}_{\text{ORPO}}: the total ORPO loss
  • λ\lambda: the coefficient weighting the odds ratio loss against the SFT loss

The hyperparameter λ\lambda controls the trade-off between learning to generate good responses and learning to distinguish good responses from bad ones. Too small a λ\lambda and the model ignores preferences; too large and the model may sacrifice fluency for preference optimization.

Why Odds Ratios Work

The odds ratio formulation gives implicit regularization without an explicit reference model. To understand why this works, consider what happens during training. The SFT loss pulls the model toward generating the chosen response. The OR loss pushes chosen odds higher relative to rejected odds. Both losses operate on the same policy, creating a coupled optimization.

The key insight is that increasing odds for chosen responses while decreasing odds for rejected responses automatically constrains how much the model can deviate from generating coherent text. If the model tried to maximize the odds ratio by assigning near-zero probability to rejected responses, the SFT loss on chosen responses would suffer because probability mass must be conserved. The model cannot simply declare everything "bad"; it must maintain a coherent probability distribution over all possible responses.

This coupling between the two loss components creates an implicit regularization effect. The SFT loss ensures the model keeps generating reasonable text, while the OR loss ensures it prefers better text to worse text. Neither loss alone would reach both objectives, but together they give a balanced training signal without requiring a separate frozen reference model.

Comparing Log-Ratio and Odds-Ratio Objectives

Comparing what DPO and ORPO optimize reveals how they prevent policy drift.

DPO optimizes:

log⁡πθ(yw∣x)πref(yw∣x)−log⁡πθ(yl∣x)πref(yl∣x)\log \frac{\pi_\theta(y_w|x)}{\pi_{\text{ref}}(y_w|x)} - \log \frac{\pi_\theta(y_l|x)}{\pi_{\text{ref}}(y_l|x)}

where:

  • πθ,πref\pi_\theta, \pi_{\text{ref}}: the policy and reference model probabilities
  • yw,yly_w, y_l: the chosen and rejected responses

ORPO optimizes:

log⁡oddsθ(yw∣x)oddsθ(yl∣x)=log⁡oddsθ(yw∣x)−log⁡oddsθ(yl∣x)\log \frac{\text{odds}_\theta(y_w|x)}{\text{odds}_\theta(y_l|x)} = \log \text{odds}_\theta(y_w|x) - \log \text{odds}_\theta(y_l|x)

where:

  • log⁡oddsθ(y∣x)\log \text{odds}_\theta(y|x): the log-odds of a response under the current policy
  • yw,yly_w, y_l: the chosen and rejected responses

DPO measures how much the policy's preference for ywy_w over yly_l has changed relative to the reference. ORPO measures the absolute odds ratio within the current policy. The reference model in DPO is an external anchor. ORPO's SFT component and odds formulation give an alternative form of internal anchoring. The two approaches reach similar regularization effects through fundamentally different mechanisms.

Out[9]:
Visualization
Diagram of the DPO architecture showing two boxes labeled Policy and Reference connected by arrows to a DPO Loss box, with a note showing 2x model memory usage.
DPO architecture requiring two models. The policy model is trained while a frozen reference model is maintained in memory to compute the probability ratios for the loss.
Diagram of the ORPO architecture showing a single Policy box connected to a combined SFT plus OR Loss box, with a note showing 1x model memory usage.
ORPO architecture using a single model. The policy is its own reference through the odds ratio formulation, combining SFT and preference optimization in one forward pass.

Conservative DPO (cDPO)

The variants discussed so far address algorithmic limitations. cDPO, introduced by Mitchell et al. (2023), addresses a data quality issue: label noise in preference annotations. While IPO and ORPO assume that the training labels are reliable (even if the optimization dynamics are problematic), cDPO starts from the recognition that human annotations are inherently uncertain and builds that uncertainty directly into the loss function.

Think of cDPO as the difference between a GPS that reports its location with false precision and one that reports a confidence interval. Standard DPO says "response A is better than response B, and that fact is certain." cDPO says "response A is probably better than response B, with probability 1−ϵ1 - \epsilon." The probabilistic framing leads to a fundamentally different loss landscape with a bounded optimal preference gap rather than an infinite one.

Label Noise in Preference Data

Human preference annotations are inherently noisy, and this noise can arise from several sources. Annotators may disagree with each other on the same comparison, make mistakes due to fatigue or inattention, apply inconsistent criteria across examples, or be influenced by surface features such as formatting and length rather than actual quality differences.

Studies of inter-annotator agreement on preference tasks typically show agreement rates of 70 to 80 percent, meaning 20 to 30 percent of labels may be "wrong" in the sense that a different annotator would have labeled them oppositely. Even gold-standard human annotations on tasks like toxicity detection or factual accuracy show substantial disagreement when the same examples are labeled by different annotators. The problem is not that annotators are careless; preferences are subjective, and the ranking of two similar responses is often not clear-cut.

Standard DPO treats all preference labels as ground truth. If the training data says yw≻yly_w \succ y_l, the model learns to prefer ywy_w with full confidence. When this label is wrong, the model learns an incorrect preference, and because DPO's unbounded optimization pushes that preference to an extreme, even a small percentage of mislabeled examples can cause significant downstream problems. The model in effect memorizes the noise in the data and tries to push it to infinity.

Label Smoothing for Preferences

cDPO applies label smoothing to preference labels. The idea behind label smoothing is familiar from classification: instead of training on hard targets (0 or 1), we train on soft targets that acknowledge uncertainty. For preferences, instead of treating each comparison as a certainty, we model the possibility that the annotation might be wrong.

Instead of treating preferences as deterministic (P(yw≻yl)=1P(y_w \succ y_l) = 1), cDPO models uncertainty:

P(yw≻yl)=1−ϵP(yl≻yw)=ϵ\begin{aligned} P(y_w \succ y_l) &= 1 - \epsilon \\ P(y_l \succ y_w) &= \epsilon \end{aligned}

where:

  • ϵ\epsilon: the label smoothing parameter, ϵ∈[0,0.5)\epsilon \in [0, 0.5), representing uncertainty about annotation correctness

The parameter ϵ\epsilon is our belief about how often the annotations are wrong. If ϵ=0\epsilon = 0, we trust all annotations completely, recovering standard DPO. If ϵ=0.3\epsilon = 0.3, we believe that 30 percent of the time the annotators got it backwards and the rejected response was better.

The cDPO loss modifies the Bradley-Terry model probability accordingly:

LcDPO=−E(x,yw,yl)[(1−ϵ)log⁡σ(β⋅hθ)+ϵlog⁡σ(−β⋅hθ)]\mathcal{L}_{\text{cDPO}} = -\mathbb{E}_{(x, y_w, y_l)} \Big[ (1-\epsilon) \log \sigma(\beta \cdot h_\theta) + \epsilon \log \sigma(-\beta \cdot h_\theta) \Big]

where:

  • LcDPO\mathcal{L}_{\text{cDPO}}: the cDPO loss function
  • E(x,yw,yl)\mathbb{E}_{(x, y_w, y_l)}: the expectation over the dataset
  • hθh_\theta: the log-ratio difference between chosen and rejected responses
  • ϵ\epsilon: the label smoothing parameter
  • σ\sigma: the sigmoid function
  • β\beta: the temperature parameter

This loss function has two terms. The first term, weighted by (1−ϵ)(1-\epsilon), is the standard DPO loss: it rewards the model for preferring the labeled chosen response. The second term, weighted by ϵ\epsilon, is the opposite: it rewards the model for preferring the labeled rejected response. The combination is the expected loss under our belief about annotation accuracy. We train the model to do well on average across both the case where the annotation is correct and the case where it is wrong.

Using the identity σ(−z)=1−σ(z)\sigma(-z) = 1 - \sigma(z), we can rewrite this loss to explicitly show the probability mass assigned to the rejected response:

LcDPO=−E[(1−ϵ)log⁡σ(β⋅hθ)+ϵlog⁡(1−σ(β⋅hθ))]\mathcal{L}_{\text{cDPO}} = -\mathbb{E} \Big[ (1-\epsilon) \log \sigma(\beta \cdot h_\theta) + \epsilon \log (1 - \sigma(\beta \cdot h_\theta)) \Big]

where:

  • 1−σ(β⋅hθ)1 - \sigma(\beta \cdot h_\theta): the probability assigned to the rejected response (equivalent to σ(−β⋅hθ)\sigma(-\beta \cdot h_\theta))
  • ϵ\epsilon: the weighting factor for the "incorrect" preference direction

This form makes the label smoothing interpretation particularly clear: we are mixing the standard DPO loss with its inverse, with the mixing coefficient ϵ\epsilon representing the probability of annotation error.

Effect on Optimization

Label smoothing fundamentally changes the loss landscape. The gradient of cDPO with respect to hθh_\theta is:

∂LcDPO∂hθ=−β[(1−ϵ)⋅σ(−β⋅hθ)−ϵ⋅σ(β⋅hθ)]\frac{\partial \mathcal{L}_{\text{cDPO}}}{\partial h_\theta} = -\beta \left[ (1-\epsilon) \cdot \sigma(-\beta \cdot h_\theta) - \epsilon \cdot \sigma(\beta \cdot h_\theta) \right]

where:

  • ∂LcDPO∂hθ\frac{\partial \mathcal{L}_{\text{cDPO}}}{\partial h_\theta}: the gradient of the loss with respect to the log-ratio difference

This gradient has two competing terms. The first term pushes toward larger hθh_\theta (stronger preference for chosen). The second term pushes toward smaller hθh_\theta (acknowledging that the rejected response might be better). These competing forces create a balance point where the gradient is zero.

Setting the gradient to zero allows us to solve for the optimal log-ratio difference hθ∗h_\theta^*:

σ(β⋅hθ∗)=1−ϵ(1−ϵ)+ϵ(rearrange terms)=1−ϵ(simplify denominator)β⋅hθ∗=log⁡(1−ϵϵ)(invert sigmoid)\begin{aligned} \sigma(\beta \cdot h_\theta^*) &= \frac{1 - \epsilon}{(1 - \epsilon) + \epsilon} && \text{(rearrange terms)} \\ &= 1 - \epsilon && \text{(simplify denominator)} \\ \beta \cdot h_\theta^* &= \log \left( \frac{1-\epsilon}{\epsilon} \right) && \text{(invert sigmoid)} \end{aligned}

where:

  • hθ∗h_\theta^*: the optimal value of the log-ratio difference
  • ϵ\epsilon: the label smoothing parameter

This means cDPO has a finite optimal gap (similar to IPO), which depends on ϵ\epsilon. With standard DPO (ϵ=0\epsilon = 0), the optimal gap is infinite. With cDPO (ϵ>0\epsilon > 0), the model converges to a bounded preference strength. The relationship between ϵ\epsilon and the optimal gap is intuitive: higher annotation uncertainty (larger ϵ\epsilon) leads to a smaller optimal gap, because the model should not be confident when the labels might be wrong.

Out[10]:
Visualization
Line plot of the cDPO equilibrium log-ratio gap as a function of label smoothing epsilon, with red scatter points at epsilon values 0.1, 0.2, and 0.3 annotating the corresponding h-star values.
cDPO equilibrium point as a function of label smoothing epsilon. Higher uncertainty leads to a smaller optimal log-ratio gap, which shows reduced confidence in annotations.
Line plot comparing normalized cDPO loss curves for epsilon values 0, 0.1, 0.2, and 0.3 across log-ratio differences, showing how increasing smoothing shifts the loss minimum toward zero.
cDPO loss landscapes for different smoothing values. As epsilon increases, the loss minimum shifts closer to zero and the slope becomes gentler, which shows increased uncertainty in the preference labels.

Choosing the Smoothing Parameter

The smoothing parameter ϵ\epsilon should reflect actual annotation noise. When inter-annotator agreement data is available, ϵ\epsilon can be estimated directly as the disagreement rate. When that data is unavailable, practical guidelines based on dataset characteristics give reasonable starting points:

  • ϵ=0.1\epsilon = 0.1: Assumes 90% annotation accuracy, appropriate for high-quality curated datasets with expert annotators.
  • ϵ=0.2\epsilon = 0.2: Assumes 80% accuracy, appropriate for crowd-sourced annotations where annotators receive limited instructions.
  • ϵ=0.3\epsilon = 0.3: Assumes 70% accuracy, appropriate for noisy or automated annotations generated by another model.

The effective training signal strength scales with (1−2ϵ)(1 - 2\epsilon). At ϵ=0.1\epsilon = 0.1, the signal is 80% as strong as standard DPO. At ϵ=0.3\epsilon = 0.3, it is only 40% as strong. This means that high smoothing values require more training data or more epochs to reach the same degree of preference learning, creating a practical trade-off between robustness to noise and data efficiency.

Worked Example: Computing Losses Step by Step

To make the differences between these methods concrete, let's walk through a specific numerical example using a single preference pair. This will illustrate how each loss function responds to the same underlying data and why their behaviors diverge.

Suppose we are training a language model to answer questions more helpfully. For the prompt "What is the boiling point of water?" we have two candidate responses:

  • ywy_w (chosen): "Water boils at 100 degrees Celsius (212 degrees Fahrenheit) at standard atmospheric pressure."
  • yly_l (rejected): "Water boils at a really high temperature."

Our reference model assigns log-probabilities:

  • log⁡πref(yw∣x)=−3.2\log \pi_{\text{ref}}(y_w|x) = -3.2 (moderately likely under the base model)
  • log⁡πref(yl∣x)=−2.8\log \pi_{\text{ref}}(y_l|x) = -2.8 (also reasonably likely, since it is short and plausible)

After some training steps, our current policy assigns:

  • log⁡πθ(yw∣x)=−2.4\log \pi_\theta(y_w|x) = -2.4 (the policy has learned to prefer the specific response)
  • log⁡πθ(yl∣x)=−4.1\log \pi_\theta(y_l|x) = -4.1 (the policy has learned to penalize the vague response)

Let's set β=0.1\beta = 0.1.

Step 1: Compute the implicit rewards.

The implicit reward for each response is the log-ratio relative to the reference:

rθ(x,yw)=log⁡πθ(yw∣x)−log⁡πref(yw∣x)=−2.4−(−3.2)=0.8\begin{aligned} r_\theta(x, y_w) &= \log \pi_\theta(y_w|x) - \log \pi_{\text{ref}}(y_w|x) \\ &= -2.4 - (-3.2) = 0.8 \end{aligned} rθ(x,yl)=log⁡πθ(yl∣x)−log⁡πref(yl∣x)=−4.1−(−2.8)=−1.3\begin{aligned} r_\theta(x, y_l) &= \log \pi_\theta(y_l|x) - \log \pi_{\text{ref}}(y_l|x) \\ &= -4.1 - (-2.8) = -1.3 \end{aligned}

The chosen response has a positive implicit reward (the policy has moved toward it relative to the reference), while the rejected response has a negative implicit reward (the policy has moved away from it).

Step 2: Compute hθh_\theta.

The log-ratio difference is:

hθ=rθ(x,yw)−rθ(x,yl)=0.8−(−1.3)=2.1h_\theta = r_\theta(x, y_w) - r_\theta(x, y_l) = 0.8 - (-1.3) = 2.1

Step 3: Compute the DPO loss.

LDPO=−log⁡σ(β⋅hθ)=−log⁡σ(0.1×2.1)=−log⁡σ(0.21)\mathcal{L}_{\text{DPO}} = -\log \sigma(\beta \cdot h_\theta) = -\log \sigma(0.1 \times 2.1) = -\log \sigma(0.21)

σ(0.21)=11+e−0.21≈11+0.811≈0.552\sigma(0.21) = \frac{1}{1 + e^{-0.21}} \approx \frac{1}{1 + 0.811} \approx 0.552

LDPO=−log⁡(0.552)≈0.594\mathcal{L}_{\text{DPO}} = -\log(0.552) \approx 0.594

This loss is still substantial. DPO would continue pushing hθh_\theta higher, even though the model already clearly prefers the chosen response.

Step 4: Compute the IPO loss with β=0.1\beta = 0.1, giving target 12β=5.0\frac{1}{2\beta} = 5.0.

LIPO=(hθ−5.0)2=(2.1−5.0)2=(−2.9)2=8.41\mathcal{L}_{\text{IPO}} = (h_\theta - 5.0)^2 = (2.1 - 5.0)^2 = (-2.9)^2 = 8.41

IPO reports a large loss because the current hθ=2.1h_\theta = 2.1 is far below the target of 5.0. The gradient is:

∂LIPO∂hθ=2(2.1−5.0)=−5.8\frac{\partial \mathcal{L}_{\text{IPO}}}{\partial h_\theta} = 2(2.1 - 5.0) = -5.8

This gradient is negative, meaning IPO is pushing hθh_\theta upward toward the target, just as DPO does in this case. If instead hθ=8.0h_\theta = 8.0, the IPO gradient would be 2(8.0−5.0)=+6.02(8.0 - 5.0) = +6.0, pushing hθh_\theta back down. DPO would never push downward.

Step 5: Compute the cDPO loss with ϵ=0.2\epsilon = 0.2.

LcDPO=−(1−0.2)log⁡σ(0.21)−0.2log⁡σ(−0.21)=−0.8log⁡(0.552)−0.2log⁡(0.448)=−0.8×(−0.594)−0.2×(−0.802)≈0.475+0.160=0.635\begin{aligned} \mathcal{L}_{\text{cDPO}} &= -(1 - 0.2) \log \sigma(0.21) - 0.2 \log \sigma(-0.21) \\ &= -0.8 \log(0.552) - 0.2 \log(0.448) \\ &= -0.8 \times (-0.594) - 0.2 \times (-0.802) \\ &\approx 0.475 + 0.160 = 0.635 \end{aligned}

The cDPO loss is slightly higher than DPO here (0.635 vs. 0.594). More importantly, the equilibrium for cDPO with ϵ=0.2\epsilon = 0.2 and β=0.1\beta = 0.1 is:

hθ∗=log⁡((1−0.2)/0.2)0.1=log⁡(4)0.1≈1.3860.1=13.86h_\theta^* = \frac{\log((1 - 0.2)/0.2)}{0.1} = \frac{\log(4)}{0.1} \approx \frac{1.386}{0.1} = 13.86

This is a finite target, far smaller than DPO's implicit infinite target.

Step 6: Compute the KTO value for the chosen response, assuming a reference point z0=0.05z_0 = 0.05.

v(x,yw)=σ(β⋅rθ(x,yw)−z0)=σ(0.1×0.8−0.05)=σ(0.03)≈0.507v(x, y_w) = \sigma(\beta \cdot r_\theta(x, y_w) - z_0) = \sigma(0.1 \times 0.8 - 0.05) = \sigma(0.03) \approx 0.507

The KTO loss contribution for this desirable example (with λd=1.0\lambda_d = 1.0) is:

λd⋅(1−v)=1.0×(1−0.507)=0.493\lambda_d \cdot (1 - v) = 1.0 \times (1 - 0.507) = 0.493

Since the implicit reward just barely exceeds the reference point, the model has done only slightly better than average for this example. The loss is close to 0.5, the maximum possible, suggesting the model should continue improving this response.

This worked example illustrates a key pattern: DPO, IPO, and cDPO all respond to the same hθh_\theta value but with different equilibria, while KTO operates on individual implicit rewards rather than their difference.

Implementation Comparison

With the theory and intuition in place, let's implement all four variants to see the differences in practice. Clear implementations help reveal the structural similarities and differences that the mathematics sometimes obscures.

In[11]:
Code
import torch
import torch.nn.functional as F


def compute_log_ratios(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    reference_chosen_logps: torch.Tensor,
    reference_rejected_logps: torch.Tensor,
) -> torch.Tensor:
    """Compute the log-ratio difference h_θ used by DPO, IPO, and cDPO."""
    chosen_ratio = policy_chosen_logps - reference_chosen_logps
    rejected_ratio = policy_rejected_logps - reference_rejected_logps
    return chosen_ratio - rejected_ratio


def dpo_loss(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    reference_chosen_logps: torch.Tensor,
    reference_rejected_logps: torch.Tensor,
    beta: float = 0.1,
) -> torch.Tensor:
    """Standard DPO loss."""
    h_theta = compute_log_ratios(
        policy_chosen_logps,
        policy_rejected_logps,
        reference_chosen_logps,
        reference_rejected_logps,
    )
    return -F.logsigmoid(beta * h_theta).mean()


def ipo_loss(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    reference_chosen_logps: torch.Tensor,
    reference_rejected_logps: torch.Tensor,
    beta: float = 0.1,
) -> torch.Tensor:
    """IPO loss with squared error objective."""
    h_theta = compute_log_ratios(
        policy_chosen_logps,
        policy_rejected_logps,
        reference_chosen_logps,
        reference_rejected_logps,
    )
    target = 1.0 / (2.0 * beta)
    return ((h_theta - target) ** 2).mean()


def cdpo_loss(
    policy_chosen_logps: torch.Tensor,
    policy_rejected_logps: torch.Tensor,
    reference_chosen_logps: torch.Tensor,
    reference_rejected_logps: torch.Tensor,
    beta: float = 0.1,
    epsilon: float = 0.1,  # label smoothing
) -> torch.Tensor:
    """Conservative DPO with label smoothing."""
    h_theta = compute_log_ratios(
        policy_chosen_logps,
        policy_rejected_logps,
        reference_chosen_logps,
        reference_rejected_logps,
    )
    # Smoothed loss: (1-ε)log σ(βh) + ε log σ(-βh)
    loss = -(1 - epsilon) * F.logsigmoid(
        beta * h_theta
    ) - epsilon * F.logsigmoid(-beta * h_theta)
    return loss.mean()

Now let's implement KTO and ORPO, which have different data requirements:

In[12]:
Code
import torch
import torch.nn.functional as F


def kto_loss(
    policy_logps: torch.Tensor,  # log probs for all examples
    reference_logps: torch.Tensor,  # reference log probs
    is_desirable: torch.Tensor,  # binary: 1=good, 0=bad
    kl_reference: float,  # running average KL
    beta: float = 0.1,
    lambda_d: float = 1.0,
    lambda_u: float = 1.0,
) -> torch.Tensor:
    """KTO loss for unpaired binary feedback."""
    implicit_reward = beta * (policy_logps - reference_logps)

    # Different value functions for desirable vs undesirable
    desirable_mask = is_desirable.bool()

    values = torch.zeros_like(policy_logps)
    values[desirable_mask] = torch.sigmoid(
        implicit_reward[desirable_mask] - kl_reference
    )
    values[~desirable_mask] = torch.sigmoid(
        kl_reference - implicit_reward[~desirable_mask]
    )

    # Weighted loss
    weights = torch.where(desirable_mask, lambda_d, lambda_u)
    return (weights * (1 - values)).mean()


def orpo_loss(
    policy_chosen_logps: torch.Tensor,  # per-token log probs, averaged
    policy_rejected_logps: torch.Tensor,
    chosen_length: torch.Tensor,  # sequence lengths
    rejected_length: torch.Tensor,
    lambda_or: float = 0.1,
) -> torch.Tensor:
    """ORPO loss combining SFT with odds ratio preference."""
    # SFT loss on chosen (negative log likelihood)
    sft_loss = -policy_chosen_logps.mean()

    # Compute log odds (approximation using average log probs)
    # log odds = log_prob - log(1 - exp(log_prob))
    # For numerical stability with small probabilities:
    def log_odds(log_probs):
        # Clamp to avoid numerical issues
        probs = torch.exp(log_probs.clamp(min=-100))
        return log_probs - torch.log1p(-probs.clamp(max=1 - 1e-7))

    chosen_log_odds = log_odds(policy_chosen_logps)
    rejected_log_odds = log_odds(policy_rejected_logps)

    # Odds ratio loss
    log_or = chosen_log_odds - rejected_log_odds
    or_loss = -F.logsigmoid(log_or).mean()

    return sft_loss + lambda_or * or_loss

Let's visualize how these losses behave differently:

In[13]:
Code
import numpy as np

# Create range of h_theta values (log-ratio differences)
h_theta = np.linspace(-2, 12, 200)
beta = 0.1

# DPO loss: -log σ(βh)
dpo = -np.log(1 / (1 + np.exp(-beta * h_theta)))

# IPO loss: (h - 1/2β)²
target = 1 / (2 * beta)
ipo = (h_theta - target) ** 2

# cDPO loss with ε=0.3
epsilon = 0.3
sigmoid_pos = 1 / (1 + np.exp(-beta * h_theta))
sigmoid_neg = 1 / (1 + np.exp(beta * h_theta))
cdpo = -(1 - epsilon) * np.log(sigmoid_pos + 1e-10) - epsilon * np.log(
    sigmoid_neg + 1e-10
)

# Normalize for visualization
dpo_norm = (dpo - dpo.min()) / (dpo.max() - dpo.min() + 1e-10)
ipo_norm = (ipo - ipo.min()) / (ipo.max() - ipo.min() + 1e-10)
cdpo_norm = (cdpo - cdpo.min()) / (cdpo.max() - cdpo.min() + 1e-10)
Out[14]:
Visualization
Line plot comparing DPO, IPO, and cDPO loss functions across log-ratio difference values.
Normalized loss functions for DPO, IPO, and cDPO across log-ratio difference values. DPO's loss decreases monotonically without bound, while IPO and cDPO exhibit stable minima that prevent unbounded reward gaps and stop optimization at principled targets.

Now let's examine the gradients:

In[15]:
Code
import numpy as np

# Compute gradients analytically
# DPO gradient: -β * σ(-βh)
dpo_grad = -beta * (1 / (1 + np.exp(beta * h_theta)))

# IPO gradient: 2(h - 1/2β)
ipo_grad = 2 * (h_theta - target)

# cDPO gradient: -β[(1-ε)σ(-βh) - ε·σ(βh)]
cdpo_grad = -beta * ((1 - epsilon) * sigmoid_neg - epsilon * sigmoid_pos)
Out[16]:
Visualization
Line plot comparing gradient magnitudes of DPO, IPO, and cDPO as functions of log-ratio difference.
Gradient magnitudes for DPO, IPO, and cDPO as functions of the log-ratio difference. DPO's gradient vanishes at high confidence values, while IPO and cDPO maintain corrective signals that drive the model toward their respective stable equilibrium points.

Comparing Alignment Methods

With multiple alignment approaches now available, selecting the right method depends on your specific constraints: data format, computational budget, annotation quality, and desired training dynamics. No single method dominates all scenarios, so understanding the trade-offs is needed for making an informed choice.

Decision Framework

The choice between alignment methods depends on several factors that you can evaluate before starting an experiment. Working through these questions systematically is more reliable than defaulting to DPO simply because it is the most familiar.

Data format availability:

  • Paired preferences (chosen vs. rejected for the same prompt): use DPO, IPO, cDPO, or ORPO.
  • Unpaired binary feedback (thumbs up/down on individual responses): use KTO.
  • Ranked lists with multiple responses per prompt: DPO can be adapted with pairwise comparisons extracted from the ranking.

Reference model constraints:

  • Can maintain frozen reference in memory: DPO, IPO, cDPO, KTO.
  • Memory-constrained single-model training: ORPO.

Annotation quality:

  • High-quality expert annotations (>90% consistency): standard DPO.
  • Moderate-quality annotations (70-90%): cDPO with appropriate ϵ\epsilon.
  • Noisy or automated labels (<70%): cDPO with higher ϵ\epsilon, or KTO.

Training stability requirements:

  • Need bounded optimization: IPO or cDPO.
  • Standard training sufficient: DPO.
Out[17]:
Visualization
Flowchart diagram with decision nodes asking about paired preferences, memory constraints, label noise, and training stability, leading to outcome nodes recommending KTO, ORPO, cDPO, DPO, or IPO.
Decision flowchart for selecting a DPO variant based on data format, memory constraints, annotation quality, and stability requirements. Following the path from top to bottom leads to the method best matched to your situation.

Empirical Performance Comparison

Published results show that no single method dominates across all benchmarks. General patterns from the literature reveal a fine-grained picture in which each method has conditions under which it performs best.

DPO remains a strong baseline. Its simplicity and well-understood behavior make it the default choice for many applications. Performance degradation typically occurs only with very long training or highly noisy data.

IPO shows advantages when training for many epochs or when the preference data has high confidence. The bounded optimization prevents the overconfident predictions that can emerge with extended DPO training. On clean benchmarks with clear preference signals, IPO often matches or exceeds DPO's performance while showing better calibration.

KTO reaches comparable performance to DPO when preference data is artificially unpaired from originally paired data. When working with naturally unpaired data, KTO materially outperforms naive approaches like randomly constructing pairs from unpaired labels, because those artificial pairs may reverse the actual preference relationship.

ORPO reduces computational overhead by 20-30% by eliminating the reference model. Quality results are competitive with DPO on standard benchmarks, though some studies report slightly lower performance on challenging alignment tasks where the stability provided by the reference model turns out to matter.

cDPO gives consistent improvements over standard DPO when annotation noise is known to be present. The gains are proportional to the actual noise level; cDPO shows little difference from DPO on clean data, which means the conservative treatment of labels does not hurt when labels are reliable.

Computational Costs

The computational characteristics differ materially across methods:

Computational requirements for DPO variants.
MethodMemoryForward PassesTraining Stages
DPO2x model2x per stepSFT, then DPO
IPO2x model2x per stepSFT, then IPO
cDPO2x model2x per stepSFT, then cDPO
KTO2x model2x per stepSFT, then KTO
ORPO1x model1x per stepSingle stage

ORPO's single-stage training can reduce wall-clock time by 40-50% compared to two-stage approaches, which makes it attractive for rapid iteration where each experiment needs to run quickly.

Training Dynamics

Let's simulate training dynamics for each method to see how they converge differently:

In[18]:
Code
import numpy as np


def simulate_training(method, n_steps=100, lr=0.1, beta=0.1, epsilon=0.1):
    """Simulate how h_theta evolves during training for different methods."""
    h = 0.0  # Start with equal preference
    history = [h]
    target = 1 / (2 * beta)

    for _ in range(n_steps):
        if method == "dpo":
            # DPO gradient: β * σ(-βh)
            grad = beta / (1 + np.exp(beta * h))
        elif method == "ipo":
            # IPO gradient: -2(h - target)
            grad = -2 * (h - target)
        elif method == "cdpo":
            # cDPO gradient
            sigma_pos = 1 / (1 + np.exp(-beta * h))
            sigma_neg = 1 - sigma_pos
            grad = beta * ((1 - epsilon) * sigma_neg - epsilon * sigma_pos)

        h = h + lr * grad
        history.append(h)

    return history


# Simulate all methods
dpo_history = simulate_training("dpo", n_steps=200, lr=3.0)
ipo_history = simulate_training("ipo", n_steps=200)
cdpo_history = simulate_training("cdpo", n_steps=200, lr=10.0)
Out[19]:
Visualization
Line plot showing h_theta evolution over training steps for DPO, IPO, and cDPO methods.
Simulated training dynamics for DPO, IPO, and cDPO. DPO's log-ratio difference grows indefinitely, while IPO converges to its target at 5.0 and cDPO converges to its finite equilibrium point, both shown as horizontal dotted lines.

Practical Guidelines

Based on the analysis above, here are practical recommendations for choosing and implementing DPO variants. These guidelines are intended as starting points; the right choices for your specific application will depend on empirical validation on held-out data.

Method Selection Checklist

When selecting an alignment method, work through these questions in order:

  1. What data do you have?

    • Paired preferences: DPO, IPO, cDPO, or ORPO.
    • Unpaired binary feedback: KTO.
  2. How much memory can you allocate?

    • Can fit two models: any method.
    • Limited to one model: ORPO.
  3. How clean are your labels?

    • High quality (>90% agreement): DPO or IPO.
    • Moderate quality (70-90%): cDPO with ϵ≈0.1\epsilon \approx 0.1-0.20.2.
    • Low quality (<70%): cDPO with ϵ≈0.2\epsilon \approx 0.2-0.30.3.
  4. How long will you train?

    • Short training (1-3 epochs): DPO.
    • Extended training: IPO or cDPO.

Hyperparameter Recommendations

Each method has specific hyperparameter considerations worth understanding before you begin.

DPO: The β\beta parameter typically works well in the range 0.1-0.5. Lower values allow more deviation from the reference; higher values keep the policy closer to the reference. Start with β=0.1\beta = 0.1 and adjust based on whether outputs are too conservative (increase β\beta) or too different from the base model (decrease β\beta).

IPO: Uses the same β\beta parameter, but interpretation differs. The target margin 12β\frac{1}{2\beta} means smaller β\beta creates larger targets. Start with β=0.1\beta = 0.1 (target 5.0) and adjust if convergence is too slow (increase β\beta) or produces weak preferences (decrease β\beta).

cDPO: Choose ϵ\epsilon based on estimated annotation noise. If unknown, start with ϵ=0.1\epsilon = 0.1 as a conservative default. The effective training signal strength is (1−2ϵ)(1 - 2\epsilon), so ϵ=0.3\epsilon = 0.3 reduces effective signal by 60%, requiring more data or longer training.

KTO: The loss aversion ratio λu/λd\lambda_u / \lambda_d is typically set to 1.0-2.0. Higher ratios emphasize avoiding bad outputs over creating good ones. The KL reference running average uses momentum 0.9-0.99. Higher momentum gives a more stable baseline but adapts more slowly to policy changes.

ORPO: The λor\lambda_{\text{or}} weight balances SFT and preference objectives. Values of 0.1-0.5 are typical. Too low ignores preferences; too high destabilizes SFT and may cause the model to sacrifice fluency for preference optimization.

Common Pitfalls

Several issues frequently arise when implementing DPO variants, and knowing them in advance can save significant debugging time.

Length bias is the most common and least obvious problem. All preference methods can develop biases toward response length if chosen responses are systematically longer or shorter than rejected ones. When annotators prefer more detailed responses, every variant will learn that length correlates with quality and may produce unnecessarily verbose outputs. Monitor average response lengths during training and consider normalizing log-probabilities by sequence length.

Mode collapse occurs when aggressive preference optimization causes the model to produce repetitive outputs. The KL penalty gives some protection, but it is insufficient when β\beta is very small. Monitoring diversity metrics like distinct n-grams alongside preference metrics helps detect this early.

Reward hacking manifests as the model learning to exploit artifacts in preference data rather than learning real quality signals. Annotators sometimes prefer responses that sound confident over responses that are accurate, or prefer formatting conventions over substance. Regular evaluation on held-out prompts with diverse evaluation criteria helps detect reward hacking before it becomes severe.

Training instability is a practical concern when combining small β\beta values with large learning rates. The coupled optimization in ORPO is particularly sensitive to learning rate choices. Starting with learning rates an order of magnitude smaller than SFT and increasing gradually is a reliable approach.

Limitations and Open Problems

The growth of DPO variants shows both the importance of alignment and the difficulty of getting it right under real-world conditions. Each variant addresses a real problem, but none fully resolves the basic tension between what alignment methods optimize and the behavior expected from deployed language models.

The most basic challenge persists across all variants: they optimize for human preferences as measured during data collection, which may differ from what humans want during deployment. Preferences are noisy, context-dependent, and subject to strategic behavior from both annotators and models. A model that perfectly optimizes collected preferences may still exhibit concerning behaviors on novel inputs, because the distribution of deployment prompts differs from the distribution used for training. This generalization gap motivates ongoing research into more reliable alignment approaches that go beyond fitting a fixed preference dataset.

The paired data assumption shared by DPO, IPO, cDPO, and ORPO creates additional problems that are easy to overlook. Constructing preference pairs requires that two responses exist for the same prompt, which either means generating multiple candidate responses during dataset construction (expensive) or collecting head-to-head comparisons from annotators (also expensive and difficult to scale). KTO partially addresses this, but it still requires a reference model and relies on the implicit reward as a surrogate for quality, which may not align with downstream task performance. The gap between proxy signals used during alignment and actual downstream metrics remains a core unsolved problem.

Computational efficiency gains from ORPO are real but come with theoretical trade-offs. Eliminating the reference model means there is no explicit constraint on how far the policy drifts from sensible behavior. The SFT loss gives implicit regularization, but this regularization is not backed by the same theoretical guarantees as the KL constraint. In practice, ORPO works well on standard benchmarks, but the conditions under which the implicit regularization fails are not yet fully understood, which makes it harder to debug when problems arise.

All of these methods also share a deeper limitation: they assume that the preference labels capture something real about alignment, but the relationship between pairwise preferences and broader safety or helpfulness is not straightforward. A model can learn to win every head-to-head comparison by generating responses that sound confident and fluent while containing subtle errors or harmful content that annotators miss on casual inspection. The field continues to develop evaluation methodologies that go beyond preference accuracy, but connecting those evaluations back to training objectives remains an open research challenge.

Summary

This chapter examined four influential DPO variants, each addressing specific limitations of the original algorithm while preserving the core insight that preference learning can bypass explicit reward modeling.

IPO reformulates preference learning as regression with a target margin 12β\frac{1}{2\beta}, preventing the unbounded preference growth that can occur with standard DPO. The squared loss creates self-correcting gradients that converge to a stable equilibrium, and the gradient changes sign when the model overshoots the target. This gives natural protection against overconfidence.

KTO lets learning from unpaired binary feedback by incorporating insights from prospect theory. The asymmetric value function and adaptive reference point allow training on thumbs-up and thumbs-down data without requiring direct comparisons between responses. The loss aversion weighting shows the empirical finding that humans are more sensitive to bad outcomes than equivalent good ones.

ORPO eliminates the reference model entirely by combining SFT with odds-ratio preference optimization. This reduces memory requirements by half and collapses two-stage training into a single pass. The coupling between the SFT loss and the odds ratio loss gives implicit regularization that substitutes for the KL constraint provided by the reference model in other methods.

cDPO handles annotation noise through label smoothing, modeling preferences probabilistically rather than deterministically. The smoothing parameter ϵ\epsilon should reflect actual annotation uncertainty and directly determines the finite optimal gap at which training stabilizes.

Method selection depends on data format, computational constraints, annotation quality, and training duration. No single method dominates across all scenarios. IPO and cDPO both produce bounded optimization with stable equilibria; they differ in how that equilibrium is derived. KTO handles unpaired data; ORPO saves memory. Understanding these trade-offs at the level of loss functions and gradients, rather than just benchmark numbers, equips you to make informed choices and to diagnose problems when experiments do not go as expected.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about DPO variants and their design principles.

DPO Variants

Question 1 of 80 of 8 completed
What is the target margin that IPO trains the model to achieve for the log-ratio difference?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026dpovariants, author = {Michael Brenndoerfer}, title = {DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment}, year = {2026}, url = {https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment. Retrieved from https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment
MLAAcademic
Michael Brenndoerfer. "DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment>.
CHICAGOAcademic
Michael Brenndoerfer. "DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment.
HARVARDAcademic
Michael Brenndoerfer (2026) 'DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment'. Available at: https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). DPO Variants: IPO, KTO, ORPO & cDPO for LLM Alignment. https://mbrenndoerfer.com/writing/dpo-variants-ipo-kto-orpo-cdpo-llm-alignment

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.