Part of Language AI Handbook
DPO variants include IPO, KTO, ORPO, and cDPO. Compare their objectives, data requirements, computational costs, and suitable alignment tasks.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
DPO Variants: IPO, KTO, ORPO, and cDPO
In the previous chapters, we derived Direct Preference Optimization from first principles and implemented it as a simpler alternative to RLHF. DPO was a breakthrough: it showed that you could align language models with human preferences using a simple classification-style loss, without ever explicitly training a reward model or running a reinforcement learning loop. The key derivation revealed that the optimal policy under the KL-constrained reward maximization objective could be expressed directly in terms of the policy's own log-probabilities, collapsing what had been a two-stage pipeline into a single supervised objective. That elegance made DPO enormously attractive.
Yet elegance often conceals hidden fragility. As practitioners began applying DPO at scale, they encountered a cluster of practical problems that the theoretical derivation had glossed over. The training loss continued to decrease even when models were creating clearly overconfident predictions. Systems trained on crowd-sourced preference data often behaved erratically, apparently because some annotators had labeled responses incorrectly. Teams working with limited GPU memory struggled to keep two copies of a large model in VRAM simultaneously. And researchers collecting feedback from deployed products found that most of their data came in the form of thumbs-up and thumbs-down ratings on individual responses, not as head-to-head comparisons that DPO's paired format requires. Each of these problems was real, each caused real engineering pain, and each motivated a distinct line of research.
This chapter explores four of the most influential responses to those problems. Identity Preference Optimization (IPO) addresses the unbounded reward growth problem by replacing DPO's monotone loss with a regression target, giving optimization a well-defined stopping point. Kahneman-Tversky Optimization (KTO) discards the paired data requirement entirely, drawing on behavioral economics to define a learning objective that works with simple binary labels. Odds Ratio Preference Optimization (ORPO) eliminates the reference model, combining supervised fine-tuning and preference learning into a single forward pass that requires only one copy of the model. Conservative DPO (cDPO) tackles annotation noise through label smoothing, acknowledging that human judgments are probabilistic rather than deterministic.
Think of these four variants as a family of tools, each sharpened for a different job. Understanding when to reach for each one requires understanding the problem each one solves and the mechanism it uses to solve it. We will work through the mathematics of all four variants carefully, compare their gradient behavior, implement them in code, and walk through a concrete numerical example that makes the differences tangible. By the end of this chapter, you will have the conceptual toolkit to select the right alignment method for your data characteristics and computational constraints, rather than defaulting to the original DPO when a better-matched alternative exists.
The DPO variants covered in this chapter appeared in rapid succession between 2023 and 2024. This shows how quickly the alignment field moved once DPO demonstrated that RL-free preference learning was viable. IPO (Azar et al., 2024) identified a theoretical gap in DPO's convergence guarantees. KTO (Ethayarajh et al., 2024) drew inspiration from the Nobel Prize-winning work of Kahneman and Tversky on human decision-making under uncertainty. This produces a method that mirrors how behavioral economists model value perception. ORPO (Hong et al., 2024) emerged partly from the practical constraints of researchers at institutions without access to large multi-GPU clusters. cDPO (Mitchell et al., 2023) applied a decades-old classification technique, label smoothing, to the preference learning setting. The fact that all four addressed different limitations of the same base algorithm within roughly twelve months illustrates both how fertile the space was and how quickly the community can respond when a tractable framework becomes available.
Identity Preference Optimization (IPO)
IPO emerged from a careful analysis of DPO's theoretical properties. Azar et al. (2024) identified a subtle but important issue: as DPO training progresses, the implicit reward gap between chosen and rejected responses can grow unboundedly, causing the policy to assign extreme probabilities that do not reflect the actual strength of human preferences. This section explains exactly why that happens, what consequences it produces, and how IPO's squared-error reformulation resolves the problem.
Think of DPO's loss as a one-sided ramp: no matter how strongly the model already prefers the chosen response, the ramp keeps tilting in the same direction, always pushing the model to make the chosen response even more probable. IPO replaces the ramp with a bowl: there is a specific level of preference the model should reach, and any deviation from that level, whether too weak or too strong, incurs a penalty. The bowl has a well-defined minimum, and once the model reaches it, training stops pushing.
The Overfitting Problem in DPO
To understand why IPO was developed, we must first examine a basic tension within DPO's optimization dynamics. Recall from our DPO derivation that the loss function is:
where:
- : the DPO loss function
- : the expectation over the dataset of preference pairs
- : the probability of response given prompt under the policy being trained
- : the probability under the frozen reference model
- : the chosen (winner) and rejected (loser) responses
- : the temperature parameter controlling the strength of the KL divergence penalty
- : the sigmoid function
To appreciate the overfitting problem, consider what this loss function is asking the model to do. The expression inside the sigmoid is the difference between two log-ratios: how much more likely the policy makes the chosen response compared to the reference, minus the same quantity for the rejected response. The negative log-sigmoid of this quantity becomes the loss. When training minimizes this loss, it maximizes what is inside the log-sigmoid.
The sigmoid function approaches 1 as its argument grows large. To minimize this loss, the model is incentivized to make the log-ratio difference as large as possible. In the limit, this pushes the model toward assigning probability approaching 1 to chosen responses and probability approaching 0 to rejected ones. There is no natural stopping point in this formulation: the gradient always points toward making the gap larger, even when the gap is already enormous.
This behavior is problematic for several reasons. First, human preferences are not absolute. A response labeled "chosen" is not infinitely better than the "rejected" alternative; it is merely preferred in that comparison. Two responses might differ only slightly in quality, yet the labeling treats one as definitively better. When the model learns to assign extreme probabilities based on these labels, it develops a false sense of certainty that does not match the underlying reality of human judgment. Second, training instability compounds as log probabilities approach for rejected responses, causing gradients to become numerically unreliable. Third, and most subtly, the model may generalize poorly: it in effect memorizes that specific responses are "infinitely good" or "infinitely bad" rather than learning the generalizable patterns that distinguish better responses from worse ones.

IPO's Squared Error Objective
IPO addresses the overfitting problem by reformulating preference learning as a regression problem with a specific target. The reason is that instead of letting the reward gap to grow without bound, we should specify how large we want that gap to be and then train the model to reach exactly that target. This turns an unbounded optimization problem into a bounded one with a clear convergence criterion.
The main point is: bounded optimization requires a target. DPO has no target for the log-ratio difference; it simply wants that difference to be as large as possible. IPO sets the target at and penalizes any deviation from it, whether the model has not learned a strong enough preference or has learned an excessively strong one. This single design choice fundamentally changes the optimization landscape.
Instead of pushing the reward gap to infinity, IPO targets a fixed margin:
where:
- : the IPO loss function
- : the expectation over the dataset of preference pairs
- : the target margin derived from the regularization strength
- : the policy and reference models
The key innovation is the squared loss with a specific target. Rather than rewarding the model for making the gap as large as possible, IPO rewards the model for making the gap equal to a specific value. Any deviation from this target, whether too small or too large, incurs a penalty. This is the essence of regression: we have a target value and we minimize the squared distance to that target.
To analyze these dynamics more precisely, let's define the log-ratio difference as:
where:
- : the difference in log-probability ratios between chosen and rejected responses
- : the "implicit reward" for a specific response
This quantity captures the essence of what preference optimization reaches. It measures how much the policy has learned to prefer the chosen response over the rejected one, relative to what the reference model would predict. A positive value means the policy has shifted probability mass toward the chosen response; a larger positive value means a stronger learned preference.
With this notation, the IPO loss becomes elegantly simple:
where:
- : the IPO loss function
- : the expectation over the dataset
- : the log-ratio difference computed by the model
- : the specific target value that the difference should converge to
This formulation makes the regression nature of IPO crystal clear. We are asking the model to make equal to for every preference pair in the dataset. The squared loss penalizes deviations in either direction: if the model has not learned a strong enough preference (), the loss is positive and gradients push toward a larger gap; if the model has learned too strong a preference (), the loss is also positive and gradients push toward a smaller gap.
The gradient with respect to reveals this self-correcting behavior:
where:
- : the gradient of the loss with respect to the log-ratio difference
- : the error term (distance from target) that drives the update
This gradient is zero precisely when , creating a stable equilibrium. Once the log-ratio difference reaches the target margin, there is no further pressure to increase it. The equilibrium is stable and attractive: regardless of where training starts, the gradients always point toward the target, and the strength of the gradient is proportional to the distance from the target.
Interpreting the Target Margin
The target has a principled interpretation that connects IPO back to the theoretical foundations of preference learning. The parameter controls the strength of the KL constraint. A larger means weaker regularization, letting larger deviations from the reference policy. With IPO's target:
- When is large (weak KL penalty), the target margin is small.
- When is small (strong KL penalty), the target margin is large.
This might seem counterintuitive until you consider a budget analogy. If you have a strict overall spending limit, you can afford to splurge on individual items because you are being careful everywhere else. Conversely, if your overall limit is loose, you must be more conservative on each purchase to avoid overshooting. Stronger KL constraints (small ) permit larger per-example margins because the overall deviation is more tightly controlled.
The target margin also has an interpretation in terms of the Bradley-Terry model that underlies preference learning. In that model, the probability that response is preferred to depends on the difference in their rewards. The target corresponds to a specific preference probability representing a moderate preference rather than absolute certainty. This aligns with the reality that human annotations express preferences, not certainties. An annotator who chooses response A over response B is not asserting that A is infinitely better; they are expressing a judgment that may be uncertain, context-dependent, and occasionally reversed.

Gradient Comparison
Different loss functions produce different gradient behaviors, and those differences explain why methods perform differently in practice. The gradient determines how the model updates its parameters at each step, so differences in gradient behavior translate directly into differences in training dynamics. Understanding these gradients is the key to predicting how a method will behave when you deploy it on your data.
Think of the gradient as the engine that drives training. DPO's engine runs at full throttle when the model is uncertain and idles as the model becomes confident, but it never turns off. IPO's engine runs proportionally to how far the model is from the target and cuts off completely once the target is reached.
The DPO gradient magnitude is:
where:
- : the temperature parameter
- : the sigmoid of the negative scaled log-ratio, acting as a weighting factor
This sigmoid term means DPO's gradient is largest when is near zero (uncertain predictions) and vanishes as (confident predictions). The intuition is that DPO pushes hardest when the model is uncertain about which response is preferred and pushes more gently when the model is already confident. While vanishing gradients prevent infinite growth in principle, they also mean learning slows sharply once the model becomes confident. This creates a problematic dynamic: the model can still drift toward extreme probabilities because the gradients, though small, remain consistently positive, never reversing direction.
The IPO gradient magnitude is:
where:
- : the absolute distance from the target margin
IPO's gradient magnitude is proportional to distance from the target. If the model overshoots the target (), the gradient reverses direction and pushes back. This self-correcting behavior is entirely absent in DPO. The gradient grows stronger as the model strays further from the target, which means that even large deviations are corrected rather than allowed to persist indefinitely.
Kahneman-Tversky Optimization (KTO)
While DPO and IPO improve upon RLHF's computational complexity, they share a basic data requirement: paired preferences. Each training example must contain a prompt together with both a chosen and rejected response. This pairing constraint creates practical challenges that are easy to underestimate if you have only worked with carefully curated benchmark datasets.
Think of the difference between a controlled taste test and a product review. In a taste test, each participant explicitly compares two options side by side and says which they prefer. Product reviews, by contrast, arrive as standalone ratings: a customer gives three stars or five stars to a single item, with no explicit comparison to an alternative. Most of the feedback you collect from deployed language models looks much more like product reviews than taste tests, and forcing it into paired-comparison format requires either discarding most of your data or constructing artificial comparisons that may not reflect real user preferences.
KTO, introduced by Ethayarajh et al. (2024), eliminates the paired data requirement by designing an objective that works directly with unpaired binary feedback. The result is a method that can absorb the full breadth of available user signal without the preprocessing overhead of converting it into an unnatural format.
The Unpaired Feedback Problem
Real-world human feedback often comes in unpaired form, and this mismatch between how feedback is collected and how preference optimization algorithms expect data creates significant friction in practical alignment pipelines. Consider the forms that user feedback naturally takes:
- Binary ratings: users click thumbs up or thumbs down on individual responses, with no stated alternative.
- Flagging systems: users report problematic outputs without giving a corrected version.
- Implicit signals: engagement metrics, session lengths, and follow-up question rates indicate whether a response was helpful, but never identify a counterfactual.
Converting this abundant unpaired feedback into DPO's paired format requires either discarding data or artificially constructing pairs. Discarding wastes useful signal. Constructing pairs introduces artifacts: if you pair a thumbs-up response to a randomly selected thumbs-down response, the pair may not reflect a real comparison because the two responses came from different conversations with different prompts and different contexts.
Consider a scenario where you have 10,000 thumbs-up ratings and 5,000 thumbs-down ratings, but these ratings come from different conversations. To use DPO, you would need to either match these into pairs somehow, losing most of your data, or generate new responses to create artificial comparisons. KTO avoids this entirely by treating each labeled response as an independent training example.
Inspiration from Prospect Theory
KTO draws inspiration from Kahneman and Tversky's prospect theory, which describes how humans make decisions under uncertainty. This connection to behavioral economics is more than a naming convention: it gives principled guidance for how to weight different types of feedback. Two key insights from behavioral economics directly inform KTO's design.
Reference dependence is the first insight. People evaluate outcomes relative to a reference point, not in absolute terms. A gain of $100 feels different depending on whether you expected $0 or $200. For alignment, this suggests that the quality of a model response should be measured relative to some baseline expectation, not in absolute terms. KTO implements this by measuring implicit rewards relative to the expected KL divergence across the training distribution, creating a dynamic baseline that adapts as training progresses.
Loss aversion is the second insight. Losses loom larger than equivalent gains. Losing $100 feels psychologically worse than gaining $100 feels good, and behavioral experiments consistently estimate that losses are weighted roughly 2x more heavily than gains in human value assessments. For alignment, this suggests that suppressing bad outputs may be more important than promoting good ones. You might forgive a bland response, but you remember a harmful or embarrassingly incorrect one. KTO incorporates this asymmetry by letting different loss weights for desirable and undesirable examples.
KTO incorporates both principles into its loss function, treating "desirable" and "undesirable" responses asymmetrically in a way that mirrors how humans experience quality differences.

The KTO Loss Function
The construction of KTO's loss function proceeds in several stages, each motivated by the behavioral economics principles described above. We begin by defining the implicit reward, which is the raw signal that KTO turns into a learning objective.
For a response to prompt with binary label , KTO defines:
where:
- : the implicit reward assigned to response given prompt
- : the probability of the response under the current policy
- : the probability of the response under the reference model
This is the same implicit reward used in DPO. It measures how much more likely the current policy makes this response compared to the reference. A positive implicit reward means the policy has learned to favor this response; a negative implicit reward means the policy has learned to disfavor it. The important property is that this quantity is well-defined for a single response, without needing a comparison partner.
KTO then defines a reference point that implements the reference dependence principle from prospect theory:
where:
- : the reference point (baseline) for evaluation
- : the policy and reference models
- : the training dataset distribution
- : the expectation over prompts sampled from
- : the Kullback-Leibler divergence measuring the drift of the policy from the reference
The reference point is the expected KL divergence between policy and reference across the training distribution. In practice, this is estimated from a running average during training. The reference point defines what counts as "above average" versus "below average" performance. Rather than using an arbitrary fixed threshold, KTO adapts the reference point to the current state of training, creating a dynamic baseline that evolves as the model improves.
With the implicit reward and reference point defined, KTO constructs a value function that differs based on whether the response is desirable:
where:
- : the value assigned to the response, bounded between 0 and 1
- : the sigmoid function
- : the label showing if the response is desirable or undesirable
- : the shifted implicit reward used for desirable examples
This asymmetric definition is the mathematical implementation of reference dependence. For desirable responses, we ask whether the implicit reward exceeds the reference point. For undesirable responses, we ask whether the implicit reward falls below the reference point. The sigmoid function squashes these comparisons into the range (0, 1), creating a smooth value measure that is easy to differentiate.
The loss weights desirable and undesirable examples differently, implementing the loss aversion principle:
where:
- : the KTO loss function
- : the expectation over the dataset of labeled examples
- : the weighting factor specific to the label type
- : the value computed by the value function
The weighting factors are and , typically with to implement loss aversion. By setting , we tell the model that failing to suppress a bad response is worse than failing to promote a good response. This asymmetry shows the empirical finding that people are more bothered by failures than impressed by equivalent successes.
Understanding the Value Function
The asymmetric value function encodes prospect theory's core insights, and understanding its behavior illuminates why KTO works. The function turns implicit rewards into values differently depending on the label. This creates distinct learning signals for positive and negative feedback.
For desirable responses, the sigmoid argument is . The model is rewarded with a low loss when the implicit reward exceeds the reference point. Intuitively, the model should push desirable responses to have higher implicit rewards than average. When , the sigmoid output is greater than 0.5, meaning the value is high and the loss is low.
For undesirable responses, the sigmoid argument is . The model is rewarded when the implicit reward falls below the reference point. When , the sigmoid output is greater than 0.5, meaning the value is high and the loss is low. The model learns to associate undesirable responses with below-average implicit rewards.
The reference point is the dividing line between "gains" and "losses" in the prospect theory sense. This relative framing means KTO does not need paired comparisons; it learns to push good responses up and bad responses down relative to an adaptive baseline. The elegance of this approach is that it converts an inherently comparative problem into a classification problem, which is exactly the form that unpaired feedback gives.


Practical Implementation Details
KTO requires tracking the reference point during training. This running average must be estimated from the current batch of data and updated incrementally as training proceeds.
import torch
def kto_loss(
policy_logps: torch.Tensor, # log π_θ(y|x) for batch
reference_logps: torch.Tensor, # log π_ref(y|x) for batch
is_desirable: torch.Tensor, # binary mask: 1 for desirable, 0 for undesirable
kl_reference: float, # z_0: running average KL
beta: float = 0.1,
lambda_d: float = 1.0, # weight for desirable
lambda_u: float = 1.0, # weight for undesirable (often > lambda_d)
) -> torch.Tensor:
"""
Compute KTO loss for a batch of (potentially unpaired) examples.
"""
# Implicit reward: log ratio
implicit_reward = policy_logps - reference_logps
# Value function differs by desirability
desirable_mask = is_desirable.bool()
values = torch.zeros_like(policy_logps)
values[desirable_mask] = torch.sigmoid(
beta * implicit_reward[desirable_mask] - kl_reference
)
values[~desirable_mask] = torch.sigmoid(
kl_reference - beta * implicit_reward[~desirable_mask]
)
# Weighted loss
weights = torch.where(desirable_mask, lambda_d, lambda_u)
loss = weights * (1 - values)
return loss.mean()The running average KL is typically computed as:
def update_kl_reference(
kl_reference: float, batch_kl: float, momentum: float = 0.99
):
"""Update running average of KL divergence."""
return momentum * kl_reference + (1 - momentum) * batch_klThe momentum parameter controls how quickly the reference point adapts to changes in the policy. A high momentum (close to 1) means the reference point changes slowly. This gives a stable baseline. A low momentum means it tracks the current policy more aggressively, which can be destabilizing if the policy is changing rapidly at the start of training.
KTO's Advantages
KTO offers several practical benefits for real-world alignment pipelines:
- Data efficiency: Uses all available binary feedback without discarding unpaired examples.
- Natural data format: Matches how most user feedback is collected in deployed products.
- Principled asymmetry: The loss-aversion weighting shows empirical findings about human judgment rather than an arbitrary design choice.
- Stable training: The reference point gives a grounding mechanism similar to IPO's target, preventing unbounded optimization.
Odds Ratio Preference Optimization (ORPO)
ORPO takes a more radical departure from the DPO framework by eliminating the reference model entirely. Introduced by Hong et al. (2024), ORPO combines supervised fine-tuning with preference optimization into a single training objective, compressing what had been a two-stage pipeline into a single pass.
Think of ORPO as designing a building that gives its own structural support rather than leaning against an adjacent building. DPO leans against the reference model: it measures every change relative to a frozen copy of the initial policy. ORPO's odds formulation gives internal stability by coupling the SFT loss and the odds ratio loss in a way that creates natural constraints on how far the policy can drift.
Motivation: The Reference Model Burden
All previous methods, including RLHF, DPO, IPO, and KTO, require maintaining a reference policy . This creates practical complications that matter most for researchers and engineers working with large models under tight resource constraints.
Memory overhead is the most immediate problem. Two copies of the model must reside in GPU memory simultaneously, or the reference model must be frequently loaded and unloaded from disk. For a 70-billion-parameter model, the reference copy alone requires roughly 140 GB of VRAM in BF16 precision, before accounting for gradients, optimizer states, and activations. This pushes many alignment experiments beyond the reach of single-node setups.
Two-stage training adds pipeline complexity. Models typically undergo supervised fine-tuning first to create a sensible reference policy, then preference training relative to that reference. Every bug, hyperparameter choice, and data quality issue must be diagnosed across two separate training runs rather than one.
Computational cost compounds the memory pressure. Every forward pass during preference training requires evaluating both policies: the current trainable policy and the frozen reference. This doubles the compute cost of every gradient step relative to supervised fine-tuning alone.
ORPO asks whether we can eliminate the reference model while still preventing the policy from drifting arbitrarily far from creating sensible outputs. The answer turns out to be yes, provided we design the objective carefully.
The Odds Ratio Approach
Instead of comparing log probabilities to a reference, ORPO uses the odds ratio between chosen and rejected responses within the current policy itself. This shift is significant. Instead of measuring change from a fixed reference, ORPO measures the relative likelihood of responses within the current policy. The odds of generating response given prompt are:
where:
- : the odds of generating response under policy
- : the probability of the response
The odds representation has a natural interpretation: it tells us how likely the response is compared to everything else. If the odds are 2:1, the response is twice as likely as all alternatives combined. The odds formulation is particularly useful because ratios of odds have clean mathematical properties that are difficult to reach with plain probability ratios.
For a language model, the probability of a specific sequence becomes numerically insignificant as length increases, making direct odds calculation unstable. Consider that even a moderately long sequence might have a probability of or smaller, which would make both the numerator and denominator of the odds calculation problematically small. ORPO instead defines the sequence-level odds as the geometric mean of the token-level odds:
where:
- : the input prompt
- : the length of the response in tokens
- : the token at step
- : the sequence of tokens preceding step (the context)
- : the probability of the next token given the prompt and previous tokens
- : the exponentiation to convert average log-odds back to the odds scale
This formulation works at the token level, where probabilities are large enough to be numerically stable, and then aggregates these token-level odds into a sequence-level measure. The averaging by sequence length ensures that longer sequences are not automatically penalized, since we take a geometric mean rather than a product.
The ratio of odds between chosen () and rejected () responses becomes:
where:
- : the odds ratio between the chosen and rejected responses
- : the chosen and rejected responses
This odds ratio captures the relative preference of the current policy for the chosen response over the rejected one. An odds ratio greater than 1 means the policy favors the chosen response; we want to train the policy to increase this ratio.
The ORPO Loss Function
ORPO combines two components that work together to give both language modeling signal and preference signal. This combination is what allows ORPO to eliminate the separate SFT stage and the reference model.
The first component is the supervised fine-tuning loss on the chosen response:
where:
- : the supervised fine-tuning loss component
- : the expectation over prompts and chosen responses
- : the likelihood of the chosen response
This component serves two purposes: it teaches the model to generate fluent, coherent text (the standard language modeling objective), and it gives an anchor that prevents the model from drifting too far from creating sensible outputs. By training on the chosen responses, the model learns what good outputs look like and maintains a probability distribution that assigns reasonable mass to those outputs.
The second component is the odds ratio loss that increases the relative odds of chosen over rejected:
where:
- : the odds ratio loss component
- : the expectation over the dataset of preference pairs
- : the chosen and rejected responses
- : the log odds ratio (log of the ratio of odds)
- : the sigmoid function
This component implements the preference learning objective. By maximizing the log-sigmoid of the log odds ratio, we push the model to make the chosen response relatively more likely than the rejected one. The structure is similar to DPO's loss, but operates on odds ratios within a single policy rather than log-probability ratios between two policies.
The combined ORPO objective is:
where:
- : the total ORPO loss
- : the coefficient weighting the odds ratio loss against the SFT loss
The hyperparameter controls the trade-off between learning to generate good responses and learning to distinguish good responses from bad ones. Too small a and the model ignores preferences; too large and the model may sacrifice fluency for preference optimization.
Why Odds Ratios Work
The odds ratio formulation gives implicit regularization without an explicit reference model. To understand why this works, consider what happens during training. The SFT loss pulls the model toward generating the chosen response. The OR loss pushes chosen odds higher relative to rejected odds. Both losses operate on the same policy, creating a coupled optimization.
The key insight is that increasing odds for chosen responses while decreasing odds for rejected responses automatically constrains how much the model can deviate from generating coherent text. If the model tried to maximize the odds ratio by assigning near-zero probability to rejected responses, the SFT loss on chosen responses would suffer because probability mass must be conserved. The model cannot simply declare everything "bad"; it must maintain a coherent probability distribution over all possible responses.
This coupling between the two loss components creates an implicit regularization effect. The SFT loss ensures the model keeps generating reasonable text, while the OR loss ensures it prefers better text to worse text. Neither loss alone would reach both objectives, but together they give a balanced training signal without requiring a separate frozen reference model.
Comparing Log-Ratio and Odds-Ratio Objectives
Comparing what DPO and ORPO optimize reveals how they prevent policy drift.
DPO optimizes:
where:
- : the policy and reference model probabilities
- : the chosen and rejected responses
ORPO optimizes:
where:
- : the log-odds of a response under the current policy
- : the chosen and rejected responses
DPO measures how much the policy's preference for over has changed relative to the reference. ORPO measures the absolute odds ratio within the current policy. The reference model in DPO is an external anchor. ORPO's SFT component and odds formulation give an alternative form of internal anchoring. The two approaches reach similar regularization effects through fundamentally different mechanisms.


Conservative DPO (cDPO)
The variants discussed so far address algorithmic limitations. cDPO, introduced by Mitchell et al. (2023), addresses a data quality issue: label noise in preference annotations. While IPO and ORPO assume that the training labels are reliable (even if the optimization dynamics are problematic), cDPO starts from the recognition that human annotations are inherently uncertain and builds that uncertainty directly into the loss function.
Think of cDPO as the difference between a GPS that reports its location with false precision and one that reports a confidence interval. Standard DPO says "response A is better than response B, and that fact is certain." cDPO says "response A is probably better than response B, with probability ." The probabilistic framing leads to a fundamentally different loss landscape with a bounded optimal preference gap rather than an infinite one.
Label Noise in Preference Data
Human preference annotations are inherently noisy, and this noise can arise from several sources. Annotators may disagree with each other on the same comparison, make mistakes due to fatigue or inattention, apply inconsistent criteria across examples, or be influenced by surface features such as formatting and length rather than actual quality differences.
Studies of inter-annotator agreement on preference tasks typically show agreement rates of 70 to 80 percent, meaning 20 to 30 percent of labels may be "wrong" in the sense that a different annotator would have labeled them oppositely. Even gold-standard human annotations on tasks like toxicity detection or factual accuracy show substantial disagreement when the same examples are labeled by different annotators. The problem is not that annotators are careless; preferences are subjective, and the ranking of two similar responses is often not clear-cut.
Standard DPO treats all preference labels as ground truth. If the training data says , the model learns to prefer with full confidence. When this label is wrong, the model learns an incorrect preference, and because DPO's unbounded optimization pushes that preference to an extreme, even a small percentage of mislabeled examples can cause significant downstream problems. The model in effect memorizes the noise in the data and tries to push it to infinity.
Label Smoothing for Preferences
cDPO applies label smoothing to preference labels. The idea behind label smoothing is familiar from classification: instead of training on hard targets (0 or 1), we train on soft targets that acknowledge uncertainty. For preferences, instead of treating each comparison as a certainty, we model the possibility that the annotation might be wrong.
Instead of treating preferences as deterministic (), cDPO models uncertainty:
where:
- : the label smoothing parameter, , representing uncertainty about annotation correctness
The parameter is our belief about how often the annotations are wrong. If , we trust all annotations completely, recovering standard DPO. If , we believe that 30 percent of the time the annotators got it backwards and the rejected response was better.
The cDPO loss modifies the Bradley-Terry model probability accordingly:
where:
- : the cDPO loss function
- : the expectation over the dataset
- : the log-ratio difference between chosen and rejected responses
- : the label smoothing parameter
- : the sigmoid function
- : the temperature parameter
This loss function has two terms. The first term, weighted by , is the standard DPO loss: it rewards the model for preferring the labeled chosen response. The second term, weighted by , is the opposite: it rewards the model for preferring the labeled rejected response. The combination is the expected loss under our belief about annotation accuracy. We train the model to do well on average across both the case where the annotation is correct and the case where it is wrong.
Using the identity , we can rewrite this loss to explicitly show the probability mass assigned to the rejected response:
where:
- : the probability assigned to the rejected response (equivalent to )
- : the weighting factor for the "incorrect" preference direction
This form makes the label smoothing interpretation particularly clear: we are mixing the standard DPO loss with its inverse, with the mixing coefficient representing the probability of annotation error.
Effect on Optimization
Label smoothing fundamentally changes the loss landscape. The gradient of cDPO with respect to is:
where:
- : the gradient of the loss with respect to the log-ratio difference
This gradient has two competing terms. The first term pushes toward larger (stronger preference for chosen). The second term pushes toward smaller (acknowledging that the rejected response might be better). These competing forces create a balance point where the gradient is zero.
Setting the gradient to zero allows us to solve for the optimal log-ratio difference :
where:
- : the optimal value of the log-ratio difference
- : the label smoothing parameter
This means cDPO has a finite optimal gap (similar to IPO), which depends on . With standard DPO (), the optimal gap is infinite. With cDPO (), the model converges to a bounded preference strength. The relationship between and the optimal gap is intuitive: higher annotation uncertainty (larger ) leads to a smaller optimal gap, because the model should not be confident when the labels might be wrong.


Choosing the Smoothing Parameter
The smoothing parameter should reflect actual annotation noise. When inter-annotator agreement data is available, can be estimated directly as the disagreement rate. When that data is unavailable, practical guidelines based on dataset characteristics give reasonable starting points:
- : Assumes 90% annotation accuracy, appropriate for high-quality curated datasets with expert annotators.
- : Assumes 80% accuracy, appropriate for crowd-sourced annotations where annotators receive limited instructions.
- : Assumes 70% accuracy, appropriate for noisy or automated annotations generated by another model.
The effective training signal strength scales with . At , the signal is 80% as strong as standard DPO. At , it is only 40% as strong. This means that high smoothing values require more training data or more epochs to reach the same degree of preference learning, creating a practical trade-off between robustness to noise and data efficiency.
Worked Example: Computing Losses Step by Step
To make the differences between these methods concrete, let's walk through a specific numerical example using a single preference pair. This will illustrate how each loss function responds to the same underlying data and why their behaviors diverge.
Suppose we are training a language model to answer questions more helpfully. For the prompt "What is the boiling point of water?" we have two candidate responses:
- (chosen): "Water boils at 100 degrees Celsius (212 degrees Fahrenheit) at standard atmospheric pressure."
- (rejected): "Water boils at a really high temperature."
Our reference model assigns log-probabilities:
- (moderately likely under the base model)
- (also reasonably likely, since it is short and plausible)
After some training steps, our current policy assigns:
- (the policy has learned to prefer the specific response)
- (the policy has learned to penalize the vague response)
Let's set .
Step 1: Compute the implicit rewards.
The implicit reward for each response is the log-ratio relative to the reference:
The chosen response has a positive implicit reward (the policy has moved toward it relative to the reference), while the rejected response has a negative implicit reward (the policy has moved away from it).
Step 2: Compute .
The log-ratio difference is:
Step 3: Compute the DPO loss.
This loss is still substantial. DPO would continue pushing higher, even though the model already clearly prefers the chosen response.
Step 4: Compute the IPO loss with , giving target .
IPO reports a large loss because the current is far below the target of 5.0. The gradient is:
This gradient is negative, meaning IPO is pushing upward toward the target, just as DPO does in this case. If instead , the IPO gradient would be , pushing back down. DPO would never push downward.
Step 5: Compute the cDPO loss with .
The cDPO loss is slightly higher than DPO here (0.635 vs. 0.594). More importantly, the equilibrium for cDPO with and is:
This is a finite target, far smaller than DPO's implicit infinite target.
Step 6: Compute the KTO value for the chosen response, assuming a reference point .
The KTO loss contribution for this desirable example (with ) is:
Since the implicit reward just barely exceeds the reference point, the model has done only slightly better than average for this example. The loss is close to 0.5, the maximum possible, suggesting the model should continue improving this response.
This worked example illustrates a key pattern: DPO, IPO, and cDPO all respond to the same value but with different equilibria, while KTO operates on individual implicit rewards rather than their difference.
Implementation Comparison
With the theory and intuition in place, let's implement all four variants to see the differences in practice. Clear implementations help reveal the structural similarities and differences that the mathematics sometimes obscures.
import torch
import torch.nn.functional as F
def compute_log_ratios(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
reference_chosen_logps: torch.Tensor,
reference_rejected_logps: torch.Tensor,
) -> torch.Tensor:
"""Compute the log-ratio difference h_θ used by DPO, IPO, and cDPO."""
chosen_ratio = policy_chosen_logps - reference_chosen_logps
rejected_ratio = policy_rejected_logps - reference_rejected_logps
return chosen_ratio - rejected_ratio
def dpo_loss(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
reference_chosen_logps: torch.Tensor,
reference_rejected_logps: torch.Tensor,
beta: float = 0.1,
) -> torch.Tensor:
"""Standard DPO loss."""
h_theta = compute_log_ratios(
policy_chosen_logps,
policy_rejected_logps,
reference_chosen_logps,
reference_rejected_logps,
)
return -F.logsigmoid(beta * h_theta).mean()
def ipo_loss(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
reference_chosen_logps: torch.Tensor,
reference_rejected_logps: torch.Tensor,
beta: float = 0.1,
) -> torch.Tensor:
"""IPO loss with squared error objective."""
h_theta = compute_log_ratios(
policy_chosen_logps,
policy_rejected_logps,
reference_chosen_logps,
reference_rejected_logps,
)
target = 1.0 / (2.0 * beta)
return ((h_theta - target) ** 2).mean()
def cdpo_loss(
policy_chosen_logps: torch.Tensor,
policy_rejected_logps: torch.Tensor,
reference_chosen_logps: torch.Tensor,
reference_rejected_logps: torch.Tensor,
beta: float = 0.1,
epsilon: float = 0.1, # label smoothing
) -> torch.Tensor:
"""Conservative DPO with label smoothing."""
h_theta = compute_log_ratios(
policy_chosen_logps,
policy_rejected_logps,
reference_chosen_logps,
reference_rejected_logps,
)
# Smoothed loss: (1-ε)log σ(βh) + ε log σ(-βh)
loss = -(1 - epsilon) * F.logsigmoid(
beta * h_theta
) - epsilon * F.logsigmoid(-beta * h_theta)
return loss.mean()Now let's implement KTO and ORPO, which have different data requirements:
import torch
import torch.nn.functional as F
def kto_loss(
policy_logps: torch.Tensor, # log probs for all examples
reference_logps: torch.Tensor, # reference log probs
is_desirable: torch.Tensor, # binary: 1=good, 0=bad
kl_reference: float, # running average KL
beta: float = 0.1,
lambda_d: float = 1.0,
lambda_u: float = 1.0,
) -> torch.Tensor:
"""KTO loss for unpaired binary feedback."""
implicit_reward = beta * (policy_logps - reference_logps)
# Different value functions for desirable vs undesirable
desirable_mask = is_desirable.bool()
values = torch.zeros_like(policy_logps)
values[desirable_mask] = torch.sigmoid(
implicit_reward[desirable_mask] - kl_reference
)
values[~desirable_mask] = torch.sigmoid(
kl_reference - implicit_reward[~desirable_mask]
)
# Weighted loss
weights = torch.where(desirable_mask, lambda_d, lambda_u)
return (weights * (1 - values)).mean()
def orpo_loss(
policy_chosen_logps: torch.Tensor, # per-token log probs, averaged
policy_rejected_logps: torch.Tensor,
chosen_length: torch.Tensor, # sequence lengths
rejected_length: torch.Tensor,
lambda_or: float = 0.1,
) -> torch.Tensor:
"""ORPO loss combining SFT with odds ratio preference."""
# SFT loss on chosen (negative log likelihood)
sft_loss = -policy_chosen_logps.mean()
# Compute log odds (approximation using average log probs)
# log odds = log_prob - log(1 - exp(log_prob))
# For numerical stability with small probabilities:
def log_odds(log_probs):
# Clamp to avoid numerical issues
probs = torch.exp(log_probs.clamp(min=-100))
return log_probs - torch.log1p(-probs.clamp(max=1 - 1e-7))
chosen_log_odds = log_odds(policy_chosen_logps)
rejected_log_odds = log_odds(policy_rejected_logps)
# Odds ratio loss
log_or = chosen_log_odds - rejected_log_odds
or_loss = -F.logsigmoid(log_or).mean()
return sft_loss + lambda_or * or_lossLet's visualize how these losses behave differently:
import numpy as np
# Create range of h_theta values (log-ratio differences)
h_theta = np.linspace(-2, 12, 200)
beta = 0.1
# DPO loss: -log σ(βh)
dpo = -np.log(1 / (1 + np.exp(-beta * h_theta)))
# IPO loss: (h - 1/2β)²
target = 1 / (2 * beta)
ipo = (h_theta - target) ** 2
# cDPO loss with ε=0.3
epsilon = 0.3
sigmoid_pos = 1 / (1 + np.exp(-beta * h_theta))
sigmoid_neg = 1 / (1 + np.exp(beta * h_theta))
cdpo = -(1 - epsilon) * np.log(sigmoid_pos + 1e-10) - epsilon * np.log(
sigmoid_neg + 1e-10
)
# Normalize for visualization
dpo_norm = (dpo - dpo.min()) / (dpo.max() - dpo.min() + 1e-10)
ipo_norm = (ipo - ipo.min()) / (ipo.max() - ipo.min() + 1e-10)
cdpo_norm = (cdpo - cdpo.min()) / (cdpo.max() - cdpo.min() + 1e-10)
Now let's examine the gradients:
import numpy as np
# Compute gradients analytically
# DPO gradient: -β * σ(-βh)
dpo_grad = -beta * (1 / (1 + np.exp(beta * h_theta)))
# IPO gradient: 2(h - 1/2β)
ipo_grad = 2 * (h_theta - target)
# cDPO gradient: -β[(1-ε)σ(-βh) - ε·σ(βh)]
cdpo_grad = -beta * ((1 - epsilon) * sigmoid_neg - epsilon * sigmoid_pos)
Comparing Alignment Methods
With multiple alignment approaches now available, selecting the right method depends on your specific constraints: data format, computational budget, annotation quality, and desired training dynamics. No single method dominates all scenarios, so understanding the trade-offs is needed for making an informed choice.
Decision Framework
The choice between alignment methods depends on several factors that you can evaluate before starting an experiment. Working through these questions systematically is more reliable than defaulting to DPO simply because it is the most familiar.
Data format availability:
- Paired preferences (chosen vs. rejected for the same prompt): use DPO, IPO, cDPO, or ORPO.
- Unpaired binary feedback (thumbs up/down on individual responses): use KTO.
- Ranked lists with multiple responses per prompt: DPO can be adapted with pairwise comparisons extracted from the ranking.
Reference model constraints:
- Can maintain frozen reference in memory: DPO, IPO, cDPO, KTO.
- Memory-constrained single-model training: ORPO.
Annotation quality:
- High-quality expert annotations (>90% consistency): standard DPO.
- Moderate-quality annotations (70-90%): cDPO with appropriate .
- Noisy or automated labels (<70%): cDPO with higher , or KTO.
Training stability requirements:
- Need bounded optimization: IPO or cDPO.
- Standard training sufficient: DPO.

Empirical Performance Comparison
Published results show that no single method dominates across all benchmarks. General patterns from the literature reveal a fine-grained picture in which each method has conditions under which it performs best.
DPO remains a strong baseline. Its simplicity and well-understood behavior make it the default choice for many applications. Performance degradation typically occurs only with very long training or highly noisy data.
IPO shows advantages when training for many epochs or when the preference data has high confidence. The bounded optimization prevents the overconfident predictions that can emerge with extended DPO training. On clean benchmarks with clear preference signals, IPO often matches or exceeds DPO's performance while showing better calibration.
KTO reaches comparable performance to DPO when preference data is artificially unpaired from originally paired data. When working with naturally unpaired data, KTO materially outperforms naive approaches like randomly constructing pairs from unpaired labels, because those artificial pairs may reverse the actual preference relationship.
ORPO reduces computational overhead by 20-30% by eliminating the reference model. Quality results are competitive with DPO on standard benchmarks, though some studies report slightly lower performance on challenging alignment tasks where the stability provided by the reference model turns out to matter.
cDPO gives consistent improvements over standard DPO when annotation noise is known to be present. The gains are proportional to the actual noise level; cDPO shows little difference from DPO on clean data, which means the conservative treatment of labels does not hurt when labels are reliable.
Computational Costs
The computational characteristics differ materially across methods:
| Method | Memory | Forward Passes | Training Stages |
|---|---|---|---|
| DPO | 2x model | 2x per step | SFT, then DPO |
| IPO | 2x model | 2x per step | SFT, then IPO |
| cDPO | 2x model | 2x per step | SFT, then cDPO |
| KTO | 2x model | 2x per step | SFT, then KTO |
| ORPO | 1x model | 1x per step | Single stage |
ORPO's single-stage training can reduce wall-clock time by 40-50% compared to two-stage approaches, which makes it attractive for rapid iteration where each experiment needs to run quickly.
Training Dynamics
Let's simulate training dynamics for each method to see how they converge differently:
import numpy as np
def simulate_training(method, n_steps=100, lr=0.1, beta=0.1, epsilon=0.1):
"""Simulate how h_theta evolves during training for different methods."""
h = 0.0 # Start with equal preference
history = [h]
target = 1 / (2 * beta)
for _ in range(n_steps):
if method == "dpo":
# DPO gradient: β * σ(-βh)
grad = beta / (1 + np.exp(beta * h))
elif method == "ipo":
# IPO gradient: -2(h - target)
grad = -2 * (h - target)
elif method == "cdpo":
# cDPO gradient
sigma_pos = 1 / (1 + np.exp(-beta * h))
sigma_neg = 1 - sigma_pos
grad = beta * ((1 - epsilon) * sigma_neg - epsilon * sigma_pos)
h = h + lr * grad
history.append(h)
return history
# Simulate all methods
dpo_history = simulate_training("dpo", n_steps=200, lr=3.0)
ipo_history = simulate_training("ipo", n_steps=200)
cdpo_history = simulate_training("cdpo", n_steps=200, lr=10.0)
Practical Guidelines
Based on the analysis above, here are practical recommendations for choosing and implementing DPO variants. These guidelines are intended as starting points; the right choices for your specific application will depend on empirical validation on held-out data.
Method Selection Checklist
When selecting an alignment method, work through these questions in order:
-
What data do you have?
- Paired preferences: DPO, IPO, cDPO, or ORPO.
- Unpaired binary feedback: KTO.
-
How much memory can you allocate?
- Can fit two models: any method.
- Limited to one model: ORPO.
-
How clean are your labels?
- High quality (>90% agreement): DPO or IPO.
- Moderate quality (70-90%): cDPO with -.
- Low quality (<70%): cDPO with -.
-
How long will you train?
- Short training (1-3 epochs): DPO.
- Extended training: IPO or cDPO.
Hyperparameter Recommendations
Each method has specific hyperparameter considerations worth understanding before you begin.
DPO: The parameter typically works well in the range 0.1-0.5. Lower values allow more deviation from the reference; higher values keep the policy closer to the reference. Start with and adjust based on whether outputs are too conservative (increase ) or too different from the base model (decrease ).
IPO: Uses the same parameter, but interpretation differs. The target margin means smaller creates larger targets. Start with (target 5.0) and adjust if convergence is too slow (increase ) or produces weak preferences (decrease ).
cDPO: Choose based on estimated annotation noise. If unknown, start with as a conservative default. The effective training signal strength is , so reduces effective signal by 60%, requiring more data or longer training.
KTO: The loss aversion ratio is typically set to 1.0-2.0. Higher ratios emphasize avoiding bad outputs over creating good ones. The KL reference running average uses momentum 0.9-0.99. Higher momentum gives a more stable baseline but adapts more slowly to policy changes.
ORPO: The weight balances SFT and preference objectives. Values of 0.1-0.5 are typical. Too low ignores preferences; too high destabilizes SFT and may cause the model to sacrifice fluency for preference optimization.
Common Pitfalls
Several issues frequently arise when implementing DPO variants, and knowing them in advance can save significant debugging time.
Length bias is the most common and least obvious problem. All preference methods can develop biases toward response length if chosen responses are systematically longer or shorter than rejected ones. When annotators prefer more detailed responses, every variant will learn that length correlates with quality and may produce unnecessarily verbose outputs. Monitor average response lengths during training and consider normalizing log-probabilities by sequence length.
Mode collapse occurs when aggressive preference optimization causes the model to produce repetitive outputs. The KL penalty gives some protection, but it is insufficient when is very small. Monitoring diversity metrics like distinct n-grams alongside preference metrics helps detect this early.
Reward hacking manifests as the model learning to exploit artifacts in preference data rather than learning real quality signals. Annotators sometimes prefer responses that sound confident over responses that are accurate, or prefer formatting conventions over substance. Regular evaluation on held-out prompts with diverse evaluation criteria helps detect reward hacking before it becomes severe.
Training instability is a practical concern when combining small values with large learning rates. The coupled optimization in ORPO is particularly sensitive to learning rate choices. Starting with learning rates an order of magnitude smaller than SFT and increasing gradually is a reliable approach.
Limitations and Open Problems
The growth of DPO variants shows both the importance of alignment and the difficulty of getting it right under real-world conditions. Each variant addresses a real problem, but none fully resolves the basic tension between what alignment methods optimize and the behavior expected from deployed language models.
The most basic challenge persists across all variants: they optimize for human preferences as measured during data collection, which may differ from what humans want during deployment. Preferences are noisy, context-dependent, and subject to strategic behavior from both annotators and models. A model that perfectly optimizes collected preferences may still exhibit concerning behaviors on novel inputs, because the distribution of deployment prompts differs from the distribution used for training. This generalization gap motivates ongoing research into more reliable alignment approaches that go beyond fitting a fixed preference dataset.
The paired data assumption shared by DPO, IPO, cDPO, and ORPO creates additional problems that are easy to overlook. Constructing preference pairs requires that two responses exist for the same prompt, which either means generating multiple candidate responses during dataset construction (expensive) or collecting head-to-head comparisons from annotators (also expensive and difficult to scale). KTO partially addresses this, but it still requires a reference model and relies on the implicit reward as a surrogate for quality, which may not align with downstream task performance. The gap between proxy signals used during alignment and actual downstream metrics remains a core unsolved problem.
Computational efficiency gains from ORPO are real but come with theoretical trade-offs. Eliminating the reference model means there is no explicit constraint on how far the policy drifts from sensible behavior. The SFT loss gives implicit regularization, but this regularization is not backed by the same theoretical guarantees as the KL constraint. In practice, ORPO works well on standard benchmarks, but the conditions under which the implicit regularization fails are not yet fully understood, which makes it harder to debug when problems arise.
All of these methods also share a deeper limitation: they assume that the preference labels capture something real about alignment, but the relationship between pairwise preferences and broader safety or helpfulness is not straightforward. A model can learn to win every head-to-head comparison by generating responses that sound confident and fluent while containing subtle errors or harmful content that annotators miss on casual inspection. The field continues to develop evaluation methodologies that go beyond preference accuracy, but connecting those evaluations back to training objectives remains an open research challenge.
Summary
This chapter examined four influential DPO variants, each addressing specific limitations of the original algorithm while preserving the core insight that preference learning can bypass explicit reward modeling.
IPO reformulates preference learning as regression with a target margin , preventing the unbounded preference growth that can occur with standard DPO. The squared loss creates self-correcting gradients that converge to a stable equilibrium, and the gradient changes sign when the model overshoots the target. This gives natural protection against overconfidence.
KTO lets learning from unpaired binary feedback by incorporating insights from prospect theory. The asymmetric value function and adaptive reference point allow training on thumbs-up and thumbs-down data without requiring direct comparisons between responses. The loss aversion weighting shows the empirical finding that humans are more sensitive to bad outcomes than equivalent good ones.
ORPO eliminates the reference model entirely by combining SFT with odds-ratio preference optimization. This reduces memory requirements by half and collapses two-stage training into a single pass. The coupling between the SFT loss and the odds ratio loss gives implicit regularization that substitutes for the KL constraint provided by the reference model in other methods.
cDPO handles annotation noise through label smoothing, modeling preferences probabilistically rather than deterministically. The smoothing parameter should reflect actual annotation uncertainty and directly determines the finite optimal gap at which training stabilizes.
Method selection depends on data format, computational constraints, annotation quality, and training duration. No single method dominates across all scenarios. IPO and cDPO both produce bounded optimization with stable equilibria; they differ in how that equilibrium is derived. KTO handles unpaired data; ORPO saves memory. Understanding these trade-offs at the level of loss functions and gradients, rather than just benchmark numbers, equips you to make informed choices and to diagnose problems when experiments do not go as expected.
DPO Variants
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!