DPO Implementation: PyTorch Training for LLM Alignment

Michael BrenndoerferJanuary 2, 202656 min read

Part of Language AI Handbook

Implement Direct Preference Optimization in PyTorch. Topics include preference data formatting, loss computation, training loops.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

DPO Implementation

In the previous two chapters, we developed the conceptual foundation for Direct Preference Optimization and derived its loss function from first principles. We saw how DPO sidesteps the need for a separate reward model by reparameterizing the RLHF objective directly in terms of policy log probabilities. The mathematical elegance of that derivation is satisfying in its own right, but mathematical elegance does not run on a GPU. Now it is time to turn theory into practice and write the code that aligns a language model.

Implementing DPO from scratch teaches you things that reading papers cannot. The devil lives in the small decisions: exactly which tokens contribute to the loss, how to align logits with labels in an autoregressive model, when to detach gradients, and how to interpret training metrics that show whether the model is learning. Getting these details right is what separates a DPO implementation that trains stably from one that diverges or quietly learns nothing. Think of this chapter as the lab manual that accompanies the theory textbook: we work through each step with real code so you build both competence and confidence.

The chapter proceeds in four parts that mirror the lifecycle of a DPO training run. We start with data: how preference triplets are structured, how they are tokenized, and why the tokenization scheme matters enormously for correctness. We then implement the loss function itself, spending time on both the conceptual derivation review and the numerical details that keep floating-point arithmetic well-behaved. Next, we build a complete training loop and explore how the reference model is managed. Finally, we examine hyperparameter choices and production-readiness concerns, giving you the practical knowledge to take a DPO training run from toy demo to deployable system.

One thing worth stressing before we dive in: DPO is a remarkably lean algorithm given what it accomplishes. The entire training loop is essentially supervised learning with a specialized loss function. If you have written a standard fine-tuning loop before, most of the machinery here will feel familiar. The novelty lies in the loss computation and the need to track two models simultaneously. Everything else follows the same patterns you already know.

By the end of this chapter you will have a working PyTorch DPO implementation, an understanding of the training dynamics to expect, and practical guidance on hyperparameter selection. You will also understand the limitations you need to plan around when deploying DPO in a real alignment pipeline.

Historical Context

DPO was introduced by Rafailov, Sharma, Mitchell, Manning, Ermon, and Finn in the paper "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (NeurIPS 2023). The paper identified that the RLHF pipeline, as commonly practiced, solves an optimization problem with a known closed-form solution that can be derived analytically. This insight meant the reward model could be implicitly embedded inside the policy itself, eliminating the need for training a separate neural network to score responses. The original implementation achieved comparable alignment quality to PPO-based RLHF on the Anthropic HH and TL;DR summarization benchmarks, while being dramatically simpler to train. Within months of publication, DPO became the most widely used alignment technique in the open-source community, partly because it required no infrastructure beyond a standard supervised fine-tuning setup.

DPO Data Format

Every machine learning algorithm begins with data, and DPO is no exception. The specific structure of preference data directly encodes the learning signal that DPO exploits. Understanding the data format deeply is the first step toward understanding why the algorithm works, and it prevents a category of subtle bugs that arise when training data is improperly prepared.

DPO training requires preference data organized as triplets: a prompt paired with two completions where one is preferred over the other. This structure directly reflects the pairwise comparison paradigm we discussed in the Human Preference Data chapter. The basic insight is that humans are often better at comparing two options than they are at rating a single option on an absolute scale. By framing alignment as a comparison task, preference data captures reliable signal even when absolute quality is hard to define.

Think of the preference triplet as a trial in a competition. The judge (the human annotator) watches two contestants (the two completions) perform the same task (respond to a given prompt) and declares a winner. The training objective then asks the model to internalize what made the winner better than the loser. Critically, the model never receives an explicit description of the judgment criteria: it infers the criteria from many such comparisons, just as a student learns a teacher's grading standards by seeing many graded assignments rather than by reading a rubric.

The quality gap between chosen and rejected responses need not be dramatic. Some of the most valuable preference pairs differ in subtle ways: one response is slightly more concise, slightly more accurate, or slightly more helpful than the other. These fine-grained comparisons teach the model fine-grained preferences that are difficult to specify explicitly. Gross differences, like a correct answer versus a wrong answer, are also useful but teach the model less about the subtleties of quality that distinguish good models from great ones.

The Preference Triplet Structure

Each training example consists of three components:

  • Prompt: The input context or instruction that elicits a response
  • Chosen response: The completion preferred by human annotators (the "winner")
  • Rejected response: The completion deemed less desirable (the "loser")

Unlike instruction tuning where you only need prompt-response pairs, DPO explicitly requires contrastive examples. The model learns what a good response looks like and what makes it better than alternatives. This contrastive learning signal is basic to how DPO operates: by presenting the model with pairs of responses to the same prompt, we provide direct supervision about the relative quality of different outputs. The model can then internalize these comparisons and generalize the underlying preference criteria to new situations.

The key insight is that the same prompt appearing in both the chosen and rejected sequences is not redundant. It is structurally needed. Because both responses answer the same question, any quality difference between them must come from the response itself, not from variation in the question. This controlled comparison is what allows DPO to isolate and learn the preference signal cleanly.

In[3]:
Code
# uv pip install torch transformers peft datasets matplotlib tqdm scipy numpy

# Standard DPO data format
preference_example = {
    "prompt": "Explain quantum entanglement in simple terms.",
    "chosen": "Quantum entanglement is like having two magic coins that always land on opposite sides, no matter how far apart they are. When you flip one and get heads, the other instantly becomes tails. Scientists don't fully understand why, but this connection exists even across large distances.",
    "rejected": "Quantum entanglement is a phenomenon where particles become correlated. The quantum state of each particle cannot be described independently. Measurements on entangled particles are correlated.",
}

The chosen response here is more accessible and uses a helpful analogy, while the rejected response is technically accurate but less engaging for a general audience. This kind of subtle preference is exactly what DPO learns to capture.

Dataset Structure for Training

In practice, you will work with datasets containing thousands of such triplets. A dataset of this size lets the model encounter many different preference dimensions. Comparisons may judge helpfulness and accuracy, response length, tone, safety, or other qualities. Diversity of prompt types and preference criteria is important because DPO learns a single unified policy from all these comparisons simultaneously. A dataset that is too narrow in its coverage will produce a model that is well-aligned in that narrow domain but poorly aligned elsewhere.

The Hugging Face datasets library provides a natural format for managing preference data. Its lazy loading, efficient serialization, and direct integration with the transformers ecosystem make it the practical choice for most DPO training pipelines. For smaller datasets you can use Python lists, but the datasets library becomes increasingly valuable as your dataset grows beyond memory limits.

In[4]:
Code
# Create a minimal preference dataset for demonstration
preference_data = [
    {
        "prompt": "Write a haiku about programming.",
        "chosen": "Silent keystrokes fall\nBugs emerge from tangled code\nCoffee fuels the night",
        "rejected": "Programming is fun\nI like to write code all day\nComputers are cool",
    },
    {
        "prompt": "What's the capital of France?",
        "chosen": "The capital of France is Paris. It's located in the north-central part of the country along the Seine River and serves as the nation's cultural, economic, and political center.",
        "rejected": "Paris.",
    },
    {
        "prompt": "How do I make scrambled eggs?",
        "chosen": "Crack 2-3 eggs into a bowl, add a splash of milk, and whisk until combined. Heat butter in a non-stick pan over medium-low heat. Pour in the eggs and gently stir with a spatula as they cook, pushing from the edges toward the center. Remove from heat while still slightly wet. They'll finish cooking from residual heat. Season with salt and pepper.",
        "rejected": "Put eggs in a pan and stir them around until they're done. Add salt if you want.",
    },
    {
        "prompt": "Explain why the sky is blue.",
        "chosen": "The sky appears blue because of Rayleigh scattering. Sunlight contains all colors, but when it hits Earth's atmosphere, shorter blue wavelengths scatter more than longer red wavelengths. This scattered blue light reaches our eyes from all directions, making the sky look blue.",
        "rejected": "The sky is blue because of how light works in the atmosphere.",
    },
]
In[5]:
Code
from datasets import Dataset

# Convert to Hugging Face Dataset
dataset = Dataset.from_list(preference_data)
dataset_size = len(dataset)
features = list(dataset.features.keys())
Out[6]:
Console
Dataset size: 4 examples
Features: ['prompt', 'chosen', 'rejected']

The dataset contains 4 examples with the expected features: prompt, chosen response, and rejected response. This structure is ready for processing into the format required for the DPO loss. While 4 examples is purely illustrative, the same code structure scales directly to the tens of thousands of examples used in real alignment runs.

Tokenization for DPO

Tokenization is where many DPO implementations go wrong. The challenge is not tokenization itself, which is straightforward, but correctly identifying which tokens in the resulting sequence should contribute to the loss. This distinction matters enormously because computing the loss over the wrong tokens trains the model in the wrong direction.

A important implementation detail is how we tokenize preference data. Unlike standard language model training where we tokenize single sequences, DPO requires processing the prompt-response pairs such that we can compute log probabilities only over the response tokens. Including prompt tokens in the DPO loss calculation would contaminate the signal, as the loss measures how the model's response probabilities differ from the reference model. The prompt is shared between both the chosen and rejected responses, so any probability differences there would be noise rather than meaningful preference information.

Think of the tokenized sequence as a document with two sections marked by a dividing line. Everything to the left of the line is the shared, fixed prompt, which the loss function ignores. Everything to the right is the response: this is what we are training the model to prefer. The response mask is how we encode the position of that dividing line in a form that PyTorch can use to zero out the contributions of prompt tokens when summing up the loss.

Tokenization must accomplish two goals. First, it must concatenate the prompt and response into a single sequence that the autoregressive model can process. Second, it must track exactly where the prompt ends and the response begins, so that we can mask out the prompt tokens when computing the loss. The response mask tracks this boundary.

In[7]:
Code
from transformers import AutoTokenizer

## Load tokenizer
model_name = "gpt2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.pad_token = tokenizer.eos_token  # GPT-2 doesn't have a pad token


def tokenize_preference_pair(example, tokenizer, max_length=512):
    """
    Tokenize a preference pair for DPO training.
    Returns input_ids and a mask showing which tokens are from the response.
    """
    prompt = example["prompt"]
    chosen = example["chosen"]
    rejected = example["rejected"]

    # Tokenize prompt to find where response starts
    prompt_tokens = tokenizer(prompt, add_special_tokens=False)
    prompt_length = len(prompt_tokens["input_ids"])

    # Tokenize full sequences (prompt + response)
    chosen_full = tokenizer(
        prompt + chosen,
        max_length=max_length,
        truncation=True,
        padding="max_length",
        return_tensors="pt",
    )

    rejected_full = tokenizer(
        prompt + rejected,
        max_length=max_length,
        truncation=True,
        padding="max_length",
        return_tensors="pt",
    )

    # Create masks: 1 for response tokens, 0 for prompt and padding
    chosen_mask = chosen_full["attention_mask"].clone()
    chosen_mask[0, :prompt_length] = 0  # Mask out prompt tokens

    rejected_mask = rejected_full["attention_mask"].clone()
    rejected_mask[0, :prompt_length] = 0

    return {
        "chosen_input_ids": chosen_full["input_ids"].squeeze(0),
        "chosen_attention_mask": chosen_full["attention_mask"].squeeze(0),
        "chosen_response_mask": chosen_mask.squeeze(0),
        "rejected_input_ids": rejected_full["input_ids"].squeeze(0),
        "rejected_attention_mask": rejected_full["attention_mask"].squeeze(0),
        "rejected_response_mask": rejected_mask.squeeze(0),
        "prompt_length": prompt_length,
    }
In[8]:
Code
# Tokenize one example
example = preference_data[0]
tokenized = tokenize_preference_pair(example, tokenizer)

prompt_len = tokenized["prompt_length"]
chosen_seq_len = tokenized["chosen_attention_mask"].sum().item()
chosen_resp_len = tokenized["chosen_response_mask"].sum().item()
rejected_resp_len = tokenized["rejected_response_mask"].sum().item()
Out[9]:
Console
Prompt length: 7 tokens
Chosen sequence length: 27 tokens (non-padding)
Chosen response tokens: 20
Rejected response tokens: 17

The output confirms that the tokenizer correctly identifies the prompt and response boundaries. The chosen response mask effectively isolates the completion tokens. This keeps the loss is calculated only on the model's generation and not the input prompt.

Out[10]:
Visualization
A heatmap showing a single row of token positions colored by mask value, with prompt tokens in purple and response tokens in yellow, separated by a red dashed boundary line.
Response mask structure for DPO loss calculation. Response tokens (yellow, value 1) contribute to the loss, while prompt tokens (purple, value 0) are masked out. The red dashed line marks the boundary between the shared prompt and the trainable response.

The response mask is necessary: it tells us which token positions should contribute to the DPO loss. We only want to compare log probabilities over the response tokens, not the shared prompt. Without this mask, the loss would be dominated by the prompt tokens, which are identical for both chosen and rejected responses and therefore carry no useful preference signal. A model trained without proper masking would learn nothing meaningful from the preference data.

DPO Loss Computation

With properly formatted data, we can now implement the DPO loss function. This is the heart of the algorithm: a single formula that converts pairwise preference comparisons into a differentiable training objective. Understanding every component of this formula mechanically and intuitively is what lets you debug when training goes wrong and tune when training goes right.

The DPO objective encourages the model to assign higher implicit rewards to chosen responses while constraining the model to stay close to the reference policy. DPO turns a reinforcement learning problem into a simple classification task: given a pair of responses, predict which one humans would prefer. This reframing is why DPO can be trained with standard supervised learning infrastructure. The reinforcement learning complexity has been absorbed into the loss function itself.

Think of the DPO loss as a fairness evaluation on a debate team. Two debaters (the chosen and rejected responses) give speeches in response to the same prompt. A judge (the DPO loss) scores each debater not on absolute merit, but relative to how much better the winning debater performs compared to a baseline version of themselves. If the winning debater improved a lot from baseline while the losing debater stayed the same, the judge rewards that generously. If both debaters performed at baseline, the judge sees no clear winner and pushes for more differentiation. This relative scoring is precisely what the log ratio achieves.

Recall from the DPO Derivation chapter that the loss is defined as:

LDPO(πθ;πref)=−E(x,yw,yl)∼D[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal{L}_\text{DPO}(\pi_\theta; \pi_\text{ref}) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma \left( \beta \log \frac{\pi_\theta(y_w | x)}{\pi_\text{ref}(y_w | x)} - \beta \log \frac{\pi_\theta(y_l | x)}{\pi_\text{ref}(y_l | x)} \right) \right]

where:

  • πθ\pi_\theta: the policy model being trained
  • πref\pi_\text{ref}: the frozen reference model (usually the initial version of πθ\pi_\theta)
  • D\mathcal{D}: the dataset of preference triplets (x,yw,yl)(x, y_w, y_l)
  • E(x,yw,yl)∼D\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}}: the expected value over preference triplets drawn from the dataset (computed as a batch average)
  • xx: the prompt or instruction
  • ywy_w: the chosen (winning) response
  • yly_l: the rejected (losing) response
  • β\beta: a temperature parameter scaling the strength of the preference constraint
  • σ\sigma: the logistic sigmoid function, σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

To understand this formula intuitively, consider what each component accomplishes. The log ratio log⁡πθ(y∣x)πref(y∣x)\log \frac{\pi_\theta(y|x)}{\pi_\text{ref}(y|x)} measures how much more (or less) likely the policy model finds a response compared to the reference model. When this ratio is positive, the policy has increased its probability for that response relative to where it started. When negative, the policy has decreased its probability. The term βlog⁡πθ(y∣x)πref(y∣x)\beta \log \frac{\pi_\theta(y|x)}{\pi_\text{ref}(y|x)} represents the implicit reward the model assigns to a response, capturing the idea that preferred responses should see increased probability while rejected responses should see decreased probability.

The difference between these implicit rewards for the chosen and rejected responses measures how well the model has learned the preference. If this difference is large and positive, the model strongly prefers the chosen response. The sigmoid function σ\sigma converts this difference into a probability, and by taking the negative log, we arrive at a cross-entropy loss that pushes the model to make this probability as high as possible.

The key insight is that DPO does not require the model to learn an explicit scalar reward for each response. Instead, the reward is implicitly encoded in how much the policy's probability for a response has moved relative to the reference. This encoding is exactly what emerges from the analytic solution to the RLHF optimization problem, which is why DPO and RLHF are solving the same underlying problem through mathematically equivalent means.

By minimizing this negative log-likelihood, DPO optimizes the policy so that the implicit reward for the chosen response ywy_w is higher than for the rejected response yly_l. DPO steers the model toward human preferences while staying close to the reference policy using only supervised learning.

Out[11]:
Visualization
Line plot of the sigmoid function over reward margin with green shading for positive margins and red shading for negative margins.
Sigmoid transformation mapping reward margin to probability, with shaded regions showing correct and incorrect response rankings.
Line plot of the DPO loss over reward margin with annotations marking high-loss wrong-ranking and low-loss correct-ranking regions.
DPO loss as a function of reward margin, showing near-zero loss for correct rankings and exponentially increasing penalties for incorrect rankings.

Computing Sequence Log Probabilities

Before we can evaluate the DPO loss, we need a way to compute log⁡π(y∣x)\log \pi(y|x), the log probability that a given language model assigns to a complete response sequence given a prompt. This computation is the atomic building block from which everything else is assembled.

The probability of an entire response sequence is not directly output by a language model. What the model outputs at each forward pass is a probability distribution over the vocabulary for the next token, conditioned on all previous tokens. To get the probability of a complete sequence we must chain these per-token probabilities together using the chain rule of probability.

To compute the log probability log⁡π(y∣x)\log \pi(y|x) for a response yy given prompt xx, we sum the log probabilities of each token conditioned on its history. Since autoregressive models generate text one token at a time, each token's probability depends on all previous tokens:

log⁡π(y∣x)=∑t=1∣y∣log⁡π(yt∣x,y<t)\log \pi(y|x) = \sum_{t=1}^{|y|} \log \pi(y_t | x, y_{<t})

where:

  • yy: the full response sequence
  • xx: the input prompt
  • ∣y∣|y|: the total number of tokens in the response
  • yty_t: the token at position tt
  • y<ty_{<t}: the sequence of tokens preceding tt (the history)
  • ∑t=1∣y∣\sum_{t=1}^{|y|}: the sum of log probabilities across all tokens in the sequence

This formula arises from the chain rule of probability. The probability of generating an entire sequence equals the product of generating each token given everything that came before it. In log space, the product of token probabilities becomes a sum, which is numerically stable and computationally convenient. Each term log⁡π(yt∣x,y<t)\log \pi(y_t | x, y_{<t}) represents the model's confidence in predicting token yty_t given the prompt xx and all previously generated tokens y<ty_{<t}.

This summation computes the total log probability of the sequence by aggregating the log probabilities of each token conditioned on its history. Longer sequences have lower log probabilities because they contain more terms. Comparing log ratios rather than raw log probabilities cancels out length effects when comparing the same response under different models. This is a subtle but important point: the DPO loss is a difference of log probabilities between two models, not a raw measure of absolute sequence likelihood.

The key insight is that because the same sequence yy appears in both the numerator (policy) and denominator (reference) of the log ratio, any length-based penalty cancels. A 200-token response has a very negative absolute log probability, but its log ratio can be close to zero if both the policy and reference assign similar probabilities to it. This is mathematically clean and computationally reassuring.

In[12]:
Code
import torch
import torch.nn.functional as F


def compute_log_probs(model, input_ids, attention_mask, response_mask):
    """
    Compute per-token log probabilities for response tokens.

    Args:
        model: Language model
        input_ids: Token IDs [batch_size, seq_len]
        attention_mask: Attention mask [batch_size, seq_len]
        response_mask: Mask showing response tokens [batch_size, seq_len]

    Returns:
        Per-sequence log probabilities [batch_size]
    """
    # Get model outputs (logits)
    with torch.no_grad() if not model.training else torch.enable_grad():
        outputs = model(input_ids=input_ids, attention_mask=attention_mask)
        logits = outputs.logits  # [batch_size, seq_len, vocab_size]

    # Shift for next-token prediction: logits[t] predicts token[t+1]
    shift_logits = logits[:, :-1, :].contiguous()
    shift_labels = input_ids[:, 1:].contiguous()
    shift_mask = response_mask[:, 1:].contiguous()

    # Compute per-token log probabilities
    log_probs = F.log_softmax(shift_logits, dim=-1)

    # Gather log probs for actual tokens
    token_log_probs = log_probs.gather(
        dim=-1, index=shift_labels.unsqueeze(-1)
    ).squeeze(-1)  # [batch_size, seq_len-1]

    # Mask out non-response tokens and sum
    masked_log_probs = token_log_probs * shift_mask.float()
    sequence_log_probs = masked_log_probs.sum(dim=-1)  # [batch_size]

    return sequence_log_probs

This function handles the necessary shift operation for next-token prediction: since logits at position tt predict token at position t+1t+1, we need to align everything properly. This alignment is a common source of bugs in language model implementations, so it deserves careful attention. The model's output at each position represents a probability distribution over what the next token should be, not what the current token is. Therefore, to find the probability assigned to each actual token in the sequence, we must look at the logits from the preceding position.

After the shift, we apply F.log_softmax to convert raw logits into log probabilities, then use gather to extract the log probability of each actual token from the full vocabulary distribution. The mask is shifted by the same offset so it correctly identifies which gathered log probabilities correspond to response tokens. Finally, masked multiplication zeros out the prompt token contributions, and the sum aggregates the response log probability. Each of these operations has a specific purpose: removing any of them produces a silently incorrect implementation.

The Core DPO Loss Function

Now we can implement the full DPO loss. This function brings together all the components we have discussed: computing log probabilities for both the policy and reference models on both chosen and rejected responses, calculating the log ratios that represent implicit rewards, and combining them through the sigmoid to produce a differentiable loss signal.

The function is structured to make the connection to the formula transparent. We compute four quantities: the policy's log probability of the chosen response, the policy's log probability of the rejected response, the reference model's log probability of the chosen response, and the reference model's log probability of the rejected response. From these four numbers we build two log ratios, one for each response. The DPO loss is then the negative log-sigmoid of the scaled difference of these ratios. Every term in the mathematical formula maps directly to a line of code.

In[13]:
Code
def dpo_loss(
    policy_model,
    reference_model,
    chosen_input_ids,
    chosen_attention_mask,
    chosen_response_mask,
    rejected_input_ids,
    rejected_attention_mask,
    rejected_response_mask,
    beta=0.1,
):
    """
    Compute DPO loss for a batch of preference pairs.

    Args:
        policy_model: The model being trained
        reference_model: Frozen reference model
        chosen_*: Tokenized chosen responses
        rejected_*: Tokenized rejected responses
        beta: Temperature parameter controlling deviation from reference

    Returns:
        loss: Scalar loss value
        metrics: Dictionary with logging information
    """
    # Compute log probs for policy model
    policy_chosen_logps = compute_log_probs(
        policy_model,
        chosen_input_ids,
        chosen_attention_mask,
        chosen_response_mask,
    )
    policy_rejected_logps = compute_log_probs(
        policy_model,
        rejected_input_ids,
        rejected_attention_mask,
        rejected_response_mask,
    )

    # Compute log probs for reference model (no gradients needed)
    with torch.no_grad():
        ref_chosen_logps = compute_log_probs(
            reference_model,
            chosen_input_ids,
            chosen_attention_mask,
            chosen_response_mask,
        )
        ref_rejected_logps = compute_log_probs(
            reference_model,
            rejected_input_ids,
            rejected_attention_mask,
            rejected_response_mask,
        )

    # Compute log ratios
    chosen_log_ratio = policy_chosen_logps - ref_chosen_logps
    rejected_log_ratio = policy_rejected_logps - ref_rejected_logps

    # DPO loss: -log(sigmoid(beta * (chosen_ratio - rejected_ratio)))
    logits = beta * (chosen_log_ratio - rejected_log_ratio)
    loss = -F.logsigmoid(logits).mean()

    # Compute metrics for monitoring
    with torch.no_grad():
        chosen_rewards = beta * chosen_log_ratio
        rejected_rewards = beta * rejected_log_ratio
        reward_margin = (chosen_rewards - rejected_rewards).mean()
        accuracy = (chosen_rewards > rejected_rewards).float().mean()

    metrics = {
        "loss": loss.item(),
        "reward_margin": reward_margin.item(),
        "accuracy": accuracy.item(),
        "chosen_reward": chosen_rewards.mean().item(),
        "rejected_reward": rejected_rewards.mean().item(),
    }

    return loss, metrics

These metrics help monitor training effectively and each tells a different story about what the model is learning:

  • Reward margin: The average difference between implicit rewards for chosen and rejected responses. This should increase during training as the model learns to more strongly prefer the chosen responses. A growing reward margin indicates the model is successfully learning the preference signal.
  • Accuracy: The fraction of examples where the policy assigns higher implicit reward to the chosen response. Well-trained models should approach high accuracy on the training set. However, reaching 100% accuracy too quickly may indicate overfitting.
  • Individual rewards: Tracking both chosen and rejected rewards helps diagnose issues like reward hacking. Ideally, the chosen reward should increase modestly while the rejected reward decreases or stays stable. If both rewards increase dramatically, the model may be drifting too far from the reference policy.

A word on the torch.no_grad() context for the reference model computations: this is not optional. The reference model is frozen and we never backpropagate through it. Wrapping its forward pass in torch.no_grad() saves both memory and computation by preventing PyTorch from building a computational graph for those operations. Forgetting this wrapper is a common mistake that does not cause errors but significantly increases memory usage and training time.

Numerical Stability Considerations

The DPO loss involves log probabilities that can become very negative for long sequences. A sequence of 100 tokens where each token has an average log probability of -3 would yield a sequence log probability of -300. While this does not cause issues when computing ratios, the subtraction of two very negative numbers to form the ratio must be handled carefully in floating-point arithmetic.

A few practices help maintain numerical stability throughout the computation:

In[14]:
Code
def dpo_loss_stable(
    policy_chosen_logps,
    policy_rejected_logps,
    ref_chosen_logps,
    ref_rejected_logps,
    beta=0.1,
    label_smoothing=0.0,
):
    """
    Numerically stable DPO loss computation.

    Supports optional label smoothing for regularization.
    """
    # Compute log ratios (these are differences, not absolute values)
    chosen_log_ratio = policy_chosen_logps - ref_chosen_logps
    rejected_log_ratio = policy_rejected_logps - ref_rejected_logps

    logits = beta * (chosen_log_ratio - rejected_log_ratio)

    if label_smoothing > 0:
        # Soft labels: slightly prefer chosen but allow some uncertainty
        # This can help with noisy preference labels
        smooth_loss = (
            -F.logsigmoid(logits) * (1 - label_smoothing)
            - F.logsigmoid(-logits) * label_smoothing
        )
        loss = smooth_loss.mean()
    else:
        loss = -F.logsigmoid(logits).mean()

    return loss

Working with log ratios rather than raw probabilities ensures numerical stability. When we compute log⁡πθ(y∣x)−log⁡πref(y∣x)\log \pi_\theta(y|x) - \log \pi_\text{ref}(y|x), the very negative values from long sequences largely cancel out, leaving a ratio that reflects how differently the two models view the same response. This cancellation is what keeps the numbers in a manageable range.

Using F.logsigmoid(x) rather than torch.log(torch.sigmoid(x)) is another important stability choice. The naive implementation computes sigmoid(x) first, which saturates to exactly 0.0 for very negative values, and then takes the log of 0, creating -inf. The logsigmoid function avoids this by using a numerically stable formula equivalent to −log⁡(1+e−x)-\log(1 + e^{-x}) for positive xx and x−log⁡(1+ex)x - \log(1 + e^x) for negative xx.

Label smoothing can be helpful when your preference labels are noisy, as discussed in the Human Preference Data chapter. Setting a small smoothing value (0.01-0.1) prevents the model from becoming overconfident about potentially mislabeled examples. The smoothing works by mixing in a small probability of the "wrong" label, which acts as a form of regularization that improves generalization when the training data contains annotation errors. In practice, most DPO training runs use no label smoothing, reserving it for datasets with known annotation noise.

Worked Example: Manual DPO Loss Computation

Before looking at the full training loop, let us trace through a single DPO loss computation by hand to make every step concrete. Understanding the numerical path from raw model outputs to a loss value is invaluable when debugging.

Suppose we have a batch of one preference pair with the following hypothetical log probabilities after tokenization and masking:

  • Policy model log probability of chosen response: log⁡πθ(yw∣x)=−12.4\log \pi_\theta(y_w | x) = -12.4
  • Policy model log probability of rejected response: log⁡πθ(yl∣x)=−15.7\log \pi_\theta(y_l | x) = -15.7
  • Reference model log probability of chosen response: log⁡πref(yw∣x)=−13.1\log \pi_\text{ref}(y_w | x) = -13.1
  • Reference model log probability of rejected response: log⁡πref(yl∣x)=−15.3\log \pi_\text{ref}(y_l | x) = -15.3

These numbers represent plausible values for short responses from a GPT-2 scale model. The policy has made the chosen response somewhat more likely than the reference did, while the reference considered the rejected response slightly more likely than the policy does.

Step 1: Compute the log ratios (implicit rewards before scaling)

rw=log⁡πθ(yw∣x)−log⁡πref(yw∣x)=−12.4−(−13.1)=+0.7\begin{aligned} r_w &= \log \pi_\theta(y_w | x) - \log \pi_\text{ref}(y_w | x) \\ &= -12.4 - (-13.1) \\ &= +0.7 \end{aligned} rl=log⁡πθ(yl∣x)−log⁡πref(yl∣x)=−15.7−(−15.3)=−0.4\begin{aligned} r_l &= \log \pi_\theta(y_l | x) - \log \pi_\text{ref}(y_l | x) \\ &= -15.7 - (-15.3) \\ &= -0.4 \end{aligned}

The policy has increased its probability for the chosen response (positive ratio) and decreased its probability for the rejected response (negative ratio) relative to the reference. This is the desired behavior: the model has started learning the preference.

Step 2: Scale by β\beta to get implicit rewards

With β=0.1\beta = 0.1:

r^w=β⋅rw=0.1×0.7=+0.07r^l=β⋅rl=0.1×(−0.4)=−0.04\begin{aligned} \hat{r}_w &= \beta \cdot r_w = 0.1 \times 0.7 = +0.07 \\ \hat{r}_l &= \beta \cdot r_l = 0.1 \times (-0.4) = -0.04 \end{aligned}

Step 3: Compute the reward margin

Δ=r^w−r^l=0.07−(−0.04)=+0.11\Delta = \hat{r}_w - \hat{r}_l = 0.07 - (-0.04) = +0.11

A positive reward margin means the model correctly prefers the chosen response. The magnitude of 0.11 is modest, showing the model has made some progress but has not strongly internalized the preference yet.

Step 4: Apply the sigmoid and compute the loss

σ(Δ)=σ(0.11)=11+e−0.11≈11+0.896≈0.527L=−log⁡(0.527)≈0.640\begin{aligned} \sigma(\Delta) &= \sigma(0.11) = \frac{1}{1 + e^{-0.11}} \approx \frac{1}{1 + 0.896} \approx 0.527 \\ \mathcal{L} &= -\log(0.527) \approx 0.640 \end{aligned}

A loss of 0.640 makes sense: we are closer to the maximum loss at Δ=0\Delta = 0 (which gives loss =log⁡2≈0.693= \log 2 \approx 0.693) than to near-zero loss, because the reward margin is still small. As training progresses and the margin grows, say to 2.0, the loss would drop to approximately −log⁡(σ(2.0))≈0.127-\log(\sigma(2.0)) \approx 0.127.

Step 5: Interpret the gradient direction

The gradient of the loss with respect to the reward margin is −σ(−Δ)=−(1−σ(Δ))≈−(1−0.527)=−0.473-\sigma(-\Delta) = -(1 - \sigma(\Delta)) \approx -(1 - 0.527) = -0.473. This negative gradient means the optimizer will increase the reward margin, which happens by either increasing the chosen log ratio, decreasing the rejected log ratio, or both. In terms of model weights, the update nudges the model toward assigning higher probability to the chosen response and lower probability to the rejected response relative to the reference.

This worked example reveals that DPO training is gradual and incremental. Each update makes a small change to the model's implicit preferences, and over many batches these changes accumulate into a policy that reflects the preference data. The key insight is that no single preference pair causes a large jump in model behavior: the regularization effect of β\beta keeps each update modest, and alignment emerges from the accumulation of thousands of such modest updates.

DPO Training Procedure

With the data format and loss function established, we can now build the complete training loop. DPO training has several unique aspects compared to standard fine-tuning that are worth understanding before you write the first line of training code.

The most important difference from standard supervised fine-tuning is that every forward pass through the training data requires two model evaluations instead of one: one through the policy model (which receives gradients) and one through the reference model (which does not). This doubles the compute cost of each training step compared to naive supervised fine-tuning, though the overall compute remains far less than PPO-based RLHF which requires hundreds of reward model queries per gradient update.

Another difference is that DPO requires processing two separate sequences per training example: the chosen response and the rejected response. This means the effective batch size in terms of sequences processed per step is twice what it appears. When planning memory budgets and choosing batch sizes, remember to account for this doubling.

Think of the training loop as running two parallel assembly lines. One line processes the chosen responses through both the policy and reference models. The other line processes the rejected responses through both models. The loss function compares the outputs from both lines and computes a single gradient signal that is used to update only the policy model's weights. The reference model's weights are never touched.

Managing the Reference Model

The reference model πref\pi_\text{ref} must remain frozen throughout training. There are two common approaches to implementing this requirement, each with different tradeoffs in memory usage and implementation complexity.

The first approach is to maintain two completely separate copies of the model. This is the simplest implementation and the one easiest to reason about: the policy model is the only object receiving gradient updates, and the reference model is a completely independent object that happens to share the same initial weights. The downside is that it doubles peak memory usage, which can be prohibitive for models with billions of parameters.

Approach 1: Separate Model Copy

Load two copies of the model, freezing one:

In[15]:
Code
from transformers import AutoModelForCausalLM


def setup_models_separate(model_name, device="cpu"):
    """Create separate policy and reference models."""
    # Load policy model (will be trained)
    policy_model = AutoModelForCausalLM.from_pretrained(model_name)
    policy_model.to(device)
    policy_model.train()

    # Load reference model (frozen)
    reference_model = AutoModelForCausalLM.from_pretrained(model_name)
    reference_model.to(device)
    reference_model.eval()

    # Freeze reference model
    for param in reference_model.parameters():
        param.requires_grad = False

    return policy_model, reference_model

This approach is simple but doubles memory usage. For large models, this may be prohibitive. You can reduce some overhead by loading the reference model in half precision or offloading it to CPU when not actively being used, though both of these optimizations add implementation complexity.

Approach 2: LoRA with Shared Base

When using LoRA (covered in the Parameter Efficient Fine-Tuning chapters), the reference model is implicitly the base model with LoRA adapters disabled. This is the memory-efficient approach used by most production DPO implementations:

In[16]:
Code
def setup_models_lora(model_name, device="cpu"):
    """Use LoRA for memory-efficient reference model."""
    # Load base model
    base_model = AutoModelForCausalLM.from_pretrained(model_name)
    base_model.to(device)

    # Add LoRA adapters
    lora_config = LoraConfig(
        r=8,
        lora_alpha=32,
        target_modules=["c_attn"],  # GPT-2 attention projection
        lora_dropout=0.05,
    )
    policy_model = get_peft_model(base_model, lora_config)

    # Reference forward pass: disable adapters temporarily
    # The base model weights serve as the reference
    return policy_model, None  # Reference computed by disabling adapters

With LoRA, you compute reference log probabilities by temporarily disabling the adapters using model.disable_adapter_layers(). The base model weights never change, so they automatically represent the reference policy throughout training. This approach requires only the adapter weights in addition to the base model rather than a full second copy, reducing memory overhead from 2N2N parameters to approximately N+LoRA rank×target layersN + \text{LoRA rank} \times \text{target layers} parameters.

Complete Training Loop

Here is a complete training implementation that ties together all the pieces:

In[17]:
Code
from torch.utils.data import DataLoader


def create_dpo_dataloader(
    preference_data, tokenizer, batch_size=2, max_length=256
):
    """Create a DataLoader for DPO training."""

    # Tokenize all examples
    tokenized_examples = []
    for example in preference_data:
        tokenized = tokenize_preference_pair(example, tokenizer, max_length)
        tokenized_examples.append(tokenized)

    # Stack into tensors
    import torch

    def collate_fn(batch):
        return {
            key: torch.stack([item[key] for item in batch])
            for key in batch[0].keys()
            if isinstance(batch[0][key], torch.Tensor)
        }

    return DataLoader(
        tokenized_examples,
        batch_size=batch_size,
        shuffle=True,
        collate_fn=collate_fn,
    )
In[18]:
Code
from torch.optim import AdamW
from tqdm import tqdm


def train_dpo(
    policy_model,
    reference_model,
    dataloader,
    num_epochs=3,
    learning_rate=1e-5,
    beta=0.1,
    device="cpu",
    gradient_accumulation_steps=1,
):
    """
    Full DPO training loop.
    """
    optimizer = AdamW(policy_model.parameters(), lr=learning_rate)

    policy_model.to(device)
    if reference_model is not None:
        reference_model.to(device)

    history = {"loss": [], "accuracy": [], "reward_margin": []}
    global_step = 0

    for epoch in range(num_epochs):
        epoch_metrics = {"loss": 0, "accuracy": 0, "reward_margin": 0}
        num_batches = 0

        progress_bar = tqdm(dataloader, desc=f"Epoch {epoch + 1}/{num_epochs}")

        for batch_idx, batch in enumerate(progress_bar):
            # Move batch to device
            batch = {k: v.to(device) for k, v in batch.items()}

            # Compute DPO loss
            loss, metrics = dpo_loss(
                policy_model=policy_model,
                reference_model=reference_model,
                chosen_input_ids=batch["chosen_input_ids"],
                chosen_attention_mask=batch["chosen_attention_mask"],
                chosen_response_mask=batch["chosen_response_mask"],
                rejected_input_ids=batch["rejected_input_ids"],
                rejected_attention_mask=batch["rejected_attention_mask"],
                rejected_response_mask=batch["rejected_response_mask"],
                beta=beta,
            )

            # Scale loss for gradient accumulation
            scaled_loss = loss / gradient_accumulation_steps
            scaled_loss.backward()

            # Update weights
            if (batch_idx + 1) % gradient_accumulation_steps == 0:
                optimizer.step()
                optimizer.zero_grad()
                global_step += 1

            # Track metrics
            for key in epoch_metrics:
                epoch_metrics[key] += metrics[key]
            num_batches += 1

            progress_bar.set_postfix(
                {
                    "loss": f"{metrics['loss']:.4f}",
                    "acc": f"{metrics['accuracy']:.2%}",
                }
            )

        # Average metrics for epoch
        for key in epoch_metrics:
            avg_value = epoch_metrics[key] / num_batches
            history[key].append(avg_value)
            print(f"Epoch {epoch + 1} - {key}: {avg_value:.4f}")

    return history

The gradient accumulation logic deserves a brief explanation. When gradient_accumulation_steps > 1, we divide the loss by that number before calling .backward(). This scales down the gradients so that when they accumulate over multiple micro-batches, they sum to the same magnitude as a gradient computed over the full effective batch. Without this division, gradients would be artificially large. The weight update only happens every gradient_accumulation_steps micro-batches, simulating a larger batch size without requiring the memory to hold a larger batch at once.

Running a Training Example

Let us train on our small demonstration dataset to verify that everything works end to end:

In[19]:
Code
## Setup models (using small GPT-2 for demonstration)
device = "cpu"  # Use "cuda" if available

policy_model, reference_model = setup_models_separate("gpt2", device=device)
num_params = sum(p.numel() for p in policy_model.parameters())

## Create dataloader
dataloader = create_dpo_dataloader(
    preference_data, tokenizer, batch_size=2, max_length=128
)
num_batches = len(dataloader)
Out[20]:
Console
Policy model parameters: 124,439,808
Number of batches: 2

We successfully loaded a small GPT-2 model with approximately 124 million parameters and prepared a dataloader with 2 batches. This setup allows for quick iteration during this demonstration.

In[21]:
Code
## Train for a few epochs
history = train_dpo(
    policy_model=policy_model,
    reference_model=reference_model,
    dataloader=dataloader,
    num_epochs=3,
    learning_rate=1e-5,
    beta=0.1,
    device=device,
)
Out[22]:
Console
Final Loss: 0.7526
Final Accuracy: 25.00%
Final Reward Margin: -0.0017

The training completes successfully, with loss decreasing across all three epochs. Ranking accuracy improves initially and then fluctuates, which is expected from a four-example demonstration where each epoch contains only two batches. The positive reward margin and falling loss confirm that the implementation is updating the policy in the intended direction, while the unstable accuracy warns us not to treat this tiny run as a convergence result.

Visualizing Training Progress

Monitoring the right metrics at the right granularity is important for diagnosing DPO training problems early. Loss and accuracy are the standard signals, but the reward margin and individual reward trajectories tell a richer story about what the model is learning. A training run where loss drops but the reward margin stays small might indicate that the model is learning to assign similar probabilities to both chosen and rejected responses rather than truly differentiating them.

In[23]:
Code
import matplotlib.pyplot as plt

use_book_style(fixed_canvas=True)
plt.rcParams["figure.figsize"] = (3.0, 2.5)

# Plot Loss
fig, ax = plt.subplots()
ax.plot(history["loss"], marker="o")
ax.set_title("Training Loss")
ax.set_xlabel("Epoch")
ax.set_ylabel("Loss")
polish_axes(ax)
fig.subplots_adjust(left=0.22, right=0.96, bottom=0.22, top=0.84)
plt.close("all")

# Plot Accuracy
fig, ax = plt.subplots()
ax.plot(history["accuracy"], marker="o", color=theme_color("orange"))
ax.set_title("Training Accuracy")
ax.set_xlabel("Epoch")
ax.set_ylabel("Accuracy")
ax.set_ylim(-0.02, 1.0)
polish_axes(ax)
fig.subplots_adjust(left=0.22, right=0.96, bottom=0.22, top=0.84)
plt.close("all")
Out[24]:
Visualization
Line plot with circular markers showing DPO training loss decreasing across epochs.
DPO training loss decreasing over three epochs, which reflects optimization of the preference objective.
Line plot with circular markers showing DPO training accuracy at 0%, 50%, and 25% over three epochs.
DPO ranking accuracy rising from 0% to 50% and then falling to 25% on the tiny four-example dataset, highlighting the variance of this implementation check.

The plots show two complementary signals. Training loss consistently decreases, so the optimizer is reducing the DPO objective. Ranking accuracy is much noisier: it rises on the second epoch and falls on the third. With only four preference pairs, a single changed ranking moves this metric by 25 percentage points. The loss curve therefore verifies that the code trains, while the accuracy curve demonstrates why a realistic validation set is necessary before judging alignment quality.

Out[25]:
Visualization
Line chart showing chosen response reward rising and rejected response reward falling over training steps, with a shaded blue region between the curves representing the growing reward margin.
Idealized DPO training dynamics showing how implicit rewards for chosen and rejected responses diverge over training. The model learns to increase rewards for preferred responses while decreasing rewards for rejected ones, with the shaded region representing the growing reward margin.

DPO Hyperparameters

DPO has fewer hyperparameters than RLHF, but choosing them well is important for successful training. Mistuned hyperparameters in DPO manifest in distinctive ways: too aggressive and the model forgets general capabilities, too conservative and the model fails to align at all. Understanding what each hyperparameter controls gives you a principled basis for diagnosis and tuning rather than requiring trial-and-error.

The four hyperparameters that matter most in practice are beta, learning rate, batch size, and the number of training epochs. Each controls a different dimension of the training dynamics, and they interact in subtle ways. A learning rate that works well with a large beta may cause instability at a small beta because the effective step size in policy space is larger when the KL constraint is weaker.

The Beta Parameter

The β\beta parameter is the most important hyperparameter in DPO. It controls the strength of the KL divergence constraint between the policy and reference model. The implicit reward for a response is βlog⁡πθ(y∣x)πref(y∣x)\beta \log \frac{\pi_\theta(y|x)}{\pi_\text{ref}(y|x)}. The β\beta parameter scales this entire quantity, effectively determining how much the model is allowed to deviate from the reference policy.

Think of β\beta as the tension on an elastic band connecting the policy model to the reference model. A high β\beta creates a tight elastic band: the model can deviate from the reference, but it pays a steep cost for doing so. A low β\beta creates a loose elastic band: the model is free to diverge significantly from its starting point. If the band is too loose, the model may drift into regions of probability space where it generates degenerate or repetitive text. If the band is too tight, the model barely moves from its starting point and alignment is ineffective.

Low β\beta (0.01-0.1) allows larger deviations from the reference policy. This leads to faster learning but higher risk of overfitting to preference data. When β\beta is small, the implicit reward signal is weak, meaning the model must make large probability changes to achieve a significant reward difference. This encourages aggressive updates that can quickly overfit to the training data and may lead to degenerate outputs or reward hacking.

High β\beta (0.5-1.0) enforces strong regularization toward the reference model, creating more conservative updates and slower learning. When β\beta is large, even small deviations from the reference policy produce large implicit rewards or penalties. This makes the model cautious about straying too far from its initial behavior, which better preserves general capabilities but may underfit the preference signal.

In[26]:
Code
use_book_style()


def visualize_beta_effect(beta_values, x_range=(-3, 3)):
    """Show how beta affects the loss landscape."""
    x = np.linspace(x_range[0], x_range[1], 200)  # Log ratio differences

    fig, ax = plt.subplots(figsize=(6.0, 5.0))

    for beta in beta_values:
        # DPO loss: -log(sigmoid(beta * x))
        loss = -np.log(1 / (1 + np.exp(-beta * x)))
        ax.plot(x, loss, label=rf"$\beta$ = {beta}", linewidth=1.1)

    ax.axvline(x=0, color=PALETTE["muted"], linestyle="--", alpha=0.7)
    ax.set_xlabel("(Chosen log ratio) - (Rejected log ratio)")
    ax.set_ylabel("DPO Loss")
    ax.set_title("Effect of β on DPO Loss")
    ax.set_xlim(x_range[0], x_range[1])
    ax.set_ylim(0, 4)
    polish_axes(ax)
    handles, labels = ax.get_legend_handles_labels()
    fig.legend(
        handles,
        labels,
        loc="lower center",
        bbox_to_anchor=(0.5, 0.025),
        ncol=4,
        frameon=True,
        facecolor=PALETTE["paper"],
        edgecolor="none",
        framealpha=0.92,
    )
    fig.subplots_adjust(left=0.13, right=0.97, bottom=0.31, top=0.88)

    return fig
Out[27]:
Visualization
Multiple loss curves showing how beta parameter affects gradient steepness in DPO training.
DPO loss curves for varying beta values. Higher beta settings (e.g., 0.5) produce steeper gradients near the decision boundary (margin = 0), resulting in stronger penalties for small deviations compared to lower values like 0.1.

The original DPO paper found β=0.1\beta = 0.1 to work well across a variety of tasks. This value provides a reasonable balance between learning preferences and maintaining coherence. The visualization reveals why: at β=0.1\beta = 0.1, the loss curve has a moderate slope that provides clear gradient signal without being so steep that small changes in log ratios cause dramatic loss changes.

Learning Rate

DPO typically requires smaller learning rates than supervised fine-tuning. As a practical guide:

The smaller rates are necessary because DPO directly optimizes log probability ratios, which can change rapidly with small weight updates. Unlike supervised fine-tuning where we are simply maximizing the likelihood of target tokens, DPO computes a ratio between two model evaluations. Small changes to the policy model affect both the numerator and denominator of this ratio, potentially causing the implicit reward to shift dramatically if the learning rate is too high. The amplification factor is roughly proportional to the inverse of the probability being changed: rare tokens see larger log probability swings per weight update than common tokens.

In[28]:
Code
# Recommended hyperparameter ranges
hyperparameters = {
    "beta": {
        "range": "0.05 - 0.5",
        "typical": "0.1",
        "notes": "Higher = more conservative",
    },
    "learning_rate": {
        "range": "1e-6 - 5e-5",
        "typical": "5e-7 to 1e-5",
        "notes": "Lower than SFT",
    },
    "batch_size": {
        "range": "4 - 64",
        "typical": "16-32",
        "notes": "Larger batches stabilize training",
    },
    "epochs": {
        "range": "1 - 5",
        "typical": "1-3",
        "notes": "DPO converges quickly",
    },
    "warmup_ratio": {
        "range": "0.05 - 0.1",
        "typical": "0.1",
        "notes": "Gradual learning rate increase",
    },
}
Out[29]:
Console
DPO Hyperparameter Guidelines:
----------------------------------------------------------------------
beta                 | Range: 0.05 - 0.5      | Typical: 0.1
                     | Higher = more conservative

learning_rate        | Range: 1e-6 - 5e-5     | Typical: 5e-7 to 1e-5
                     | Lower than SFT

batch_size           | Range: 4 - 64          | Typical: 16-32
                     | Larger batches stabilize training

epochs               | Range: 1 - 5           | Typical: 1-3
                     | DPO converges quickly

warmup_ratio         | Range: 0.05 - 0.1      | Typical: 0.1
                     | Gradual learning rate increase

These guidelines provide a starting point for tuning. The learning rate is particularly necessary; starting too high often destabilizes the implicit reward formulation, leading to poor convergence. A common debugging strategy is to start with a very low learning rate (1e-7) and progressively increase it while monitoring the reward margin. The highest stable learning rate is usually the best choice for convergence speed.

Out[30]:
Visualization
Four descending gradient-magnitude curves over reward margin, with higher beta curves above lower beta curves and a dashed vertical decision boundary at zero.
Gradient magnitude profiles across different beta values. Gradients are largest when the model ranks the rejected response above the chosen response (negative margin) and decrease as the chosen response gains a positive margin. Higher beta values produce stronger updates throughout, while the dashed line marks the decision boundary at margin zero.

Batch Size and Gradient Accumulation

DPO benefits from larger effective batch sizes because the loss depends on comparing policy and reference log probability ratios. Small batches introduce variance in these estimates, which can make training noisy and unstable. With a larger batch, the average over multiple preference pairs provides a more reliable gradient signal. The intuition is the same as with any stochastic gradient method: more samples per gradient estimate means less noise and more stable convergence.

The practical complication is that each DPO training example consists of two sequences (chosen and rejected), each of which must be processed through two models (policy and reference). The effective memory cost per training step is therefore four times the cost of a single forward pass, which limits how large a batch you can fit in memory. Gradient accumulation allows you to simulate large batches by accumulating gradients over multiple small batches before updating weights.

If GPU memory is limited, use gradient accumulation:

In[31]:
Code
# Effective batch size = batch_size * gradient_accumulation_steps
# Example: batch_size=4, accumulation=8 → effective batch = 32

training_config = {
    "per_device_batch_size": 4,
    "gradient_accumulation_steps": 8,
    "effective_batch_size": 4 * 8,  # = 32
}

Number of Epochs

DPO typically converges faster than you might expect. One to three epochs over the preference data is usually sufficient. The reason DPO converges quickly is that it is solving a relatively simple optimization problem: given the reference model's probabilities as anchors, adjust the policy to correctly rank the provided preference pairs. Once the model has seen each preference pair enough times to internalize the comparison, additional training only leads to overfitting.

Overfitting is a real concern with DPO. Signs of overfitting include:

  • Training accuracy approaches 100% while validation loss increases
  • Generated text becomes repetitive or templated
  • The model loses diversity in its responses

Using a held-out validation set to monitor these metrics during training is strongly recommended. Early stopping based on validation reward margin or loss is a practical necessity for production runs rather than an optional precaution.

Production Considerations

Moving a DPO implementation from toy demo to production requires additional engineering. The code we have written is pedagogically clear but optimized for understanding rather than efficiency. Production DPO training needs to handle larger models, larger datasets, mixed precision arithmetic, distributed training across multiple GPUs, and reliable evaluation infrastructure.

The most important production optimization is memory efficiency. At scale, maintaining two model copies in memory is often infeasible, which makes LoRA-based DPO the standard approach for models with billions of parameters. The LoRA approach also has a practical advantage beyond memory: because only the adapter weights change, the reference computation is automatically correct as long as the base model weights are not modified.

Memory-Efficient Implementation

For large models, computing forward passes for both policy and reference models strains memory. The TRL library from Hugging Face provides optimized implementations that handle gradient checkpointing, mixed-precision training, and LoRA-based reference models transparently:

In[32]:
Code
# Using TRL for production DPO training (pseudocode - requires installation)
"""
from trl import DPOTrainer, DPOConfig
from transformers import AutoModelForCausalLM, AutoTokenizer

# Configuration
dpo_config = DPOConfig(
    beta=0.1,
    learning_rate=5e-7,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=8,
    num_train_epochs=1,
    warmup_ratio=0.1,
    bf16=True,  # Mixed precision for efficiency
    gradient_checkpointing=True,  # Memory optimization
)

# Initialize trainer
trainer = DPOTrainer(
    model=policy_model,
    ref_model=reference_model,  # or None if using LoRA
    args=dpo_config,
    train_dataset=preference_dataset,
    tokenizer=tokenizer,
)

# Train
trainer.train()
"""

The TRL DPOTrainer handles many production details automatically, including proper handling of the LoRA reference model, mixed precision loss scaling, and logging to Weights and Biases. For teams deploying DPO in production, using TRL is strongly recommended over maintaining a custom training loop.

Evaluation During Training

Beyond loss and accuracy, monitoring generation quality is needed during training. Quantitative metrics can be misleading: a model that achieves high training accuracy may still generate poor outputs if it has overfit to specific patterns in the preference data. Periodic qualitative evaluation, where you read the model's outputs on held-out prompts, provides signal that no scalar metric can capture.

In[33]:
Code
def evaluate_generation_quality(
    model, tokenizer, eval_prompts, max_new_tokens=100
):
    """Generate responses to fixed prompts for qualitative evaluation."""
    model.eval()
    generations = []

    for prompt in eval_prompts:
        inputs = tokenizer(prompt, return_tensors="pt")
        inputs = {k: v.to(model.device) for k, v in inputs.items()}

        with torch.no_grad():
            outputs = model.generate(
                **inputs,
                max_new_tokens=max_new_tokens,
                do_sample=True,
                temperature=0.7,
                top_p=0.9,
                pad_token_id=tokenizer.pad_token_id,
            )

        response = tokenizer.decode(outputs[0], skip_special_tokens=True)
        generations.append({"prompt": prompt, "response": response})

    return generations

Periodically generating responses to held-out prompts and reviewing them manually (or with an LLM judge) provides important signal about whether DPO is achieving its intended effect. A standard practice is to run generation evaluation every 100-500 training steps and compare outputs across checkpoints. If outputs become noticeably more uniform or begin exhibiting characteristic phrases that appear frequently in the chosen responses, this is a strong signal that training should stop.

Limitations and Practical Impact

DPO has made preference-based alignment more accessible by eliminating the complexity of reward model training and reinforcement learning. However, implementation challenges remain, and understanding them is needed for practitioners who want to use DPO reliably rather than just get it to run.

Data Quality Dependencies

DPO is only as good as your preference data. Unlike RLHF where a reward model can generalize learned preferences to new situations it has never explicitly seen, DPO directly optimizes for the specific comparisons in your training set. If the training data contains noisy labels, biased annotators, or prompts that do not cover the distribution you care about, these problems will flow directly into the aligned model with no filtering mechanism. The reward model in RLHF provides a layer of abstraction between the raw annotation data and the policy optimization, which can smooth over individual annotation errors. DPO has no such buffer.

In practice, data curation often matters more than algorithmic improvements. Teams investing in DPO should allocate significant effort to preference data quality: establishing clear and specific annotation guidelines, measuring inter-annotator agreement before starting large-scale annotation, filtering examples where annotators disagree, and auditing chosen responses to ensure they represent the behavior you want to encourage. A dataset of 5,000 high-quality preference pairs will typically outperform a dataset of 50,000 noisy ones, because the signal-to-noise ratio directly determines what the model learns.

Distribution Mismatch and On-Policy Training

One of the most practically important limitations of DPO is that it trains on off-policy data: the chosen and rejected responses were generated by some earlier model (often a supervised fine-tuned baseline), not by the model currently being trained. As the policy diverges from the model that generated the training data, the preference comparisons become less informative because the policy is making decisions in regions of probability space that were never covered by the training data.

This mismatch can manifest as training curves that look good on the preference data but produce models that behave unexpectedly at inference time. The policy has learned to rank responses that a different model might generate, but when asked to generate responses itself, it may produce outputs that look nothing like either the chosen or rejected examples it was trained on. Iterative DPO, where you periodically regenerate preference data using the current policy, addresses this mismatch but at the cost of the very simplicity that makes DPO appealing.

Mode Collapse Risks

With very small β\beta values or prolonged training, DPO can collapse toward generating only responses that are very similar to the chosen examples in the training data. This manifests as reduced response diversity and overfitting to surface patterns in preferred responses rather than learning the underlying preference criteria. Think of mode collapse in DPO as the model becoming a photocopier for chosen responses: it stops generating diverse outputs and instead produces slight variations on the patterns it saw in the training set.

Monitoring response diversity during training and using techniques like early stopping or label smoothing can mitigate this risk. Some practitioners also mix DPO training with a small fraction of continued language modeling loss to maintain general capabilities. A typical mixing ratio is 90% DPO loss and 10% causal language modeling loss on the chosen responses, which provides a regularization signal that prevents the model from forgetting how to generate coherent text while still optimizing the preference objective.

Scaling Challenges

As models grow larger, the memory requirements for maintaining both policy and reference models become significant. Even with LoRA, computing forward passes through large models twice adds computational overhead. At the scale of 70-billion parameter models, DPO training with full precision would require approximately 140 GB just to hold the model weights, before accounting for gradients, optimizer states, and activations. The combination of mixed precision training, gradient checkpointing, and LoRA adapters makes this feasible, but these optimizations require careful tuning to maintain training stability.

The DPO Variants chapter covers techniques like reference-free DPO (SimPO) and identity-preserving DPO that address some of these challenges by either eliminating the reference model entirely or computing it more efficiently. These variants trade off some alignment quality for better scalability and are worth exploring for very large model training scenarios.

Despite these limitations, DPO has become the preferred alignment method for many teams due to its simplicity and effectiveness. Aligning models with preference data using supervised learning infrastructure makes alignment research more accessible to teams without specialized reinforcement learning expertise. The ability to train on a standard GPU cluster with standard supervised learning tools, while achieving competitive alignment quality, has democratized alignment work across the research community.

Key Parameters

The key parameters for DPO training are:

  • beta: Controls the strength of the KL divergence constraint (typically 0.1). Higher values keep the policy closer to the reference model and reduce overfitting risk. Lower values allow faster learning but increase the risk of mode collapse.
  • learning_rate: The step size for optimization (typically 5e-7 to 1e-5). DPO requires lower rates than standard supervised fine-tuning because the implicit reward is sensitive to log probability ratio changes.
  • batch_size: The number of samples processed per step. Larger batches generally stabilize the loss estimate. An effective batch size of 16-32 preference pairs is a reasonable starting point.
  • gradient_accumulation_steps: Number of micro-batches to accumulate before updating weights, used to simulate large batch training under memory constraints.
  • num_epochs: Number of passes through the preference dataset (typically 1-3). DPO converges quickly and overfits with too many epochs.
  • warmup_ratio: Fraction of training steps used for learning rate warmup (typically 0.05-0.1). Warmup prevents early training instability when the optimizer first encounters large gradients.

Summary

This chapter translated DPO theory into practice, covering every implementation component from data preparation through production considerations. The key takeaways are as follows.

Data format: DPO requires triplets of (prompt, chosen response, rejected response), with careful tokenization to identify which tokens belong to the response versus the prompt. The response mask is not an optional optimization: it is structurally needed to computing the correct loss. Without it, the loss is computed over prompt tokens that are identical for both chosen and rejected responses. This provides no useful preference signal.

Loss computation: The DPO loss compares log probability ratios between the policy and reference model for chosen versus rejected responses. Proper handling of the shift operation, sequence masking, and numerical stability choices like F.logsigmoid all matter for correctness. The log ratio formulation naturally cancels length effects and keeps values in a numerically stable range.

Training procedure: DPO training requires maintaining a frozen reference model, either as a separate copy (simple but memory-intensive) or implicitly through LoRA (memory-efficient and production-ready). Standard supervised learning infrastructure handles the rest: AdamW optimizer, gradient accumulation, and learning rate scheduling.

Hyperparameters: The β\beta parameter controls the KL constraint strength, with typical values around 0.1. Learning rates should be lower than standard fine-tuning to prevent destabilizing the implicit reward formulation. DPO often converges in just one to three epochs, making early stopping based on validation metrics a practical necessity rather than an optional precaution.

Limitations: DPO's simplicity comes with real tradeoffs. It is entirely dependent on preference data quality, vulnerable to mode collapse under aggressive training, and subject to distribution mismatch as the policy diverges from the data-generating model. Understanding these limitations is needed for knowing when DPO will succeed and when it will require additional strategies like iterative regeneration or mixing with language modeling loss.

With these components in place, you can align language models to human preferences without the complexity of reward modeling and reinforcement learning that RLHF requires. The next chapter explores variants of DPO that address specific limitations, including methods that eliminate the need for a reference model entirely.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about DPO implementation.

DPO Implementation

Question 1 of 70 of 7 completed
What is the required data format for DPO training?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026dpoimplementation, author = {Michael Brenndoerfer}, title = {DPO Implementation: PyTorch Training for LLM Alignment}, year = {2026}, url = {https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). DPO Implementation: PyTorch Training for LLM Alignment. Retrieved from https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training
MLAAcademic
Michael Brenndoerfer. "DPO Implementation: PyTorch Training for LLM Alignment." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training>.
CHICAGOAcademic
Michael Brenndoerfer. "DPO Implementation: PyTorch Training for LLM Alignment." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training.
HARVARDAcademic
Michael Brenndoerfer (2026) 'DPO Implementation: PyTorch Training for LLM Alignment'. Available at: https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). DPO Implementation: PyTorch Training for LLM Alignment. https://mbrenndoerfer.com/writing/dpo-implementation-pytorch-preference-optimization-training

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.