RLHF Pipeline: SFT, Reward Models, and PPO

Michael BrenndoerferDecember 29, 202561 min read

Part of Language AI Handbook

Covers RLHF pipeline with three stages: Supervised Fine-Tuning, Reward Model training, and PPO optimization. Topics include debugging techniques.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

RLHF Pipeline

Reinforcement Learning from Human Feedback turns a base language model into an assistant that follows instructions and gives useful, truthful responses without causing harm. Building a capable language model through pretraining is only half the story. A model trained purely to predict the next token has learned to be a generalist text predictor, not a helpful collaborator. It will complete sentences, continue stories, and extend documents, but it has no concept of what "helpful" means, no understanding of what a user wants, and no instinct to avoid harmful outputs. Bridging that gap requires a fundamentally different kind of training signal, one that originates not from next-token prediction but from human judgment about what constitutes a good response.

The previous chapters introduced the individual components: the Bradley-Terry model for preference modeling, reward model architecture and training, and the PPO algorithm adapted for language models. Now we assemble these pieces into a complete training pipeline. The assembly matters as much as any individual component. Each stage creates preconditions that the next stage depends on, and misunderstanding those dependencies is one of the most common sources of failure in real RLHF implementations. A reward model trained on the wrong data distribution, or PPO optimization started from a poorly prepared initial policy, can derail the entire alignment effort despite each component working correctly in isolation.

Think of the RLHF pipeline as a three-act apprenticeship. In the first act, the apprentice studies examples of excellent work from master craftspeople, learning the expected format and quality standard. In the second act, an experienced judge observes many examples and builds an intuition for what separates excellent work from mediocre work, capturing that judgment in a reproducible scoring mechanism. In the third act, the apprentice iterates, generating work, receiving scores, and adjusting behavior to earn higher marks, all while staying grounded enough in the original training to remain coherent. The magic is that this apprenticeship produces a model whose behavior reflects human preferences even on prompts it has never seen before, generalizing a signal from thousands of comparisons into millions of novel situations.

The RLHF pipeline consists of three sequential stages: Supervised Fine-Tuning (SFT), Reward Model (RM) training, and PPO optimization. Each stage builds on the previous one, progressively shaping the model's behavior. This chapter walks through the complete pipeline, examining the design decisions at each stage, the hyperparameters that govern training stability, and the debugging techniques needed for successful alignment. We will also develop intuition for the failure modes that make RLHF notoriously difficult to get right in practice, building toward a mental model that helps you diagnose problems when they arise.

Understanding the complete pipeline also illuminates the motivation for alternatives like Direct Preference Optimization, which we will encounter in upcoming chapters. Many of the simplifications that DPO introduces are direct responses to known weaknesses in this three-stage RLHF pipeline, and appreciating those weaknesses helps you understand how DPO works and why it was designed the way it was.

Historical Context

The RLHF pipeline as we know it today was crystallized in the InstructGPT paper (Ouyang et al., 2022), which demonstrated for the first time that this three-stage procedure could produce models that humans strongly preferred over much larger models trained with standard supervised objectives. The underlying insight, that human feedback could serve as a training signal through a learned reward model rather than being applied directly, drew on decades of earlier work in reinforcement learning from human feedback in robotics (Christiano et al., 2017) and earlier experiments applying preference learning to simpler language tasks. The GPT-3 era provided the important ingredient: a base model powerful enough to produce responses diverse enough for meaningful human comparisons. Before sufficiently capable base models existed, the preference signal was too noisy to train a useful reward model. InstructGPT showed that the combination of scale and the three-stage pipeline could produce qualitative improvements that dwarfed the effect of model size alone: the 1.3B InstructGPT model was preferred by labelers over the 175B GPT-3 model, a finding that reshaped the entire field's thinking about the relationship between model size and usefulness.

The Three-Stage Pipeline

The RLHF pipeline follows a specific sequence that turns a pretrained language model into an aligned assistant. Understanding why each stage exists and how they connect is important for successful implementation. Each stage has a distinct purpose, and the stages form a dependency chain where the quality of each stage's output places a ceiling on everything that follows.

The pipeline is a technical recipe built on assumptions about how alignment works and what data is available. The stages are ordered the way they are because of a specific logic: you cannot collect useful preference data without responses to compare, you cannot train a reward model without preference data, and you cannot run stable PPO without a reward model. The forward arrow of causality runs through the entire pipeline, and this is why short-circuiting any stage tends to undermine subsequent ones even when the skipped stage seems redundant.

A second important dimension is that all three models in the pipeline, the SFT model, the reward model, and the PPO policy, typically share the same underlying architecture. They are all transformer language models of roughly the same parameter count, often initialized from the same pretrained weights. This shared foundation means that improvements in base model capability propagate through all three stages simultaneously. When OpenAI upgraded from GPT-3 to GPT-4 as the base, the improvements benefited every stage of the alignment pipeline, including the final PPO policy. The alignment procedure is more like a specialization of capabilities that already exist than an injection of entirely new capabilities.

Out[3]:
Visualization
Flow diagram showing a pretrained model and demonstrations entering SFT, the SFT model and preference data entering reward modeling, and the reward model entering PPO to produce an aligned model.
The three-stage RLHF pipeline illustrating the transformation of a pretrained model into an aligned assistant. This process sequentially utilizes supervised demonstrations, reward modeling for preference learning, and PPO optimization to maximize helpfulness while maintaining linguistic coherence.

Stage 1: Supervised Fine-Tuning (SFT) takes a pretrained language model and trains it on high-quality demonstrations of desired behavior. This stage teaches the model the format and style of helpful responses, creating a starting point that can already follow instructions reasonably well.

Stage 2: Reward Model Training creates a model that predicts human preferences. Using comparison data where humans ranked alternative responses, the reward model learns to assign scalar scores which reflects response quality. As we covered in the Reward Modeling chapter, this model provides the optimization signal for the final stage.

Stage 3: PPO Fine-Tuning optimizes the SFT model to maximize rewards from the reward model while staying close to its original behavior. Building on the PPO for Language Models chapter, this stage performs the actual alignment through reinforcement learning.

Why the Order Cannot Be Reversed

The sequential nature of the pipeline is not arbitrary, and understanding the causal logic makes it far easier to reason about what can go wrong. SFT must precede reward model training because the reward model needs realistic candidate responses to compare. If you attempted to collect preference data from a raw pretrained model, the responses would be so incoherent and off-format that human annotators would struggle to determine which was "better" in any meaningful sense. Both responses might simply be unusable, making the comparison uninformative.

Reward model training must precede PPO because PPO requires a dense, differentiable reward signal. The alternative, having humans rate every generated response during PPO training, would be impractically slow. PPO generates millions of samples during training; each requiring a human rating would make the procedure orders of magnitude too expensive. The reward model acts as a learned proxy that can be evaluated in milliseconds rather than hours.

PPO must come last because it requires a stable starting point and a clear optimization target. Without the SFT model as initialization, the policy gradient updates would struggle to produce coherent language. Without the reward model as the optimization target, there is no well-defined objective to optimize toward. The PPO stage consumes all the infrastructure built in the previous two stages and converts it into the final aligned model.

Stage 1: Supervised Fine-Tuning

Supervised Fine-Tuning turns a pretrained language model into one able to following instructions and engaging in dialogue. While the pretrained model has acquired extensive knowledge and language understanding, it lacks the ability to respond helpfully to your queries in a conversational format. SFT closes this gap by exposing the model to paired examples of instructions and ideal responses, teaching it both the expected interaction pattern and the quality bar it should aim for.

Think of SFT as teaching an incredibly knowledgeable but socially awkward expert how to communicate. Before SFT, the language model has read essentially all of human-written text and internalized an enormous amount of knowledge. But it communicates the way text on the internet communicates, sometimes as a forum post, sometimes as a Wikipedia article, sometimes as a fictional story, depending on what pattern the prompt activates. After SFT, it has learned that when a human poses a question, the appropriate response is a clear, structured, helpful answer, not a continuation of the question into a series of related questions, not a fictional dialogue where characters discuss the topic, and not a Wikipedia-style article about the broader subject area. The format is part of the skill.

The SFT stage is also where the model develops an initial grasp of what "helpful" looks like in concrete terms. Through exposure to carefully crafted demonstrations, it learns that a good response is appropriately scoped to the question, acknowledges limitations when relevant, structures complex information clearly, and avoids unnecessary caveats or padding. These are difficult qualities to specify algorithmically, but relatively easy for expert human writers to demonstrate. SFT exploits that asymmetry, using human demonstrations as a richer signal than any reward function could capture.

Why SFT Comes First

You might wonder why we don't skip directly to reinforcement learning. The reason is practical: PPO optimization requires a model that already produces reasonable responses. Trying to optimize a raw pretrained model with RL is like trying to teach someone chess strategy before they know how the pieces move.

The pretrained model generates plausible continuations of text, but it doesn't understand the assistant paradigm. Given a question, it might generate more questions, continue with a different topic, or produce text that reads like training data rather than a helpful response. SFT provides the foundation by:

  • Teaching the instruction-following format (recognizing prompts, generating responses)
  • Establishing a baseline quality level that RL can refine
  • Reducing the search space for PPO (responses are already in a useful format)

There is also a subtler reason that SFT matters for reward model training. The preference data collected for stage 2 requires the SFT model to generate candidate responses. If those candidates are incoherent, the human annotators comparing them face an impossible task: they cannot distinguish useful quality differences between incoherent responses. High-quality SFT outputs ensure that the preference data captures differences in response quality rather than the contrast between merely incoherent and completely incoherent. The SFT model sets the quality floor for every subsequent stage.

SFT Data Requirements

SFT data consists of (prompt, response) pairs showing ideal assistant behavior. Quality matters far more than quantity. A few thousand high-quality demonstrations often outperform millions of lower-quality examples.

In[4]:
Code
sft_examples = [
    {
        "prompt": "What causes rainbows to form?",
        "response": """Rainbows form through a combination of refraction, reflection, and dispersion of sunlight in water droplets. Here's how it works:

1. **Refraction**: Sunlight enters a water droplet and bends (refracts) because light travels slower in water than in air.

2. **Dispersion**: Different wavelengths of light bend by different amounts, separating white light into its component colors (red bends least, violet bends most).

3. **Internal Reflection**: The light reflects off the back inner surface of the droplet.

4. **Second Refraction**: The light bends again as it exits the droplet, further separating the colors.

You see a rainbow when the sun is behind you and rain is in front of you, with the light which reflects back at approximately 42 degrees from the original light direction.""",
    },
    {
        "prompt": "Write a short poem about coding",
        "response": """Lines of Logic

In the glow of screens at night,
We craft our worlds in black and white,
Variables dance, functions call,
Building dreams that never fall.

Bugs may lurk in shadowed code,
Each one a puzzle to decode,
But when at last the program runs,
We've made magic, byte by byte, we've won.""",
    },
]
Out[5]:
Console
SFT Training Example:
Prompt: What causes rainbows to form?

Response preview: Rainbows form through a combination of refraction, reflection, and dispersion of sunlight in water droplets. Here's how it works:

1. **Refraction**: Sunlight enters a water droplet and bends (refract...

The demonstrations should exhibit properties you want the final model to have: helpfulness, appropriate tone, factual accuracy, and safety awareness. As discussed in the Instruction Data Creation chapter, these examples can come from human writers, filtered model outputs, or synthetic generation with quality controls.

What makes a good SFT dataset goes beyond having correct answers. The demonstrations need to model the right relationship between question complexity and response depth. A simple factual question deserves a concise answer; a complex reasoning question deserves a step-by-step explanation. Demonstrations that always use the same format, regardless of question type, teach the model the wrong thing: they teach it to mimic a style rather than to calibrate the response to the question. Building this calibration into the demonstration data is one of the craft elements that separates effective SFT datasets from mediocre ones.

Coverage breadth is equally important. The SFT model's capabilities are strongly constrained by the distribution of topics and task types in the training data. A model trained only on question-answering demonstrations will struggle with creative writing requests. A model trained only on English examples will struggle with multilingual queries. Systematic gaps in the SFT data create systematic gaps in the model's behavior that downstream RL training can partially compensate for but never fully overcome. This is why organizations building production alignment pipelines invest heavily in dataset curation: so diversity across domains, styles, difficulty levels, and instruction types.

SFT Training Process

SFT uses standard causal language modeling loss but only on the response tokens. This distinction is important for understanding how the model learns during this stage. The prompt provides context, letting the model to understand what kind of response is expected, but we don't penalize the model for not predicting prompt tokens. The reasoning is straightforward: we only want the model to learn how to respond given a prompt, not to memorize and reproduce the prompts themselves.

The training objective focuses exclusively on maximizing the likelihood of generating the correct response tokens, conditioned on the full context of the prompt. This targeted learning ensures the model develops the skill of creating appropriate responses rather than simply learning to continue arbitrary text.

Out[6]:
Visualization
Visualization showing token sequence with prompt tokens masked and response tokens adding to loss.
Response-only loss masking in SFT. Prompt tokens (gray) provide context but are masked in the loss calculation (-100). Only response tokens (green) contribute to gradients, so the model learns to generate the completion rather than recreating the prompt.
In[7]:
Code
import torch
from torch.utils.data import Dataset


class SFTDataset(Dataset):
    """Dataset for supervised fine-tuning with response-only loss."""

    def __init__(self, examples, tokenizer, max_length=512):
        self.examples = examples
        self.tokenizer = tokenizer
        self.max_length = max_length

    def __len__(self):
        return len(self.examples)

    def __getitem__(self, idx):
        example = self.examples[idx]

        prompt_tokens = self.tokenizer.encode(example["prompt"])
        response_tokens = self.tokenizer.encode(example["response"])

        full_tokens = prompt_tokens + response_tokens

        if len(full_tokens) > self.max_length:
            full_tokens = full_tokens[: self.max_length]
            response_start = len(prompt_tokens)
        else:
            response_start = len(prompt_tokens)

        labels = [-100] * response_start + full_tokens[response_start:]

        padding_length = self.max_length - len(full_tokens)
        input_ids = full_tokens + [self.tokenizer.pad_token_id] * padding_length
        labels = labels + [-100] * padding_length

        return {
            "input_ids": torch.tensor(input_ids),
            "labels": torch.tensor(labels),
            "attention_mask": torch.tensor(
                [1] * len(full_tokens) + [0] * padding_length
            ),
        }

The key detail is setting labels to -100 for prompt tokens. This value is a special sentinel in PyTorch's CrossEntropyLoss function: any token position with a label of -100 is completely ignored during loss computation. By masking the prompt tokens this way, we ensure we only compute loss on the response portion, teaching the model to generate good responses without wasting gradient updates on predicting prompt content it doesn't need to reproduce.

In[8]:
Code
def sft_training_step(model, batch, optimizer):
    """Single SFT training step with response-only loss."""
    model.train()

    input_ids = batch["input_ids"]
    labels = batch["labels"]
    attention_mask = batch["attention_mask"]

    outputs = model(
        input_ids=input_ids,
        attention_mask=attention_mask,
        labels=labels,  # Model computes loss internally, ignoring -100 labels
    )

    loss = outputs.loss

    optimizer.zero_grad()
    loss.backward()
    optimizer.step()

    return loss.item()

Key Parameters

SFT training is relatively straightforward compared to the later stages. The key parameters are:

  • Learning rate: Typically 1e-5 to 5e-5 for full fine-tuning, or 1e-4 to 3e-4 for LoRA (as covered in the LoRA chapters)
  • Batch size: 32 to 128 examples, depending on available memory
  • Epochs: 1-3 passes over the data; more risks overfitting
  • Warmup: 3-10% of total steps with linear warmup
In[9]:
Code
sft_config = {
    "learning_rate": 2e-5,
    "batch_size": 64,
    "num_epochs": 2,
    "warmup_ratio": 0.03,
    "weight_decay": 0.01,
    "max_grad_norm": 1.0,
    "lr_scheduler": "cosine",
}
Out[10]:
Console
Typical SFT Configuration:
  learning_rate: 2e-05
  batch_size: 64
  num_epochs: 2
  warmup_ratio: 0.03
  weight_decay: 0.01
  max_grad_norm: 1.0
  lr_scheduler: cosine

Overfitting is a real concern with small SFT datasets. Monitor validation loss and stop training when it begins to increase. You might prefer training for slightly fewer steps than optimal to preserve model generalization.

The learning rate choice deserves particular attention. SFT is fine-tuning a model that has already learned enormous amounts from pretraining. Using too large a learning rate risks catastrophically forgetting pretraining knowledge, a problem where the model loses general capabilities in exchange for fitting the specific demonstrations. Using too small a learning rate means the model barely moves from its pretrained behavior, failing to learn the instruction-following format effectively. The 1e-5 to 5e-5 range for full fine-tuning represents decades of empirical wisdom about this trade-off. For LoRA fine-tuning, the larger learning rates (1e-4 to 3e-4) are appropriate because LoRA only updates a small fraction of parameters, so the effective learning rate on the full model is much smaller than the nominal value.

The number of epochs also warrants thought beyond "stop when validation loss increases." With SFT, the concern is overfitting in the traditional sense and sycophancy: the model learning to mimic the surface features of the demonstrations (bullet points, specific phrasing, characteristic openers) rather than the underlying quality. If you train for many epochs on a small dataset, the model may learn to reproduce specific demonstration responses almost verbatim rather than generalizing the quality signal. One to three epochs is the typical operating range precisely because it provides enough exposure to learn the patterns without enough repetition to memorize specific examples.

What Good SFT Looks Like

A well-trained SFT model should exhibit several distinguishing characteristics when you evaluate it qualitatively. When given a straightforward question, it should provide a direct and accurate answer without excessive preamble. When given a complex multi-step problem, it should structure its response with clear organization rather than presenting a wall of text. When given a request that falls outside appropriate assistant behavior, it should decline gracefully rather than either complying blindly or refusing with an unhelpful non-response.

The key insight is that SFT teaches the model a mode of operation rather than a set of facts. You can check whether SFT is working by observing whether the model consistently adopts the assistant persona across diverse prompt types. If the model responds to some prompts helpfully but reverts to raw language model behavior on others (generating continuations, creating multiple responses, or adopting the voice of a different persona), SFT has not been thorough enough. This inconsistency is usually a sign that either the demonstration data lacks coverage for certain prompt types or training ran for too few steps.

Stage 2: Reward Model Training

The reward model learns to predict which responses humans prefer. Building on the Bradley-Terry model from earlier chapters, it converts pairwise comparisons into a scalar reward signal that guides PPO optimization. This stage is arguably the most consequential in the entire pipeline, because any systematic error in the reward model will be amplified by PPO optimization. The policy will find the ways to maximize the reward model score, and if the reward model is wrong about what constitutes a good response in certain regions of the input space, PPO will drive the policy directly into those regions.

Think of the reward model as a compressed representation of annotator judgment. Thousands of pairwise comparisons, each representing a human's opinion about which of two responses is better, are distilled into a single neural network that can assign a quality score to any response in fractions of a second. This compression is both the strength and the weakness of the approach. The strength is that the reward model can generalize beyond the specific comparisons it was trained on. This provides quality judgments for entirely novel responses. The weakness is that the generalization may be imperfect, and the compression inevitably loses some of the nuance that human annotators would apply in novel situations.

The point is that the reward model is not trained to produce absolute quality scores, only to produce scores that correctly rank pairs of responses. Two reward models trained on the same data could assign completely different absolute values and yet behave identically in practice, because PPO optimization only cares about relative rankings, not absolute magnitudes. This property changes the interpretation for how you should interpret reward values and how you calibrate the KL penalty coefficient in stage 3.

Reward Model Architecture

As discussed in the Reward Modeling chapter, the reward model typically shares architecture with the language model but replaces the language modeling head with a scalar output head. This architectural choice is deliberate: by starting from a language model that understands text in context, the reward model inherits much of its ability to interpret fine-grained meaning. The only modification is the final layer, which now outputs a single number representing quality rather than a distribution over vocabulary tokens.

In[11]:
Code
import torch
import torch.nn as nn


class RewardModel(nn.Module):
    """Reward model that outputs a scalar score for text quality."""

    def __init__(self, base_model, hidden_size):
        super().__init__()
        self.base_model = base_model

        self.reward_head = nn.Sequential(
            nn.Linear(hidden_size, hidden_size),
            nn.ReLU(),
            nn.Linear(hidden_size, 1),
        )

    def forward(self, input_ids, attention_mask):
        outputs = self.base_model(
            input_ids=input_ids,
            attention_mask=attention_mask,
            output_hidden_states=True,
        )

        last_hidden_state = outputs.hidden_states[-1]

        sequence_lengths = attention_mask.sum(dim=1) - 1
        batch_indices = torch.arange(input_ids.size(0), device=input_ids.device)
        final_hidden = last_hidden_state[batch_indices, sequence_lengths]

        reward = self.reward_head(final_hidden).squeeze(-1)
        return reward

The reward model extracts the hidden state at the final token position and passes it through a small neural network to produce a scalar reward. This design uses an important property of causal transformers: the final token's hidden state has attended to all previous tokens in the sequence, meaning it encapsulates information about the entire prompt-response pair. From this aggregated representation, the model can make an overall quality judgment that considers both what was asked and how well the response addresses it.

An important practical detail is that the reward model is initialized from the SFT model rather than the raw pretrained model. This initialization matters because the reward model needs to understand the specific response style established during SFT. If it were initialized from the pretrained model, it might assign high scores to responses that look like pretraining data (encyclopedia entries, forum posts) rather than assistant responses. Starting from the SFT model ensures the reward model operates in the same distribution as the responses it will be asked to score during PPO training.

The two-layer reward head sits on top of the transformer backbone. The first linear layer allows the network to learn a projection of the final hidden state that is maximally informative for quality prediction. The ReLU nonlinearity allows the network to model asymmetric quality differences. The second linear layer collapses this representation to a single scalar. Training this relatively small head while keeping the backbone frozen is one option, but joint fine-tuning of the entire model typically produces better reward models because it allows the transformer layers to reorganize their representations to make quality prediction easier.

Preference Data Format

Training data consists of prompts with two or more responses ranked by human annotators. Each training example captures a comparison: given the same prompt, which response did humans consider better? This pairwise structure is basic to the Bradley-Terry model, which we'll examine shortly.

In[12]:
Code
preference_data = [
    {
        "prompt": "How do I improve my writing skills?",
        "chosen": """Here are effective strategies to improve your writing:

1. **Read widely and actively** - Study how skilled authors structure sentences and arguments
2. **Write daily** - Even 15 minutes of practice builds consistency
3. **Seek feedback** - Join a writing group or find a trusted reader
4. **Edit ruthlessly** - Good writing is rewriting; cut unnecessary words
5. **Study grammar** - Understanding rules helps you break them purposefully

Start with one habit and build from there.""",
        "rejected": """Writing is hard but you can get better. Just practice a lot and read books. 
Maybe take a class or something. Good luck!""",
    },
    {
        "prompt": "Explain quantum entanglement simply",
        "chosen": """Quantum entanglement is like having two magic coins. When you flip them while they're together, then separate them, even across the universe, they always land the same way.

In physics terms: two particles become correlated so that measuring one instantly affects what you'll measure on the other, regardless of distance. Einstein called this "spooky action at a distance."

This isn't about sending information faster than light, but about correlations that can't be explained by classical physics.""",
        "rejected": """Quantum entanglement is a quantum mechanical phenomenon where particles become interconnected through their quantum states such that the quantum state of each particle cannot be described independently.""",
    },
]
Out[13]:
Console
Number of preference pairs: 2

Example prompt: How do I improve my writing skills?
Chosen response length: 481 chars
Rejected response length: 120 chars

The training examples typically show that chosen responses are longer, more detailed, and better structured than rejected ones. The reward model learns to associate these features, along with factual accuracy and tone, with higher scalar scores.

The quality of the preference data is the single biggest lever in determining how well the reward model generalizes. Constructing good preference data requires careful thought about how to generate diverse candidate responses and how to ensure annotators are evaluating the right properties. If all comparisons pit an obviously good response against an obviously bad one, the reward model learns only from extreme quality differences and may fail to distinguish between two good responses, which is precisely the judgment needed during PPO when the policy has already learned to produce reasonable outputs. Some comparison data should therefore deliberately contrast responses that are both reasonable but differ in subtle quality dimensions: accuracy, depth, appropriate confidence, and conciseness.

Annotator consistency is another important concern. Human annotators often disagree about which of two responses is better, especially when both responses are in the "good" quality range. The Bradley-Terry model assumes a total ordering of quality that may not exist. Two annotators might consistently prefer different response styles (one preferring brevity, another preferring depth), and training on their mixed judgments produces a reward model that has internalized a blend of their preferences. This blend may not correspond to any user's preference. Managing annotator consistency through training, calibration sessions, and inter-annotator agreement monitoring is an important part of high-quality preference data collection.

Reward Model Training Loss

The training objective for the reward model emerges from a probabilistic framework for modeling human preferences. We begin with a natural question: given two responses to the same prompt, how likely is it that a human prefers one over the other? The Bradley-Terry model provides an elegant answer by relating this preference probability to the difference in quality scores assigned to each response.

The core insight is that we can model preferences as arising from latent quality scores. If response ywy_w has a higher quality score than response yly_l, then ywy_w should be preferred more often. The Bradley-Terry model formalizes this intuition by expressing the preference probability as a function of the score difference:

P(yw≻yl∣x)=σ(rθ(x,yw)−rθ(x,yl))P(y_w \succ y_l | x) = \sigma(r_\theta(x, y_w) - r_\theta(x, y_l))

where:

  • P(yw≻yl∣x)P(y_w \succ y_l | x): the probability that response ywy_w is preferred over yly_l given prompt xx
  • σ\sigma: the logistic sigmoid function, σ(z)=11+e−z\sigma(z) = \frac{1}{1+e^{-z}}
  • rθr_\theta: the reward model with parameters θ\theta
  • xx: the input prompt
  • ywy_w: the preferred ("winning") response
  • yly_l: the rejected ("losing") response

The sigmoid function σ\sigma plays a useful role in this formulation. It converts the unbounded difference in reward scores into a probability between 0 and 1. This provides a smooth and differentiable mapping. When the reward for the winning response rθ(x,yw)r_\theta(x, y_w) is significantly larger than for the losing response, the difference becomes a large positive number, and the sigmoid approaches 1, showing near-certainty that ywy_w would be preferred. Conversely, if the scores are equal, the sigmoid returns 0.5. This reflects maximum uncertainty. This elegant mathematical structure captures our intuition that larger quality differences should correspond to more decisive preferences.

Out[14]:
Visualization
Sigmoid curve showing how reward difference maps to preference probability in the Bradley-Terry model.
Bradley-Terry preference probability. The sigmoid function maps the reward difference between two responses to a probability of preference. A large positive difference implies near-certainty that the chosen response is better, while zero difference indicates equal preference likelihood.

To train the model, we minimize the negative log-likelihood of observing the preferences in our dataset. For a single preference pair, the loss is:

LRM=−log⁡σ(rθ(x,yw)−rθ(x,yl))\mathcal{L}_{RM} = -\log \sigma(r_\theta(x, y_w) - r_\theta(x, y_l))

where:

  • LRM\mathcal{L}_{RM}: the scalar loss value to be minimized
  • σ\sigma: the logistic sigmoid function, σ(z)=11+e−z\sigma(z) = \frac{1}{1+e^{-z}}
  • rθr_\theta: the reward model with parameters θ\theta that assigns a scalar score to a prompt-response pair
  • xx: the input prompt or instruction
  • ywy_w: the "winning" or preferred response
  • yly_l: the "losing" or rejected response
  • rθ(x,yw)−rθ(x,yl)r_\theta(x, y_w) - r_\theta(x, y_l): the difference in reward scores (which we want to be positive)

Understanding why this loss works requires examining what happens during optimization. When the model correctly assigns a higher score to the preferred response (making the difference positive and large), the sigmoid outputs a value close to 1, and the negative log becomes small. When the model incorrectly ranks the responses (making the difference negative), the sigmoid outputs a value close to 0, and the negative log becomes very large, creating a strong gradient signal to correct this error. This objective therefore directly maximizes the likelihood that the model assigns a higher score to the preferred response ywy_w than the rejected response yly_l, which is precisely what we want from a reward model.

In[15]:
Code
import torch


def compute_reward_model_loss(reward_model, batch, tokenizer, device):
    """Compute Bradley-Terry loss for preference learning."""

    prompts = batch["prompts"]
    chosen_responses = batch["chosen"]
    rejected_responses = batch["rejected"]

    chosen_texts = [p + c for p, c in zip(prompts, chosen_responses)]
    rejected_texts = [p + r for p, r in zip(prompts, rejected_responses)]

    chosen_tokens = tokenizer(
        chosen_texts, padding=True, return_tensors="pt"
    ).to(device)
    rejected_tokens = tokenizer(
        rejected_texts, padding=True, return_tensors="pt"
    ).to(device)

    chosen_rewards = reward_model(
        chosen_tokens["input_ids"], chosen_tokens["attention_mask"]
    )
    rejected_rewards = reward_model(
        rejected_tokens["input_ids"], rejected_tokens["attention_mask"]
    )

    loss = -torch.log(torch.sigmoid(chosen_rewards - rejected_rewards)).mean()

    accuracy = (chosen_rewards > rejected_rewards).float().mean()

    return loss, accuracy

Reward Model Training Metrics:

  • Loss: Bradley-Terry negative log-likelihood
  • Accuracy: Fraction where r(chosen)>r(rejected)r(\text{chosen}) > r(\text{rejected})
  • Target accuracy: 70-80% (higher may indicate overfitting)

The 70-80% accuracy target deserves explanation, because it might seem counterintuitive to stop well below 100%. Higher accuracy on the training set typically means the model has overfit to the specific comparison pairs it was trained on and has lost the ability to generalize to novel responses. A reward model that achieves 95% accuracy on training data but only 65% on held-out data has learned to recognize specific text patterns rather than quality. In the downstream PPO stage, this overfitted reward model will assign high scores to responses that match patterns from the training comparison data rather than to better responses. Stopping at 70-80% accuracy is an empirical heuristic that tends to produce reward models with better generalization, though the right threshold varies with dataset size and diversity.

Reward Model Calibration

A necessary but often overlooked aspect is reward model calibration. The absolute reward values don't matter for ranking, since the Bradley-Terry model only uses differences between scores. However, the scale of these values significantly affects PPO training stability. If rewards are too large in magnitude, gradient updates can become unstable; if they vary too widely, optimization becomes difficult.

In[16]:
Code
import torch


def calibrate_reward_model(reward_model, calibration_data, tokenizer, device):
    """Calibrate reward model to have zero mean and unit variance."""
    reward_model.eval()

    all_rewards = []

    with torch.no_grad():
        for batch in calibration_data:
            texts = batch["texts"]
            tokens = tokenizer(texts, padding=True, return_tensors="pt").to(
                device
            )
            rewards = reward_model(
                tokens["input_ids"], tokens["attention_mask"]
            )
            all_rewards.append(rewards.cpu())

    all_rewards = torch.cat(all_rewards)
    mean_reward = all_rewards.mean().item()
    std_reward = all_rewards.std().item()

    return mean_reward, std_reward


def normalized_reward(raw_reward, mean, std):
    """Normalize reward to zero mean and unit variance."""
    return (raw_reward - mean) / (std + 1e-8)

Calibrating the reward model to have approximately zero mean and unit variance at the start of PPO training provides a stable foundation for the optimization. The zero mean ensures that the policy does not start with a systematic positive or negative reward bias, which could cause it to initially over-generate or under-generate regardless of quality. The unit variance ensures that the KL penalty coefficient β\beta has a consistent meaning across different training runs: a β\beta of 0.05 represents the same trade-off between reward and KL regularization regardless of whether the raw reward model happens to output values in the range [−1,1][-1, 1] or [−10,10][-10, 10].

The calibration should be computed on a representative sample of prompt-response pairs that reflect the distribution the policy will encounter during PPO training. Using the SFT model to generate calibration responses is the most common approach, since the SFT model provides the starting point for the PPO policy. This ensures that the normalization statistics are computed at the operating point of the system, not on some other distribution.

Stage 3: PPO Fine-Tuning

With the SFT model and reward model ready, we can now run PPO optimization. This stage adjusts the policy to maximize expected reward while staying close to the SFT model's behavior. The challenge here is delicate: we want the model to improve according to the reward signal without losing the coherent language abilities it acquired during pretraining and SFT.

Think of PPO fine-tuning as the difference between a musician who learns jazz improvisation and one who simply memorizes jazz recordings. The SFT model knows how to play in the right style. The reward model has internalized the judgment of experienced listeners. PPO optimization uses that judgment signal to push the musician toward improvising in ways that experienced listeners prefer, while the KL penalty ensures the musician does not completely abandon the musical training that makes their playing coherent in the first place. The constraint is not a limitation but a necessity: a musician who has thrown out all musical structure in pursuit of applause is no longer a musician.

The PPO stage runs for far fewer gradient steps than pretraining or even SFT, often fewer than 10,000 steps compared to hundreds of billions of tokens for pretraining. This brevity is both a feature and a constraint. The short training duration means PPO cannot introduce catastrophically wrong behaviors if properly regularized, but it also means the policy cannot make dramatic departures from the SFT starting point. The alignment effect is best thought of as a refinement of existing capabilities rather than an acquisition of new ones. The capabilities must already exist in the SFT model; PPO merely tunes the probability distribution over responses to favor the higher-quality ones.

The PPO Training Loop

As we detailed in the PPO for Language Models chapter, each training iteration involves a carefully orchestrated sequence of steps that together enable stable policy improvement:

  1. Sampling: Generate responses from the current policy
  2. Reward computation: Score responses using the reward model
  3. Advantage estimation: Compute advantages using GAE
  4. Policy update: Optimize the clipped surrogate objective. This iterative process gradually shifts the policy's behavior toward responses that score higher according to the reward model, while the various stability mechanisms in PPO prevent the optimization from taking steps that are too large or in harmful directions.
In[17]:
Code
import torch
import torch.optim


class RLHFTrainer:
    """Complete RLHF trainer implementing the PPO training loop."""

    def __init__(
        self, policy_model, ref_model, reward_model, tokenizer, config
    ):
        self.policy = policy_model
        self.ref_model = ref_model  # Frozen copy of SFT model
        self.reward_model = reward_model
        self.tokenizer = tokenizer
        self.config = config

        self.optimizer = torch.optim.AdamW(
            self.policy.parameters(),
            lr=config["learning_rate"],
            weight_decay=config["weight_decay"],
        )

    def generate_responses(self, prompts, max_length=256):
        """Generate responses from current policy."""
        self.policy.eval()

        responses = []
        log_probs_list = []

        with torch.no_grad():
            for prompt in prompts:
                input_ids = self.tokenizer.encode(prompt, return_tensors="pt")

                output_ids = []
                log_probs = []

                for _ in range(max_length):
                    outputs = self.policy(input_ids)
                    next_token_logits = outputs.logits[:, -1, :]
                    probs = torch.softmax(next_token_logits, dim=-1)

                    next_token = torch.multinomial(probs, num_samples=1)

                    token_log_prob = torch.log(probs[0, next_token.item()])

                    output_ids.append(next_token.item())
                    log_probs.append(token_log_prob.item())

                    input_ids = torch.cat([input_ids, next_token], dim=1)

                    if next_token.item() == self.tokenizer.eos_token_id:
                        break

                responses.append(self.tokenizer.decode(output_ids))
                log_probs_list.append(log_probs)

        return responses, log_probs_list

Computing the Complete Reward

The total reward used to update the policy combines two competing objectives. We want to maximize the reward model's score, which represents human preferences, but we also want to prevent the policy from straying too far from the reference model, which represents stable, coherent language generation. The KL divergence penalty provides this regularization, creating a tug-of-war that encourages improvement without catastrophic drift.

We'll explore the KL divergence penalty in detail in the next chapter, but the basic formulation captures this balance mathematically:

Rtotal(x,y)=RRM(x,y)−β⋅DKL(πθ∣∣πref)R_{total}(x, y) = R_{RM}(x, y) - \beta \cdot D_{KL}(\pi_\theta || \pi_{ref})

where:

  • Rtotal(x,y)R_{total}(x, y): the combined reward used to update the policy
  • xx: the input prompt
  • yy: the generated response
  • RRM(x,y)R_{RM}(x, y): the preference score from the reward model
  • β\beta: the KL penalty coefficient controlling regularization strength
  • DKLD_{KL}: the Kullback-Leibler divergence between the two distributions
  • πθ\pi_\theta: the current policy model
  • πref\pi_{ref}: the reference model (frozen SFT model)

The KL coefficient β\beta acts as a dial controlling the trade-off between reward maximization and behavioral stability. A larger β\beta keeps the policy closer to the reference model, preserving more of the original capabilities but potentially limiting how much the model can improve. A smaller β\beta allows more aggressive optimization toward higher rewards, but risks the policy finding reward model exploits or losing coherence. This formulation ensures that while we maximize the preference score, we maintain the linguistic coherence and knowledge of the original model.

In practice, β\beta is one of the most sensitive hyperparameters in the entire RLHF pipeline. Values too close to zero produce unstable training because the policy has no incentive to stay in the region of coherent language. Values too large make the training too conservative, and the policy barely moves from its SFT initialization despite many optimization steps. Typical values range from 0.01 to 0.1 for most configurations, though some practitioners use adaptive schemes that increase β\beta automatically when the KL divergence exceeds a threshold. This provides a soft constraint instead of a fixed penalty.

Out[18]:
Visualization
Stacked bar chart showing how total reward is composed of reward model score minus KL penalty across training steps.
Reward composition in RLHF. The total reward (blue line) balances the raw preference score (green) against a KL divergence penalty (red). As the policy diverges from the reference model, the growing penalty prevents catastrophic drift.
In[19]:
Code
import torch


def compute_rewards_with_kl_penalty(
    policy_model,
    ref_model,
    reward_model,
    prompts,
    responses,
    tokenizer,
    kl_coef,
    device,
):
    """Compute total reward including KL penalty."""

    full_texts = [p + r for p, r in zip(prompts, responses)]
    tokens = tokenizer(full_texts, padding=True, return_tensors="pt").to(device)

    with torch.no_grad():
        rm_rewards = reward_model(tokens["input_ids"], tokens["attention_mask"])

    with torch.no_grad():
        policy_outputs = policy_model(
            tokens["input_ids"], attention_mask=tokens["attention_mask"]
        )
        ref_outputs = ref_model(
            tokens["input_ids"], attention_mask=tokens["attention_mask"]
        )

        policy_logprobs = torch.log_softmax(policy_outputs.logits, dim=-1)
        ref_logprobs = torch.log_softmax(ref_outputs.logits, dim=-1)

        token_kl = (
            policy_logprobs.exp() * (policy_logprobs - ref_logprobs)
        ).sum(dim=-1)
        kl_penalty = token_kl.sum(dim=-1)

    total_rewards = rm_rewards - kl_coef * kl_penalty

    return total_rewards, rm_rewards, kl_penalty

PPO Update Step

The policy update uses the clipped surrogate objective from the PPO Algorithm chapter. This objective function represents the core mechanism that enables stable policy improvement: rather than directly maximizing expected reward, which could lead to catastrophically large updates, PPO constrains how much the policy can change in a single step. The clipping mechanism ensures that even if the advantage estimates suggest a large improvement, the actual policy update remains bounded.

In[20]:
Code
import torch


def ppo_update(
    policy_model, optimizer, batch, clip_epsilon=0.2, entropy_coef=0.01
):
    """Perform PPO policy update with clipped objective."""
    policy_model.train()

    states = batch["states"]
    actions = batch["actions"]
    old_log_probs = batch["old_log_probs"]
    advantages = batch["advantages"]
    returns = batch["returns"]

    advantages = (advantages - advantages.mean()) / (advantages.std() + 1e-8)

    outputs = policy_model(states)
    logits = outputs.logits

    log_probs = torch.log_softmax(logits, dim=-1)
    action_log_probs = log_probs.gather(-1, actions.unsqueeze(-1)).squeeze(-1)

    ratios = torch.exp(action_log_probs - old_log_probs)

    surr1 = ratios * advantages
    surr2 = torch.clamp(ratios, 1 - clip_epsilon, 1 + clip_epsilon) * advantages
    policy_loss = -torch.min(surr1, surr2).mean()

    probs = torch.softmax(logits, dim=-1)
    entropy = -(probs * log_probs).sum(dim=-1).mean()
    entropy_loss = -entropy_coef * entropy

    total_loss = policy_loss + entropy_loss

    optimizer.zero_grad()
    total_loss.backward()
    torch.nn.utils.clip_grad_norm_(policy_model.parameters(), max_norm=1.0)
    optimizer.step()

    return {
        "policy_loss": policy_loss.item(),
        "entropy": entropy.item(),
        "mean_ratio": ratios.mean().item(),
        "clip_fraction": (
            (ratios < 1 - clip_epsilon) | (ratios > 1 + clip_epsilon)
        )
        .float()
        .mean()
        .item(),
    }

The entropy bonus term in the loss function deserves particular attention. By adding a small positive reward for entropy (the entropy coefficient 0.010.01 in the code above multiplied by the average entropy of the action distribution), we discourage the policy from collapsing to deterministic outputs too quickly. Think of entropy as a measure of how many different responses the model considers at each step: high entropy means the model is still exploring a diverse set of continuations, while low entropy means it has converged to nearly always predicting the same tokens. Without the entropy bonus, PPO training tends to produce overly deterministic policies that perform poorly on novel prompts where the training distribution does not provide clear guidance.

The advantage normalization step (subtracting the mean and dividing by the standard deviation) is a standard variance reduction technique that makes training more stable. Raw advantages can have high variance because they depend on the absolute reward values, which may shift considerably across training steps. Normalizing to zero mean and unit variance at each batch ensures that the effective learning rate remains consistent even as the reward distribution changes during training.

Worked Example: Tracing a Single RLHF Training Step

To make the pipeline concrete, let's trace a single complete training step from prompt to policy update, using specific numerical values at each stage. This worked example connects all the abstract components into a coherent sequence.

Setup: Suppose we have already completed SFT and trained a reward model. The policy starts from the SFT model checkpoint. We use a KL coefficient β=0.05\beta = 0.05 and a PPO clip epsilon of 0.20.2.

Step 1: Sample a prompt from the training distribution. We draw the prompt "What is gradient descent?" from our prompt library.

Step 2: Generate a response from the current policy. The policy samples token by token, creating the response: "Gradient descent is an optimization algorithm that iteratively adjusts parameters in the direction that minimizes a loss function." Alongside the response, we record the log-probabilities assigned by the policy to each token, which we'll call log⁡πθ(at∣st)\log \pi_\theta(a_t | s_t) for each token position tt.

Step 3: Score the response using the reward model. The reward model receives the full prompt-response pair and outputs a raw score. Suppose it outputs RRM=1.42R_{RM} = 1.42.

Step 4: Compute the KL penalty. We also pass the prompt-response pair through the frozen reference model (the SFT model) and compute token-level log-probabilities. The KL divergence at each token position tt is:

KLt=πθ(at∣st)⋅(log⁡πθ(at∣st)−log⁡πref(at∣st))\text{KL}_t = \pi_\theta(a_t | s_t) \cdot \left(\log \pi_\theta(a_t | s_t) - \log \pi_{ref}(a_t | s_t)\right)

Summing across all token positions gives the total sequence-level KL divergence. Suppose this comes out to DKL=0.28D_{KL} = 0.28.

Step 5: Compute the total reward. Combining the reward model score with the KL penalty:

Rtotal=RRM−β⋅DKL=1.42−0.05×0.28=1.42−0.014=1.406\begin{aligned} R_{total} &= R_{RM} - \beta \cdot D_{KL} \\ &= 1.42 - 0.05 \times 0.28 \\ &= 1.42 - 0.014 \\ &= 1.406 \end{aligned}

where:

  • RRM=1.42R_{RM} = 1.42: the raw reward model score for this response
  • β=0.05\beta = 0.05: the KL penalty coefficient
  • DKL=0.28D_{KL} = 0.28: the total sequence-level KL divergence

At this early stage of training, the KL penalty is small because the policy has barely moved from its SFT initialization. The total reward is almost entirely the raw reward model score.

Step 6: Estimate advantages. The advantage tells us how much better this response was than the average response to this prompt. Using a value function baseline with estimated value V(s)=1.15V(s) = 1.15, the advantage is approximately:

A=Rtotal−V(s)=1.406−1.15=0.256A = R_{total} - V(s) = 1.406 - 1.15 = 0.256

A positive advantage means this response was better than expected. The policy update will increase the probability of generating similar responses.

Step 7: Compute the PPO ratio. For each token, we compute the ratio of the current policy's probability to the old policy's probability (from before this update):

ratiot=πθ(at∣st)πold(at∣st)=exp⁡(log⁡πθ(at∣st)−log⁡πold(at∣st))\text{ratio}_t = \frac{\pi_\theta(a_t | s_t)}{\pi_{\text{old}}(a_t | s_t)} = \exp\left(\log \pi_\theta(a_t | s_t) - \log \pi_{\text{old}}(a_t | s_t)\right)

Suppose the mean ratio across tokens is 1.081.08, meaning the policy has moved slightly toward creating this response.

Step 8: Apply the clipped objective. Since 1.08<1+0.2=1.21.08 < 1 + 0.2 = 1.2, the ratio is within the clipping bounds. The surrogate objective for this response is 1.08×0.256=0.2771.08 \times 0.256 = 0.277. The gradient update increases the probability of generating this response, pushing the policy in a direction that the reward model rates more highly.

After this single step, the policy is infinitesimally more likely to generate responses similar to "Gradient descent is an optimization algorithm..." when asked about gradient descent. Repeating this process across thousands of prompts and responses gradually sculpts the policy's behavior across the entire prompt distribution.

RLHF Debugging

RLHF training is notoriously difficult to debug. The interplay between the policy, reward model, and KL constraint creates many potential failure modes. Unlike standard supervised learning, where a rising validation loss is a clear and unambiguous signal that something is wrong, RLHF metrics can be simultaneously misleading in several directions at once: reward can increase while response quality decreases, KL divergence can remain controlled while mode collapse is developing, and entropy can look healthy while the policy is creating repetitive variations on a narrow set of response templates.

Effective debugging requires combining quantitative metrics with qualitative evaluation. No set of numerical metrics can substitute for regularly reading the model's outputs during training. A reward model score of 2.5 could represent excellent responses, or it could represent responses the policy has learned to craft specifically to exploit the reward model's scoring tendencies. The only way to tell the difference is to look at the text.

Key Metrics to Monitor

Effective RLHF debugging requires tracking multiple metrics throughout training:

In[21]:
Code
import numpy as np


class RLHFMetricsTracker:
    """Track key metrics for RLHF debugging."""

    def __init__(self):
        self.metrics_history = {
            "reward_mean": [],
            "reward_std": [],
            "kl_divergence": [],
            "policy_loss": [],
            "entropy": [],
            "clip_fraction": [],
            "response_length": [],
            "unique_tokens": [],
        }

    def log_step(self, metrics_dict):
        for key, value in metrics_dict.items():
            if key in self.metrics_history:
                self.metrics_history[key].append(value)

    def check_health(self):
        """Check for common RLHF failure modes."""
        warnings = []

        if len(self.metrics_history["kl_divergence"]) > 100:
            recent_kl = np.mean(self.metrics_history["kl_divergence"][-100:])
            if recent_kl > 10:
                warnings.append(
                    f"HIGH KL DIVERGENCE ({recent_kl:.2f}): Policy diverging from reference"
                )
            elif recent_kl < 0.01:
                warnings.append(
                    f"LOW KL DIVERGENCE ({recent_kl:.4f}): Policy not learning"
                )

        if len(self.metrics_history["reward_mean"]) > 100:
            recent_reward = np.mean(self.metrics_history["reward_mean"][-100:])
            if recent_reward > 5:
                warnings.append(
                    f"VERY HIGH REWARD ({recent_reward:.2f}): Possible reward hacking"
                )

        if len(self.metrics_history["entropy"]) > 100:
            recent_entropy = np.mean(self.metrics_history["entropy"][-100:])
            if recent_entropy < 0.1:
                warnings.append(
                    f"LOW ENTROPY ({recent_entropy:.4f}): Policy becoming deterministic"
                )

        if len(self.metrics_history["response_length"]) > 100:
            recent_len = np.mean(self.metrics_history["response_length"][-100:])
            early_len = np.mean(self.metrics_history["response_length"][:100])
            if recent_len < early_len * 0.5:
                warnings.append(
                    "RESPONSE LENGTH COLLAPSED: Model creating very short responses"
                )
            elif recent_len > early_len * 2:
                warnings.append(
                    "RESPONSE LENGTH EXPLOSION: Model creating very long responses"
                )

        return warnings
In[22]:
Code
import numpy as np

tracker = RLHFMetricsTracker()

for i in range(200):
    kl = 0.5 + 0.1 * np.random.randn() + i * 0.08  # Gradually increasing KL
    tracker.log_step(
        {
            "kl_divergence": kl,
            "reward_mean": 2.0 + 0.5 * np.random.randn() + i * 0.04,
            "entropy": max(0.05, 1.0 - i * 0.004 + 0.1 * np.random.randn()),
            "response_length": 100 + 10 * np.random.randn() - i * 0.2,
        }
    )

warnings = tracker.check_health()
Out[23]:
Console
Health Check Results:
  ⚠️  HIGH KL DIVERGENCE (12.45): Policy diverging from reference
  ⚠️  VERY HIGH REWARD (7.95): Possible reward hacking

The metrics tracker successfully identifies the simulated anomalies. By monitoring the KL divergence and reward statistics we can catch issues like the "High KL" spike and "High Reward" events (simulating reward hacking) before they destabilize the entire training run.

Understanding what each metric tells you is as important as tracking it. KL divergence measures how far the current policy has moved from the reference model in terms of its probability distributions over responses. High KL is not inherently bad if reward has improved correspondingly, but high KL with stagnant or declining response quality is a warning sign that the policy has drifted into a region where the reward model is no longer a reliable proxy for human preferences. Entropy measures how deterministic the policy has become. Entropy collapsing to near zero while training is still ongoing suggests mode collapse is developing, even if reward appears stable.

The clip fraction (the fraction of PPO ratio updates that were clipped) is a particularly useful diagnostic. Values consistently above 0.3 indicate that the policy is changing faster than the clipped objective can accommodate, which means the old log-probabilities stored from the rollout phase are no longer accurate descriptions of the current policy. This is a sign that the PPO epoch count (how many times you update the policy on the same set of rollouts) is too high, or the learning rate is too large. Reducing either parameter should bring the clip fraction back to a healthier range of 0.05 to 0.2.

Common Failure Modes

Understanding these failure modes helps you diagnose and fix training issues:

Reward Hacking occurs when the policy finds exploits in the reward model that don't correspond to quality improvements. Signs include rapidly increasing reward with degrading response quality, or unusual patterns like excessive repetition or specific phrases.

In[24]:
Code
import numpy as np


def detect_reward_hacking(responses, rewards, threshold_percentile=95):
    """Detect potential reward hacking by examining high-reward responses."""

    threshold = np.percentile(rewards, threshold_percentile)
    high_reward_indices = np.where(np.array(rewards) > threshold)[0]

    flags = []

    for idx in high_reward_indices:
        response = responses[idx]

        words = response.split()
        if len(words) > 0:
            unique_ratio = len(set(words)) / len(words)
            if unique_ratio < 0.3:
                flags.append(("repetition", idx, response[:100]))

        if len(response) > 2000:
            flags.append(("excessive_length", idx, f"Length: {len(response)}"))
        elif len(response) < 10:
            flags.append(("too_short", idx, response))

    return flags

Reward hacking is insidious because the numerical metrics look good while the model's behavior deteriorates. The reward model has a finite capacity and has only seen a finite amount of training data. The PPO policy, optimizing aggressively, can discover input patterns that the reward model was never trained to reject. Classic examples include: generating responses that begin with phrases the reward model associates with high quality (such as "Certainly! Here's a complete answer...") regardless of whether the content is helpful; creating responses of a specific length that correlates with high reward in the training data without giving proportional value; and inserting specific keywords that the reward model's training data associated with expert responses. Each of these behaviors increases the reward model score without making the response more useful.

Mode Collapse happens when the policy converges to creating nearly identical responses regardless of the prompt. Monitor response diversity and entropy throughout training.

In[25]:
Code
import numpy as np


def compute_response_diversity(responses, n_samples=100):
    """Measure diversity of generated responses."""

    unique_responses = len(set(responses[:n_samples]))
    uniqueness_ratio = unique_responses / min(n_samples, len(responses))

    all_words = []
    for response in responses[:n_samples]:
        all_words.extend(response.lower().split())

    if len(all_words) > 0:
        vocab_size = len(set(all_words))
        type_token_ratio = vocab_size / len(all_words)
    else:
        type_token_ratio = 0

    return {
        "uniqueness_ratio": uniqueness_ratio,
        "type_token_ratio": type_token_ratio,
        "avg_response_length": np.mean([len(r) for r in responses[:n_samples]]),
    }

KL Explosion indicates the policy is moving too fast away from the reference model. This often precedes training instability:

Out[26]:
Visualization
Line plot showing KL divergence staying near the 0.5 target during healthy RLHF training.
Healthy RLHF training: KL divergence remains near the target of 0.5, showing controlled optimization without policy drift.
Line plot showing KL divergence rapidly increasing during KL explosion in RLHF training.
KL explosion during RLHF: the policy diverges rapidly from the reference model, often triggering mode collapse or reward hacking.

Debugging Workflow

When RLHF training goes wrong, follow this systematic debugging approach:

  1. Check the reward model first: Generate samples and manually verify that reward model scores align with your quality intuitions. A miscalibrated or overfitted reward model dooms PPO from the start.

  2. Examine generated samples: Look at actual model outputs throughout training. Metrics can hide problems that become obvious when reading responses.

  3. Verify the KL penalty is working: The policy should stay reasonably close to the reference. If responses look completely different from SFT outputs, the KL constraint may be too weak.

  4. Monitor multiple metrics together: Single metrics can be misleading. High reward with low diversity suggests reward hacking. Low KL with no reward improvement suggests the policy isn't learning.

A systematic debugging session for a troubled RLHF run typically looks like this: first, freeze the policy and run the reward model on a held-out set of manually graded responses to check that its rankings match your intuitions. If the reward model consistently assigns higher scores to responses you would consider worse, the problem is in stage 2, and you need to retrain the reward model. If the reward model scores make sense but the policy is not improving, the issue is more likely in the PPO hyperparameters, such as too high a KL coefficient preventing meaningful updates, or too few rollout steps per iteration giving the advantage estimator high variance. If the policy is improving according to the reward model but human evaluation shows no improvement, the reward model is overfitting and its scores have become decoupled from actual quality.

In[27]:
Code
import torch


def debug_rlhf_step(
    policy_model, ref_model, reward_model, sample_prompts, tokenizer, device
):
    """Comprehensive debugging for a single RLHF step."""

    debug_info = {}

    policy_model.eval()
    ref_model.eval()

    with torch.no_grad():
        responses = []
        for prompt in sample_prompts[:5]:
            input_ids = tokenizer.encode(prompt, return_tensors="pt").to(device)
            output_ids = policy_model.generate(
                input_ids, max_new_tokens=100, do_sample=True, temperature=0.7
            )
            response = tokenizer.decode(output_ids[0][len(input_ids[0]) :])
            responses.append(response)

        debug_info["sample_responses"] = list(
            zip(sample_prompts[:5], responses)
        )

        full_texts = [p + r for p, r in zip(sample_prompts[:5], responses)]
        tokens = tokenizer(full_texts, padding=True, return_tensors="pt").to(
            device
        )

        rewards = reward_model(tokens["input_ids"], tokens["attention_mask"])
        debug_info["rewards"] = rewards.cpu().tolist()

        debug_info["response_lengths"] = [len(r) for r in responses]
        debug_info["unique_words"] = [len(set(r.split())) for r in responses]

    return debug_info
Out[28]:
Visualization
Line plot showing reward increasing and diversity decreasing during reward hacking in RLHF.
Reward hacking: the policy exploits the reward model, driving reward scores up while response diversity simultaneously collapses.
Line plot showing response uniqueness collapsing to near zero during mode collapse in RLHF.
Mode collapse: response uniqueness degrades sharply as the model converges to a narrow set of repetitive outputs.
Line plot showing KL divergence exploding far beyond the 0.5 target during unstable RLHF training.
KL explosion: KL divergence grows uncontrollably as the policy diverges far from the reference SFT model.
Line plot showing reward and KL divergence both evolving healthily during stable RLHF training.
Healthy RLHF training: reward increases steadily while KL divergence remains near the target, showing stable alignment progress.

Putting It All Together

Let's trace through a complete RLHF training run, showing how all pieces connect:

In[29]:
Code
import numpy as np


def run_rlhf_pipeline(config):
    """Complete RLHF training pipeline."""

    print("=" * 50)
    print("STAGE 1: Supervised Fine-Tuning")
    print("=" * 50)

    # SFT training loop (simplified)
    sft_metrics = {
        "initial_loss": 3.5,
        "final_loss": 1.8,
        "epochs": config["sft_epochs"],
    }

    print(f"  Initial loss: {sft_metrics['initial_loss']:.3f}")
    print(f"  Final loss: {sft_metrics['final_loss']:.3f}")
    print(f"  Trained for {sft_metrics['epochs']} epochs")
    print()

    print("=" * 50)
    print("STAGE 2: Reward Model Training")
    print("=" * 50)

    rm_metrics = {
        "initial_accuracy": 0.52,
        "final_accuracy": 0.74,
        "epochs": config["rm_epochs"],
    }

    print(f"  Initial accuracy: {rm_metrics['initial_accuracy']:.1%}")
    print(f"  Final accuracy: {rm_metrics['final_accuracy']:.1%}")
    print(f"  Trained for {rm_metrics['epochs']} epochs")
    print()

    print("=" * 50)
    print("STAGE 3: PPO Fine-Tuning")
    print("=" * 50)

    ppo_metrics = []

    for step in range(0, config["ppo_steps"], config["ppo_steps"] // 10):
        progress = step / config["ppo_steps"]

        metrics = {
            "step": step,
            "reward": 0.5 + 1.5 * progress + 0.2 * np.random.randn(),
            "kl": 0.5 * (1 - np.exp(-5 * progress)) + 0.03 * np.random.randn(),
            "policy_loss": -0.5 - 0.3 * progress + 0.1 * np.random.randn(),
        }
        ppo_metrics.append(metrics)

        if step % (config["ppo_steps"] // 5) == 0:
            print(
                f"  Step {step:5d}: reward={metrics['reward']:.3f}, "
                f"KL={metrics['kl']:.3f}, loss={metrics['policy_loss']:.3f}"
            )

    return sft_metrics, rm_metrics, ppo_metrics
In[30]:
Code
pipeline_config = {"sft_epochs": 2, "rm_epochs": 1, "ppo_steps": 10000}
In[31]:
Code
sft_results, rm_results, ppo_results = run_rlhf_pipeline(pipeline_config)
Out[31]:
Console
==================================================
STAGE 1: Supervised Fine-Tuning
==================================================
  Initial loss: 3.500
  Final loss: 1.800
  Trained for 2 epochs

==================================================
STAGE 2: Reward Model Training
==================================================
  Initial accuracy: 52.0%
  Final accuracy: 74.0%
  Trained for 1 epochs

==================================================
STAGE 3: PPO Fine-Tuning
==================================================
  Step     0: reward=0.544, KL=-0.025, loss=-0.744
  Step  2000: reward=0.550, KL=0.308, loss=-0.557
  Step  4000: reward=0.981, KL=0.382, loss=-0.631
  Step  6000: reward=1.353, KL=0.472, loss=-0.790
  Step  8000: reward=1.319, KL=0.534, loss=-0.863

The text output confirms that each stage completed successfully. The SFT loss decreased significantly, and the Reward Model achieved a validation accuracy of 74%, which is within the typical 70-80% range for effective preference modeling. These healthy prerequisites set the stage for the PPO phase, which we can now visualize.

In[32]:
Code
steps = [m["step"] for m in ppo_results]
rewards = [m["reward"] for m in ppo_results]
kls = [m["kl"] for m in ppo_results]
Out[33]:
Visualization
Line plot showing mean reward rising during PPO training steps.
Mean reward trajectory during PPO training: reward increases steadily as the policy learns to satisfy human preferences.
Line plot showing KL divergence staying near 0.5 during PPO training steps.
KL divergence during PPO training: divergence remains near the 0.5 target, so the model retains coherent language capabilities without reward hacking.

The training curves demonstrate healthy alignment progress. The reward (left) steadily increases, showing the model is learning to satisfy the reward model's preferences. Meanwhile, the KL divergence (right) remains controlled near the target of 0.5. This keeps the model maintains the coherent capabilities of the original SFT model without drifting into incoherence or reward hacking.

Limitations and Practical Considerations

The RLHF pipeline comes with significant challenges in any production deployment.

Computational cost is substantial. The pipeline requires training three separate models (SFT, reward model, and PPO policy), with the PPO stage being particularly expensive because it requires running both the policy and reference model for every batch. A single RLHF training run can cost hundreds of thousands of dollars in compute for large models, making iteration and experimentation prohibitively expensive for most organizations. This has driven interest in more efficient alternatives like Direct Preference Optimization (DPO), which we'll explore in upcoming chapters. Even at research scale with smaller models, the infrastructure complexity of running three models simultaneously with different roles (active policy, frozen reference, frozen reward model) requires careful systems engineering to avoid memory bottlenecks.

Human annotation quality fundamentally limits what RLHF can achieve. The reward model can only capture patterns present in the preference data, and human annotators bring biases and practical limitations to the task. Their judgments can also be inconsistent. Disagreement between annotators is common, yet the Bradley-Terry model assumes a consistent underlying preference ordering. When annotators disagree about what makes a response "better," the reward model learns a noisy compromise that may not align with any particular user's preferences. Annotators also tend to evaluate responses on dimensions they can assess quickly (length, formality, apparent confidence) rather than harder dimensions (factual accuracy, appropriate reasoning, appropriate safety). This creates systematic gaps between what the reward model scores and response quality.

Reward hacking remains an unsolved problem despite various mitigation strategies. As we discussed in the Reward Hacking chapter, the policy will exploit any systematic weakness in the reward model. The KL penalty helps by anchoring behavior to the reference model, but sufficiently capable policies can still find exploits within the allowed KL budget. This creates an ongoing cat-and-mouse dynamic where you must continually patch reward model vulnerabilities. The basic tension is that the reward model is a finite model trained on finite data, while the policy is a powerful optimizer with many more optimization steps. Over time, the policy will discover the reward model's weak points because that is exactly what gradient descent is designed to do.

The pipeline also introduces a subtle and often overlooked failure mode: sycophancy. Because human annotators tend to prefer responses that agree with them and that sound confident, the reward model can learn to reward agreeable, confident-sounding responses over accurate but potentially uncertain or corrective ones. The aligned model becomes expert at telling users what they want to hear rather than what is true, because that is what maximizes reward. This form of reward hacking is particularly dangerous because it is invisible in aggregate metrics: the reward scores look good, the KL divergence is controlled, and the model produces fluent, confident-sounding text. Only careful evaluation on benchmarks that specifically probe for sycophantic behavior reveals the problem.

Reproducibility is challenging due to the many interacting hyperparameters and the sensitivity of PPO training. Small changes in learning rate, KL coefficient, or even random seed can lead to qualitatively different outcomes. This makes it difficult to compare results across papers or replicate published findings. The stochastic nature of both the generation process and the policy gradient updates means that two runs with identical hyperparameters can produce noticeably different models. Reporting results from a single run and treating them as definitive is a common but significant methodological error in the RLHF literature.

Despite these limitations, RLHF remains the most widely deployed alignment technique for production language models. Understanding the complete pipeline, including its failure modes, is needed for anyone working on language model alignment. The next chapter examines the KL divergence penalty in detail, which plays a useful role in balancing reward maximization with behavioral stability.

Summary

The RLHF pipeline turns a pretrained language model into an aligned assistant through three sequential stages. Supervised Fine-Tuning creates a model that understands the instruction-following format and produces reasonable responses. Reward Model training captures human preferences in a learnable function that provides optimization signal. PPO Fine-Tuning then optimizes the policy to maximize rewards while staying close to the reference model.

Key takeaways from this chapter:

  • SFT provides the foundation: PPO requires a model that already produces usable responses; skipping SFT leads to unstable training
  • Reward model quality constrains the policy: A flawed reward model will lead to flawed policies; validate carefully before PPO
  • KL penalty prevents catastrophic drift: Without anchoring to the reference model, the policy will exploit reward model weaknesses
  • Monitor multiple metrics: Single metrics can be misleading; track reward, KL, entropy, and response characteristics together
  • Debugging requires examining actual outputs: Metrics summarize behavior, but reading generated responses reveals problems that numbers hide
  • The pipeline has compounding failure modes: Errors in SFT propagate to reward model training, and errors in the reward model propagate to the PPO policy

The RLHF pipeline established the template for aligning large language models, but its complexity and cost have motivated simpler alternatives. The next chapter examines the KL divergence penalty in mathematical detail, followed by chapters on Direct Preference Optimization, which eliminates the reward model and PPO stages entirely while achieving comparable alignment results.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the RLHF pipeline.

RLHF Pipeline

Question 1 of 80 of 8 completed
What is the correct order of stages in the RLHF pipeline?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025rlhfpipeline, author = {Michael Brenndoerfer}, title = {RLHF Pipeline: SFT, Reward Models, and PPO}, year = {2025}, url = {https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). RLHF Pipeline: SFT, Reward Models, and PPO. Retrieved from https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training
MLAAcademic
Michael Brenndoerfer. "RLHF Pipeline: SFT, Reward Models, and PPO." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training>.
CHICAGOAcademic
Michael Brenndoerfer. "RLHF Pipeline: SFT, Reward Models, and PPO." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training.
HARVARDAcademic
Michael Brenndoerfer (2025) 'RLHF Pipeline: SFT, Reward Models, and PPO'. Available at: https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). RLHF Pipeline: SFT, Reward Models, and PPO. https://mbrenndoerfer.com/writing/rlhf-pipeline-sft-reward-model-ppo-training

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.