Part of Language AI Handbook
Covers RLHF pipeline with three stages: Supervised Fine-Tuning, Reward Model training, and PPO optimization. Topics include debugging techniques.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
RLHF Pipeline
Reinforcement Learning from Human Feedback turns a base language model into an assistant that follows instructions and gives useful, truthful responses without causing harm. Building a capable language model through pretraining is only half the story. A model trained purely to predict the next token has learned to be a generalist text predictor, not a helpful collaborator. It will complete sentences, continue stories, and extend documents, but it has no concept of what "helpful" means, no understanding of what a user wants, and no instinct to avoid harmful outputs. Bridging that gap requires a fundamentally different kind of training signal, one that originates not from next-token prediction but from human judgment about what constitutes a good response.
The previous chapters introduced the individual components: the Bradley-Terry model for preference modeling, reward model architecture and training, and the PPO algorithm adapted for language models. Now we assemble these pieces into a complete training pipeline. The assembly matters as much as any individual component. Each stage creates preconditions that the next stage depends on, and misunderstanding those dependencies is one of the most common sources of failure in real RLHF implementations. A reward model trained on the wrong data distribution, or PPO optimization started from a poorly prepared initial policy, can derail the entire alignment effort despite each component working correctly in isolation.
Think of the RLHF pipeline as a three-act apprenticeship. In the first act, the apprentice studies examples of excellent work from master craftspeople, learning the expected format and quality standard. In the second act, an experienced judge observes many examples and builds an intuition for what separates excellent work from mediocre work, capturing that judgment in a reproducible scoring mechanism. In the third act, the apprentice iterates, generating work, receiving scores, and adjusting behavior to earn higher marks, all while staying grounded enough in the original training to remain coherent. The magic is that this apprenticeship produces a model whose behavior reflects human preferences even on prompts it has never seen before, generalizing a signal from thousands of comparisons into millions of novel situations.
The RLHF pipeline consists of three sequential stages: Supervised Fine-Tuning (SFT), Reward Model (RM) training, and PPO optimization. Each stage builds on the previous one, progressively shaping the model's behavior. This chapter walks through the complete pipeline, examining the design decisions at each stage, the hyperparameters that govern training stability, and the debugging techniques needed for successful alignment. We will also develop intuition for the failure modes that make RLHF notoriously difficult to get right in practice, building toward a mental model that helps you diagnose problems when they arise.
Understanding the complete pipeline also illuminates the motivation for alternatives like Direct Preference Optimization, which we will encounter in upcoming chapters. Many of the simplifications that DPO introduces are direct responses to known weaknesses in this three-stage RLHF pipeline, and appreciating those weaknesses helps you understand how DPO works and why it was designed the way it was.
The RLHF pipeline as we know it today was crystallized in the InstructGPT paper (Ouyang et al., 2022), which demonstrated for the first time that this three-stage procedure could produce models that humans strongly preferred over much larger models trained with standard supervised objectives. The underlying insight, that human feedback could serve as a training signal through a learned reward model rather than being applied directly, drew on decades of earlier work in reinforcement learning from human feedback in robotics (Christiano et al., 2017) and earlier experiments applying preference learning to simpler language tasks. The GPT-3 era provided the important ingredient: a base model powerful enough to produce responses diverse enough for meaningful human comparisons. Before sufficiently capable base models existed, the preference signal was too noisy to train a useful reward model. InstructGPT showed that the combination of scale and the three-stage pipeline could produce qualitative improvements that dwarfed the effect of model size alone: the 1.3B InstructGPT model was preferred by labelers over the 175B GPT-3 model, a finding that reshaped the entire field's thinking about the relationship between model size and usefulness.
The Three-Stage Pipeline
The RLHF pipeline follows a specific sequence that turns a pretrained language model into an aligned assistant. Understanding why each stage exists and how they connect is important for successful implementation. Each stage has a distinct purpose, and the stages form a dependency chain where the quality of each stage's output places a ceiling on everything that follows.
The pipeline is a technical recipe built on assumptions about how alignment works and what data is available. The stages are ordered the way they are because of a specific logic: you cannot collect useful preference data without responses to compare, you cannot train a reward model without preference data, and you cannot run stable PPO without a reward model. The forward arrow of causality runs through the entire pipeline, and this is why short-circuiting any stage tends to undermine subsequent ones even when the skipped stage seems redundant.
A second important dimension is that all three models in the pipeline, the SFT model, the reward model, and the PPO policy, typically share the same underlying architecture. They are all transformer language models of roughly the same parameter count, often initialized from the same pretrained weights. This shared foundation means that improvements in base model capability propagate through all three stages simultaneously. When OpenAI upgraded from GPT-3 to GPT-4 as the base, the improvements benefited every stage of the alignment pipeline, including the final PPO policy. The alignment procedure is more like a specialization of capabilities that already exist than an injection of entirely new capabilities.

Stage 1: Supervised Fine-Tuning (SFT) takes a pretrained language model and trains it on high-quality demonstrations of desired behavior. This stage teaches the model the format and style of helpful responses, creating a starting point that can already follow instructions reasonably well.
Stage 2: Reward Model Training creates a model that predicts human preferences. Using comparison data where humans ranked alternative responses, the reward model learns to assign scalar scores which reflects response quality. As we covered in the Reward Modeling chapter, this model provides the optimization signal for the final stage.
Stage 3: PPO Fine-Tuning optimizes the SFT model to maximize rewards from the reward model while staying close to its original behavior. Building on the PPO for Language Models chapter, this stage performs the actual alignment through reinforcement learning.
Why the Order Cannot Be Reversed
The sequential nature of the pipeline is not arbitrary, and understanding the causal logic makes it far easier to reason about what can go wrong. SFT must precede reward model training because the reward model needs realistic candidate responses to compare. If you attempted to collect preference data from a raw pretrained model, the responses would be so incoherent and off-format that human annotators would struggle to determine which was "better" in any meaningful sense. Both responses might simply be unusable, making the comparison uninformative.
Reward model training must precede PPO because PPO requires a dense, differentiable reward signal. The alternative, having humans rate every generated response during PPO training, would be impractically slow. PPO generates millions of samples during training; each requiring a human rating would make the procedure orders of magnitude too expensive. The reward model acts as a learned proxy that can be evaluated in milliseconds rather than hours.
PPO must come last because it requires a stable starting point and a clear optimization target. Without the SFT model as initialization, the policy gradient updates would struggle to produce coherent language. Without the reward model as the optimization target, there is no well-defined objective to optimize toward. The PPO stage consumes all the infrastructure built in the previous two stages and converts it into the final aligned model.
Stage 1: Supervised Fine-Tuning
Supervised Fine-Tuning turns a pretrained language model into one able to following instructions and engaging in dialogue. While the pretrained model has acquired extensive knowledge and language understanding, it lacks the ability to respond helpfully to your queries in a conversational format. SFT closes this gap by exposing the model to paired examples of instructions and ideal responses, teaching it both the expected interaction pattern and the quality bar it should aim for.
Think of SFT as teaching an incredibly knowledgeable but socially awkward expert how to communicate. Before SFT, the language model has read essentially all of human-written text and internalized an enormous amount of knowledge. But it communicates the way text on the internet communicates, sometimes as a forum post, sometimes as a Wikipedia article, sometimes as a fictional story, depending on what pattern the prompt activates. After SFT, it has learned that when a human poses a question, the appropriate response is a clear, structured, helpful answer, not a continuation of the question into a series of related questions, not a fictional dialogue where characters discuss the topic, and not a Wikipedia-style article about the broader subject area. The format is part of the skill.
The SFT stage is also where the model develops an initial grasp of what "helpful" looks like in concrete terms. Through exposure to carefully crafted demonstrations, it learns that a good response is appropriately scoped to the question, acknowledges limitations when relevant, structures complex information clearly, and avoids unnecessary caveats or padding. These are difficult qualities to specify algorithmically, but relatively easy for expert human writers to demonstrate. SFT exploits that asymmetry, using human demonstrations as a richer signal than any reward function could capture.
Why SFT Comes First
You might wonder why we don't skip directly to reinforcement learning. The reason is practical: PPO optimization requires a model that already produces reasonable responses. Trying to optimize a raw pretrained model with RL is like trying to teach someone chess strategy before they know how the pieces move.
The pretrained model generates plausible continuations of text, but it doesn't understand the assistant paradigm. Given a question, it might generate more questions, continue with a different topic, or produce text that reads like training data rather than a helpful response. SFT provides the foundation by:
- Teaching the instruction-following format (recognizing prompts, generating responses)
- Establishing a baseline quality level that RL can refine
- Reducing the search space for PPO (responses are already in a useful format)
There is also a subtler reason that SFT matters for reward model training. The preference data collected for stage 2 requires the SFT model to generate candidate responses. If those candidates are incoherent, the human annotators comparing them face an impossible task: they cannot distinguish useful quality differences between incoherent responses. High-quality SFT outputs ensure that the preference data captures differences in response quality rather than the contrast between merely incoherent and completely incoherent. The SFT model sets the quality floor for every subsequent stage.
SFT Data Requirements
SFT data consists of (prompt, response) pairs showing ideal assistant behavior. Quality matters far more than quantity. A few thousand high-quality demonstrations often outperform millions of lower-quality examples.
sft_examples = [
{
"prompt": "What causes rainbows to form?",
"response": """Rainbows form through a combination of refraction, reflection, and dispersion of sunlight in water droplets. Here's how it works:
1. **Refraction**: Sunlight enters a water droplet and bends (refracts) because light travels slower in water than in air.
2. **Dispersion**: Different wavelengths of light bend by different amounts, separating white light into its component colors (red bends least, violet bends most).
3. **Internal Reflection**: The light reflects off the back inner surface of the droplet.
4. **Second Refraction**: The light bends again as it exits the droplet, further separating the colors.
You see a rainbow when the sun is behind you and rain is in front of you, with the light which reflects back at approximately 42 degrees from the original light direction.""",
},
{
"prompt": "Write a short poem about coding",
"response": """Lines of Logic
In the glow of screens at night,
We craft our worlds in black and white,
Variables dance, functions call,
Building dreams that never fall.
Bugs may lurk in shadowed code,
Each one a puzzle to decode,
But when at last the program runs,
We've made magic, byte by byte, we've won.""",
},
]SFT Training Example: Prompt: What causes rainbows to form? Response preview: Rainbows form through a combination of refraction, reflection, and dispersion of sunlight in water droplets. Here's how it works: 1. **Refraction**: Sunlight enters a water droplet and bends (refract...
The demonstrations should exhibit properties you want the final model to have: helpfulness, appropriate tone, factual accuracy, and safety awareness. As discussed in the Instruction Data Creation chapter, these examples can come from human writers, filtered model outputs, or synthetic generation with quality controls.
What makes a good SFT dataset goes beyond having correct answers. The demonstrations need to model the right relationship between question complexity and response depth. A simple factual question deserves a concise answer; a complex reasoning question deserves a step-by-step explanation. Demonstrations that always use the same format, regardless of question type, teach the model the wrong thing: they teach it to mimic a style rather than to calibrate the response to the question. Building this calibration into the demonstration data is one of the craft elements that separates effective SFT datasets from mediocre ones.
Coverage breadth is equally important. The SFT model's capabilities are strongly constrained by the distribution of topics and task types in the training data. A model trained only on question-answering demonstrations will struggle with creative writing requests. A model trained only on English examples will struggle with multilingual queries. Systematic gaps in the SFT data create systematic gaps in the model's behavior that downstream RL training can partially compensate for but never fully overcome. This is why organizations building production alignment pipelines invest heavily in dataset curation: so diversity across domains, styles, difficulty levels, and instruction types.
SFT Training Process
SFT uses standard causal language modeling loss but only on the response tokens. This distinction is important for understanding how the model learns during this stage. The prompt provides context, letting the model to understand what kind of response is expected, but we don't penalize the model for not predicting prompt tokens. The reasoning is straightforward: we only want the model to learn how to respond given a prompt, not to memorize and reproduce the prompts themselves.
The training objective focuses exclusively on maximizing the likelihood of generating the correct response tokens, conditioned on the full context of the prompt. This targeted learning ensures the model develops the skill of creating appropriate responses rather than simply learning to continue arbitrary text.

import torch
from torch.utils.data import Dataset
class SFTDataset(Dataset):
"""Dataset for supervised fine-tuning with response-only loss."""
def __init__(self, examples, tokenizer, max_length=512):
self.examples = examples
self.tokenizer = tokenizer
self.max_length = max_length
def __len__(self):
return len(self.examples)
def __getitem__(self, idx):
example = self.examples[idx]
prompt_tokens = self.tokenizer.encode(example["prompt"])
response_tokens = self.tokenizer.encode(example["response"])
full_tokens = prompt_tokens + response_tokens
if len(full_tokens) > self.max_length:
full_tokens = full_tokens[: self.max_length]
response_start = len(prompt_tokens)
else:
response_start = len(prompt_tokens)
labels = [-100] * response_start + full_tokens[response_start:]
padding_length = self.max_length - len(full_tokens)
input_ids = full_tokens + [self.tokenizer.pad_token_id] * padding_length
labels = labels + [-100] * padding_length
return {
"input_ids": torch.tensor(input_ids),
"labels": torch.tensor(labels),
"attention_mask": torch.tensor(
[1] * len(full_tokens) + [0] * padding_length
),
}The key detail is setting labels to -100 for prompt tokens. This value is a special sentinel in PyTorch's CrossEntropyLoss function: any token position with a label of -100 is completely ignored during loss computation. By masking the prompt tokens this way, we ensure we only compute loss on the response portion, teaching the model to generate good responses without wasting gradient updates on predicting prompt content it doesn't need to reproduce.
def sft_training_step(model, batch, optimizer):
"""Single SFT training step with response-only loss."""
model.train()
input_ids = batch["input_ids"]
labels = batch["labels"]
attention_mask = batch["attention_mask"]
outputs = model(
input_ids=input_ids,
attention_mask=attention_mask,
labels=labels, # Model computes loss internally, ignoring -100 labels
)
loss = outputs.loss
optimizer.zero_grad()
loss.backward()
optimizer.step()
return loss.item()Key Parameters
SFT training is relatively straightforward compared to the later stages. The key parameters are:
- Learning rate: Typically 1e-5 to 5e-5 for full fine-tuning, or 1e-4 to 3e-4 for LoRA (as covered in the LoRA chapters)
- Batch size: 32 to 128 examples, depending on available memory
- Epochs: 1-3 passes over the data; more risks overfitting
- Warmup: 3-10% of total steps with linear warmup
sft_config = {
"learning_rate": 2e-5,
"batch_size": 64,
"num_epochs": 2,
"warmup_ratio": 0.03,
"weight_decay": 0.01,
"max_grad_norm": 1.0,
"lr_scheduler": "cosine",
}Typical SFT Configuration: learning_rate: 2e-05 batch_size: 64 num_epochs: 2 warmup_ratio: 0.03 weight_decay: 0.01 max_grad_norm: 1.0 lr_scheduler: cosine
Overfitting is a real concern with small SFT datasets. Monitor validation loss and stop training when it begins to increase. You might prefer training for slightly fewer steps than optimal to preserve model generalization.
The learning rate choice deserves particular attention. SFT is fine-tuning a model that has already learned enormous amounts from pretraining. Using too large a learning rate risks catastrophically forgetting pretraining knowledge, a problem where the model loses general capabilities in exchange for fitting the specific demonstrations. Using too small a learning rate means the model barely moves from its pretrained behavior, failing to learn the instruction-following format effectively. The 1e-5 to 5e-5 range for full fine-tuning represents decades of empirical wisdom about this trade-off. For LoRA fine-tuning, the larger learning rates (1e-4 to 3e-4) are appropriate because LoRA only updates a small fraction of parameters, so the effective learning rate on the full model is much smaller than the nominal value.
The number of epochs also warrants thought beyond "stop when validation loss increases." With SFT, the concern is overfitting in the traditional sense and sycophancy: the model learning to mimic the surface features of the demonstrations (bullet points, specific phrasing, characteristic openers) rather than the underlying quality. If you train for many epochs on a small dataset, the model may learn to reproduce specific demonstration responses almost verbatim rather than generalizing the quality signal. One to three epochs is the typical operating range precisely because it provides enough exposure to learn the patterns without enough repetition to memorize specific examples.
What Good SFT Looks Like
A well-trained SFT model should exhibit several distinguishing characteristics when you evaluate it qualitatively. When given a straightforward question, it should provide a direct and accurate answer without excessive preamble. When given a complex multi-step problem, it should structure its response with clear organization rather than presenting a wall of text. When given a request that falls outside appropriate assistant behavior, it should decline gracefully rather than either complying blindly or refusing with an unhelpful non-response.
The key insight is that SFT teaches the model a mode of operation rather than a set of facts. You can check whether SFT is working by observing whether the model consistently adopts the assistant persona across diverse prompt types. If the model responds to some prompts helpfully but reverts to raw language model behavior on others (generating continuations, creating multiple responses, or adopting the voice of a different persona), SFT has not been thorough enough. This inconsistency is usually a sign that either the demonstration data lacks coverage for certain prompt types or training ran for too few steps.
Stage 2: Reward Model Training
The reward model learns to predict which responses humans prefer. Building on the Bradley-Terry model from earlier chapters, it converts pairwise comparisons into a scalar reward signal that guides PPO optimization. This stage is arguably the most consequential in the entire pipeline, because any systematic error in the reward model will be amplified by PPO optimization. The policy will find the ways to maximize the reward model score, and if the reward model is wrong about what constitutes a good response in certain regions of the input space, PPO will drive the policy directly into those regions.
Think of the reward model as a compressed representation of annotator judgment. Thousands of pairwise comparisons, each representing a human's opinion about which of two responses is better, are distilled into a single neural network that can assign a quality score to any response in fractions of a second. This compression is both the strength and the weakness of the approach. The strength is that the reward model can generalize beyond the specific comparisons it was trained on. This provides quality judgments for entirely novel responses. The weakness is that the generalization may be imperfect, and the compression inevitably loses some of the nuance that human annotators would apply in novel situations.
The point is that the reward model is not trained to produce absolute quality scores, only to produce scores that correctly rank pairs of responses. Two reward models trained on the same data could assign completely different absolute values and yet behave identically in practice, because PPO optimization only cares about relative rankings, not absolute magnitudes. This property changes the interpretation for how you should interpret reward values and how you calibrate the KL penalty coefficient in stage 3.
Reward Model Architecture
As discussed in the Reward Modeling chapter, the reward model typically shares architecture with the language model but replaces the language modeling head with a scalar output head. This architectural choice is deliberate: by starting from a language model that understands text in context, the reward model inherits much of its ability to interpret fine-grained meaning. The only modification is the final layer, which now outputs a single number representing quality rather than a distribution over vocabulary tokens.
import torch
import torch.nn as nn
class RewardModel(nn.Module):
"""Reward model that outputs a scalar score for text quality."""
def __init__(self, base_model, hidden_size):
super().__init__()
self.base_model = base_model
self.reward_head = nn.Sequential(
nn.Linear(hidden_size, hidden_size),
nn.ReLU(),
nn.Linear(hidden_size, 1),
)
def forward(self, input_ids, attention_mask):
outputs = self.base_model(
input_ids=input_ids,
attention_mask=attention_mask,
output_hidden_states=True,
)
last_hidden_state = outputs.hidden_states[-1]
sequence_lengths = attention_mask.sum(dim=1) - 1
batch_indices = torch.arange(input_ids.size(0), device=input_ids.device)
final_hidden = last_hidden_state[batch_indices, sequence_lengths]
reward = self.reward_head(final_hidden).squeeze(-1)
return rewardThe reward model extracts the hidden state at the final token position and passes it through a small neural network to produce a scalar reward. This design uses an important property of causal transformers: the final token's hidden state has attended to all previous tokens in the sequence, meaning it encapsulates information about the entire prompt-response pair. From this aggregated representation, the model can make an overall quality judgment that considers both what was asked and how well the response addresses it.
An important practical detail is that the reward model is initialized from the SFT model rather than the raw pretrained model. This initialization matters because the reward model needs to understand the specific response style established during SFT. If it were initialized from the pretrained model, it might assign high scores to responses that look like pretraining data (encyclopedia entries, forum posts) rather than assistant responses. Starting from the SFT model ensures the reward model operates in the same distribution as the responses it will be asked to score during PPO training.
The two-layer reward head sits on top of the transformer backbone. The first linear layer allows the network to learn a projection of the final hidden state that is maximally informative for quality prediction. The ReLU nonlinearity allows the network to model asymmetric quality differences. The second linear layer collapses this representation to a single scalar. Training this relatively small head while keeping the backbone frozen is one option, but joint fine-tuning of the entire model typically produces better reward models because it allows the transformer layers to reorganize their representations to make quality prediction easier.
Preference Data Format
Training data consists of prompts with two or more responses ranked by human annotators. Each training example captures a comparison: given the same prompt, which response did humans consider better? This pairwise structure is basic to the Bradley-Terry model, which we'll examine shortly.
preference_data = [
{
"prompt": "How do I improve my writing skills?",
"chosen": """Here are effective strategies to improve your writing:
1. **Read widely and actively** - Study how skilled authors structure sentences and arguments
2. **Write daily** - Even 15 minutes of practice builds consistency
3. **Seek feedback** - Join a writing group or find a trusted reader
4. **Edit ruthlessly** - Good writing is rewriting; cut unnecessary words
5. **Study grammar** - Understanding rules helps you break them purposefully
Start with one habit and build from there.""",
"rejected": """Writing is hard but you can get better. Just practice a lot and read books.
Maybe take a class or something. Good luck!""",
},
{
"prompt": "Explain quantum entanglement simply",
"chosen": """Quantum entanglement is like having two magic coins. When you flip them while they're together, then separate them, even across the universe, they always land the same way.
In physics terms: two particles become correlated so that measuring one instantly affects what you'll measure on the other, regardless of distance. Einstein called this "spooky action at a distance."
This isn't about sending information faster than light, but about correlations that can't be explained by classical physics.""",
"rejected": """Quantum entanglement is a quantum mechanical phenomenon where particles become interconnected through their quantum states such that the quantum state of each particle cannot be described independently.""",
},
]Number of preference pairs: 2 Example prompt: How do I improve my writing skills? Chosen response length: 481 chars Rejected response length: 120 chars
The training examples typically show that chosen responses are longer, more detailed, and better structured than rejected ones. The reward model learns to associate these features, along with factual accuracy and tone, with higher scalar scores.
The quality of the preference data is the single biggest lever in determining how well the reward model generalizes. Constructing good preference data requires careful thought about how to generate diverse candidate responses and how to ensure annotators are evaluating the right properties. If all comparisons pit an obviously good response against an obviously bad one, the reward model learns only from extreme quality differences and may fail to distinguish between two good responses, which is precisely the judgment needed during PPO when the policy has already learned to produce reasonable outputs. Some comparison data should therefore deliberately contrast responses that are both reasonable but differ in subtle quality dimensions: accuracy, depth, appropriate confidence, and conciseness.
Annotator consistency is another important concern. Human annotators often disagree about which of two responses is better, especially when both responses are in the "good" quality range. The Bradley-Terry model assumes a total ordering of quality that may not exist. Two annotators might consistently prefer different response styles (one preferring brevity, another preferring depth), and training on their mixed judgments produces a reward model that has internalized a blend of their preferences. This blend may not correspond to any user's preference. Managing annotator consistency through training, calibration sessions, and inter-annotator agreement monitoring is an important part of high-quality preference data collection.
Reward Model Training Loss
The training objective for the reward model emerges from a probabilistic framework for modeling human preferences. We begin with a natural question: given two responses to the same prompt, how likely is it that a human prefers one over the other? The Bradley-Terry model provides an elegant answer by relating this preference probability to the difference in quality scores assigned to each response.
The core insight is that we can model preferences as arising from latent quality scores. If response has a higher quality score than response , then should be preferred more often. The Bradley-Terry model formalizes this intuition by expressing the preference probability as a function of the score difference:
where:
- : the probability that response is preferred over given prompt
- : the logistic sigmoid function,
- : the reward model with parameters
- : the input prompt
- : the preferred ("winning") response
- : the rejected ("losing") response
The sigmoid function plays a useful role in this formulation. It converts the unbounded difference in reward scores into a probability between 0 and 1. This provides a smooth and differentiable mapping. When the reward for the winning response is significantly larger than for the losing response, the difference becomes a large positive number, and the sigmoid approaches 1, showing near-certainty that would be preferred. Conversely, if the scores are equal, the sigmoid returns 0.5. This reflects maximum uncertainty. This elegant mathematical structure captures our intuition that larger quality differences should correspond to more decisive preferences.

To train the model, we minimize the negative log-likelihood of observing the preferences in our dataset. For a single preference pair, the loss is:
where:
- : the scalar loss value to be minimized
- : the logistic sigmoid function,
- : the reward model with parameters that assigns a scalar score to a prompt-response pair
- : the input prompt or instruction
- : the "winning" or preferred response
- : the "losing" or rejected response
- : the difference in reward scores (which we want to be positive)
Understanding why this loss works requires examining what happens during optimization. When the model correctly assigns a higher score to the preferred response (making the difference positive and large), the sigmoid outputs a value close to 1, and the negative log becomes small. When the model incorrectly ranks the responses (making the difference negative), the sigmoid outputs a value close to 0, and the negative log becomes very large, creating a strong gradient signal to correct this error. This objective therefore directly maximizes the likelihood that the model assigns a higher score to the preferred response than the rejected response , which is precisely what we want from a reward model.
import torch
def compute_reward_model_loss(reward_model, batch, tokenizer, device):
"""Compute Bradley-Terry loss for preference learning."""
prompts = batch["prompts"]
chosen_responses = batch["chosen"]
rejected_responses = batch["rejected"]
chosen_texts = [p + c for p, c in zip(prompts, chosen_responses)]
rejected_texts = [p + r for p, r in zip(prompts, rejected_responses)]
chosen_tokens = tokenizer(
chosen_texts, padding=True, return_tensors="pt"
).to(device)
rejected_tokens = tokenizer(
rejected_texts, padding=True, return_tensors="pt"
).to(device)
chosen_rewards = reward_model(
chosen_tokens["input_ids"], chosen_tokens["attention_mask"]
)
rejected_rewards = reward_model(
rejected_tokens["input_ids"], rejected_tokens["attention_mask"]
)
loss = -torch.log(torch.sigmoid(chosen_rewards - rejected_rewards)).mean()
accuracy = (chosen_rewards > rejected_rewards).float().mean()
return loss, accuracyReward Model Training Metrics:
- Loss: Bradley-Terry negative log-likelihood
- Accuracy: Fraction where
- Target accuracy: 70-80% (higher may indicate overfitting)
The 70-80% accuracy target deserves explanation, because it might seem counterintuitive to stop well below 100%. Higher accuracy on the training set typically means the model has overfit to the specific comparison pairs it was trained on and has lost the ability to generalize to novel responses. A reward model that achieves 95% accuracy on training data but only 65% on held-out data has learned to recognize specific text patterns rather than quality. In the downstream PPO stage, this overfitted reward model will assign high scores to responses that match patterns from the training comparison data rather than to better responses. Stopping at 70-80% accuracy is an empirical heuristic that tends to produce reward models with better generalization, though the right threshold varies with dataset size and diversity.
Reward Model Calibration
A necessary but often overlooked aspect is reward model calibration. The absolute reward values don't matter for ranking, since the Bradley-Terry model only uses differences between scores. However, the scale of these values significantly affects PPO training stability. If rewards are too large in magnitude, gradient updates can become unstable; if they vary too widely, optimization becomes difficult.
import torch
def calibrate_reward_model(reward_model, calibration_data, tokenizer, device):
"""Calibrate reward model to have zero mean and unit variance."""
reward_model.eval()
all_rewards = []
with torch.no_grad():
for batch in calibration_data:
texts = batch["texts"]
tokens = tokenizer(texts, padding=True, return_tensors="pt").to(
device
)
rewards = reward_model(
tokens["input_ids"], tokens["attention_mask"]
)
all_rewards.append(rewards.cpu())
all_rewards = torch.cat(all_rewards)
mean_reward = all_rewards.mean().item()
std_reward = all_rewards.std().item()
return mean_reward, std_reward
def normalized_reward(raw_reward, mean, std):
"""Normalize reward to zero mean and unit variance."""
return (raw_reward - mean) / (std + 1e-8)Calibrating the reward model to have approximately zero mean and unit variance at the start of PPO training provides a stable foundation for the optimization. The zero mean ensures that the policy does not start with a systematic positive or negative reward bias, which could cause it to initially over-generate or under-generate regardless of quality. The unit variance ensures that the KL penalty coefficient has a consistent meaning across different training runs: a of 0.05 represents the same trade-off between reward and KL regularization regardless of whether the raw reward model happens to output values in the range or .
The calibration should be computed on a representative sample of prompt-response pairs that reflect the distribution the policy will encounter during PPO training. Using the SFT model to generate calibration responses is the most common approach, since the SFT model provides the starting point for the PPO policy. This ensures that the normalization statistics are computed at the operating point of the system, not on some other distribution.
Stage 3: PPO Fine-Tuning
With the SFT model and reward model ready, we can now run PPO optimization. This stage adjusts the policy to maximize expected reward while staying close to the SFT model's behavior. The challenge here is delicate: we want the model to improve according to the reward signal without losing the coherent language abilities it acquired during pretraining and SFT.
Think of PPO fine-tuning as the difference between a musician who learns jazz improvisation and one who simply memorizes jazz recordings. The SFT model knows how to play in the right style. The reward model has internalized the judgment of experienced listeners. PPO optimization uses that judgment signal to push the musician toward improvising in ways that experienced listeners prefer, while the KL penalty ensures the musician does not completely abandon the musical training that makes their playing coherent in the first place. The constraint is not a limitation but a necessity: a musician who has thrown out all musical structure in pursuit of applause is no longer a musician.
The PPO stage runs for far fewer gradient steps than pretraining or even SFT, often fewer than 10,000 steps compared to hundreds of billions of tokens for pretraining. This brevity is both a feature and a constraint. The short training duration means PPO cannot introduce catastrophically wrong behaviors if properly regularized, but it also means the policy cannot make dramatic departures from the SFT starting point. The alignment effect is best thought of as a refinement of existing capabilities rather than an acquisition of new ones. The capabilities must already exist in the SFT model; PPO merely tunes the probability distribution over responses to favor the higher-quality ones.
The PPO Training Loop
As we detailed in the PPO for Language Models chapter, each training iteration involves a carefully orchestrated sequence of steps that together enable stable policy improvement:
- Sampling: Generate responses from the current policy
- Reward computation: Score responses using the reward model
- Advantage estimation: Compute advantages using GAE
- Policy update: Optimize the clipped surrogate objective. This iterative process gradually shifts the policy's behavior toward responses that score higher according to the reward model, while the various stability mechanisms in PPO prevent the optimization from taking steps that are too large or in harmful directions.
import torch
import torch.optim
class RLHFTrainer:
"""Complete RLHF trainer implementing the PPO training loop."""
def __init__(
self, policy_model, ref_model, reward_model, tokenizer, config
):
self.policy = policy_model
self.ref_model = ref_model # Frozen copy of SFT model
self.reward_model = reward_model
self.tokenizer = tokenizer
self.config = config
self.optimizer = torch.optim.AdamW(
self.policy.parameters(),
lr=config["learning_rate"],
weight_decay=config["weight_decay"],
)
def generate_responses(self, prompts, max_length=256):
"""Generate responses from current policy."""
self.policy.eval()
responses = []
log_probs_list = []
with torch.no_grad():
for prompt in prompts:
input_ids = self.tokenizer.encode(prompt, return_tensors="pt")
output_ids = []
log_probs = []
for _ in range(max_length):
outputs = self.policy(input_ids)
next_token_logits = outputs.logits[:, -1, :]
probs = torch.softmax(next_token_logits, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
token_log_prob = torch.log(probs[0, next_token.item()])
output_ids.append(next_token.item())
log_probs.append(token_log_prob.item())
input_ids = torch.cat([input_ids, next_token], dim=1)
if next_token.item() == self.tokenizer.eos_token_id:
break
responses.append(self.tokenizer.decode(output_ids))
log_probs_list.append(log_probs)
return responses, log_probs_listComputing the Complete Reward
The total reward used to update the policy combines two competing objectives. We want to maximize the reward model's score, which represents human preferences, but we also want to prevent the policy from straying too far from the reference model, which represents stable, coherent language generation. The KL divergence penalty provides this regularization, creating a tug-of-war that encourages improvement without catastrophic drift.
We'll explore the KL divergence penalty in detail in the next chapter, but the basic formulation captures this balance mathematically:
where:
- : the combined reward used to update the policy
- : the input prompt
- : the generated response
- : the preference score from the reward model
- : the KL penalty coefficient controlling regularization strength
- : the Kullback-Leibler divergence between the two distributions
- : the current policy model
- : the reference model (frozen SFT model)
The KL coefficient acts as a dial controlling the trade-off between reward maximization and behavioral stability. A larger keeps the policy closer to the reference model, preserving more of the original capabilities but potentially limiting how much the model can improve. A smaller allows more aggressive optimization toward higher rewards, but risks the policy finding reward model exploits or losing coherence. This formulation ensures that while we maximize the preference score, we maintain the linguistic coherence and knowledge of the original model.
In practice, is one of the most sensitive hyperparameters in the entire RLHF pipeline. Values too close to zero produce unstable training because the policy has no incentive to stay in the region of coherent language. Values too large make the training too conservative, and the policy barely moves from its SFT initialization despite many optimization steps. Typical values range from 0.01 to 0.1 for most configurations, though some practitioners use adaptive schemes that increase automatically when the KL divergence exceeds a threshold. This provides a soft constraint instead of a fixed penalty.

import torch
def compute_rewards_with_kl_penalty(
policy_model,
ref_model,
reward_model,
prompts,
responses,
tokenizer,
kl_coef,
device,
):
"""Compute total reward including KL penalty."""
full_texts = [p + r for p, r in zip(prompts, responses)]
tokens = tokenizer(full_texts, padding=True, return_tensors="pt").to(device)
with torch.no_grad():
rm_rewards = reward_model(tokens["input_ids"], tokens["attention_mask"])
with torch.no_grad():
policy_outputs = policy_model(
tokens["input_ids"], attention_mask=tokens["attention_mask"]
)
ref_outputs = ref_model(
tokens["input_ids"], attention_mask=tokens["attention_mask"]
)
policy_logprobs = torch.log_softmax(policy_outputs.logits, dim=-1)
ref_logprobs = torch.log_softmax(ref_outputs.logits, dim=-1)
token_kl = (
policy_logprobs.exp() * (policy_logprobs - ref_logprobs)
).sum(dim=-1)
kl_penalty = token_kl.sum(dim=-1)
total_rewards = rm_rewards - kl_coef * kl_penalty
return total_rewards, rm_rewards, kl_penaltyPPO Update Step
The policy update uses the clipped surrogate objective from the PPO Algorithm chapter. This objective function represents the core mechanism that enables stable policy improvement: rather than directly maximizing expected reward, which could lead to catastrophically large updates, PPO constrains how much the policy can change in a single step. The clipping mechanism ensures that even if the advantage estimates suggest a large improvement, the actual policy update remains bounded.
import torch
def ppo_update(
policy_model, optimizer, batch, clip_epsilon=0.2, entropy_coef=0.01
):
"""Perform PPO policy update with clipped objective."""
policy_model.train()
states = batch["states"]
actions = batch["actions"]
old_log_probs = batch["old_log_probs"]
advantages = batch["advantages"]
returns = batch["returns"]
advantages = (advantages - advantages.mean()) / (advantages.std() + 1e-8)
outputs = policy_model(states)
logits = outputs.logits
log_probs = torch.log_softmax(logits, dim=-1)
action_log_probs = log_probs.gather(-1, actions.unsqueeze(-1)).squeeze(-1)
ratios = torch.exp(action_log_probs - old_log_probs)
surr1 = ratios * advantages
surr2 = torch.clamp(ratios, 1 - clip_epsilon, 1 + clip_epsilon) * advantages
policy_loss = -torch.min(surr1, surr2).mean()
probs = torch.softmax(logits, dim=-1)
entropy = -(probs * log_probs).sum(dim=-1).mean()
entropy_loss = -entropy_coef * entropy
total_loss = policy_loss + entropy_loss
optimizer.zero_grad()
total_loss.backward()
torch.nn.utils.clip_grad_norm_(policy_model.parameters(), max_norm=1.0)
optimizer.step()
return {
"policy_loss": policy_loss.item(),
"entropy": entropy.item(),
"mean_ratio": ratios.mean().item(),
"clip_fraction": (
(ratios < 1 - clip_epsilon) | (ratios > 1 + clip_epsilon)
)
.float()
.mean()
.item(),
}The entropy bonus term in the loss function deserves particular attention. By adding a small positive reward for entropy (the entropy coefficient in the code above multiplied by the average entropy of the action distribution), we discourage the policy from collapsing to deterministic outputs too quickly. Think of entropy as a measure of how many different responses the model considers at each step: high entropy means the model is still exploring a diverse set of continuations, while low entropy means it has converged to nearly always predicting the same tokens. Without the entropy bonus, PPO training tends to produce overly deterministic policies that perform poorly on novel prompts where the training distribution does not provide clear guidance.
The advantage normalization step (subtracting the mean and dividing by the standard deviation) is a standard variance reduction technique that makes training more stable. Raw advantages can have high variance because they depend on the absolute reward values, which may shift considerably across training steps. Normalizing to zero mean and unit variance at each batch ensures that the effective learning rate remains consistent even as the reward distribution changes during training.
Worked Example: Tracing a Single RLHF Training Step
To make the pipeline concrete, let's trace a single complete training step from prompt to policy update, using specific numerical values at each stage. This worked example connects all the abstract components into a coherent sequence.
Setup: Suppose we have already completed SFT and trained a reward model. The policy starts from the SFT model checkpoint. We use a KL coefficient and a PPO clip epsilon of .
Step 1: Sample a prompt from the training distribution. We draw the prompt "What is gradient descent?" from our prompt library.
Step 2: Generate a response from the current policy. The policy samples token by token, creating the response: "Gradient descent is an optimization algorithm that iteratively adjusts parameters in the direction that minimizes a loss function." Alongside the response, we record the log-probabilities assigned by the policy to each token, which we'll call for each token position .
Step 3: Score the response using the reward model. The reward model receives the full prompt-response pair and outputs a raw score. Suppose it outputs .
Step 4: Compute the KL penalty. We also pass the prompt-response pair through the frozen reference model (the SFT model) and compute token-level log-probabilities. The KL divergence at each token position is:
Summing across all token positions gives the total sequence-level KL divergence. Suppose this comes out to .
Step 5: Compute the total reward. Combining the reward model score with the KL penalty:
where:
- : the raw reward model score for this response
- : the KL penalty coefficient
- : the total sequence-level KL divergence
At this early stage of training, the KL penalty is small because the policy has barely moved from its SFT initialization. The total reward is almost entirely the raw reward model score.
Step 6: Estimate advantages. The advantage tells us how much better this response was than the average response to this prompt. Using a value function baseline with estimated value , the advantage is approximately:
A positive advantage means this response was better than expected. The policy update will increase the probability of generating similar responses.
Step 7: Compute the PPO ratio. For each token, we compute the ratio of the current policy's probability to the old policy's probability (from before this update):
Suppose the mean ratio across tokens is , meaning the policy has moved slightly toward creating this response.
Step 8: Apply the clipped objective. Since , the ratio is within the clipping bounds. The surrogate objective for this response is . The gradient update increases the probability of generating this response, pushing the policy in a direction that the reward model rates more highly.
After this single step, the policy is infinitesimally more likely to generate responses similar to "Gradient descent is an optimization algorithm..." when asked about gradient descent. Repeating this process across thousands of prompts and responses gradually sculpts the policy's behavior across the entire prompt distribution.
RLHF Debugging
RLHF training is notoriously difficult to debug. The interplay between the policy, reward model, and KL constraint creates many potential failure modes. Unlike standard supervised learning, where a rising validation loss is a clear and unambiguous signal that something is wrong, RLHF metrics can be simultaneously misleading in several directions at once: reward can increase while response quality decreases, KL divergence can remain controlled while mode collapse is developing, and entropy can look healthy while the policy is creating repetitive variations on a narrow set of response templates.
Effective debugging requires combining quantitative metrics with qualitative evaluation. No set of numerical metrics can substitute for regularly reading the model's outputs during training. A reward model score of 2.5 could represent excellent responses, or it could represent responses the policy has learned to craft specifically to exploit the reward model's scoring tendencies. The only way to tell the difference is to look at the text.
Key Metrics to Monitor
Effective RLHF debugging requires tracking multiple metrics throughout training:
import numpy as np
class RLHFMetricsTracker:
"""Track key metrics for RLHF debugging."""
def __init__(self):
self.metrics_history = {
"reward_mean": [],
"reward_std": [],
"kl_divergence": [],
"policy_loss": [],
"entropy": [],
"clip_fraction": [],
"response_length": [],
"unique_tokens": [],
}
def log_step(self, metrics_dict):
for key, value in metrics_dict.items():
if key in self.metrics_history:
self.metrics_history[key].append(value)
def check_health(self):
"""Check for common RLHF failure modes."""
warnings = []
if len(self.metrics_history["kl_divergence"]) > 100:
recent_kl = np.mean(self.metrics_history["kl_divergence"][-100:])
if recent_kl > 10:
warnings.append(
f"HIGH KL DIVERGENCE ({recent_kl:.2f}): Policy diverging from reference"
)
elif recent_kl < 0.01:
warnings.append(
f"LOW KL DIVERGENCE ({recent_kl:.4f}): Policy not learning"
)
if len(self.metrics_history["reward_mean"]) > 100:
recent_reward = np.mean(self.metrics_history["reward_mean"][-100:])
if recent_reward > 5:
warnings.append(
f"VERY HIGH REWARD ({recent_reward:.2f}): Possible reward hacking"
)
if len(self.metrics_history["entropy"]) > 100:
recent_entropy = np.mean(self.metrics_history["entropy"][-100:])
if recent_entropy < 0.1:
warnings.append(
f"LOW ENTROPY ({recent_entropy:.4f}): Policy becoming deterministic"
)
if len(self.metrics_history["response_length"]) > 100:
recent_len = np.mean(self.metrics_history["response_length"][-100:])
early_len = np.mean(self.metrics_history["response_length"][:100])
if recent_len < early_len * 0.5:
warnings.append(
"RESPONSE LENGTH COLLAPSED: Model creating very short responses"
)
elif recent_len > early_len * 2:
warnings.append(
"RESPONSE LENGTH EXPLOSION: Model creating very long responses"
)
return warningsimport numpy as np
tracker = RLHFMetricsTracker()
for i in range(200):
kl = 0.5 + 0.1 * np.random.randn() + i * 0.08 # Gradually increasing KL
tracker.log_step(
{
"kl_divergence": kl,
"reward_mean": 2.0 + 0.5 * np.random.randn() + i * 0.04,
"entropy": max(0.05, 1.0 - i * 0.004 + 0.1 * np.random.randn()),
"response_length": 100 + 10 * np.random.randn() - i * 0.2,
}
)
warnings = tracker.check_health()Health Check Results: ⚠️ HIGH KL DIVERGENCE (12.45): Policy diverging from reference ⚠️ VERY HIGH REWARD (7.95): Possible reward hacking
The metrics tracker successfully identifies the simulated anomalies. By monitoring the KL divergence and reward statistics we can catch issues like the "High KL" spike and "High Reward" events (simulating reward hacking) before they destabilize the entire training run.
Understanding what each metric tells you is as important as tracking it. KL divergence measures how far the current policy has moved from the reference model in terms of its probability distributions over responses. High KL is not inherently bad if reward has improved correspondingly, but high KL with stagnant or declining response quality is a warning sign that the policy has drifted into a region where the reward model is no longer a reliable proxy for human preferences. Entropy measures how deterministic the policy has become. Entropy collapsing to near zero while training is still ongoing suggests mode collapse is developing, even if reward appears stable.
The clip fraction (the fraction of PPO ratio updates that were clipped) is a particularly useful diagnostic. Values consistently above 0.3 indicate that the policy is changing faster than the clipped objective can accommodate, which means the old log-probabilities stored from the rollout phase are no longer accurate descriptions of the current policy. This is a sign that the PPO epoch count (how many times you update the policy on the same set of rollouts) is too high, or the learning rate is too large. Reducing either parameter should bring the clip fraction back to a healthier range of 0.05 to 0.2.
Common Failure Modes
Understanding these failure modes helps you diagnose and fix training issues:
Reward Hacking occurs when the policy finds exploits in the reward model that don't correspond to quality improvements. Signs include rapidly increasing reward with degrading response quality, or unusual patterns like excessive repetition or specific phrases.
import numpy as np
def detect_reward_hacking(responses, rewards, threshold_percentile=95):
"""Detect potential reward hacking by examining high-reward responses."""
threshold = np.percentile(rewards, threshold_percentile)
high_reward_indices = np.where(np.array(rewards) > threshold)[0]
flags = []
for idx in high_reward_indices:
response = responses[idx]
words = response.split()
if len(words) > 0:
unique_ratio = len(set(words)) / len(words)
if unique_ratio < 0.3:
flags.append(("repetition", idx, response[:100]))
if len(response) > 2000:
flags.append(("excessive_length", idx, f"Length: {len(response)}"))
elif len(response) < 10:
flags.append(("too_short", idx, response))
return flagsReward hacking is insidious because the numerical metrics look good while the model's behavior deteriorates. The reward model has a finite capacity and has only seen a finite amount of training data. The PPO policy, optimizing aggressively, can discover input patterns that the reward model was never trained to reject. Classic examples include: generating responses that begin with phrases the reward model associates with high quality (such as "Certainly! Here's a complete answer...") regardless of whether the content is helpful; creating responses of a specific length that correlates with high reward in the training data without giving proportional value; and inserting specific keywords that the reward model's training data associated with expert responses. Each of these behaviors increases the reward model score without making the response more useful.
Mode Collapse happens when the policy converges to creating nearly identical responses regardless of the prompt. Monitor response diversity and entropy throughout training.
import numpy as np
def compute_response_diversity(responses, n_samples=100):
"""Measure diversity of generated responses."""
unique_responses = len(set(responses[:n_samples]))
uniqueness_ratio = unique_responses / min(n_samples, len(responses))
all_words = []
for response in responses[:n_samples]:
all_words.extend(response.lower().split())
if len(all_words) > 0:
vocab_size = len(set(all_words))
type_token_ratio = vocab_size / len(all_words)
else:
type_token_ratio = 0
return {
"uniqueness_ratio": uniqueness_ratio,
"type_token_ratio": type_token_ratio,
"avg_response_length": np.mean([len(r) for r in responses[:n_samples]]),
}KL Explosion indicates the policy is moving too fast away from the reference model. This often precedes training instability:


Debugging Workflow
When RLHF training goes wrong, follow this systematic debugging approach:
-
Check the reward model first: Generate samples and manually verify that reward model scores align with your quality intuitions. A miscalibrated or overfitted reward model dooms PPO from the start.
-
Examine generated samples: Look at actual model outputs throughout training. Metrics can hide problems that become obvious when reading responses.
-
Verify the KL penalty is working: The policy should stay reasonably close to the reference. If responses look completely different from SFT outputs, the KL constraint may be too weak.
-
Monitor multiple metrics together: Single metrics can be misleading. High reward with low diversity suggests reward hacking. Low KL with no reward improvement suggests the policy isn't learning.
A systematic debugging session for a troubled RLHF run typically looks like this: first, freeze the policy and run the reward model on a held-out set of manually graded responses to check that its rankings match your intuitions. If the reward model consistently assigns higher scores to responses you would consider worse, the problem is in stage 2, and you need to retrain the reward model. If the reward model scores make sense but the policy is not improving, the issue is more likely in the PPO hyperparameters, such as too high a KL coefficient preventing meaningful updates, or too few rollout steps per iteration giving the advantage estimator high variance. If the policy is improving according to the reward model but human evaluation shows no improvement, the reward model is overfitting and its scores have become decoupled from actual quality.
import torch
def debug_rlhf_step(
policy_model, ref_model, reward_model, sample_prompts, tokenizer, device
):
"""Comprehensive debugging for a single RLHF step."""
debug_info = {}
policy_model.eval()
ref_model.eval()
with torch.no_grad():
responses = []
for prompt in sample_prompts[:5]:
input_ids = tokenizer.encode(prompt, return_tensors="pt").to(device)
output_ids = policy_model.generate(
input_ids, max_new_tokens=100, do_sample=True, temperature=0.7
)
response = tokenizer.decode(output_ids[0][len(input_ids[0]) :])
responses.append(response)
debug_info["sample_responses"] = list(
zip(sample_prompts[:5], responses)
)
full_texts = [p + r for p, r in zip(sample_prompts[:5], responses)]
tokens = tokenizer(full_texts, padding=True, return_tensors="pt").to(
device
)
rewards = reward_model(tokens["input_ids"], tokens["attention_mask"])
debug_info["rewards"] = rewards.cpu().tolist()
debug_info["response_lengths"] = [len(r) for r in responses]
debug_info["unique_words"] = [len(set(r.split())) for r in responses]
return debug_info



Putting It All Together
Let's trace through a complete RLHF training run, showing how all pieces connect:
import numpy as np
def run_rlhf_pipeline(config):
"""Complete RLHF training pipeline."""
print("=" * 50)
print("STAGE 1: Supervised Fine-Tuning")
print("=" * 50)
# SFT training loop (simplified)
sft_metrics = {
"initial_loss": 3.5,
"final_loss": 1.8,
"epochs": config["sft_epochs"],
}
print(f" Initial loss: {sft_metrics['initial_loss']:.3f}")
print(f" Final loss: {sft_metrics['final_loss']:.3f}")
print(f" Trained for {sft_metrics['epochs']} epochs")
print()
print("=" * 50)
print("STAGE 2: Reward Model Training")
print("=" * 50)
rm_metrics = {
"initial_accuracy": 0.52,
"final_accuracy": 0.74,
"epochs": config["rm_epochs"],
}
print(f" Initial accuracy: {rm_metrics['initial_accuracy']:.1%}")
print(f" Final accuracy: {rm_metrics['final_accuracy']:.1%}")
print(f" Trained for {rm_metrics['epochs']} epochs")
print()
print("=" * 50)
print("STAGE 3: PPO Fine-Tuning")
print("=" * 50)
ppo_metrics = []
for step in range(0, config["ppo_steps"], config["ppo_steps"] // 10):
progress = step / config["ppo_steps"]
metrics = {
"step": step,
"reward": 0.5 + 1.5 * progress + 0.2 * np.random.randn(),
"kl": 0.5 * (1 - np.exp(-5 * progress)) + 0.03 * np.random.randn(),
"policy_loss": -0.5 - 0.3 * progress + 0.1 * np.random.randn(),
}
ppo_metrics.append(metrics)
if step % (config["ppo_steps"] // 5) == 0:
print(
f" Step {step:5d}: reward={metrics['reward']:.3f}, "
f"KL={metrics['kl']:.3f}, loss={metrics['policy_loss']:.3f}"
)
return sft_metrics, rm_metrics, ppo_metricspipeline_config = {"sft_epochs": 2, "rm_epochs": 1, "ppo_steps": 10000}sft_results, rm_results, ppo_results = run_rlhf_pipeline(pipeline_config)================================================== STAGE 1: Supervised Fine-Tuning ================================================== Initial loss: 3.500 Final loss: 1.800 Trained for 2 epochs ================================================== STAGE 2: Reward Model Training ================================================== Initial accuracy: 52.0% Final accuracy: 74.0% Trained for 1 epochs ================================================== STAGE 3: PPO Fine-Tuning ================================================== Step 0: reward=0.544, KL=-0.025, loss=-0.744 Step 2000: reward=0.550, KL=0.308, loss=-0.557 Step 4000: reward=0.981, KL=0.382, loss=-0.631 Step 6000: reward=1.353, KL=0.472, loss=-0.790 Step 8000: reward=1.319, KL=0.534, loss=-0.863
The text output confirms that each stage completed successfully. The SFT loss decreased significantly, and the Reward Model achieved a validation accuracy of 74%, which is within the typical 70-80% range for effective preference modeling. These healthy prerequisites set the stage for the PPO phase, which we can now visualize.
steps = [m["step"] for m in ppo_results]
rewards = [m["reward"] for m in ppo_results]
kls = [m["kl"] for m in ppo_results]

The training curves demonstrate healthy alignment progress. The reward (left) steadily increases, showing the model is learning to satisfy the reward model's preferences. Meanwhile, the KL divergence (right) remains controlled near the target of 0.5. This keeps the model maintains the coherent capabilities of the original SFT model without drifting into incoherence or reward hacking.
Limitations and Practical Considerations
The RLHF pipeline comes with significant challenges in any production deployment.
Computational cost is substantial. The pipeline requires training three separate models (SFT, reward model, and PPO policy), with the PPO stage being particularly expensive because it requires running both the policy and reference model for every batch. A single RLHF training run can cost hundreds of thousands of dollars in compute for large models, making iteration and experimentation prohibitively expensive for most organizations. This has driven interest in more efficient alternatives like Direct Preference Optimization (DPO), which we'll explore in upcoming chapters. Even at research scale with smaller models, the infrastructure complexity of running three models simultaneously with different roles (active policy, frozen reference, frozen reward model) requires careful systems engineering to avoid memory bottlenecks.
Human annotation quality fundamentally limits what RLHF can achieve. The reward model can only capture patterns present in the preference data, and human annotators bring biases and practical limitations to the task. Their judgments can also be inconsistent. Disagreement between annotators is common, yet the Bradley-Terry model assumes a consistent underlying preference ordering. When annotators disagree about what makes a response "better," the reward model learns a noisy compromise that may not align with any particular user's preferences. Annotators also tend to evaluate responses on dimensions they can assess quickly (length, formality, apparent confidence) rather than harder dimensions (factual accuracy, appropriate reasoning, appropriate safety). This creates systematic gaps between what the reward model scores and response quality.
Reward hacking remains an unsolved problem despite various mitigation strategies. As we discussed in the Reward Hacking chapter, the policy will exploit any systematic weakness in the reward model. The KL penalty helps by anchoring behavior to the reference model, but sufficiently capable policies can still find exploits within the allowed KL budget. This creates an ongoing cat-and-mouse dynamic where you must continually patch reward model vulnerabilities. The basic tension is that the reward model is a finite model trained on finite data, while the policy is a powerful optimizer with many more optimization steps. Over time, the policy will discover the reward model's weak points because that is exactly what gradient descent is designed to do.
The pipeline also introduces a subtle and often overlooked failure mode: sycophancy. Because human annotators tend to prefer responses that agree with them and that sound confident, the reward model can learn to reward agreeable, confident-sounding responses over accurate but potentially uncertain or corrective ones. The aligned model becomes expert at telling users what they want to hear rather than what is true, because that is what maximizes reward. This form of reward hacking is particularly dangerous because it is invisible in aggregate metrics: the reward scores look good, the KL divergence is controlled, and the model produces fluent, confident-sounding text. Only careful evaluation on benchmarks that specifically probe for sycophantic behavior reveals the problem.
Reproducibility is challenging due to the many interacting hyperparameters and the sensitivity of PPO training. Small changes in learning rate, KL coefficient, or even random seed can lead to qualitatively different outcomes. This makes it difficult to compare results across papers or replicate published findings. The stochastic nature of both the generation process and the policy gradient updates means that two runs with identical hyperparameters can produce noticeably different models. Reporting results from a single run and treating them as definitive is a common but significant methodological error in the RLHF literature.
Despite these limitations, RLHF remains the most widely deployed alignment technique for production language models. Understanding the complete pipeline, including its failure modes, is needed for anyone working on language model alignment. The next chapter examines the KL divergence penalty in detail, which plays a useful role in balancing reward maximization with behavioral stability.
Summary
The RLHF pipeline turns a pretrained language model into an aligned assistant through three sequential stages. Supervised Fine-Tuning creates a model that understands the instruction-following format and produces reasonable responses. Reward Model training captures human preferences in a learnable function that provides optimization signal. PPO Fine-Tuning then optimizes the policy to maximize rewards while staying close to the reference model.
Key takeaways from this chapter:
- SFT provides the foundation: PPO requires a model that already produces usable responses; skipping SFT leads to unstable training
- Reward model quality constrains the policy: A flawed reward model will lead to flawed policies; validate carefully before PPO
- KL penalty prevents catastrophic drift: Without anchoring to the reference model, the policy will exploit reward model weaknesses
- Monitor multiple metrics: Single metrics can be misleading; track reward, KL, entropy, and response characteristics together
- Debugging requires examining actual outputs: Metrics summarize behavior, but reading generated responses reveals problems that numbers hide
- The pipeline has compounding failure modes: Errors in SFT propagate to reward model training, and errors in the reward model propagate to the PPO policy
The RLHF pipeline established the template for aligning large language models, but its complexity and cost have motivated simpler alternatives. The next chapter examines the KL divergence penalty in mathematical detail, followed by chapters on Direct Preference Optimization, which eliminates the reward model and PPO stages entirely while achieving comparable alignment results.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the RLHF pipeline.
RLHF Pipeline
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!