Part of Language AI Handbook
Build neural networks that learn human preferences from pairwise comparisons. Topics include reward model architecture, Bradley-Terry loss.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Reward Modeling
In the previous chapter, we explored the Bradley-Terry model and how it provides a probabilistic framework for converting pairwise human preferences into a consistent scoring system. Now we turn to the practical question: how do we build a neural network that learns to predict these preferences at scale?
A reward model is a neural network that takes a prompt and response as input and outputs a scalar score showing how "good" that response is according to human preferences. This model is a proxy for human judgment. This allows us to provide dense feedback signals during reinforcement learning without requiring a human to evaluate every generated response. The reward model sits at the heart of RLHF: it translates sparse, noisy human preferences into a continuous signal that can guide policy optimization. Think of the reward model as a learned critic that has internalized the aesthetic sensibilities of your annotator pool. Once trained, it can evaluate thousands of responses per second, standing in for the human evaluators who could only score a small fraction of that volume.
Building an effective reward model requires careful consideration of architecture choices, loss functions, and evaluation methods. The model must generalize from a limited set of human comparisons to accurately score responses it has never seen before, including responses generated by future versions of the language model being trained. This generalization challenge is more subtle than typical supervised learning. We are not just asking the model to recognize patterns in text. We are asking it to internalize a subjective, context-dependent value system from a collection of binary comparison labels, then extrapolate that value system to an unbounded space of possible outputs. The better the model generalizes, the more reliably the downstream RL process will produce helpful behavior rather than exploiting statistical artifacts.
Consider the practical scale of the challenge. A human annotator reviewing responses might evaluate fifty to one hundred pairs per hour. A language model undergoing RLHF training might generate millions of candidate responses across the training run, each requiring a quality signal. Without an automated proxy, RLHF would be computationally intractable. The reward model bridges this gap, but only if it has learned a sufficiently faithful representation of human values. Every imperfection in the reward model becomes an adversarial opportunity for the policy optimizer, which will relentlessly probe for responses that exploit whatever inconsistencies exist in the learned scoring function.
The degree to which reward models succeed in this role explains much of the variation in quality between different RLHF-trained systems. Two language models trained with the same base architecture but different reward models can exhibit dramatically different behavior. A reward model that overweights surface stylistic features (confident tone, formatted lists, certain vocabulary) will steer the policy toward responses that look helpful without necessarily being helpful. A reward model that captures the substance of human preferences will instead favor accurate, informative responses suited to the request. Understanding reward-model construction and training, including how to evaluate the result, therefore gives you direct insight into why aligned language models behave the way they do.
This chapter covers the complete pipeline from architecture design through training and evaluation. We begin with the architectural choices that determine how a pretrained transformer is adapted into an evaluative model. We then derive the preference loss function from first principles and examine its mathematical properties. The practical training section covers data preparation and optimization, followed by the regularization choices that control overfitting. We cover evaluation methodology in depth, including the metrics that matter and the failure modes to watch for. A numerical worked example makes the loss computation concrete. The code section implements a complete reward model from scratch. We close by examining the basic limitations that make reward modeling one of the most active research areas in alignment.
Reward modeling as a component of RLHF was popularized by the InstructGPT paper from OpenAI (Ouyang et al., 2022), which demonstrated that a 1.3B parameter model trained with RLHF could outperform a 175B model on human preference benchmarks. The reward model in that system was trained on approximately 33,000 comparison pairs. The underlying idea of using a learned model to proxy for human judgment has roots in earlier preference-based reinforcement learning work, including Christiano et al. (2017), who demonstrated that a reward model could learn to play Atari games from human preference comparisons without access to the game score. Anthropic's work on Constitutional AI (Bai et al., 2022) extended the paradigm by using AI-generated comparisons to supplement human annotations, significantly scaling the amount of preference data available for training. The subsequent DPO line of work (Rafailov et al., 2023) showed that in some settings the explicit reward model can be bypassed entirely, folding preference learning directly into the policy update. However, explicit reward models remain the dominant approach for large-scale systems because they offer interpretability, reusability across policy versions, and the ability to apply separate quality controls to the reward signal before using it for optimization.
Reward Model Architecture
The standard approach to building a reward model starts with a pretrained language model and adds a simple regression head that maps the final hidden states to a scalar reward. This design philosophy reflects a key insight: rather than learning to understand language from scratch, we can use the rich representations already encoded in models that have been trained on large text corpora. The pretrained model provides the linguistic foundation, while the new regression head learns to interpret those representations through the lens of human preferences.
Think of the architecture as dividing the problem into two parts. The pretrained backbone is responsible for language understanding: it reads the prompt and response, parses the argument structure, identifies factual claims, and builds rich contextual representations of every token. The value head is responsible for quality assessment: given those rich representations, it extracts the signal most relevant to human preferences and collapses it to a single number. The backbone does the hard language work; the value head does the preference judgment. This division of labor is what makes the approach so sample efficient. Training the value head from scratch would require enormous amounts of preference data to simultaneously learn language understanding and preference assessment. By reusing a pretrained backbone, we only need the preference data to fine-tune the judgment layer.
The key insight is that the backbone's representations already encode many of the features that correlate with response quality: coherence, factual consistency, appropriate register, completeness of explanation. The value head's task is not to discover these features from scratch but to learn which combinations of already-discovered features predict human preference in the training domain.
The pretrained backbone is typically frozen or lightly fine-tuned during reward model training. Freezing the backbone is more computationally efficient and preserves the rich general language understanding. Light fine-tuning (using a very small learning rate for the backbone and a larger one for the value head) can improve performance when the training domain is significantly different from the pretraining corpus. In practice, most implementations fine-tune the entire model with a conservative learning rate, letting the backbone to gradually adapt while preserving its core representations.
Base Model Selection
Reward models typically use the same architecture family as the policy model they will train. If you're using RLHF to fine-tune a 7B parameter LLaMA model, your reward model might be initialized from the same pretrained checkpoint or a similar model in the family. This architectural alignment ensures the reward model can process the same input representations and has similar language understanding capabilities. The choice is not arbitrary: when the reward model shares the same "vocabulary" of internal representations as the policy model, it can more accurately evaluate the subtle qualities of responses that the policy might generate.
There is a meaningful tradeoff around reward model size. A larger reward model generally has better language understanding and can detect more subtle quality distinctions. However, it is also more expensive to run during the RL training loop, where it must be queried for every sampled response. In practice, teams often use a reward model that is the same size as or smaller than the policy, accepting some quality loss for the sake of training throughput. Some systems use reward model distillation to compress a high-quality large reward model into a smaller one that can be evaluated more cheaply during RL.
The key modification is replacing or augmenting the language modeling head. Instead of predicting the next token, we need to output a single scalar value that represents the quality of the entire response. This transformation converts a generative model into an evaluative one, shifting from the question "what comes next?" to "how good is this?" Because the parameters encoding language understanding are already in place, this transformation requires relatively little new learning. The backbone simply needs to route its existing representations through a new output projection.
One important architectural consideration is whether to initialize the reward model from the same checkpoint used to start the supervised fine-tuning (SFT) stage, or from the SFT model itself. In most RLHF pipelines, the SFT stage produces a model that has already learned to follow instructions and produce helpful responses. Initializing the reward model from this SFT checkpoint means the backbone already understands the response format expected by the system, potentially which makes it easier to learn quality distinctions within that format.
Value Head Design
The core architectural question for reward modeling is this: how do we collapse an entire sequence of hidden states, one for each token in the prompt and response, into a single number that captures the overall quality? The solution involves selecting a representative hidden state and projecting it down to a scalar through a learned transformation.
Think of the sequence of hidden states as a series of snapshots taken at each position as the model reads through the text. Each snapshot captures what the model "knows" at that point in the sequence, accumulated from all prior tokens via the attention mechanism. The final snapshot, taken after the model has processed the entire sequence, is the most information-rich because it has had access to everything that came before. Using this final snapshot as the input to the value head is like asking someone to score a piece of writing after they have read the whole thing, rather than after each sentence.
The reward model architecture can be expressed as:
where:
- : the scalar reward score output by the model
- : the prompt text input to the model
- : the response text to be scored
- : the hidden state vector at the final token position
- : the learned function (value head) mapping the hidden state to a scalar
This formulation captures the needed transformation at the heart of reward modeling. The input is a high-dimensional hidden state vector, perhaps 768 or 4096 dimensions depending on the model size, and the output is a single real number. The function must learn to extract and weigh the relevant features from this rich representation to produce a meaningful quality score.
This projection is typically implemented as a linear layer:
where:
- : the scalar output of the value head
- : the learned weight vector for the value head
- : the transpose of , letting the dot product with
- : the input hidden state vector
- : the learned bias term
- : the hidden dimension size of the base transformer model
The linear layer performs a weighted sum over all dimensions of the hidden state. Each component learns to assign importance to the corresponding dimension of the hidden representation. Positive weights mean that feature contributes positively to the reward, negative weights indicate a negative contribution, and weights near zero suggest the feature is irrelevant for quality assessment. The bias term shifts the overall reward scale, letting the model to center its predictions appropriately.
Some implementations use a two-layer MLP for the value head rather than a single linear layer:
where:
- : the first projection matrix, mapping the transformer hidden state to an intermediate dimension
- : the bias for the first linear layer
- : the rectified linear activation, introducing nonlinearity between the two layers
- : the final projection vector collapsing the intermediate representation to a scalar
- : the output bias
The nonlinear value head allows the model to learn more complex combinations of the backbone's features. In practice, the single linear layer often performs comparably to the MLP, especially when the backbone is large and already captures high-level quality features. The additional parameters of the MLP also increase the risk of overfitting on small preference datasets.
Why use the final token position? In autoregressive models, information flows from left to right through causal attention. Each token can only attend to tokens that came before it, creating a natural accumulation of information as we move through the sequence. The final token's hidden state has "seen" all preceding tokens in both the prompt and response, which makes it a natural summary of the entire sequence. By the time we reach the end, the model has processed every word, every argument, and every nuance. This is analogous to using the [CLS] token representation in BERT-style models, which we covered in Part XXIV, where a special token is positioned to aggregate information from the entire input.
Some implementations average over all response token positions instead:
where:
- : the mean pooled hidden representation
- : the number of tokens in the response
- : the hidden state at token position
- : a summation over all token positions belonging to the response
This mean pooling approach treats all positions as equally important and computes their centroid in hidden space. The intuition is that every part of the response matters, and averaging captures the "typical" representation across the sequence. However, using the final token is more common in practice because it naturally captures the complete context and requires no additional computation. The causal attention mechanism has already done the work of aggregating information, so the final position provides a ready-made summary. A potential advantage of mean pooling is greater stability when the most important content appears in the middle rather than the end, but in practice this benefit is small relative to the conceptual simplicity of the final-token approach.
A third variant, used in some systems, extracts the hidden state at the position corresponding to a special end-of-response token explicitly appended to every input. This ensures there is always a dedicated "summary" token regardless of what appears last in the tokenized sequence, which can improve consistency when response lengths vary widely.
Handling Variable-Length Inputs
The reward model must handle prompt-response pairs of varying lengths. Some prompts are brief questions, others are lengthy instructions with context. Similarly, responses range from terse answers to elaborate explanations. The architecture must gracefully accommodate this variability while maintaining consistent scoring semantics.
Think of the variable-length input problem as managing a sliding measuring tape. The reward always needs to be extracted from the right end of the tape, but the tape itself is a different length every time. The solution is to track where the tape ends and always read the final marker, regardless of the overall length.
The input is formatted as a concatenation of the prompt and response:
[prompt tokens] [response tokens] [EOS]
The model processes this sequence through the transformer layers, and we extract the hidden state at the EOS (end-of-sequence) token position for the value head. This approach uses the causal attention mechanism to ensure the reward is computed based on the complete context. The EOS token is a natural boundary marker, signaling where the response ends and giving a consistent extraction point regardless of the sequence length.
In practice, variable-length inputs require careful handling of attention masks during batching. When sequences in a batch have different lengths, shorter sequences are padded to match the length of the longest sequence in the batch. Padding tokens should not contribute to the final representation, so we must identify the position of the actual last token (not the last padding token) for each sequence. This is typically implemented by summing the attention mask to find the sequence length, then using that length to index into the hidden states.
The maximum sequence length is another practical constraint. Most transformer architectures have a fixed context window (2048 tokens for earlier models, 4096 or more for recent ones). Prompt-response pairs that exceed this limit must be truncated. The standard approach truncates from the right, which means very long responses may be partially cut off. Some implementations preferentially truncate the prompt rather than the response, reasoning that the response quality is what we are scoring. Choosing where to truncate involves a tradeoff: truncating responses can cause the model to miss important quality signals, while truncating prompts can deprive the model of context needed to judge appropriateness.
In reinforcement learning terminology, a reward function gives the immediate reward for taking action in state . A value function estimates the expected cumulative future reward from state . Our reward model acts as a reward function. This provides a scalar score for the completed response. The value function used during PPO training is a separate component, which we'll discuss in upcoming chapters on policy optimization.
Preference Loss Function
The reward model is trained on human preference data, where annotators indicate which of two responses they prefer for a given prompt. We need a loss function that encourages the model to assign higher rewards to preferred responses. The key insight is that we do not need absolute quality labels. Instead, we only need relative comparisons, and from these pairwise signals, the model can learn a consistent scoring function.
Think of this like training a judge to rank competitors. You do not need to know that competitor A deserves a score of 7.4 out of 10. You only need to know that A is better than B, and from enough of these comparisons, the judge gradually develops an internal scale. The Bradley-Terry preference loss formalizes this intuition: each comparison example provides a gradient signal that nudges the scoring function to separate the preferred response from the rejected one, and the accumulated effect of many such nudges produces a globally consistent ranking.
The preference loss operates entirely in the space of reward differences, not absolute values. This shift-invariance has a practical implication: you never need to define what a "good" absolute reward score looks like, and you never need to normalize rewards during training. The model learns to separate preferred from rejected responses regardless of where on the number line those scores happen to fall. Two reward models that produce identical rankings but different absolute values are functionally equivalent for the purpose of RL training, and the loss function makes no distinction between them.
This shift-invariance also means that the absolute magnitude of reward scores is not meaningful on its own. A reward of 3.7 says nothing about quality in isolation. What matters is whether 3.7 is higher or lower than the reward assigned to alternative responses for the same prompt. The reward model's job is to learn an ordering, not a cardinal scale. However, the downstream RL algorithm often benefits from reward signals that are normalized to a consistent range, which is why reward normalization is applied as a post-processing step after training.
Deriving the Loss from Bradley-Terry
The mathematical foundation for our loss function comes from the Bradley-Terry model, which provides a principled way to convert scalar scores into preference probabilities. As we established in the Bradley-Terry chapter, the probability that response (the winner) is preferred over response (the loser) given their reward scores is:
where:
- : the probability that response is preferred over given prompt
- : the prompt text input
- : the logistic sigmoid function,
- : the scalar score output by the reward model for a given input pair
- : the winning and losing responses, respectively
This equation has a beautiful interpretation. The preference probability depends only on the difference between rewards, not their absolute values. If one response scores 10 points higher than another, the preference probability is the same whether the scores are (15, 5) or (105, 95). This shift-invariance is both mathematically convenient and practically useful: it means the model only needs to learn relative quality, not calibrate to some arbitrary absolute scale.
The sigmoid function turns this difference into a valid probability between 0 and 1. When the reward difference is zero, both responses are equally preferred with probability 0.5. As the difference grows positive (the winner scores much higher), the probability approaches 1. As it grows negative (the winner scores lower, showing a model error), the probability approaches 0.
Using Python 3.11.14 environment at: /private/tmp/mb-language-ai-modern-plots/books/_quarto_language-ai-handbook/.venv Checked 2 packages in 4ms

To train the model, we maximize the log-likelihood of the observed preferences. This corresponds to minimizing the negative log-likelihood, which is equivalent to the binary cross-entropy loss applied to the reward difference. The derivation proceeds through several algebraic steps, each illuminating a different aspect of the loss function. We can derive the final loss form step-by-step:
where:
- : the preference loss to be minimized
- : the probability that response is preferred
- : the logistic sigmoid function
- : the prompt text input
- : the winning and losing responses
- : the reward difference (negative when the preferred response scores higher, as desired)
- : the exponential function
- : the natural logarithm
The final form of the loss, , is known as the softplus function applied to the negative reward margin. This form is numerically stable and commonly implemented directly in deep learning frameworks as F.logsigmoid(reward_diff).neg() or equivalently via the binary cross-entropy function with logits.
The main point behind this loss is that it turns a ranking problem into a maximum likelihood problem. Instead of asking "is A ranked above B?", it asks "what is the probability that A is preferred over B, and how do we maximize that probability?" This probabilistic framing enables smooth gradient-based optimization and naturally handles noisy or inconsistent annotations by letting the model to express uncertainty through reward differences that are neither very large nor very small.
Intuition Behind the Loss
This equation has a clear interpretation that connects directly to how we want the model to behave. Consider what happens in different scenarios.
When the reward model correctly assigns a much higher score to the preferred response (), the difference is a large positive number, and of that difference approaches 1. Taking the negative log gives a loss close to zero. The model has confidently made the right prediction, so there is little left to learn from this example.
Conversely, when the model incorrectly assigns a higher score to the rejected response, the difference is negative, outputs a value near 0, and the negative log produces a large loss. This large loss creates strong gradients that push the model away from its incorrect belief.

The gradient of this loss pushes the model to:
- Increase (the preferred response's reward)
- Decrease (the rejected response's reward)
The magnitude of these updates depends on how confident the model currently is. If the model already strongly prefers the correct response, gradients are small. If it's uncertain or wrong, gradients are larger. This is the standard behavior of log-loss functions and provides natural calibration during training. The model learns most from examples where it is wrong or uncertain, while examples it already handles well contribute minimally to parameter updates.
One subtle property of the preference loss is that it has no natural floor for the margin. Even if the model correctly ranks a pair with a margin of 100, the loss is not exactly zero; it approaches zero but never reaches it. This means training never "stops" in the sense of achieving exactly zero loss on the training data, which acts as a mild form of implicit regularization. The model is always being pushed to increase the margin, which encourages confident predictions, but the diminishing gradient ensures this pressure decreases as confidence grows.
Batch Loss Formulation
In practice, we train on batches of preference pairs rather than individual examples. For a dataset of preference pairs , the full training objective is:
where:
- : the total loss over the dataset parameterized by model weights
- : the total number of preference pairs in the batch
- : summation over all examples in the batch
- : the prompt for the -th example
- : the preferred and rejected responses for the -th example
- : the reward model function parameterized by
- : the sigmoid function converting the score difference into a probability
- : the natural logarithm
The averaging by normalizes the loss across different batch sizes. This keeps learning rates behave consistently regardless of batch configuration. This is the standard reward modeling loss used in systems like InstructGPT, Anthropic's Constitutional AI, and most open-source RLHF implementations. Its simplicity and effectiveness have made it the default choice for training reward models from pairwise preferences.
From a statistical perspective, minimizing this loss is equivalent to maximum likelihood estimation of the Bradley-Terry model parameters. We are finding the reward function that maximizes the joint probability of observing all the preference comparisons in our dataset, under the assumption that preferences follow the Bradley-Terry model. This principled probabilistic foundation distinguishes the preference loss from ad hoc margin-based objectives and ensures the resulting reward model has well-defined calibration properties.
Margin-Based Variants
Some implementations add a margin term to encourage larger reward differences between preferred and rejected responses:
where:
- : the margin-based preference loss
- : the prompt text
- : the preferred and rejected responses
- : a fixed margin hyperparameter so the winner's score exceeds the loser's by at least
- : the raw reward difference
- : the sigmoid function
- : the natural logarithm
The margin acts as a buffer zone. With a margin of, say, 0.5, the model is penalized unless the preferred response scores at least 0.5 points higher than the rejected one. This prevents the model from being satisfied with arbitrarily small reward differences, even when the ranking is technically correct.
This penalizes the model even when it correctly ranks preferences but with a small margin, encouraging more confident predictions. The idea is that a reliable reward model should produce clearly separated scores. This makes it easier for downstream RL algorithms to distinguish good from bad responses. However, this can hurt calibration and is not universally used. The margin effectively changes the decision boundary from zero to , which may not align with the true underlying preference probabilities.
An adaptive variant uses a learned or example-specific margin based on annotator confidence. If multiple annotators rated a comparison and they agreed strongly, the margin is larger. If they disagreed or the comparison was rated as close, the margin is smaller. This approach more accurately reflects the difficulty of each comparison and can lead to better-calibrated reward models, but requires richer annotation metadata that is not always available.
Reward Model Training
Training a reward model involves several practical considerations beyond the loss function, including data preparation, optimization settings, and regularization strategies. Getting these details right matters as much as the architectural and loss function choices. A perfectly designed reward model architecture can still produce poor results if the training procedure is poorly configured, especially when working with the relatively small preference datasets typical in RLHF pipelines.
Think of reward model training as sitting in a particularly challenging part of the spectrum between pretraining and fine-tuning. Unlike pretraining, where we have abundant data and are learning from scratch, reward model training uses a much smaller dataset of human comparisons. Unlike standard fine-tuning, where we have labeled examples of the correct output, reward model training has only relative judgments with no absolute quality labels. This combination of data scarcity and label ambiguity requires careful handling to produce a model that generalizes rather than memorizes.
The main point governing most reward model training decisions is that we are trying to learn a smooth, generalizable preference function from a small number of noisy comparisons. Every design choice should be evaluated against this goal. Conservative learning rates prevent catastrophic forgetting of useful pretrained representations. Early stopping prevents memorization of specific training pairs. Regularization smooths the learned reward function. And careful data preprocessing ensures the comparisons we train on reflect human values rather than annotation artifacts.
Data Preparation
Each training example consists of a prompt and two responses with a preference label. The standard format is:
{
"prompt": "Explain quantum entanglement to a 10-year-old.",
"chosen": "Imagine you have two magic coins...",
"rejected": "Quantum entanglement is a phenomenon...",
}{'prompt': 'Explain quantum entanglement to a 10-year-old.',
'chosen': 'Imagine you have two magic coins...',
'rejected': 'Quantum entanglement is a phenomenon...'}During training, we need to compute rewards for both responses. A common approach processes both responses in a single forward pass by concatenating them:
[prompt] [chosen response] [EOS] [PAD] ... [prompt] [rejected response] [EOS]
The model computes hidden states for both sequences, extracts the final token representations for each, passes them through the value head, and computes the loss on their difference.
Data quality matters enormously at this stage. Before training, the dataset should be audited for several issues. First, check for duplicate or near-duplicate pairs, which can cause the model to overfit to specific examples. Second, examine the distribution of preference pairs across different prompt types, since a dataset dominated by one domain (say, coding questions) will produce a reward model that generalizes poorly to other domains. Third, filter out pairs where the quality difference between chosen and rejected is very small or where the human label may be unreliable. Pairs where annotators showed high disagreement are informative but should be handled carefully, either by upweighting them during training or by assigning softer labels.
The formatting of prompt-response pairs also deserves attention. In production RLHF systems, the prompt and response are typically formatted using the same chat template that the policy model was fine-tuned on. This ensures the reward model sees the same input format it will encounter during RL training. Mismatches between the format seen during reward model training and the format seen during RL training can cause systematic scoring errors that are difficult to diagnose.
Optimization Configuration
Reward model training typically uses conservative hyperparameters. These choices reflect the need to adapt a pretrained model to a new task while preserving its core language understanding capabilities.
The standard configuration includes:
- Learning rate: to , lower than instruction tuning to preserve pretrained knowledge
- Batch size: Large batches (32-128 pairs) help with gradient stability
- Epochs: 1-3 epochs over the preference data; more risks overfitting
- Optimizer: AdamW with weight decay 0.01-0.1
- Learning rate schedule: Linear warmup for the first 5-10% of training steps, followed by linear decay
The learning rate is kept low because we're building on a pretrained model that already has strong language understanding. We want to learn the preference structure without disturbing the underlying representations too much. Think of the pretrained backbone as a carefully calibrated instrument: we need to tune it slightly for our specific measurement task, but we should be careful not to disturb its basic sensitivity. A large learning rate would be like recalibrating an instrument with a blunt screwdriver, potentially destroying the precision that makes it useful.
Different learning rates for the backbone and the value head can improve performance. The value head is initialized randomly and needs to learn from scratch, so a higher learning rate (e.g., ) helps it converge quickly. The backbone is already well-trained and only needs minor adjustments, so a lower rate (e.g., ) prevents destructive updates. This pattern, called "differential learning rates" or "layerwise learning rate decay," is common in fine-tuning practice and applies especially well to reward modeling.
Regularization Considerations
Reward models are susceptible to overfitting, especially when trained on limited preference data. Several techniques help prevent the model from memorizing training pairs rather than learning the underlying preference structure.
Early stopping based on validation accuracy is needed. The model should generalize to held-out preferences, not memorize training pairs. Hold out at least 10% of your preference data as a validation set and monitor performance after each epoch. Stop training when validation accuracy plateaus or begins to decline, even if training accuracy continues to improve. This is a clearer signal of overfitting for reward models than it is for standard classification tasks, because the training distribution of human comparisons is often narrower than the distribution of responses the model will encounter during RL training.
Dropout in the value head (though not typically in the pretrained layers) adds regularization without affecting the language model's representations. A dropout rate of 0.1 to 0.2 on the value head prevents the model from becoming overly reliant on specific features that happen to correlate with the training labels. The pretrained backbone layers should generally not have their dropout rates modified, as those were set during pretraining for different purposes.
Label smoothing can help with noisy labels. If some human preferences are inconsistent or reflect borderline cases, treating them as probabilistic rather than hard labels reduces sensitivity to noise. With label smoothing parameter , a preference label of 1 (chosen is better) becomes , and the model is trained to predict a probability of rather than 1. This discourages overconfident predictions and makes the model more reliable to annotation noise. A typical value is .
Training Stability
The reward scale is arbitrary, unlike classification where outputs are bounded probabilities. This can cause optimization instability, particularly in the early stages of training when the value head is freshly initialized and may produce extreme reward values. Several techniques help maintain stable optimization throughout training.
Reward normalization after training scales rewards to have zero mean and unit variance over a reference set. This makes the reward signal easier to use in downstream RL training. The normalization is computed on a held-out reference set rather than the training data. This keeps it represents the typical distribution of responses the reward model will encounter during RL. Specifically, we collect a large set of prompt-response pairs (using the SFT model to generate responses to a diverse set of prompts), compute the rewards for all of them, and record the mean and standard deviation. During RL training, rewards are normalized using these statistics: .
Gradient clipping prevents large updates from outlier examples. A max gradient norm of 1.0 is typical. Without clipping, a single training example with an unusually large loss can cause a large gradient update that disrupts the model's existing knowledge. Clipping ensures that no single example has an outsized influence on the parameters, making training more reliable to label noise and distribution outliers.
Learning rate warmup over the first 5-10% of training helps stabilize early optimization. During warmup, the learning rate gradually increases from a very small value (e.g., 0) to its target value. This prevents the early training steps from making large, potentially destructive updates when the value head is still randomly initialized and creating arbitrary rewards.
Reward Model Evaluation
Evaluating reward models is important but challenging. Unlike language models where we can measure perplexity on held-out text, reward models are evaluated on their ability to predict human preferences. The evaluation must assess whether the learned reward function captures the underlying preference structure, not just memorizes the training comparisons.
Think of reward model evaluation as answering three distinct questions simultaneously. First, does the model correctly rank responses on held-out examples from the same distribution as the training data? Second, does it correctly rank responses from distributions it has not seen, including responses generated by the policy during RL training? Third, are there systematic biases or shortcuts in the model's scoring that could be exploited during RL optimization? A model can pass the first test while failing the second and third, which is the evaluation trap that allows reward hacking to occur. A thorough evaluation probes all three dimensions.
The difficulty of complete evaluation is one of the basic challenges in reward modeling research. We typically only have access to held-out human comparison data from the same annotation protocol used for training, which addresses only the first question. Measuring out-of-distribution generalization requires collecting new annotations for distribution-shifted samples, which is expensive. Detecting exploitable biases requires adversarial testing: constructing responses specifically designed to exploit suspected weaknesses and checking whether the reward model scores them inappropriately. This kind of adversarial red-teaming is resource-intensive but invaluable before deploying a reward model in RL training.
Primary Metrics
Pairwise accuracy is the most direct evaluation. On a held-out set of preference pairs, we measure how often the reward model assigns a higher score to the human-preferred response. This metric directly measures the model's ability to perform its intended function: distinguishing better responses from worse ones.
where:
- : the fraction of correctly ranked pairs
- : the total number of examples in the test set
- : summation over all test examples
- : the prompt for the -th test example
- : the preferred and rejected responses for the -th test example
- : the indicator function, evaluating to 1 if the condition is true and 0 otherwise
- : the predicted reward for the -th pair's responses
The indicator function returns 1 when the condition inside is true (the chosen response receives a higher reward) and 0 otherwise. Summing these indicators and dividing by the total count gives us the proportion of correctly ranked pairs.
A random model achieves 50% accuracy, while human inter-annotator agreement typically ranges from 65-80% depending on the task difficulty. A good reward model should approach but not necessarily exceed human agreement levels, since disagreements in the training data cap achievable performance. If humans themselves only agree 75% of the time on which response is better, we cannot expect the model to exceed this ceiling. In fact, a model achieving much higher accuracy might be exploiting artifacts in the data rather than capturing preferences. The ceiling imposed by inter-annotator agreement is not a failure of reward modeling but a reflection of ambiguity in human values, and reward models that appear to exceed this ceiling should be scrutinized carefully for data leakage or shortcut learning.
Calibration measures whether the model's confidence matches its accuracy. If the model predicts a preference probability of 0.8, it should be correct about 80% of the time for similar-confidence predictions. Calibration is important because the reward differences will be used as optimization signals. A well-calibrated model produces reliable gradients: large reward differences indicate strong preferences, while small differences reflect uncertainty. Poor calibration in either direction is problematic. Overconfidence means the model assigns large margins to pairs it only weakly prefers, which can cause the RL optimizer to over-exploit noisy reward signals. Underconfidence means the model assigns small margins even to clear quality differences, which can make the RL optimization slow or unstable.
Agreement with Human Evaluators
Beyond automatic metrics, direct comparison with human evaluations provides insight into model quality. These evaluations are more expensive but can reveal failure modes that automatic metrics miss.
Correlation with human scores offers a continuous measure of alignment. If human evaluators rate responses on a Likert scale (1-5), we can measure Spearman correlation between model rewards and human ratings. Spearman correlation is appropriate here because we care about the ranking order rather than the absolute values of the scores. A correlation of 0.7 or higher with experienced human evaluators is generally considered good performance, though the threshold depends on the task difficulty and the consistency of the human ratings themselves.
Head-to-head win rates provide a complementary evaluation. Show humans new responses ranked by the reward model and ask them to validate the rankings. This catches cases where the model has learned spurious correlations. For example, if the model systematically overscores responses that use bullet points, head-to-head evaluation with human judges will reveal this: the judges will find that the model's highly-scored bulleted responses are often no better than the lower-scored prose responses.
It is also valuable to assess the reward model against the specific human annotators whose preferences it was trained on. If the model achieves high correlation with annotator A's preferences but low correlation with annotator B's, this reveals that the model has implicitly learned to weight different annotators' judgments differently. This can be a problem or a feature, depending on the relative quality of different annotators, but it should be understood rather than ignored.
Detecting Reward Model Weaknesses
Reward models can learn shortcuts that do not align with true quality. Identifying these shortcuts before RL training begins can prevent the policy from being steered in undesirable directions.
Length bias is one of the most common shortcuts. Models often prefer longer responses, even when brevity is more appropriate. This occurs because annotators, when uncertain about which response is better, often prefer the more detailed-seeming one, and response length correlates imperfectly with detail. Test by constructing pairs where the shorter response is clearly superior (e.g., a concise correct answer vs. a verbose but partially incorrect one) and checking whether the reward model correctly scores the concise response higher.
Style over substance is another pervasive issue. Models may prefer responses with confident tone, specific formatting (headers, numbered lists), or polished writing regardless of accuracy. Test with factually incorrect but confidently-written responses against accurate but plainly-written ones. If the reward model systematically prefers the confident-but-wrong response, style has contaminated the quality signal.
Prompt sensitivity reveals inconsistencies that could be exploited. Check if reward differences are consistent across paraphrased prompts asking the same question. If response A is preferred over response B when the prompt is "What is photosynthesis?" but B is preferred when the prompt is "Explain how plants make food," the reward model has learned prompt-specific patterns rather than a general quality function. This type of inconsistency is especially dangerous because the policy can learn to exploit it by generating responses that happen to work well with specific prompt phrasings.
Sycophancy is a particular concern for reward models trained on human feedback. Human annotators often prefer responses that agree with them or praise their question, even when such agreement is not warranted. A sycophantic reward model will score flattering or agreeable responses higher than honest but potentially unwelcome responses, which can steer the policy toward telling users what they want to hear rather than what is accurate.
These weaknesses become necessary during RL training, when the policy model can exploit them. We'll explore this issue in depth in the upcoming chapter on reward hacking.
Worked Example: Computing Preference Loss
Let's trace through the loss computation for a single preference pair to solidify understanding. Walking through a concrete numerical example makes the abstract formulas tangible and reveals how the loss function behaves in practice across different scoring scenarios.
Suppose we have a prompt asking for a simple explanation of gravity, with two candidate responses. Response A (the chosen response) gives a simple, accessible explanation using an everyday analogy. Response B (the rejected response) gives a technically correct but overly complex description using field equations. After training, our reward model produces the following scores:
- for the preferred response (A, the accessible explanation)
- for the rejected response (B, the technical description)
Step 1: Compute the reward difference.
where:
- : the reward difference between the chosen and rejected responses
- : the reward score for the preferred response ()
- : the reward score for the rejected response ()
This positive difference of 1.3 indicates the model correctly believes the preferred response is better. The question is: how confident is this prediction, and how much loss does it incur?
Step 2: Convert to a preference probability via the sigmoid.
where:
- : the calculated probability that the model assigns to the correct preference
- : the sigmoid function applied to the reward difference
- : the resulting probability (approximately 78.6%)
The model assigns about 78.6% probability to the correct preference. This is a reasonably confident prediction. This reflects the moderately large reward margin of 1.3 points. The model is more likely than not to get this comparison right, but it is not highly certain.
Step 3: Compute the negative log-likelihood loss.
where:
- : the computed preference loss value
- : the negative log-likelihood of the correct preference
The loss of 0.241 is relatively small, showing the model is performing well on this example. For reference, a model that assigns exactly 50% probability (pure uncertainty) would incur a loss of .
Step 4: Compare with an incorrect ranking scenario.
Now consider if the model had incorrectly scored the responses, assigning higher reward to the technical description:
- (accessible explanation, incorrectly ranked lower)
- (technical description, incorrectly ranked higher)
Following the same steps: the reward difference is . The sigmoid gives . The loss is .
Step 5: Compare the two loss values.
The loss of 1.54 for the incorrect ranking is more than six times larger than the 0.241 for the correct ranking. This asymmetry demonstrates how the loss heavily penalizes incorrect rankings while giving smaller gradients when the model is already correct. The gradient signal is strongest where the model needs improvement, naturally focusing the learning process on mistakes.
Step 6: Consider the gradient implications.
The gradient of the loss with respect to the reward difference is . For the correct-ranking case (), this gradient has magnitude . For the incorrect-ranking case (), the gradient magnitude is . While the gradient magnitudes are the same (by symmetry of the sigmoid), the direction is reversed: for the correct ranking, the gradient pushes toward a larger margin; for the incorrect ranking, it pushes toward flipping the order entirely. The large loss difference between the two cases reflects the log transformation, which compresses the scale near 1 and expands it near 0, creating the six-fold difference in loss from equal-magnitude gradient contributions.


This demonstration illustrates that the preference loss is a ranking objective and a calibrated probabilistic one. The gradients it produces reflect both the direction of the error and the model's current confidence, creating a training signal that naturally focuses effort where the model most needs to improve.
Code Implementation
Let's build a reward model from scratch using a small pretrained transformer. We'll implement the architecture, loss function, and training loop, then assess the model's ability to generalize to new preference pairs.
Setting Up the Environment
!uv pip install transformers torch matplotlib numpy
import torch
import torch.nn as nn
import torch.nn.functional as F
from torch.utils.data import Dataset, DataLoader
import torch.optim
from transformers import AutoModel, AutoTokenizer
import numpy as np
import matplotlib.pyplot as plt
from dataclasses import dataclass
from typing import List, Tuple, Optional
import warnings
warnings.filterwarnings('ignore')Reward Model Architecture
We'll build a reward model by adding a value head to a pretrained transformer. The value head is a simple linear layer that projects the final hidden state to a scalar.
class RewardModel(nn.Module):
"""
Reward model built on a pretrained transformer.
Takes (prompt, response) pairs and outputs scalar rewards.
"""
def __init__(self, model_name: str, dropout: float = 0.1):
super().__init__()
# Load pretrained transformer as the backbone
self.backbone = AutoModel.from_pretrained(model_name)
hidden_size = self.backbone.config.hidden_size
# Value head: projects final hidden state to scalar reward
self.value_head = nn.Sequential(
nn.Dropout(dropout), nn.Linear(hidden_size, 1)
)
def forward(
self,
input_ids: torch.Tensor,
attention_mask: torch.Tensor,
) -> torch.Tensor:
"""
Compute reward for input sequences.
Args:
input_ids: Token IDs [batch_size, seq_len]
attention_mask: Attention mask [batch_size, seq_len]
Returns:
Tensor: Scalar rewards [batch_size]
"""
# Get hidden states from backbone
outputs = self.backbone(
input_ids=input_ids, attention_mask=attention_mask
)
hidden_states = outputs.last_hidden_state # [batch, seq, hidden]
# Find position of last non-padding token for each sequence
# Sum attention mask to get sequence lengths
seq_lengths = attention_mask.sum(dim=1) - 1 # -1 for 0-indexing
batch_indices = torch.arange(
hidden_states.size(0), device=hidden_states.device
)
# Extract final token hidden state for each sequence
final_hidden = hidden_states[
batch_indices, seq_lengths
] # [batch, hidden]
# Project to scalar reward
rewards = self.value_head(final_hidden).squeeze(-1) # [batch]
return rewardsLet's verify the architecture outputs the expected shapes.
# Initialize model and tokenizer
model_name = "distilbert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
reward_model = RewardModel(model_name)
# Create a sample input
sample_text = "What is machine learning? Machine learning is a field of AI."
inputs = tokenizer(sample_text, return_tensors="pt", padding=True)
# Forward pass
with torch.no_grad():
reward = reward_model(inputs["input_ids"], inputs["attention_mask"])Input shape: torch.Size([1, 15]) Reward shape: torch.Size([1]) Reward value: 0.2864
The output confirms that the model processes the tokenized sequence and produces a single scalar score (batch size 1, output dimension 1). This scalar represents the reward for the prompt-response pair, which will be used to rank different responses against each other.
Preference Dataset
Now let's create a dataset class that handles preference pairs. Each example contains a prompt with a chosen (preferred) and rejected response.
@dataclass
class PreferencePair:
prompt: str
chosen: str
rejected: str
class PreferenceDataset(Dataset):
"""Dataset of preference pairs for reward model training."""
def __init__(
self, pairs: List[PreferencePair], tokenizer, max_length: int = 256
):
self.pairs = pairs
self.tokenizer = tokenizer
self.max_length = max_length
def __len__(self):
return len(self.pairs)
def __getitem__(self, idx: int):
pair = self.pairs[idx]
# Tokenize prompt and chosen response
# Passing two strings automatically adds the separator token
chosen_enc = self.tokenizer(
pair.prompt,
pair.chosen,
max_length=self.max_length,
padding="max_length",
truncation=True,
return_tensors="pt",
)
# Tokenize prompt and rejected response
rejected_enc = self.tokenizer(
pair.prompt,
pair.rejected,
max_length=self.max_length,
padding="max_length",
truncation=True,
return_tensors="pt",
)
return {
"chosen_ids": chosen_enc["input_ids"].squeeze(0),
"chosen_mask": chosen_enc["attention_mask"].squeeze(0),
"rejected_ids": rejected_enc["input_ids"].squeeze(0),
"rejected_mask": rejected_enc["attention_mask"].squeeze(0),
}Let's create some synthetic preference data for demonstration.
# Synthetic preference pairs for demonstration
# 20 distinct pairs covering diverse topics so train and val sets are independent
all_preference_pairs = [
PreferencePair(
prompt="Explain gravity simply.",
chosen="Gravity is the force that pulls objects toward each other. The Earth pulls you down, which is why you stay on the ground.",
rejected="Gravity is described by Einstein's field equations relating the curvature of spacetime to energy-momentum.",
),
PreferencePair(
prompt="What is photosynthesis?",
chosen="Photosynthesis is how plants make food from sunlight, water, and carbon dioxide, creating oxygen as a byproduct.",
rejected="It's a plant thing.",
),
PreferencePair(
prompt="How do computers work?",
chosen="Computers process information using tiny electronic switches called transistors that can be on or off, representing 1s and 0s.",
rejected="Computers work by executing machine code instructions on a von Neumann architecture with fetch-decode-execute cycles.",
),
PreferencePair(
prompt="Why is the sky blue?",
chosen="Sunlight contains all colors. Blue light scatters more than other colors when it hits air molecules, so we see blue when we look up.",
rejected="The sky is blue.",
),
PreferencePair(
prompt="What causes rain?",
chosen="Water evaporates from oceans and lakes, rises as vapor, cools in clouds, and falls as rain when droplets get heavy enough.",
rejected="Rain is caused by the condensation of atmospheric water vapor into droplets when air masses cool below the dew point temperature.",
),
PreferencePair(
prompt="What is the immune system?",
chosen="Your immune system is your body's defense network. White blood cells recognize and destroy germs before they make you sick.",
rejected="The immune system is a complex network of cells and proteins.",
),
PreferencePair(
prompt="How does electricity work?",
chosen="Electricity is the flow of electrons through a wire, like water flowing through a pipe. A battery pushes those electrons to power devices.",
rejected="Electricity involves electromagnetic fields described by Maxwell's equations governing charge motion.",
),
PreferencePair(
prompt="What is DNA?",
chosen="DNA is a molecule shaped like a twisted ladder that holds the instructions for building every part of your body, stored in nearly every cell.",
rejected="dna stuff",
),
PreferencePair(
prompt="Why do we need sleep?",
chosen="Sleep lets your brain consolidate memories, repair cells, and clear out waste products that build up during the day.",
rejected="Sleep is a periodic state of rest involving complex neurological oscillations.",
),
PreferencePair(
prompt="What is climate change?",
chosen="Climate change refers to long-term shifts in global temperatures, largely driven by humans releasing greenhouse gases like CO2 into the atmosphere.",
rejected="It's when the weather changes.",
),
PreferencePair(
prompt="How do vaccines work?",
chosen="Vaccines teach your immune system to recognize a pathogen by exposing it to a harmless piece of the germ, so it's ready to fight the real thing.",
rejected="Vaccines involve antigen presentation and adaptive humoral immunity mediated by B-lymphocytes.",
),
PreferencePair(
prompt="What is inflation?",
chosen="Inflation is when prices rise over time, so the same amount of money buys less than it used to. Central banks try to keep it low and stable.",
rejected="inflation",
),
PreferencePair(
prompt="How does the internet work?",
chosen="Computers send data in small packets across a global network of cables and routers. Each packet finds its own path to the destination.",
rejected="The internet operates on TCP/IP protocols letting packet-switched communication across distributed nodes.",
),
PreferencePair(
prompt="What is evolution?",
chosen="Evolution is the process by which living things gradually change over generations. Individuals with helpful traits survive and pass them on.",
rejected="Evolution refers to allele frequency changes in populations across successive generations.",
),
PreferencePair(
prompt="Why do stars shine?",
chosen="Stars shine because they fuse hydrogen atoms into helium in their cores, releasing enormous amounts of energy as light and heat.",
rejected="stars are bright",
),
PreferencePair(
prompt="What is machine learning?",
chosen="Machine learning is when computers learn patterns from data to make predictions without being explicitly programmed for each task.",
rejected="ML utilizes gradient descent optimization on parameterized function approximators.",
),
PreferencePair(
prompt="How does a car engine work?",
chosen="A car engine burns fuel in cylinders, creating small explosions that push pistons down and turn the wheels through a series of gears.",
rejected="Internal combustion engines operate on thermodynamic cycles converting chemical energy to mechanical work via reciprocating pistons.",
),
PreferencePair(
prompt="What is democracy?",
chosen="Democracy is a system where citizens vote to choose their leaders and influence the laws that govern them.",
rejected="democracy",
),
PreferencePair(
prompt="How does sound travel?",
chosen="Sound travels as vibrations through air, like ripples on a pond. The vibrations reach your ears and your brain interprets them as sound.",
rejected="Sound propagates as longitudinal pressure waves through elastic media with frequency-dependent attenuation.",
),
PreferencePair(
prompt="What is gravity on the moon?",
chosen="The moon has about one-sixth of Earth's gravity because it is much less massive. That is why astronauts bounce when they walk there.",
rejected="Lunar surface gravity is approximately 1.62 m/s due to the moon's lower mass and radius.",
),
]
# Use 15 pairs for training and 5 held-out pairs for validation
train_preference_pairs = all_preference_pairs[:15]
val_preference_pairs = all_preference_pairs[15:]
# Duplicate training pairs to give enough batches for visible learning curves
extended_pairs = train_preference_pairs * 4 # 60 training examplesPreference Loss Function
The loss function implements the Bradley-Terry preference model. We compute rewards for both responses and maximize the probability of preferring the chosen response.
def compute_preference_loss(
reward_model: nn.Module,
chosen_ids: torch.Tensor,
chosen_mask: torch.Tensor,
rejected_ids: torch.Tensor,
rejected_mask: torch.Tensor,
) -> Tuple[torch.Tensor, dict]:
"""
Compute the Bradley-Terry preference loss.
Args:
reward_model: The reward model
chosen_ids: Token IDs for chosen responses [batch, seq]
chosen_mask: Attention mask for chosen [batch, seq]
rejected_ids: Token IDs for rejected responses [batch, seq]
rejected_mask: Attention mask for rejected [batch, seq]
Returns:
loss: Scalar loss value
metrics: Dictionary with additional metrics
"""
# Compute rewards for both responses
chosen_rewards = reward_model(chosen_ids, chosen_mask)
rejected_rewards = reward_model(rejected_ids, rejected_mask)
# Preference loss: -log(sigmoid(r_chosen - r_rejected))
# Equivalent to: log(1 + exp(r_rejected - r_chosen))
reward_diff = chosen_rewards - rejected_rewards
loss = -F.logsigmoid(reward_diff).mean()
# Compute accuracy (how often chosen reward > rejected reward)
accuracy = (reward_diff > 0).float().mean()
# Average reward margin
margin = reward_diff.mean()
metrics = {
"loss": loss.item(),
"accuracy": accuracy.item(),
"margin": margin.item(),
"chosen_reward_mean": chosen_rewards.mean().item(),
"rejected_reward_mean": rejected_rewards.mean().item(),
}
return loss, metricsTraining Loop
Now we implement the complete training loop with logging and validation.
def train_reward_model(
model: nn.Module,
train_loader: DataLoader,
val_loader: Optional[DataLoader],
num_epochs: int = 3,
learning_rate: float = 2e-5,
device: str = "cpu",
) -> dict:
"""
Train the reward model on preference data.
Args:
model: Reward model to train
train_loader: Training data loader
val_loader: Validation data loader (optional)
num_epochs: Number of training epochs
learning_rate: Learning rate for optimizer
device: Device to train on
Returns:
history: Dictionary with training metrics
"""
model = model.to(device)
optimizer = torch.optim.AdamW(
model.parameters(), lr=learning_rate, weight_decay=0.01
)
history = {
"train_loss": [],
"train_accuracy": [],
"val_loss": [],
"val_accuracy": [],
}
for epoch in range(num_epochs):
# Training phase
model.train()
train_metrics = {"loss": 0, "accuracy": 0, "count": 0}
for batch in train_loader:
chosen_ids = batch["chosen_ids"].to(device)
chosen_mask = batch["chosen_mask"].to(device)
rejected_ids = batch["rejected_ids"].to(device)
rejected_mask = batch["rejected_mask"].to(device)
optimizer.zero_grad()
loss, metrics = compute_preference_loss(
model, chosen_ids, chosen_mask, rejected_ids, rejected_mask
)
loss.backward()
# Gradient clipping for stability
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
train_metrics["loss"] += metrics["loss"] * len(chosen_ids)
train_metrics["accuracy"] += metrics["accuracy"] * len(chosen_ids)
train_metrics["count"] += len(chosen_ids)
# Compute epoch averages
train_loss = train_metrics["loss"] / train_metrics["count"]
train_acc = train_metrics["accuracy"] / train_metrics["count"]
history["train_loss"].append(train_loss)
history["train_accuracy"].append(train_acc)
# Validation phase
if val_loader is not None:
model.eval()
val_metrics = {"loss": 0, "accuracy": 0, "count": 0}
with torch.no_grad():
for batch in val_loader:
chosen_ids = batch["chosen_ids"].to(device)
chosen_mask = batch["chosen_mask"].to(device)
rejected_ids = batch["rejected_ids"].to(device)
rejected_mask = batch["rejected_mask"].to(device)
_, metrics = compute_preference_loss(
model,
chosen_ids,
chosen_mask,
rejected_ids,
rejected_mask,
)
val_metrics["loss"] += metrics["loss"] * len(chosen_ids)
val_metrics["accuracy"] += metrics["accuracy"] * len(
chosen_ids
)
val_metrics["count"] += len(chosen_ids)
val_loss = val_metrics["loss"] / val_metrics["count"]
val_acc = val_metrics["accuracy"] / val_metrics["count"]
history["val_loss"].append(val_loss)
history["val_accuracy"].append(val_acc)
return historyLet's train the model on our synthetic preference data.
# Create datasets -- val_preference_pairs are held-out pairs not seen during training
train_pairs = extended_pairs
val_pairs = val_preference_pairs
train_dataset = PreferenceDataset(train_pairs, tokenizer, max_length=128)
val_dataset = PreferenceDataset(val_pairs, tokenizer, max_length=128)
train_loader = DataLoader(train_dataset, batch_size=8, shuffle=True)
val_loader = DataLoader(val_dataset, batch_size=8)
# Initialize fresh model for training
reward_model = RewardModel(model_name)
# Train
history = train_reward_model(
reward_model, train_loader, val_loader, num_epochs=5, learning_rate=2e-5
)Training Results: ---------------------------------------- Epoch 1: Train Loss: 0.5974, Accuracy: 75.00% Val Loss: 0.3318, Accuracy: 100.00% Epoch 2: Train Loss: 0.1549, Accuracy: 100.00% Val Loss: 0.1605, Accuracy: 100.00% Epoch 3: Train Loss: 0.0074, Accuracy: 100.00% Val Loss: 0.0942, Accuracy: 100.00% Epoch 4: Train Loss: 0.0007, Accuracy: 100.00% Val Loss: 0.0688, Accuracy: 100.00% Epoch 5: Train Loss: 0.0002, Accuracy: 100.00% Val Loss: 0.0617, Accuracy: 100.00%
The model quickly learns to distinguish between preferred and rejected responses on this small dataset.
Visualizing Training Progress


The training curves demonstrate that the model effectively learns the preference ranking task, with validation accuracy tracking closely with training accuracy. This indicates the model is generalizing well to unseen preference pairs without significant overfitting.
Evaluating on New Examples
Let's assess the trained model on examples it hasn't seen during training.
def score_response(
model, tokenizer, prompt: str, response: str, device: str = "cpu"
) -> float:
"""Score a single prompt-response pair."""
model.eval()
inputs = tokenizer(
prompt,
response,
return_tensors="pt",
padding=True,
truncation=True,
max_length=128,
)
with torch.no_grad():
reward = model(
inputs["input_ids"].to(device), inputs["attention_mask"].to(device)
)
return reward.item()
# Test on new examples
test_prompt = "What is machine learning?"
test_responses = [
(
"Good explanation",
"Machine learning is when computers learn patterns from data to make predictions without being explicitly programmed.",
),
(
"Too technical",
"ML utilizes gradient descent optimization on parameterized function approximators.",
),
("Too brief", "It's AI stuff."),
]
# Compute scores for all responses
scores = []
for label, response in test_responses:
score = score_response(reward_model, tokenizer, test_prompt, response)
scores.append((score, label, response))
# Sort by score (higher is better)
scores.sort(reverse=True)Prompt: What is machine learning? Response Rankings: ------------------------------------------------------------ 1. [Good explanation] Score: -3.8208 "Machine learning is when computers learn patterns from data ..." 2. [Too technical] Score: -4.4391 "ML utilizes gradient descent optimization on parameterized f..." 3. [Too brief] Score: -4.6247 "It's AI stuff."
The model has learned to rank responses in a way that aligns with our training preferences: clear explanations over technical jargon or overly brief responses.
Analyzing Reward Distributions
A well-calibrated reward model should produce meaningful score separations between good and bad responses. Let's analyze the distribution of rewards on our validation set.
# Collect rewards from validation set
chosen_rewards = []
rejected_rewards = []
# Ensure model is on the correct device
device = next(reward_model.parameters()).device
reward_model.eval()
with torch.no_grad():
for batch in val_loader:
# Move inputs to the same device as the model
chosen_ids = batch["chosen_ids"].to(device)
chosen_mask = batch["chosen_mask"].to(device)
rejected_ids = batch["rejected_ids"].to(device)
rejected_mask = batch["rejected_mask"].to(device)
chosen_r = reward_model(chosen_ids, chosen_mask)
rejected_r = reward_model(rejected_ids, rejected_mask)
# Move to CPU for logging/plotting
chosen_rewards.extend(chosen_r.cpu().numpy().tolist())
rejected_rewards.extend(rejected_r.cpu().numpy().tolist())
The clear separation between distributions indicates the model has learned meaningful preference distinctions.
Key Parameters
The key parameters for the Reward Model implementation are:
- model_name: The pretrained backbone (e.g.,
"distilbert-base-uncased"). Smaller models allow for faster iteration during experimentation. - dropout: Regularization applied to the value head (set to
0.1) to prevent overfitting on the small preference dataset. - learning_rate: A conservative rate (
2e-5) is used to fine-tune the backbone without destroying pretrained features. - batch_size: Set to
8for this demonstration, though larger batches are preferred for stability in full-scale training. - num_epochs: Training is limited to
5epochs to avoid overfitting on the small synthetic dataset.
Limitations and Impact
Reward modeling is a powerful technique but comes with significant challenges that you must understand before deploying it in production RLHF systems. These limitations are not merely theoretical: they have caused real quality problems in deployed language models and have motivated major lines of research in the alignment community. Understanding where reward models can fail, and why, is as important as understanding how to build them.
The Proxy Problem
The basic limitation of reward models is that they are proxies for human preferences, not perfect representations. This is the key tension at the heart of RLHF: we want to optimize for human values, but the RL optimizer sees only the reward model's approximation of those values. Every error in the approximation becomes an opportunity for the optimizer. The model learns from a finite set of comparisons made by a specific group of annotators under particular conditions. It cannot generalize perfectly to all possible responses or capture the full complexity of human values. When the policy model optimizes against this learned reward, it may find responses that score highly according to the proxy but do not satisfy human preferences.
Think of the reward model as a map and human preferences as the territory. A map can be useful for navigation even when it is not perfectly accurate. But under severe optimization pressure, the map's imperfections start to matter: a path that looks short on the map may be impassable in reality. Similarly, a policy that is trained for just a few RL steps will likely improve. A policy that is trained for many RL steps under high optimization pressure will start exploiting the gaps between the map and the territory. This phenomenon, known as reward hacking or Goodhart's Law in action, becomes more severe as optimization pressure increases. We'll explore this challenge in depth in the next chapter.
Annotation Quality and Consistency
Reward model quality is bounded by the quality of the underlying preference data. Human annotators disagree, make mistakes, and have biases. If 70% of annotators prefer response A over B, the "correct" label is somewhat arbitrary. The reward model learns from these noisy, inconsistent signals, and this uncertainty propagates into the learned reward function.
Different annotator pools may have systematically different preferences based on cultural background, expertise, or task understanding. A reward model trained on one population may not generalize to another. This is particularly concerning when the intended users of the deployed model are demographically or culturally different from the annotator pool. The reward model may have learned preferences that are appropriate for one group but misaligned with another.
The annotation interface and instructions also matter. Subtle differences in how the comparison task is framed can lead to systematic biases in the annotations. If annotators are asked to select "the better response" without further guidance, some may focus on helpfulness, others on tone, others on accuracy, and still others on length. The resulting reward model will reflect this mixture of criteria, which may not align with any single coherent value system. Clear annotation guidelines that specify exactly what "better" means for the intended application reduce this problem but cannot eliminate it entirely.
Distribution Shift
During RLHF training, the policy model generates responses that may differ substantially from those in the reward model's training set. The reward model must extrapolate to these out-of-distribution samples, and its predictions become less reliable as the policy's outputs diverge from the training distribution. This creates a feedback loop: the policy learns to generate responses that score well according to the reward model's potentially incorrect extrapolations, which can lead to degraded actual quality even as measured rewards increase.
The KL divergence penalty used in most RLHF implementations (we'll cover this in detail in the PPO chapter) is designed to limit this distribution shift. By penalizing the policy for deviating too far from the SFT initialization, the KL penalty keeps the policy's outputs in a region where the reward model's extrapolations are more likely to be reliable. However, this is a blunt instrument: it slows policy improvement overall, not just in areas where the reward model is unreliable.
Research on reward model ensembles addresses this limitation by using multiple reward models trained on different subsets of the data or with different architectures. When the ensemble members disagree on a response, that response is likely out-of-distribution for at least some models, and the disagreement can be used as an uncertainty signal. Conservative RL algorithms can then be penalized for exploring uncertain regions, focusing optimization on areas where the reward is reliably estimated.
Computational Costs
Training reward models requires significant computational resources, particularly when using large base models to ensure the reward model has sufficient language understanding. A 70B parameter reward model requires more GPU memory and compute than most teams can afford for a component that serves only as an auxiliary signal. During RL training, the reward model must evaluate every generated response, adding substantial inference costs to an already expensive training procedure.
There are several strategies for managing these costs. First, use a reward model that is smaller than the policy, accepting some quality loss for throughput. Second, apply reward model distillation: train a large, high-quality reward model once, then distill it into a smaller model for use during RL training. Third, cache reward model evaluations when the policy generates the same or similar responses repeatedly. Fourth, use reward model parallelism: if you have multiple GPUs, run the reward model on a dedicated subset while the policy runs on others.
The trend toward larger models creates an escalating cost dynamic. As policies grow larger, more capable reward models are needed to evaluate them accurately, which in turn increases inference costs. This scaling challenge is one of the motivations for DPO and other reward-model-free approaches to preference learning.
Reward Hacking and Specification Gaming
Beyond the proxy problem, reward models are vulnerable to specification gaming: the policy may find valid-seeming ways to achieve high rewards that nonetheless violate the intent of the human preferences. Common examples include verbosity (padding responses with filler content to appear thorough), hedging (adding excessive caveats to avoid making claims that could be wrong), sycophancy (agreeing with the user even when the user is incorrect), and false confidence (asserting uncertain facts with unwarranted certainty).
These behaviors can be hard to detect from reward model scores alone because the reward model was trained on the same distribution that generated the problematic preference labels. If human annotators preferred verbose responses during training, the reward model will score verbose responses highly, and the policy will produce verbose responses. The problem only becomes visible when humans assess the RL-trained policy and notice that it produces unnecessarily long, formulaic answers. By that point, the policy may already be deeply trained on the problematic reward signal, and retraining requires collecting new annotations specifically targeted at the identified bias.
The best defense against specification gaming is proactive adversarial assessment of the reward model before RL training begins. Construct response pairs specifically designed to test hypothesized weaknesses, collect human judgments on those pairs, and verify that the reward model's scores align with human judgments. This kind of red-teaming is time-consuming but far cheaper than discovering specification gaming problems after a full RL training run.
Impact on RLHF Systems
Despite these limitations, reward models have enabled major advances in language model alignment. They provide the important bridge between sparse human feedback and dense training signals. The InstructGPT system that powers ChatGPT, Anthropic's Claude models, and many open-source chat models all rely on reward models as a core component. The quality improvements these systems demonstrate over base language models are directly attributable to the reward model's ability to encode human preferences and feed them into the RL training loop.
The reward modeling approach has also influenced research directions, spurring work on direct preference optimization (DPO) methods that eliminate the need for explicit reward models, as well as techniques for reward model ensembles, uncertainty quantification, and reliable optimization. Understanding reward modeling deeply is needed for grasping both current RLHF systems and the alternatives being developed to address its limitations. The basic challenge of learning human values from comparisons, and building a representation of those values reliable enough to withstand adversarial optimization, remains one of the central open problems in the field.
Summary
This chapter covered the complete pipeline for building reward models that learn to predict human preferences and serve as the optimization target for RLHF.
Architecture: Reward models add a value head to pretrained transformers, projecting the final token's hidden state to a scalar reward. Using the same architecture family as the policy model ensures compatible representations. The value head is typically a single linear layer, though more complex heads can be used. The final token's position is preferred because causal attention accumulates information from the entire sequence.
Loss function: The Bradley-Terry preference loss maximizes the probability of correctly ranking preference pairs. Gradients naturally emphasize uncertain or incorrect predictions. The loss is shift-invariant: only reward differences matter, not absolute values. Margin-based variants encourage larger score separations but can hurt calibration.
Training: Conservative hyperparameters preserve pretrained knowledge while learning preference structure. Early stopping, gradient clipping, label smoothing, and appropriate regularization prevent overfitting to limited preference data. Post-training reward normalization ensures the reward signal is in a consistent range for downstream RL optimization.
Evaluation: Pairwise accuracy measures ranking performance, with human agreement giving an upper bound. Calibration testing, agreement with human evaluators, and adversarial probing for biases like length preference or style over substance are all needed components of complete evaluation. The evaluation must address in-distribution accuracy, out-of-distribution generalization, and susceptibility to exploitation.
Limitations: Reward models are proxies for human preferences and are bounded by annotation quality, susceptible to distribution shift, and vulnerable to reward hacking. These limitations motivate both the KL penalty in RLHF and alternative preference learning methods. Understanding these failure modes is as important as understanding the training procedure.
The reward model is the necessary interface between human preferences and policy optimization. In the following chapters, we'll examine how reward hacking can undermine this proxy relationship, and then explore how policy gradient methods and PPO use reward signals to improve language models.
Reward Modeling
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!