RLHF Foundations: Learning from Human Preferences in RL

Michael BrenndoerferJune 9, 202513 min read

Part of History of Language AI

Covers preference-based learning, the framework developed by Christiano et al. in 2017.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

2017: RLHF Foundations

In 2017, researchers at OpenAI and DeepMind studied a problem in reinforcement learning: how could agents learn complex behaviors when engineers could not specify an adequate reward function? For tasks involving natural language or robotics, desired behavior often depended on human preferences that were difficult to express mathematically. Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei developed a method that allowed reinforcement-learning agents to learn from human comparisons rather than only predefined rewards.

The paper, "Deep Reinforcement Learning from Human Preferences," described a framework in which evaluators compared pairs of agent behaviors. A reward model learned from those comparisons, then supplied the signal used for reinforcement learning. This replaced an explicitly programmed reward function with preferences expressed through comparative judgments. In the paper's experiments, agents learned behaviors that matched evaluator preferences even when those preferences were difficult to formalize.

The framework was later adapted beyond the reinforcement-learning tasks in the original paper. Language models presented a related problem: maximum-likelihood training could produce fluent text without directly optimizing for evaluator preferences about its content or behavior. Preference learning supplied one method for adding that signal during model fine-tuning.

The method reduced the need to define rewards for every possible scenario. Evaluators supplied a sample of comparisons, and the learned reward model estimated preferences for new trajectories. This property made the approach applicable to language models, whose possible outputs cannot be enumerated in advance, although collecting comparison data still required substantial human effort.

The Problem

Traditional reinforcement learning relies on a reward function that provides a scalar reward signal to the agent after each action or at the end of an episode. This reward function is the primary learning signal, guiding the agent to discover behaviors that maximize expected cumulative reward. For many applications, this approach works well when the reward function can be precisely specified. In game-playing scenarios, for example, the reward might be the game score, which is clearly defined and measurable. In control tasks, the reward might be based on measurable quantities like distance traveled, energy consumed, or task completion metrics.

However, many important applications involve objectives that are difficult or impossible to encode as precise reward functions. Consider a robotic agent learning to assist with household tasks. An engineer might try to define a reward based on metrics like task completion time or the number of objects moved, but these metrics might miss important aspects of what makes assistance helpful: being gentle with fragile items, cleaning up afterward, or understanding the user's unstated preferences. Similarly, for language tasks, a model trained to maximize likelihood might generate text that is grammatically correct but unhelpful, verbose, or unsafe. What makes text "good" often involves subtle qualities that resist simple mathematical formulation.

Another challenge arises when a reward function does not fully capture the desired behavior. Optimization can exploit gaps in such a function, producing high reward without the intended result. Examples include game-playing agents that pause indefinitely to avoid negative rewards and cleaning robots that hide messes instead of removing them. A reward that appears reasonable during design can therefore encourage unintended behavior.

Human preferences also involve trade-offs that are difficult to capture in one scalar function. An evaluator might prefer a response that balances information with a friendly, concise presentation over one optimized only for informativeness. Encoding such judgments requires engineers to choose weights and assumptions about qualities whose relative importance changes across users or contexts.

Reward-function engineering also becomes more expensive as task variety grows. A new scenario can require a revised reward structure followed by testing and adjustment based on observed behavior. Systems that must handle varied situations or user preferences can therefore spend substantial effort on reward design.

Evaluation posed a related problem. Engineers still needed human judgments to determine whether a hand-written reward function produced acceptable behavior. When evaluation exposed a mismatch, they revised the function and retrained the agent. Incorporating comparisons into the learning process moved those judgments into the reward model instead of using them only after training.

The Solution

The preference-based learning framework addresses these challenges by separating reward specification from reward learning. Instead of requiring engineers to design reward functions, the method learns a reward model from human feedback, then uses this learned reward model to train reinforcement learning agents. The key insight is that humans can more easily compare behaviors than they can specify exact reward values or functions. By collecting pairwise comparisons from human evaluators, the system can learn what humans value without requiring them to formalize their preferences mathematically.

The process works in three main stages. First, a reinforcement learning agent interacts with the environment, generating trajectories of behavior. These trajectories are presented to human evaluators in pairs, and the evaluators indicate which trajectory they prefer. Second, a reward model is trained to predict human preferences by learning to assign higher rewards to trajectories that humans prefer. This reward model learns from the pairwise comparison data, developing an internal representation of what makes behaviors desirable according to human judgment. Third, the reinforcement learning agent is trained using the learned reward model as the source of rewards, optimizing its policy to maximize expected reward according to the reward model's predictions.

The reward model is typically implemented as a neural network that takes a trajectory as input and outputs a scalar reward value. During training, the model learns to assign higher rewards to preferred trajectories and lower rewards to non-preferred ones. The training objective encourages the reward model to correctly rank trajectories according to human preferences. Specifically, the model is trained using a ranking loss that maximizes the difference in predicted rewards between preferred and non-preferred trajectories. This approach allows the reward model to generalize beyond the specific comparisons it was trained on, learning general principles about what makes behaviors desirable.

The key technical innovation involves how trajectories are compared. Rather than requiring humans to evaluate complete trajectories, which might be long or complex, the method allows comparisons of trajectory segments. Human evaluators might compare specific segments of behavior, focusing attention on the most relevant parts. This segmentation makes human evaluation more efficient and allows the reward model to learn fine-grained preferences about specific aspects of behavior. The learned reward model can then aggregate these segment-level preferences when evaluating complete trajectories.

Another important aspect is how the reward model handles uncertainty. When the reward model is uncertain about which trajectory is preferred, it should express this uncertainty rather than making confident but potentially incorrect predictions. The framework incorporates this through the reward model's training, which learns point estimates of reward together with confidence in those estimates. This uncertainty can be used to guide further human evaluation, focusing attention on cases where the reward model is least confident and where additional human feedback would be most informative.

The integration with reinforcement learning uses the learned reward model as a drop-in replacement for a traditional reward function. Standard reinforcement learning algorithms, such as policy gradient methods or actor-critic approaches, can be applied without modification. The agent receives rewards from the reward model rather than from a hand-coded function, but from the agent's perspective, the learning process is identical. This compatibility meant that existing reinforcement learning infrastructure and algorithms could be used with minimal changes, making the approach practical to implement and experiment with.

The method also supports iterative updates. As the policy changes, it generates trajectories that differ from those used to train the initial reward model. Evaluators can compare these new trajectories, supplying additional data for another reward-model update. Repeating this process updates both the learned reward and the policy against behaviors the current policy can produce.

Applications and Impact

The initial experiments covered simulated robotic manipulation and Atari games. In the robotics tasks, agents learned behavior that evaluators judged safe and efficient. In Atari, comparison feedback could favor behavior evaluators considered interesting or skillful rather than behavior selected only by game score. These experiments tested preferences that were difficult to express through simple reward functions.

The same approach could be applied where user judgments mattered but were difficult to formalize. A recommendation system, for example, could learn from comparative user behavior rather than a single hand-written quality metric. An autonomous system could also learn preferences about safe operation and appropriate interaction without an exhaustive list of rules.

Several years later, researchers adapted the preference-learning framework to large language models. Standard language-model objectives optimized next-token prediction rather than evaluator judgments about a response. The resulting alignment problem resembled the earlier task of fitting agent behavior to human comparisons.

For language models, evaluators compared responses to the same prompt instead of trajectories in an environment. A reward model learned from those comparisons and guided reinforcement-learning fine-tuning, often with Proximal Policy Optimization. This pipeline became known as Reinforcement Learning from Human Feedback, or RLHF.

Developers have used RLHF or related preference-learning methods to adjust models such as GPT-3.5, GPT-4, and Claude toward evaluator-selected behavior. These methods supplement language-model training with judgments about response usefulness and acceptable conduct. They have become a common stage in post-training at major AI labs.

Preference learning also applies beyond language models. Image generators can learn from comparative judgments about outputs, while code generators can use programmer rankings of proposed code. The approach is relevant when direct human comparison captures desired behavior better than a fixed metric.

Limitations

While preference-based learning addresses many challenges, it also introduces new limitations and considerations. One fundamental issue is that human preferences may be inconsistent, context-dependent, or difficult to express through pairwise comparisons. Different humans might have different preferences, and the same human might express different preferences at different times or in different contexts. The method assumes that preferences are reasonably stable and consistent enough to learn from, but in practice, this assumption may not always hold.

The quality of the learned reward model depends critically on the quality and quantity of human preference data. If human evaluators provide noisy, inconsistent, or biased comparisons, the reward model will learn to reflect those issues. Poor quality preference data can lead to reward models that don't accurately capture human values, which then guide agents or models toward misaligned behaviors. Collecting high-quality preference data requires careful attention to evaluation protocols, training for evaluators, and quality assurance measures.

The scalability of human evaluation remains a challenge. While preference-based learning reduces the need for humans to specify reward functions, it still requires significant human effort to provide comparisons. For large-scale applications, this can become expensive and time-consuming. The method improves efficiency by learning generalizable reward models from relatively small amounts of comparison data, but human evaluation remains a bottleneck that limits how quickly systems can be improved or adapted to new domains.

Reward hacking remains possible when an agent finds behavior that scores highly under the learned reward model but conflicts with evaluator preferences. The reward model is an approximation and may contain blind spots or systematic errors. Optimization can exploit those gaps, just as it can exploit a hand-written reward function, while the learned model may make the failure harder to diagnose.

The method also assumes that pairwise comparisons provide sufficient information to learn good reward models. However, some preferences might be inherently difficult to express through comparisons, especially when preferences involve multiple dimensions that can't be easily reduced to a single ranking. For example, a conversation might be preferred in some ways but not others, and forcing a single preference judgment might lose important nuance. More sophisticated preference elicitation methods might be needed for complex multi-dimensional preferences.

There are also questions about whose preferences are being learned. The method learns from the preferences of the humans who provide comparisons, but these humans might not be representative of all users or stakeholders. Preferences might vary across cultures, contexts, or individuals, and learning from one group's preferences might not produce behaviors that align with other groups' values. This concern becomes especially important for language models and other systems with broad user bases, where diverse preferences must be considered.

Finally, the iterative improvement process, while valuable, can be slow. Each cycle of generating behaviors, collecting human feedback, updating the reward model, and retraining the agent takes time and resources. For applications requiring rapid adaptation or frequent updates, this iterative process might not be fast enough. The method works well when there is time for careful alignment, but may be less suitable for scenarios requiring quick deployment or frequent changes.

Legacy and Looking Forward

The 2017 framework established a reusable pattern: learn a reward model from human comparisons, then optimize a policy against that model. Researchers later adapted this pattern to new environments and to language-model post-training. The work demonstrated a practical way to incorporate comparative human feedback into a machine-learning objective.

The connection is most visible in language-model post-training. Systems including GPT-3.5, GPT-4, and Claude have used variants of preference learning to adjust outputs beyond the next-token prediction objective. Implementations differ, but each uses evaluator judgments as a training signal for model behavior.

Research continues to improve upon the original framework. Recent work has explored alternatives to the three-stage RLHF process, such as Direct Preference Optimization, which fine-tunes language models directly from preference data without training a separate reward model. Other research has investigated methods for making human evaluation more efficient, such as using AI assistants to help with preference collection or developing better protocols for training human evaluators. These improvements build on the core insights from 2017 while addressing some of the method's limitations.

The framework also contributed to safety and alignment research focused on objectives derived from human judgments rather than simple proxy metrics. It made the collection and evaluation of human feedback part of the training pipeline while also exposing questions about whose preferences the data represents.

Current research seeks to reduce the cost of preference data, represent disagreement among evaluators, and limit exploitation of learned reward models. These problems matter whenever a system treats a finite set of comparisons as a proxy for broader human judgments.

The 2017 work showed how comparative judgments could replace part of the hand-written reward-design process. Its later adaptation to language models linked reinforcement-learning research with modern post-training practice. The method does not remove the alignment problem: it changes the source of the objective and introduces new dependencies on evaluator data and reward-model accuracy.

Quiz

Test your understanding of preference data, learned reward models, and reinforcement-learning updates in the RLHF pipeline.

RLHF Foundations Quiz

Question 1 of 60 of 6 completed
What fundamental problem did preference-based learning solve in reinforcement learning?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025rlhffoundations, author = {Michael Brenndoerfer}, title = {RLHF Foundations: Learning from Human Preferences in RL}, year = {2025}, url = {https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2025). RLHF Foundations: Learning from Human Preferences in RL. Retrieved from https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences
MLAAcademic
Michael Brenndoerfer. "RLHF Foundations: Learning from Human Preferences in RL." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences>.
CHICAGOAcademic
Michael Brenndoerfer. "RLHF Foundations: Learning from Human Preferences in RL." Accessed September 27, 2026. https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences.
HARVARDAcademic
Michael Brenndoerfer (2025) 'RLHF Foundations: Learning from Human Preferences in RL'. Available at: https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2025). RLHF Foundations: Learning from Human Preferences in RL. https://mbrenndoerfer.com/writing/rlhf-foundations-reinforcement-learning-human-preferences

About the author

Continue with the full handbook

This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore History of Language AI
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.