Human Preference Data: Collection for LLM Alignment

Michael BrenndoerferDecember 22, 202567 min read

Part of Language AI Handbook

Collect and process human preference data for RLHF. Topics include pairwise comparisons, annotator guidelines, quality metrics, and interface design.

Human Preference Data

The alignment problem, as we discussed in the previous chapter, fundamentally comes down to a question of measurement: how do we capture what humans want from language models? Abstract concepts like helpfulness, harmlessness, and honesty resist simple formalization. You cannot write a loss function that directly optimizes for "be helpful without being sycophantic." The space of possible model behaviors is vast, the space of desirable behaviors is hard to characterize, and models tend to fail in the gap between what a rule captures and the intended behavior.

Think of trying to specify "good teaching" through a rubric. You might write rules like "explains concepts clearly," "uses examples," "checks for understanding," and "adapts to the student's level." But any experienced teacher knows these rules capture only a fraction of what makes teaching effective. A teacher who follows all the rules mechanically but fails to engage students or respond intuitively to their needs still falls short in ways the rubric doesn't measure. The same problem plagues rule-based alignment: the rules capture the surface structure of good behavior while missing the deeper patterns that matter.

Human preference data is the empirical foundation for alignment precisely because it sidesteps this specification problem. Rather than trying to formalize desired behavior through rules or heuristics, we collect examples of humans choosing between model outputs. These choices encode implicit knowledge about what makes a response good: whether it answers the question correctly, whether the tone is appropriate, whether it avoids harmful content, whether it acknowledges uncertainty when uncertain. The approach can capture subtle preferences that resist explicit specification. When a human annotator chooses Response A over Response B, they are expressing something richer than any rule could capture: an integrated judgment across multiple quality dimensions.

The main point is that preference data turns an ill-specified optimization problem into a well-specified one. Instead of optimizing "be helpful," we optimize "produce responses that humans prefer," and we have direct empirical measurements of those preferences. This reframing is powerful even though it introduces new challenges: preferences vary across humans, reflect biases and context, and may not always capture what we want in the long run.

This chapter covers the practical machinery of preference data collection in depth: how to design comparison interfaces, what makes a good preference dataset, how to write annotator guidelines that produce consistent judgments, how to measure and maintain data quality, and how to detect and correct systematic biases. These decisions affect the entire alignment pipeline, shaping the reward models we will train and ultimately determining how our models behave in deployment. The quality of your preference data is the ceiling on the quality of your alignment.

Historical Context

Human preference data for language model alignment has its roots in classical social science methodology. Pairwise comparison methods were developed by psychologists in the early twentieth century to measure preferences that resist direct quantification. Thurstone's Law of Comparative Judgment (1927) provided the first formal framework for deriving scale values from pairwise choices. The Bradley-Terry model, developed in 1952, gave statisticians a principled probabilistic framework for the same problem. These ideas lay dormant in NLP until the RLHF paper by Christiano et al. (2017) demonstrated that pairwise human preferences could train reward models for deep reinforcement learning agents. InstructGPT (Ouyang et al., 2022) then applied this approach at scale to language models, showing that a relatively small amount of preference data could dramatically improve model behavior. The Anthropic Constitutional AI work and the subsequent open-source RLHF efforts further demonstrated the scalability of preference data collection. Today, essentially every major language model uses some form of human preference data in its training pipeline.

Preference Collection Paradigms

Human preferences can be collected in several formats, each with different tradeoffs. Some formats scale easily but produce less reliable or less information-dense labels. The format we choose shapes the annotation process and the mathematical models we can apply to learn from the data. We must consider both the cognitive demands on annotators and the statistical properties of the resulting data. A format that is cognitively demanding produces fewer annotations per hour and shows more noise; a format that is statistically inefficient requires more data to learn the same signal. Good preference collection design balances these tradeoffs.

The central challenge is that quality is a multidimensional concept. A response can be helpful but inaccurate, concise but incomplete, appropriately safe but overly cautious. Different collection formats make different choices about how to handle this multidimensionality. Some collapse it into a single comparison; others try to capture it explicitly through structured dimensions. Understanding these tradeoffs helps you choose the right format for your alignment goals and design collection workflows that minimize systematic distortions.

Pairwise Comparisons

The dominant paradigm in RLHF presents annotators with two model responses to the same prompt and asks which is better. This format has become the standard for preference collection because it addresses both human and mathematical needs simultaneously.

From the annotator's perspective, pairwise comparisons reduce a complex quality judgment to its simplest possible form. Instead of asking "how good is this response?" we ask "which of these two responses is better?" This reframing has several advantages that compound to make the data much more reliable.

First, consider cognitive simplicity. Comparing two options is much easier than rating on an absolute scale. The annotator need only decide which response they would rather receive, not quantify exactly how much better one is than the other. Humans are generally better at making relative judgments than absolute ones. Think of asking a wine critic to rate a wine from 1 to 100 versus asking them which of two wines they prefer: the comparative task is less ambiguous and less subject to individual scale interpretation.

Second, pairwise comparisons are calibration-free. Annotators don't need to agree on what "4 out of 5" means. When two annotators both prefer response A over response B, they have communicated the same information regardless of whether they would have assigned different numerical ratings. This removes a significant source of noise in absolute rating approaches, where annotators with different internal scales produce incomparable scores even when their underlying preferences are identical.

Third, pairwise data aligns naturally with the Bradley-Terry model we'll cover in the next chapter. This mathematical framework was specifically designed to learn underlying quality scores from pairwise comparisons, and it provides theoretically grounded aggregation of preferences across annotators and comparisons. The data format and the mathematical model are matched to each other.

A typical pairwise comparison looks like:

Prompt: "Explain photosynthesis to a 10-year-old." Response A: "Photosynthesis is how plants make food! Plants have special parts called chloroplasts that capture sunlight..." Response B: "Photosynthesis is the biochemical process by which photoautotrophs convert electromagnetic radiation..." Which response is better? [A] [B] [Tie]

The annotator selects which response better satisfies the implicit quality criteria for the given prompt and audience. Most preference datasets discard or minimize tie votes, as they provide less signal for training reward models. This reflects a main point: the most valuable information comes from cases where annotators can make a clear distinction between options. Ties represent uncertainty in the signal, and while they contain some information (the annotator could not distinguish between the responses), they are harder to use effectively in training.

Some annotation schemes offer a more granular response set, such as "A is much better," "A is slightly better," "Tie," "B is slightly better," "B is much better." This captures strength of preference, not just direction, and can provide richer training signal. The tradeoff is added cognitive complexity: annotators must calibrate what "slightly" versus "much" means, reintroducing some of the calibration problems that pairwise comparison was designed to avoid.

Likert Ratings

An alternative approach asks annotators to rate individual responses on a numerical scale (e.g., 1-5 or 1-7). This captures absolute quality judgments:

Response: "Photosynthesis is how plants make food using sunlight..." Rate this response: [1 - Very Poor] [2] [3 - Adequate] [4] [5 - Excellent]

Likert ratings provide more information per annotation, but they suffer from calibration issues that can be severe in practice. Different annotators interpret the scale differently, and even the same annotator may shift their standards over time. One annotator's "4" might correspond to another annotator's "3," making it difficult to aggregate ratings across annotators without careful normalization procedures.

The calibration problem also interacts with domain: annotators may have different standards for different types of tasks. An annotator with deep coding expertise might rate a mediocre code explanation as "2," while a non-programmer rates the same response as "4." These differences reflect different standards rather than annotator error, but they are difficult to disentangle from quality differences.

Despite these challenges, Likert ratings have one significant advantage: they allow you to evaluate a response without needing another response to compare it against. This matters when you have a large pool of responses and cannot enumerate all possible pairs. Ratings can be converted to implied preferences, but the conversion requires assumptions about scale interpretation that introduce additional noise.

Ranking Multiple Responses

Some collection schemes present more than two responses and ask for a full ranking:

Rank these responses from best (1) to worst (4): [ ] Response A [ ] Response B [ ] Response C [ ] Response D

Rankings generate multiple pairwise comparisons from a single annotation task. Specifically, ranking kk items yields (k2)=k(k−1)/2\binom{k}{2} = k(k-1)/2 pairwise comparisons, since we learn the relative ordering of every pair. For k=4k = 4 items, this yields 6 pairwise comparisons; for k=8k = 8 items, 28 comparisons. In theory, this is highly efficient.

In practice, however, the efficiency gain is smaller than it appears. Ranking becomes cognitively demanding as the number of items increases. Annotators may carefully compare the top two or three responses but make increasingly arbitrary decisions among lower-ranked items. This produces high-quality signal for the top comparisons and noisy signal for the rest. Studies of annotation quality in ranking tasks confirm this pattern: inter-annotator agreement drops substantially for the lower-ranked comparisons.

There is also a practical concern about ranking consistency. Ranking requires the annotator to hold multiple items in working memory simultaneously and make transitively consistent judgments. Human judgment is not always transitive: an annotator might prefer A to B, B to C, but prefer C to A when compared directly. Rankings force transitivity by construction, which can introduce distortions when the annotator's actual preferences are cyclic or context-dependent.

Out[3]:
Visualization
Grouped bar chart comparing pairwise comparisons and cognitive load for ranking sets of 2 to 8 items, showing cognitive load rising faster than the number of comparisons.
Comparison of information yield versus cognitive load for ranking tasks of increasing size. While the number of pairwise comparisons grows quadratically with the number of items ($k(k-1)/2$), the cognitive load increases at a faster rate due to working memory constraints and the need for global consistency, illustrating the diminishing returns of ranking large sets.

Constitutional and Criteria-Based Evaluation

A newer approach, developed largely by Anthropic, asks annotators to evaluate responses against an explicit set of principles or a "constitution." Rather than just asking "which is better?" the interface presents the constitution and asks "which response better follows these principles?"

This approach has the advantage of making the evaluation criteria explicit and consistent. Annotators are not relying on personal intuition about what "better" means; they are applying a shared framework. The disadvantage is that the constitution must be carefully designed, and annotators must understand and apply it correctly. If the constitution is ambiguous, different annotators will interpret it differently, reintroducing the calibration problem.

Constitutional evaluation is also useful for multi-dimensional quality assessment. A constitution might specify separate criteria for helpfulness, harmlessness, and honesty, allowing annotators to mark which criterion a response fails. This produces richer signal than a single preference judgment, though it requires more annotation time per comparison.

Why Pairwise Comparisons Dominate

Despite the information efficiency of rankings and the richness of ratings, pairwise comparisons remain the standard for several reasons that reinforce each other in practice.

The first reason is reliability. Binary choices show higher inter-annotator agreement than multi-point scales. When forced to make a simple choice between two options, annotators tend to agree with each other more consistently than when asked to assign numerical scores or produce full rankings. This higher agreement translates directly into cleaner training signal for reward models. If two annotators looking at the same pair both prefer Response A, we have strong evidence that Response A is better. If they assign different ratings (say, 3 and 5), we have much weaker signal.

The second reason is theoretical grounding. The Bradley-Terry model provides a principled way to aggregate pairwise comparisons into reward scores. This model assumes that each response has some underlying "quality" score and that the probability of preferring one response over another follows a logistic function of the difference in their quality scores. We will explore this model in depth in the next chapter, but the key point is that pairwise data fits naturally into this framework, enabling principled learning from diverse comparison data.

The third reason is practical simplicity. The interface is straightforward to build, the data format is simple to process, and the annotation task is easy to explain to new annotators. These practical considerations matter enormously when scaling to thousands of comparisons across many annotators. Simpler collection pipelines have fewer points of failure and are easier to quality-control.

Most major preference datasets, including those used to train InstructGPT, Claude, and Llama 2, rely primarily on pairwise comparisons, which validates this choice empirically across multiple organizations and use cases.

Collection Interface Design

The interface through which annotators provide preferences significantly impacts data quality. Small design decisions can introduce systematic biases or reduce annotator efficiency in ways that are difficult to detect after the fact. A well-designed interface makes the annotator's job easier and produces cleaner data; a poorly designed one can corrupt the entire dataset with subtle biases that persist through all downstream training.

Think of the interface as a measuring instrument. A thermometer that reads one degree too high does not become more accurate just because you use it many times. Similarly, an interface that systematically biases annotators toward longer responses produces biased data no matter how many annotations you collect. The systematic error does not average out with scale; it compounds. This makes interface design one of the highest-impact decisions in the entire preference collection process.

Layout Considerations

The physical arrangement of responses matters more than one might initially expect. Standard practice places the two responses side-by-side or stacked vertically with consistent formatting. Several key principles have emerged from studies of annotation bias.

Randomized position is essential: the position of responses (left versus right, or top versus bottom) must be randomized across comparisons. Without randomization, position bias can corrupt the data. Research on human choice behavior consistently finds that people show position preferences, often favoring the first or last option presented. In web interfaces, users sometimes show a tendency to click the left option or the top option. If one model's responses always appear in position A, any systematic position preference will masquerade as a quality difference between models.

Consistent formatting is equally important. Both responses should use identical font sizes, paragraph spacing, and code formatting. Visual differences can bias judgments toward the more polished-looking response, even if content quality is equal. A response that happens to render with better code highlighting or cleaner bullet points may receive unwarranted preference because it simply looks more professional. This is particularly relevant for technical responses that include code snippets, mathematical notation, or structured lists. The interface should render all formatting identically regardless of which model generated the response.

No metadata leakage is a critical requirement. The interface should hide information about which model generated which response. If annotators can guess that "Response A" comes from the newer model, they may unconsciously favor it, introducing a systematic bias toward the model whose outputs are more recognizable. This is particularly important when comparing responses from different model versions or different model families, where stylistic differences may make model identity identifiable to experienced annotators. Some annotation platforms go so far as to rephrase responses slightly to remove identifying stylistic signatures.

Response Length Considerations

Length is a major confounder in preference data, and managing it requires deliberate effort throughout the collection process. Annotators often prefer longer responses, even when the additional content is filler. This creates a problematic incentive structure: models learn that verbosity correlates with preference, independent of actual quality. The model effectively learns to pad responses with tangents and hedges or simply elaborate beyond what the question requires.

This phenomenon arises from a reasonable heuristic that often fails in practice. Longer responses frequently contain more information, and more information is often helpful. However, this correlation breaks down when models learn to pad responses with redundant content or tangential details. The annotator's implicit preference for thoroughness gets exploited by the model, producing responses that appear complete but contain little useful content per word.

Studies on preference data consistently find that response length is one of the strongest predictors of preference, sometimes exceeding actual content quality in predictive power. Research from Anthropic and other labs has documented this length bias extensively. In several public preference datasets, the longer response was preferred in over 65% of comparisons, a far higher rate than would be expected if length and quality were uncorrelated.

Mitigation strategies include:

  • Training annotators to explicitly consider length-quality tradeoffs and include examples in guidelines showing that concise, accurate responses are preferred over verbose, padded ones.
  • Including explicit guidelines that state "longer is not automatically better" and providing examples where the shorter response is clearly superior.
  • Generating response pairs of similar length when possible, removing length as a confounding factor.
  • Post-hoc analysis to detect and correct length bias by examining whether annotators systematically prefer longer responses and adjusting accordingly.

None of these strategies fully eliminates length bias, but together they can substantially reduce it. The most effective approach is combining explicit annotator training with pair generation strategies that control for length.

Annotation Time Tracking

Recording how long annotators spend on each comparison provides valuable quality signals that are otherwise invisible. Suspiciously fast annotations (under 10 seconds for long responses) may indicate satisficing or random selection: the annotator is clicking through without reading carefully. This behavior can be rational from the annotator's perspective if they are paid per annotation and accuracy is not enforced, but it produces low-quality data.

Time data also helps identify difficult comparisons. Very slow annotations may indicate cases where the annotator is uncertain, which suggests these are close comparisons that may need clearer guidelines or additional annotators. A comparison that takes an average annotator two minutes to decide is a very different kind of data point than one that takes ten seconds.

Time data enables several downstream quality control mechanisms. You can flag and review all annotations below a minimum time threshold, calibrate expected annotation times for different response lengths, and monitor for annotators who progressively speed up over time (a sign of fatigue or gaming). Some platforms automatically reject annotations below a minimum time threshold and require the annotator to redo the comparison.

Forced Attention Mechanisms

Beyond tracking time, some annotation platforms use forced attention mechanisms to ensure annotators read the responses. These include:

Forced scrolling, which requires annotators to scroll through the full response before the comparison buttons become active. This ensures annotators cannot skip to the end without reading. Attention checks, which embed short questions about the content of the responses that annotators must answer correctly to submit their preference. These questions catch annotators who are selecting randomly without reading. Progress indicators that show annotators how far through each response they have read, creating a social norm of full engagement.

These mechanisms add friction to the annotation process, which can reduce throughput, but they improve quality by ensuring that preferences reflect actual reading of the responses rather than cursory inspection. The right balance depends on your annotator population and the consequences of noisy data.

Comparison Design

The prompts and responses used in preference collection determine what behaviors the resulting model will learn. Poor comparison design limits what alignment can achieve, no matter how good your annotators are. This section explores how to construct comparisons that provide maximal information about the quality distinctions we care about.

The central principle of comparison design is that you get the model behavior you measure for. If your preference dataset consists entirely of factual questions about science, your aligned model will learn to produce good science answers but may not generalize to other domains. If your comparisons only contrast clearly good responses against clearly bad ones, your model will learn to avoid obvious failures but won't learn to make fine-grained quality distinctions in the middle of the quality distribution. Intentional, principled comparison design produces alignment that generalizes in the directions you intend.

Prompt Selection

The prompts used for preference collection should cover the distribution of tasks the model will encounter in deployment. This requires careful consideration of multiple dimensions that together determine the coverage and informativeness of your dataset.

Domain coverage is the most obvious consideration. Instructions should span different use cases. Creative writing and coding assistance require different behavior from factual questions or advice requests. Analysis and summarization add further demands, as does translation. A dataset heavily weighted toward one domain will produce a model with uneven capabilities, excellent in the covered domain but poorly calibrated in others. The distribution of prompts in your preference dataset should match, as closely as possible, the distribution of requests the model will receive in production.

Difficulty distribution is equally important but often neglected. Including both simple and complex prompts ensures the model learns preferences across difficulty levels. Easy prompts, like basic factual questions with clear answers, may have nearly universal agreement among annotators and provide clear but shallow signal. Harder prompts, like subtle ethical dilemmas or ambiguous requests, generate more disagreement but provide richer signal about what good judgment looks like in difficult cases. A good preference dataset includes prompts across the full difficulty spectrum.

Edge cases deserve deliberate inclusion. Prompts that probe potential failure modes should be specifically added to the dataset: ambiguous requests where the intended meaning is unclear, requests with hidden assumptions that may not hold, requests that might elicit harmful content if handled poorly, and requests where the appropriate response involves acknowledging uncertainty or declining. These adversarial prompts help the model learn appropriate boundaries and exception handling.

Prompt diversity can be achieved through several strategies. Organic collection gathers prompts from actual user interactions with early model versions, capturing the true distribution of real usage. Templated generation creates prompts by filling templates with varied content. This keeps systematic coverage of topic and difficulty combinations. Adversarial generation deliberately crafts prompts that probe edge cases, using techniques like red-teaming to find prompts that expose model weaknesses.

Response Generation

The responses being compared typically come from the model being aligned, or an earlier version of it, but the generation strategy affects what the model can learn. Different generation strategies expose different aspects of the model's behavior distribution.

Temperature diversity is a standard technique. Generating responses at different temperatures produces variety in style and quality. High-temperature samples (temperature above 1.0) produce creative but sometimes risky or incoherent outputs. Low-temperature samples (temperature near 0) produce safe but generic outputs that may not show useful variation. Comparing responses generated at different temperatures helps the model learn when creativity versus caution is appropriate for different types of prompts.

Using multiple model versions or models is another common strategy. Comparing outputs from the current model against a previous version creates informative contrasts that help the model learn from its own past mistakes. Comparing against a different model entirely, such as using GPT-4 outputs to create positive examples for training a smaller model, creates strong directional signal about what ideal outputs look like. If annotators consistently prefer GPT-4's responses over the model being trained, those comparisons teach the model to emulate GPT-4's style and quality.

Intentional degradation involves deliberately including low-quality responses to establish clear negative examples. These might be truncated responses, factually wrong responses, stylistically poor responses, or responses that fail to follow instructions. Easy comparisons where one response is clearly much better than the other provide strong, low-noise signal about basic quality expectations. However, these easy comparisons are less informative than close comparisons between two reasonable responses, because they don't teach the model to make fine-grained distinctions. A good dataset includes both easy and hard comparisons.

Contrastive Pair Generation

The most informative comparisons involve responses that differ in specific, meaningful ways. Rather than comparing a clearly good response to a clearly bad one, effective datasets include pairs where both responses are reasonable but differ along dimensions that matter for quality. These contrastive pairs force annotators, and later the reward model, to make fine-grained distinctions.

Effective contrastive pairs include situations where:

  • Both responses are reasonable but differ in style or emphasis, teaching the model which style is preferred for different contexts.
  • One response is more complete while the other is more concise, teaching the model to calibrate length to the complexity and nature of the request.
  • One response directly answers the question while the other provides useful context first, teaching the model appropriate ordering of information.
  • Both responses are correct but one is more clearly explained, teaching the model the value of clarity over technical correctness alone.
  • One response acknowledges uncertainty while the other states the same information confidently, teaching the model appropriate epistemic calibration.

The key insight is that we learn most from close comparisons where the distinction is subtle, not from obvious cases where any reasonable annotator would agree. A dataset composed entirely of easy comparisons produces a reward model that can identify poor responses but cannot distinguish between good and excellent ones, which is often the more important distinction in practice.

The Informativeness-Agreement Tradeoff

There is a fundamental tension in preference data design between informativeness and agreement. Easy comparisons (clearly good vs. clearly bad) produce high inter-annotator agreement but low information about fine-grained quality differences. Hard comparisons (two reasonable responses) provide high information about fine-grained distinctions but produce lower inter-annotator agreement. The optimal dataset includes both: enough easy comparisons to anchor the quality scale and enough hard comparisons to teach fine-grained distinctions. Most successful preference datasets target an overall agreement rate of 70-80%, which typically requires a mix of easy and challenging comparisons.

Annotator Guidelines

Clear guidelines are essential for consistent preference judgments. Without explicit criteria, different annotators will apply different standards, introducing noise into the dataset that cannot be easily distinguished from legitimate disagreement about quality. Think of annotator guidelines as the shared vocabulary of quality: they ensure that when two annotators say "this response is better," they mean something similar by "better."

Guidelines serve multiple functions simultaneously. They communicate what quality dimensions matter and how to weight them. They provide examples that calibrate annotators' internal standards to a common reference point. They address ambiguous cases before annotators encounter them, preventing inconsistent ad-hoc decisions. And they create a feedback loop: as annotation proceeds, new edge cases emerge that can be documented and added to the guidelines, progressively improving consistency over time.

The investment in guidelines pays off disproportionately. A poorly specified guideline leads to systematic biases that corrupt thousands of annotations before they are detected. A well-specified guideline, by contrast, reduces noise across the entire dataset and makes downstream reward model training more efficient by providing cleaner signal.

Defining Quality Criteria

Guidelines should specify what annotators should consider when comparing responses. These criteria operationalize our abstract notion of "quality" into concrete factors that annotators can evaluate independently. Good criteria are specific enough to guide consistent judgment but not so mechanical that they prevent an overall evaluation.

Common criteria in RLHF guidelines include:

  • Helpfulness: Does the response help you accomplish your goal? Does it answer the question asked, or does it dodge the question with tangentially related information? Does it provide actionable guidance rather than vague generalities? A response that sounds impressive but doesn't help the user accomplish their task fails on helpfulness even if it is technically accurate.
  • Accuracy: Is the information correct? For factual claims, does the evidence support the assertion? For reasoning tasks, are the logical steps valid and is the conclusion supported by the premises? For code, does the code work as described? Accuracy is often the highest-priority criterion because inaccurate information actively harms users.
  • Clarity: Is the response easy to understand? Is it well-organized? Does it explain technical concepts at an appropriate level for the apparent audience? A response that is technically correct but incomprehensible to the intended reader fails on clarity. Good clarity means the reader spends cognitive effort understanding the topic, not deciphering the response.
  • Harmlessness: Does the response avoid providing dangerous or harmful information? Does it refuse inappropriate requests appropriately without being unnecessarily restrictive? A response that correctly refuses to help with a clearly harmful request should be preferred over one that complies. But a response that refuses a legitimate request out of excessive caution is also penalized on this dimension.
  • Honesty: Does the response acknowledge uncertainty when appropriate? Does it avoid claiming capabilities the model doesn't have? Does it represent the state of knowledge accurately rather than fabricating confident-sounding answers? Epistemic honesty is particularly important for preventing harmful overconfidence.
  • Conciseness: Does the response communicate what it needs to without unnecessary padding? Is it as long as necessary but no longer? This criterion explicitly counteracts length bias by making conciseness a quality dimension in its own right.

Guidelines should specify how to weigh these criteria when they conflict. For example: "If one response is more helpful but slightly less accurate, prefer the accurate response. Accuracy takes priority over helpfulness." These explicit priority rules reduce the arbitrariness of difficult comparisons and help ensure that different annotators resolve the same conflict the same way.

The criteria themselves should also be domain-adapted where appropriate. The most important quality dimensions for a coding question differ from those for a creative writing request. Some annotation systems use different guideline versions for different task types. This keeps the quality framework is always appropriate for the type of comparison being made.

Handling Ambiguous Cases

Some comparisons have no clear winner, and guidelines must address this directly. Leaving ambiguous cases unspecified leads annotators to develop their own inconsistent resolution strategies, introducing noise that is particularly harmful because it concentrates in exactly the hard cases where clear signal is most valuable.

For ties, guidelines should specify when annotators may select "tie" or "no preference" and when they should push through to a decision. Some annotation projects discourage ties almost entirely, on the grounds that even a small preference difference provides training signal, while forcing annotators to decide pushes them to examine the responses more carefully. Others allow ties freely, preferring not to force arbitrary choices that add noise. The guideline should be explicit about which policy applies, and the policy should be chosen based on how ties will be handled in downstream training.

For cases where responses take different valid approaches, guidelines should specify resolution strategies. Two responses might both correctly answer a programming question but use different algorithms, or both correctly explain a concept but use different analogies. If both responses would fully satisfy a reasonable user, ties may be appropriate. If one approach is more pedagogically effective or more efficient in practice, annotators should prefer it and the guidelines should specify how to recognize that.

For cases with partial quality differences, where one response is better on some criteria and worse on others, guidelines should provide an explicit priority ordering or aggregation rule. Without this, annotators will apply their personal intuition about the relative importance of different criteria, introducing systematic variation based on annotator characteristics rather than response quality.

Calibration Examples

Including worked examples in the guidelines is one of the most effective investments in annotation quality. Calibration examples transform abstract criteria into concrete demonstrations that anchor annotators' internal quality standards. They show rather than tell, which is particularly important for subtle quality distinctions that resist pure verbal description.

Each calibration example should show the full context: the prompt, both responses, the correct choice, and a detailed explanation of why that choice is correct. The explanation should explicitly reference the quality criteria in the guidelines. This shows how to apply them. This creates a direct bridge between the abstract criteria and the concrete practice of annotation.

Effective calibration examples serve specific pedagogical functions:

  • Clear cases establish the baseline: an obviously good response versus an obviously poor one. This shows what the extremes of the quality scale look like and making the criteria feel concrete.
  • Close calls illustrate subtle distinctions: two good responses where one is slightly better on a specific dimension. This shows how to make fine-grained judgments when both options are reasonable.
  • Traps catch common mistakes: cases where surface features like length, formality, or confidence might mislead annotators away from the better response. For example, a confident-sounding but factually wrong response compared to a hesitant but accurate one, where the guidelines should make clear that accuracy takes priority over apparent confidence.
  • Edge cases document hard problems: ambiguities that the guidelines resolve in a specific way, even if the resolution is somewhat arbitrary, so that different annotators handle the same type of case consistently.

The calibration examples should be expanded over time as new patterns emerge. Every time a project manager reviews a batch of annotations and notices systematic inconsistency, the team should diagnose and document it. They can then add a calibration example or clarify the existing guidelines.

Documenting Edge Cases

As annotation proceeds, edge cases emerge that the original guidelines didn't anticipate. Maintaining a living FAQ or edge case document helps ensure consistent handling across annotators and over time. This document should be version-controlled and accessible to all annotators, with new entries added as decisions are made.

Q: What if one response is correct but uses informal language while the other is wrong but sounds professional? A: Prefer correctness over tone. Factual accuracy always takes priority over stylistic preferences. Q: What if the prompt asks for something unethical? A: Prefer the response that appropriately declines while still being respectful to the user. Q: What if one response is much more concise but the other covers an additional relevant point? A: If the additional point is genuinely relevant and useful, prefer the more complete response. If the additional point is tangential or the user clearly did not ask for it, prefer the concise response.

The act of documenting edge cases as they arise forces explicit decisions that improve consistency. Without documentation, different annotators will make different decisions for the same case type, introducing systematic variation that looks like random noise but reflects policy inconsistency.

Quality Control and Agreement

Preference data quality directly impacts alignment quality, and the relationship is not linear. Low-quality data is not just less efficient than high-quality data; it actively trains reward models in wrong directions. A reward model trained on data where annotators systematically prefer verbose responses will assign higher rewards to verbosity, and the resulting aligned model will be verbose even when conciseness would serve users better. Quality control is not optional overhead; it is core to alignment.

Multiple mechanisms help ensure data reliability, and they work best in combination. No single quality control mechanism is sufficient on its own. Inter-annotator agreement metrics tell you whether annotators are consistent but not whether they are consistently correct. Gold questions catch random clicking but not systematic biases. Statistical outlier detection flags individual annotators but cannot catch collective biases shared by all annotators in your pool. Effective quality control requires layered defenses.

Inter-Annotator Agreement

Having multiple annotators label the same comparison allows measurement of agreement, which provides a direct proxy for data quality. If annotators frequently disagree about which response is better, the data is noisy. If they almost always agree, the comparisons may be too easy to provide useful training signal. The target range of 70-80% agreement reflects a sweet spot where comparisons are challenging enough to be informative but consistent enough to be reliable.

Raw agreement rate is the simplest metric: the percentage of comparisons where annotators chose the same response. For binary choices (ignoring ties), random agreement is 50%, so rates should substantially exceed this. A raw agreement rate of 65% means annotators are agreeing at a rate only slightly better than chance, which suggests the data is very noisy. A rate of 90% might suggest the comparisons are too easy.

Cohen's Kappa addresses the key limitation of raw agreement rate: it does not account for the possibility that annotators might agree purely by chance. If two annotators are each selecting randomly, they would still agree about half the time on binary choices. We need a metric that tells us how much better than chance our annotators are agreeing.

Cohen's Kappa measures inter-annotator agreement while accounting for the probability of chance agreement. The core insight is that we want to measure the agreement that goes beyond what we would expect from random labeling. To compute Kappa, we first determine what fraction of agreement is due to chance by considering the marginal distributions of annotator choices, then measure how much actual agreement exceeds that baseline.

Formally, if we let pop_o be the observed agreement and pep_e be the expected agreement by chance, Cohen's Kappa is:

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

where:

  • κ\kappa: the Cohen's Kappa coefficient, ranging from negative infinity to 1
  • pop_o: the observed agreement (the proportion of items where annotators chose the same option)
  • pep_e: the expected agreement by chance (derived from the marginal distributions of annotator choices, computed as the sum over all labels of the product of each annotator's marginal probability of choosing that label)
  • po−pep_o - p_e: the agreement achieved beyond chance, the numerator of the correction
  • 1−pe1 - p_e: the maximum possible agreement beyond chance, used to normalize the metric to a fixed scale

To understand this formula, consider a concrete example. Suppose two annotators each label 100 comparisons. Annotator 1 prefers Response A in 60 cases and Response B in 40 cases. Annotator 2 prefers Response A in 55 cases and Response B in 45 cases. They agree in 70 cases (po=0.70p_o = 0.70).

The expected chance agreement is computed from their marginal distributions. The probability that both choose A by chance is 0.60×0.55=0.330.60 \times 0.55 = 0.33. The probability that both choose B by chance is 0.40×0.45=0.180.40 \times 0.45 = 0.18. So pe=0.33+0.18=0.51p_e = 0.33 + 0.18 = 0.51.

Applying the formula: κ=(0.70−0.51)/(1−0.51)=0.19/0.49≈0.39\kappa = (0.70 - 0.51) / (1 - 0.51) = 0.19 / 0.49 \approx 0.39, which falls in the "fair" agreement range despite a raw agreement rate of 70%. This illustrates why Cohen's Kappa can paint a more sobering picture than raw agreement: it corrects for the substantial chance agreement that exists in binary choice tasks.

The numerator po−pep_o - p_e represents how much agreement we observed beyond what chance would predict. The denominator 1−pe1 - p_e represents the maximum possible improvement over chance, which normalizes the metric. A value of 0 implies agreement no better than chance, while 1 implies perfect agreement. Negative values indicate agreement worse than chance, which might occur if annotators are systematically misunderstanding the task or applying opposite criteria.

Out[4]:
Visualization
Horizontal color-coded bar chart showing Cohen's Kappa interpretation ranges from Poor (below 0) through Slight, Fair, Moderate, Substantial, to Almost Perfect agreement (above 0.8).
Interpretation scale for Cohen's Kappa inter-annotator agreement values. The color-coded bars illustrate standard thresholds used in annotation quality assessment. Values above 0.6 indicate substantial agreement that is typically sufficient for reward model training, while values below 0.2 suggest guidelines need clarification or annotators need retraining.

Krippendorff's Alpha is a more flexible agreement measure that handles multiple annotators simultaneously and deals gracefully with missing data. Unlike Cohen's Kappa, which is defined for exactly two annotators, Krippendorff's Alpha generalizes to any number of annotators and accommodates the situation where different subsets of annotators label different items. This makes it more practical for large annotation projects where not every annotator labels every comparison. Alpha values are interpreted similarly to Kappa: above 0.8 indicates high reliability, 0.67-0.8 is acceptable for tentative conclusions, and below 0.67 is considered unreliable.

Typical preference datasets achieve 70-80% pairwise agreement, corresponding to Cohen's Kappa values in the moderate to substantial range. Perfect agreement is neither expected nor desirable: it would suggest the comparisons are so easy they provide little training signal. A Kappa of 0.5-0.7 is typically considered acceptable for preference data collection in practice.

A Worked Example: Computing Cohen's Kappa by Hand

Let's work through a concrete numerical example to make the Kappa calculation concrete. Suppose we have two annotators, each labeling the same 10 preference comparisons with choices A, B, or Tie.

Their labels are:

ComparisonAnnotator 1Annotator 2
1AA
2AB
3BB
4AA
5BB
6TieA
7AA
8BB
9AA
10BTie

Step 1: Compute observed agreement pop_o. The annotators agree on comparisons 1, 3, 4, 5, 7, 8, 9, for a total of 7 agreements out of 10. So po=7/10=0.70p_o = 7/10 = 0.70.

Step 2: Compute marginal distributions. Annotator 1 chose A on items 1,2,4,6,7,9 (count 6) and B on items 3,5,8,10 (count 4). That count mistakenly treats item 6 as A, but it is a Tie. The corrected counts are: Annotator 1 chose A on items 1,2,4,7,9 (5 times), B on items 3,5,8,10 (4 times), and Tie on item 6 (1 time). Annotator 2 chose A on items 1,4,6,7,9 (5 times), B on items 2,3,5,8 (4 times), and Tie on item 10 (1 time).

Step 3: Compute expected chance agreement pep_e. For each label, multiply the proportion chosen by Annotator 1 by the proportion chosen by Annotator 2, then sum:

  • For label A: (5/10)×(5/10)=0.25(5/10) \times (5/10) = 0.25
  • For label B: (4/10)×(4/10)=0.16(4/10) \times (4/10) = 0.16
  • For label Tie: (1/10)×(1/10)=0.01(1/10) \times (1/10) = 0.01
  • pe=0.25+0.16+0.01=0.42p_e = 0.25 + 0.16 + 0.01 = 0.42

Step 4: Compute Kappa:

κ=po−pe1−pe=0.70−0.421−0.42=0.280.58≈0.48\kappa = \frac{p_o - p_e}{1 - p_e} = \frac{0.70 - 0.42}{1 - 0.42} = \frac{0.28}{0.58} \approx 0.48

This falls in the "moderate" agreement range, which is typical for preference annotation tasks involving subtle quality judgments. The raw agreement of 70% looked more impressive than the Kappa of 0.48 indicates, because a substantial fraction of that agreement is attributable to both annotators frequently choosing A regardless of the comparison.

Disagreement Analysis

When annotators disagree, understanding why is as important as measuring how much. Different types of disagreement call for different responses, and categorizing disagreements systematically guides guideline improvements.

Ambiguous prompts create disagreement when different annotators interpret the prompt differently. If annotators disagree about what a prompt is asking for, they may evaluate responses against different implicit standards and produce inconsistent preferences. This type of disagreement calls for prompt revision or clarification rather than annotator retraining.

Subjective criteria generate disagreement that reflects real diversity in preferences. Some quality dimensions, like formality of tone or level of detail, have no objectively correct answer. Annotators with different backgrounds or preferences may prefer different styles without either being wrong. Documenting and quantifying this type of disagreement helps calibrate expectations and may motivate explicitly stratifying annotators by relevant characteristics.

Criteria conflicts occur when annotators weigh quality dimensions differently. One annotator might prioritize helpfulness while another prioritizes harmlessness, leading to different preferences for the same pair. This type of disagreement often responds well to clearer guidelines that specify priority orderings.

Annotation errors include cases where annotators misread a response, were confused by the interface, or made a mechanical mistake. These errors are often identifiable through follow-up review and can sometimes be corrected retrospectively. Some annotation platforms allow annotators to flag their own uncertainty or request clarification on specific comparisons.

Systematically categorizing disagreements across a sample of disputed comparisons can reveal patterns that guide targeted improvements to guidelines, training, or comparison design. The analysis pays for itself when it prevents the same type of disagreement from repeating across thousands of future annotations.

Quality Assurance Mechanisms

Several mechanisms work together to maintain data quality throughout collection:

Gold questions are comparisons with known correct answers embedded in the annotation stream. For example, one response might be clearly helpful and the other clearly harmful, where any reasonable annotator should prefer the helpful response. Annotators who miss gold questions can be flagged for retraining, additional supervision, or removal from the project. Gold questions also provide a continuous quality signal: if gold question accuracy drops over time for an annotator, it may indicate fatigue or decreasing engagement.

Consistency checks show the same comparison twice (sometimes with response order flipped) and verify whether the annotator's choice is consistent. Inconsistency on the same comparison indicates noise: the annotator is not making a stable quality judgment but rather clicking somewhat randomly. A small rate of inconsistency is expected (humans are not perfectly consistent), but a high rate indicates either a very difficult comparison or an unreliable annotator.

Feedback loops involve regular communication between project managers and annotators to discuss difficult cases, clarify guidelines, and address systematic problems. Annotation is an iterative process, and the people doing the annotation often have valuable insights about where the guidelines are unclear or missing important cases. Creating channels for this feedback and acting on it improves quality throughout the project.

Statistical outlier detection identifies annotators whose judgments systematically differ from the majority. If one annotator consistently prefers the shorter response while all others prefer the longer response, or consistently prefers Response A regardless of prompt, these patterns may indicate systematic bias, confusion about the task, or gaming. Some divergence is expected and may even be desirable (you want diverse perspectives), but extreme outliers that are statistically inconsistent with the annotator pool warrant investigation.

Implementing Preference Data Collection

Let's implement a simple preference data processing pipeline that demonstrates key concepts in code. This implementation will illustrate the data structures, quality metrics, and analysis workflows you would use in a real preference collection project.

The goal here is not to show production-ready code but to make the concepts concrete. A real preference collection platform would require a web interface, database backend, authentication system, and many other components. What we show here is the core data processing and analysis logic.

In[5]:
Code
from dataclasses import dataclass
from typing import Literal


@dataclass
class PreferenceSample:
    """A single preference comparison sample."""

    prompt: str
    response_a: str
    response_b: str
    preference: Literal["a", "b", "tie"]
    annotator_id: str
    annotation_time_seconds: float
    metadata: dict = None

    def __post_init__(self):
        if self.metadata is None:
            self.metadata = {}


@dataclass
class PreferenceDataset:
    """Collection of preference samples with quality metrics."""

    samples: list[PreferenceSample]

    def compute_agreement(self, other_samples: list[PreferenceSample]) -> float:
        """Compute agreement rate between two sets of annotations."""
        # Match samples by prompt
        prompt_to_self = {s.prompt: s for s in self.samples}

        agreements = 0
        comparisons = 0

        for other in other_samples:
            if other.prompt in prompt_to_self:
                self_pref = prompt_to_self[other.prompt].preference
                if self_pref == other.preference:
                    agreements += 1
                comparisons += 1

        return agreements / comparisons if comparisons > 0 else 0.0

The PreferenceSample dataclass captures all the information we need about a single preference comparison: the prompt, both responses, the annotator's choice, the annotator's identity (for quality tracking), the time spent on the annotation, and a flexible metadata field for additional tracking information. The PreferenceDataset wraps a collection of samples and provides analysis utilities.

Now let's create some example preference data that illustrates common patterns:

In[6]:
Code
import random

# Simulate realistic preference data
example_prompts = [
    "Explain quantum computing to a beginner.",
    "Write a haiku about artificial intelligence.",
    "What are the pros and cons of remote work?",
    "How do I debug a Python memory leak?",
    "Summarize the French Revolution in 3 sentences.",
]


def generate_mock_responses(prompt: str) -> tuple[str, str]:
    """Generate mock response pairs for demonstration."""
    responses = {
        "Explain quantum computing to a beginner.": (
            "Quantum computing uses qubits instead of regular bits. While normal bits are either 0 or 1, qubits can be both at once through superposition. This lets quantum computers solve certain problems much faster than regular computers.",
            "Quantum computing is a revolutionary paradigm in computational theory that leverages the principles of quantum mechanics, specifically superposition and entanglement, to perform calculations that would be intractable for classical computing architectures. The fundamental unit of quantum information is the qubit, which exists in a probabilistic superposition of states until measurement collapses the wave function.",
        ),
        "Write a haiku about artificial intelligence.": (
            "Silicon mind wakes\nPatterns emerge from chaos\nLearning never ends",
            "Artificial\nIntelligence is really cool\nRobots are the best",
        ),
        "What are the pros and cons of remote work?": (
            "Remote work offers flexibility and eliminates commuting, but can lead to isolation and blurred work-life boundaries. Success depends on individual work style and team communication practices.",
            "Pros: flexibility, no commute, comfortable environment. Cons: isolation, distractions, communication challenges. Remote work has both advantages and disadvantages that vary by person.",
        ),
        "How do I debug a Python memory leak?": (
            "Use the tracemalloc module to track memory allocations. Start with tracemalloc.start(), run your code, then use tracemalloc.take_snapshot() to see what's using memory. Look for objects that grow over time.",
            "Memory leaks in Python are rare due to garbage collection, but can occur with circular references or C extensions. You should probably just restart your program periodically to clear memory.",
        ),
        "Summarize the French Revolution in 3 sentences.": (
            "The French Revolution (1789-1799) overthrew the monarchy and established a republic based on Enlightenment ideals. It began with the storming of the Bastille and led to radical social change, including the abolition of feudalism. The revolution ended with Napoleon's rise to power.",
            "The French Revolution was when French people got mad at the king. There was lots of fighting and the king got his head cut off. Then Napoleon became the leader.",
        ),
    }
    return responses.get(prompt, ("Response A", "Response B"))


def simulate_annotation(
    prompt: str, response_a: str, response_b: str, annotator_id: str
) -> PreferenceSample:
    """Simulate an annotator making a preference judgment."""
    # Simple heuristics to simulate realistic annotation patterns
    len_diff = len(response_b) - len(response_a)

    # Simulate some annotators preferring verbose responses
    if "verbose" in annotator_id:
        preference = "b" if len(response_b) > len(response_a) else "a"
    # Others prefer concise, clear responses
    elif "concise" in annotator_id:
        preference = "a" if len(response_a) < len(response_b) else "b"
    # Most annotators have mixed preferences
    else:
        # Prefer response A for prompts about technical topics (our mock data has better A responses)
        if (
            "debug" in prompt.lower()
            or "quantum" in prompt.lower()
            or "revolution" in prompt.lower()
        ):
            preference = "a"
        else:
            preference = random.choice(["a", "b"])

    # Add some noise to simulate disagreement
    if random.random() < 0.15:
        preference = "b" if preference == "a" else "a"

    # Simulate annotation time (longer responses take longer to read)
    base_time = 15.0
    time_per_char = 0.02
    total_chars = len(response_a) + len(response_b)
    annotation_time = (
        base_time + total_chars * time_per_char + random.gauss(0, 5)
    )
    annotation_time = max(5.0, annotation_time)  # Minimum 5 seconds

    return PreferenceSample(
        prompt=prompt,
        response_a=response_a,
        response_b=response_b,
        preference=preference,
        annotator_id=annotator_id,
        annotation_time_seconds=annotation_time,
        metadata={"len_a": len(response_a), "len_b": len(response_b)},
    )

The simulation models three types of annotators: those who systematically prefer longer responses ("verbose"), those who systematically prefer shorter responses ("concise"), and neutral annotators who tend to prefer the better response but with some random noise. This mirrors the heterogeneity you observe in real annotator pools.

In[7]:
Code
# Generate preference data from multiple annotators
annotator_ids = ["ann_1", "ann_2", "ann_3_verbose", "ann_4_concise", "ann_5"]

all_samples = []
for prompt in example_prompts:
    response_a, response_b = generate_mock_responses(prompt)
    for annotator in annotator_ids:
        sample = simulate_annotation(prompt, response_a, response_b, annotator)
        all_samples.append(sample)

dataset = PreferenceDataset(samples=all_samples)
Out[8]:
Console
Total samples collected: 25
Unique prompts: 5
Annotators: 5

The dataset contains 25 samples distributed across 5 prompts and 5 annotators. This scale is sufficient to demonstrate the data structures and analysis pipelines, though real-world datasets typically contain thousands to hundreds of thousands of examples. The InstructGPT dataset used approximately 40,000 preference comparisons; Anthropic's HH-RLHF dataset contains roughly 169,000 comparisons. Scale matters because reward models are being trained on this data and need sufficient coverage of the prompt distribution.

Let's analyze the preference distribution and look for potential biases:

In[9]:
Code
from collections import defaultdict


def analyze_preferences(dataset: PreferenceDataset) -> dict:
    """Analyze preference patterns in the dataset."""
    analysis = {
        "preference_distribution": defaultdict(int),
        "annotator_preferences": defaultdict(lambda: defaultdict(int)),
        "length_correlation": {
            "preferred_longer": 0,
            "preferred_shorter": 0,
            "same_length": 0,
        },
        "avg_annotation_time": 0.0,
        "fast_annotations": 0,  # Under 10 seconds
    }

    total_time = 0.0
    for sample in dataset.samples:
        # Preference distribution
        analysis["preference_distribution"][sample.preference] += 1

        # Per-annotator preferences
        analysis["annotator_preferences"][sample.annotator_id][
            sample.preference
        ] += 1

        # Length correlation
        len_a = sample.metadata.get("len_a", len(sample.response_a))
        len_b = sample.metadata.get("len_b", len(sample.response_b))

        if sample.preference == "a":
            if len_a > len_b:
                analysis["length_correlation"]["preferred_longer"] += 1
            elif len_a < len_b:
                analysis["length_correlation"]["preferred_shorter"] += 1
            else:
                analysis["length_correlation"]["same_length"] += 1
        elif sample.preference == "b":
            if len_b > len_a:
                analysis["length_correlation"]["preferred_longer"] += 1
            elif len_b < len_a:
                analysis["length_correlation"]["preferred_shorter"] += 1
            else:
                analysis["length_correlation"]["same_length"] += 1

        # Annotation time
        total_time += sample.annotation_time_seconds
        if sample.annotation_time_seconds < 10.0:
            analysis["fast_annotations"] += 1

    analysis["avg_annotation_time"] = total_time / len(dataset.samples)

    return dict(analysis)


analysis = analyze_preferences(dataset)
Out[10]:
Console
Preference Distribution:
  a: 15 (60.0%)
  b: 10 (40.0%)

Average annotation time: 21.5 seconds
Fast annotations (<10s): 1

Length correlation (excluding ties):
  Preferred longer response: 56.0%
  Preferred shorter response: 44.0%

The distribution reveals a mix of preferences, with some annotators showing distinct biases. The length correlation metrics help identify whether annotators systematically favor longer responses, a common quality issue where verbosity is mistaken for helpfulness. In a real dataset, finding that 65% of preferences go to the longer response would be a red flag requiring investigation and possible guideline revision.

Out[11]:
Visualization
Scatter plot of response length differences (preferred minus rejected character count) across annotation samples, colored by annotator type, with verbose-preferring annotators clustering above zero and concise-preferring annotators clustering below.
Response length differences between preferred and rejected outputs, colored by annotator type. Positive values indicate the preferred response was longer than the rejected one. The distinct clustering of verbose-preferring annotators (red) above the zero line visually confirms their systematic length bias, while concise-preferring annotators (blue) cluster below. Neutral annotators (gray) show a mixed pattern reflecting genuine quality judgments.

Now let's compute inter-annotator agreement using Cohen's Kappa:

In[12]:
Code
def compute_cohens_kappa(annotations_1: list, annotations_2: list) -> float:
    """
    Compute Cohen's Kappa for two sets of annotations.

    Kappa = (p_o - p_e) / (1 - p_e)
    where p_o is observed agreement and p_e is expected agreement by chance.
    """
    assert len(annotations_1) == len(annotations_2)
    n = len(annotations_1)

    if n == 0:
        return 0.0

    # Get unique labels
    labels = sorted(set(annotations_1) | set(annotations_2))

    # Compute observed agreement
    agreements = sum(
        1 for a1, a2 in zip(annotations_1, annotations_2) if a1 == a2
    )
    p_o = agreements / n

    # Compute expected agreement by chance
    p_e = 0.0
    for label in labels:
        # Proportion of times annotator 1 chose this label
        p1 = sum(1 for a in annotations_1 if a == label) / n
        # Proportion of times annotator 2 chose this label
        p2 = sum(1 for a in annotations_2 if a == label) / n
        p_e += p1 * p2

    # Compute kappa
    if p_e == 1.0:
        return 1.0  # Perfect agreement

    kappa = (p_o - p_e) / (1 - p_e)
    return kappa


def compute_pairwise_agreement(dataset: PreferenceDataset) -> dict:
    """Compute agreement between all pairs of annotators."""
    # Group samples by prompt and annotator
    prompt_annotations = defaultdict(dict)
    for sample in dataset.samples:
        prompt_annotations[sample.prompt][sample.annotator_id] = (
            sample.preference
        )

    # Get all annotators
    annotators = sorted(set(s.annotator_id for s in dataset.samples))

    # Compute pairwise kappa
    agreement_matrix = {}
    for i, ann1 in enumerate(annotators):
        for ann2 in annotators[i + 1 :]:
            # Get annotations from both annotators for shared prompts
            ann1_labels = []
            ann2_labels = []
            for prompt, annotations in prompt_annotations.items():
                if ann1 in annotations and ann2 in annotations:
                    ann1_labels.append(annotations[ann1])
                    ann2_labels.append(annotations[ann2])

            if ann1_labels:
                kappa = compute_cohens_kappa(ann1_labels, ann2_labels)
                agreement_matrix[(ann1, ann2)] = kappa

    return agreement_matrix


agreement = compute_pairwise_agreement(dataset)
Out[13]:
Console
Pairwise Inter-Annotator Agreement (Cohen's Kappa):
  ann_1 vs ann_2: κ = 0.000 (slight)
  ann_1 vs ann_3_verbose: κ = 0.000 (slight)
  ann_1 vs ann_4_concise: κ = 0.000 (slight)
  ann_1 vs ann_5: κ = 0.000 (slight)
  ann_2 vs ann_3_verbose: κ = 0.167 (slight)
  ann_2 vs ann_4_concise: κ = 0.000 (slight)
  ann_2 vs ann_5: κ = -0.364 (slight)
  ann_3_verbose vs ann_4_concise: κ = 0.000 (slight)
  ann_3_verbose vs ann_5: κ = -0.364 (slight)
  ann_4_concise vs ann_5: κ = 0.000 (slight)

Average Kappa: -0.056

The agreement analysis reveals important patterns. Annotators with different biases (verbose versus concise preferences) show lower agreement with each other than they do with annotators who share their biases. The neutral annotators show moderate agreement with each other and mixed agreement with the biased annotators, depending on whether their quality judgments happen to align with the biased annotator's length-based heuristic. In real preference collection, observing this pattern would motivate investigating what drives the low-agreement annotator pairs and potentially removing or retraining the most extreme outliers.

Out[14]:
Visualization
Heatmap grid of pairwise Cohen's Kappa scores between five annotators, with green diagonal cells at 1.00 and mostly orange or red off-diagonal cells at low or negative agreement.
Pairwise Cohen's Kappa agreement matrix for five annotators. Apart from perfect self-agreement on the diagonal, most pairwise scores are near zero or negative. The low agreement involving differently biased annotators shows how inconsistent internal criteria can degrade the training signal if they are not detected and corrected.

Let's visualize the preference patterns across annotators to understand how their biases manifest in the data:

In[15]:
Code
annotators = sorted(set(s.annotator_id for s in dataset.samples))
preferences = ["a", "b", "tie"]

# Build data matrix
data = np.zeros((len(annotators), len(preferences)))
for i, ann in enumerate(annotators):
    ann_samples = [s for s in dataset.samples if s.annotator_id == ann]
    for j, pref in enumerate(preferences):
        data[i, j] = sum(1 for s in ann_samples if s.preference == pref)

# Normalize to percentages
data_pct = data / data.sum(axis=1, keepdims=True) * 100
Out[16]:
Visualization
Stacked bar chart of five annotators'' Response A and Response B selections: ann_1 is 100 percent A, ann_2 and ann_3_verbose are 60 percent A, ann_4_concise is 100 percent B, and ann_5 is 80 percent A.
Preference distributions for five annotators showing the proportion of A, B, and Tie selections in this small seeded example. The concise-preferring annotator selects Response B throughout, while the verbose-preferring annotator has a more modest 60/40 split. The neutral annotators vary substantially, illustrating why per-annotator diagnostics need enough samples before a pattern is treated as a stable bias.
Out[17]:
Visualization
Histogram of annotation completion times in seconds, with a dashed red vertical line at 10 seconds indicating the minimum reliable annotation threshold, and a solid green line showing the mean annotation time.
Distribution of annotation times in seconds across all 25 samples. The vertical dashed red line at 10 seconds marks the quality control threshold; annotations to the left of this line are flagged as potentially unreliable and subject to review. The distribution is right-skewed because longer responses require more reading time, with the mean typically falling between 20 and 40 seconds for the response lengths in this dataset.

Converting to Training Format

Preference data must be converted to a format suitable for training reward models before it can be used. The standard format pairs each prompt with its chosen and rejected responses, creating a direct learning signal: the reward model should assign higher scores to chosen than to rejected.

This conversion step also provides an opportunity for final quality filtering. You can exclude ties (which provide no directional signal), remove annotations with suspicious timing, weight annotations by annotator reliability if you have quality estimates, or deduplicate prompts to avoid over-representing certain topics in the training data.

In[18]:
Code
def convert_to_training_format(
    dataset: PreferenceDataset, exclude_ties: bool = True
) -> list[dict]:
    """
    Convert preference samples to the standard training format.

    Returns list of dicts with keys: prompt, chosen, rejected
    """
    training_data = []

    for sample in dataset.samples:
        if exclude_ties and sample.preference == "tie":
            continue

        if sample.preference == "a":
            chosen = sample.response_a
            rejected = sample.response_b
        else:
            chosen = sample.response_b
            rejected = sample.response_a

        training_data.append(
            {
                "prompt": sample.prompt,
                "chosen": chosen,
                "rejected": rejected,
                "annotator_id": sample.annotator_id,
            }
        )

    return training_data


training_data = convert_to_training_format(dataset)
Out[19]:
Console
Training samples (excluding ties): 25

Sample training instance:
Prompt: Explain quantum computing to a beginner....
Chosen: Quantum computing uses qubits instead of regular bits. While normal bits are eit...
Rejected: Quantum computing is a revolutionary paradigm in computational theory that lever...

This format directly supports the reward modeling objective we'll cover in upcoming chapters: the reward model learns to assign higher scalar scores to chosen responses than rejected ones for the same prompt. The training loss penalizes the model whenever its score for the chosen response is lower than its score for the rejected response, driving the model to internalize the quality distinctions captured by the preference data.

The annotator_id field in the training format is valuable for advanced training strategies. You can filter training data to include only high-agreement annotations, weight samples by annotator reliability, or train separate reward models for different annotator subgroups and then ensemble them. Preserving annotator identity in the training format keeps these options open.

Key Parameters

The key parameters for the preference data structures are:

  • prompt: The instruction or question provided to the model. The distribution of prompts determines which behaviors the aligned model learns to optimize.
  • response_a / response_b: The two model outputs presented for comparison. The generation strategy and response quality distribution determine the signal available in the data.
  • preference: The annotator's judgment ("a", "b", or "tie"). This is the primary learning signal for the reward model.
  • annotator_id: Unique identifier for the annotator, used to track individual biases and agreement patterns. Essential for quality control.
  • annotation_time_seconds: Time spent on the annotation, used to flag potentially unreliable fast annotations and identify difficult comparisons.
  • metadata: Flexible field for additional tracking information such as response lengths, annotation round, or task category. Useful for bias analysis and quality control.

Real-World Preference Datasets

Several public preference datasets demonstrate these principles at scale, and studying them reveals the design choices that different organizations have made and their consequences.

Anthropic HH-RLHF contains approximately 169,000 conversations rated for helpfulness and harmlessness. Annotators compared assistant responses in multi-turn dialogues, focusing on whether responses were helpful without being harmful. The dataset explicitly separates the helpfulness and harmlessness dimensions. This provides separate labeled subsets for each. This design choice reflects the tension between helpfulness and safety: responses that are maximally helpful may sometimes be harmful, and perfectly safe responses may refuse to be helpful. Studying where this tension appears is itself informative.

OpenAssistant is a crowdsourced dataset with multi-level rankings. Annotators ranked multiple responses per prompt, and the data includes quality flags, annotator metadata, and multi-turn conversation structure. The crowdsourced nature provides diverse annotator perspectives but introduces more variability in annotation quality than professionally managed annotation projects. OpenAssistant demonstrates both the potential and the challenges of large-scale community annotation.

Stanford Human Preferences (SHP) is derived from Reddit, where upvotes serve as implicit preference signals. When a user's comment receives more upvotes than another comment on the same post, we can interpret this as a collective preference for the higher-upvoted response. This represents a fundamentally different collection paradigm: instead of explicit annotation by paid workers, we harvest naturally occurring human feedback from existing social data. The advantages are scale (millions of comparisons) and authenticity (real human judgments in natural context). The disadvantages include noisiness (upvotes reflect many factors beyond response quality), demographic skew (Reddit users are not representative of all users), and the inability to control what types of comparisons are available.

UltraFeedback uses GPT-4 to provide preference judgments, representing the AI feedback approach where a stronger model labels data for training weaker models. This is sometimes called RLAIF (Reinforcement Learning from AI Feedback) rather than RLHF. The approach is far more scalable than human annotation: GPT-4 can evaluate thousands of comparisons per hour at low cost. The quality of the resulting signal depends on the quality of the judge model, and GPT-4 has its own biases and blind spots that transfer into the training data. Research has shown that RLAIF can produce competitive results to RLHF in many settings, but the two approaches may have different failure modes.

Each dataset makes different design choices that affect what behaviors the resulting models learn. Understanding these choices helps practitioners select appropriate data for their alignment goals and diagnose why models trained on different datasets exhibit different behavioral patterns.

Limitations and Impact

Human preference data, while powerful, has fundamental limitations that shape what alignment can achieve. Understanding these limitations is essential for using preference data responsibly and for interpreting the behavior of RLHF-trained models.

The most significant challenge is that preferences are not ground truth about what is objectively good. Annotators bring different values and backgrounds to the task, so they may interpret the same response differently. A response one annotator finds helpful might strike another as condescending. A response deemed harmless by one might seem problematic to another based on cultural context. When we aggregate preferences, we construct a statistical summary of diverse opinions, not discover objective facts about quality. This means aligned models reflect a particular distribution of human views, not some universal standard. The model is aligned to the preferences of the annotator pool, which may differ substantially from the preferences of the eventual user population.

Annotator demographics matter enormously, and most preference datasets do not adequately represent global diversity. Most preference datasets are labeled by workers in specific geographic regions, typically English-speaking and with access to the labeling platforms. These annotators may not represent the global user base of language models. Preferences about appropriate humor, cultural references, discussion of sensitive topics, and interpersonal communication styles vary across populations. Models aligned to one population's preferences may behave inappropriately or paternalistically for users from different cultural backgrounds. This concern is documented: models trained on primarily Western preference data can exhibit cultural assumptions that surprise or offend users from other backgrounds.

The comparison format itself limits what preferences can be expressed. Pairwise comparisons work well for choosing between similar responses but struggle to capture the full structure of human preferences. An annotator might think both responses are terrible, or both are excellent, but the interface forces a choice between them. Expressed preference in such cases tells us about relative quality within the pair but nothing about absolute quality. Some preference relationships are also intransitive: an annotator might prefer A to B, prefer B to C, but also prefer C to A when compared directly. This can happen when different comparisons make different quality dimensions salient, and the Bradley-Terry model cannot represent intransitive preferences correctly. The model assumes a total ordering exists, which forces it to approximate intransitive preference structures with the closest consistent ordering.

Scale and cost remain persistent challenges. Collecting high-quality preference data is expensive and slow. Annotators must read lengthy responses carefully, consider multiple quality dimensions, and make difficult judgments. Rushing this process degrades quality, but thorough annotation limits dataset size. This tension between quality and quantity affects all preference datasets and constrains how much alignment can be achieved. The RLHF datasets that produced the most impressive models, like those used for InstructGPT and early Claude versions, required substantial investment in both annotator training and quality control infrastructure.

The preference labeling process can also introduce its own systematic distortions. Annotators may be influenced by factors that should not affect quality judgments: the visual length of the response, the presence of formatting like bullet points, the apparent confidence of the writing style, or even subtle cues about which model produced which response. These biases propagate through the reward model into the aligned model, teaching the model to optimize for features that correlate with annotator preference rather than with response quality. Length bias is the most well-documented example, but similar effects likely exist for many other surface features.

Reward model overfitting is another fundamental concern. The reward model learns from a finite sample of human preferences and may generalize imperfectly to new cases. When the language model is optimized against the reward model through RLHF, it may find responses that exploit gaps in the reward model's generalization, producing outputs that score high on the reward model but would not be preferred by humans. This phenomenon, sometimes called reward hacking or specification gaming, is a fundamental challenge for any learned reward signal. Detecting and preventing it requires ongoing human evaluation beyond the initial preference collection process.

Despite these limitations, human preference data has enabled substantial progress in alignment. The RLHF pipeline, built on preference data, transformed language models from capable but unreliable systems into assistants that follow instructions, acknowledge uncertainty, and refuse clearly harmful requests with some consistency. The preference data approach succeeded where explicit rules failed, precisely because preferences encode implicit knowledge that resists formalization. A model trained on millions of examples of humans preferring clear explanations to jargon-laden ones learns something about clarity that no rule fully captures. The limitations are real and must be understood, but so is the progress.

Summary

Human preference data provides the empirical foundation for language model alignment. Rather than specifying desired behavior through rules, we collect examples of humans choosing between model outputs, then use these choices to train reward models that guide model behavior. The simplicity of this idea belies the substantial engineering and judgment required to execute it well.

The key elements of preference data collection include:

  • Collection format: Pairwise comparisons dominate due to cognitive simplicity and theoretical grounding with the Bradley-Terry model, though rankings and ratings offer alternatives with different tradeoffs in information density and annotator cognitive load.
  • Interface design: Randomized response positions, consistent formatting, hidden model identity, and forced attention mechanisms prevent systematic biases from corrupting the data. Small interface decisions compound into large dataset-level effects at scale.
  • Comparison design: Prompts should cover expected use cases across domains and difficulty levels, including edge cases that probe failure modes. Response pairs should include both easy comparisons that anchor the quality scale and hard, contrastive comparisons that teach fine-grained distinctions.
  • Annotator guidelines: Clear, specific criteria for quality dimensions (helpfulness, accuracy, harmlessness, honesty, conciseness) with explicit priority orderings for conflicts and worked calibration examples help ensure consistent judgments across annotators and over time.
  • Quality control: Inter-annotator agreement metrics (Cohen's Kappa, Krippendorff's Alpha), gold questions, consistency checks, annotation time tracking, and statistical outlier detection together maintain data reliability through the collection process. No single mechanism is sufficient; layered defenses are necessary.
  • Bias analysis: Length bias, position bias, and demographic bias in the annotator pool must be actively monitored and mitigated throughout collection, not just at the end.

The decisions made during preference collection cascade through the entire alignment pipeline. In the next chapter, we'll see how the Bradley-Terry model provides a principled framework for converting pairwise preference data into the reward scores that drive RLHF training.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about human preference data collection for language model alignment.

Human Preference Data

Question 1 of 70 of 7 completed
Why have pairwise comparisons become the dominant paradigm for collecting human preference data in RLHF?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025humanpreference, author = {Michael Brenndoerfer}, title = {Human Preference Data: Collection for LLM Alignment}, year = {2025}, url = {https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Human Preference Data: Collection for LLM Alignment. Retrieved from https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment
MLAAcademic
Michael Brenndoerfer. "Human Preference Data: Collection for LLM Alignment." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment>.
CHICAGOAcademic
Michael Brenndoerfer. "Human Preference Data: Collection for LLM Alignment." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Human Preference Data: Collection for LLM Alignment'. Available at: https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Human Preference Data: Collection for LLM Alignment. https://mbrenndoerfer.com/writing/human-preference-data-collection-rlhf-alignment

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.