Part of Language AI Handbook
Explains how position bias, verbosity bias, and sycophancy distort LLM evaluation. Measure swap consistency, detect length effects.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Position Bias in LLM Judges
In the previous chapter on LLM-as-Judge, we established that language models can serve as scalable, human-correlated evaluators: cheaper than annotation, faster than benchmarks, and surprisingly accurate on a wide range of quality dimensions. We saw how to design judge prompts, pick judge models, and calibrate outputs against human ratings. But we ended with a warning: LLM judges inherit the same failure modes as the models they evaluate. They can be manipulated or fooled, producing systematic errors in ways that are invisible unless you specifically look for them.
This chapter is about one of the most pervasive of those failure modes: position bias. When an LLM judge reads two candidate answers and decides which is better, it should base that decision entirely on quality. But it often does not. The order in which responses appear, the length of each response, and whether one response happens to agree with what the judge already believes can all tip the verdict, independently of actual quality. These biases distort evaluation results, compromise model comparisons, and silently mislead practitioners who trust LLM judges without measuring their reliability.
Understanding position bias is needed for anyone building production evaluation pipelines. We will cover what the bias is, why it occurs, how to measure it quantitatively, and what mitigation strategies work and when they fail. We will also look at two closely related biases, verbosity bias and sycophancy, which frequently compound position effects and require their own treatment.
Position bias in LLM evaluation refers to the tendency of a judge model to systematically prefer responses that appear in a particular position in the prompt (for example, first or last), regardless of the actual quality of those responses.
Why Position Bias Exists
To understand why position bias occurs in LLM judges, it helps to think about what happens inside an autoregressive language model when it reads a long prompt containing two candidate answers. The model processes tokens left to right, building up a contextual representation as it goes. By the time it reaches the instruction to compare and score, the representations of Response A and Response B are not equally salient in the model's attention pattern. The attention mechanism, trained on data where position in a sequence carries real meaning, treats earlier and later content differently in ways that do not simply cancel out.
There are two primary mechanisms that contribute to position bias. The first is recency bias: models often weight recent tokens more heavily because recency is a reliable signal in natural language. In a news article, the most important conclusions often appear near the end. In an argument, the final rebuttal carries more rhetorical weight. In narrative writing, the resolution comes last. When applied to an evaluation prompt, recency manifests as a preference for whichever response appeared last. The second mechanism is primacy bias: in some contexts, the first item in a list or the first argument in a debate receives outsized weight simply because it sets the reference point for everything that follows. Human psychology exhibits a similar split depending on the task, and models trained on human-generated text absorb these contradictory patterns.
The key insight is that neither bias operates alone, and which direction dominates depends on the judge model's architecture, the nature of its training data, the length of the prompt, and even the specific comparison task. Some models show strong first-position preference. Others show strong last-position preference. Some flip direction depending on whether the evaluation prompt uses markdown headers, whether responses are separated by explicit delimiters, or whether there is a brief summarizing sentence before the final verdict. The only way to know which pattern your judge exhibits is to measure it empirically.
Swap consistency measures how often a judge gives the same relative verdict when the order of two responses is reversed. A perfectly unbiased judge would achieve 100% swap consistency. Values below 100% reveal position-dependent judgments.
Position bias was first rigorously documented in the context of LLM evaluation by the MT-Bench paper (Zheng et al., 2023) and in the LMSYS Chatbot Arena analysis. Researchers found that GPT-4, when used as a judge, showed a statistically significant preference for the first response in roughly 60 to 70 percent of comparisons where it changed its verdict when the order was swapped. This was not a small or ignorable effect. It meant that the outcome of an LLM evaluation depended substantially on which team happened to submit their model's response first in the prompt template, a variable with no connection to quality whatsoever.
The Attention Mechanism and Positional Salience
To build an even deeper intuition for why position bias exists, consider how a transformer's attention mechanism operates during a long-form evaluation prompt. When the model reads a prompt containing 800 tokens of context and reaches the final instruction "Which response is better? Answer A or B," the attention patterns over that context are not uniform. The model attends more strongly to some tokens than others, and both position within the sequence and proximity to the current generation step influence those weights.
Research on attention patterns in long contexts has shown that many transformer models exhibit a characteristic attention shape sometimes called the "U-curve": tokens at the very beginning and very end of a long sequence receive disproportionately high attention, while tokens in the middle tend to be underweighted. This is sometimes called the "lost in the middle" phenomenon. For an evaluation prompt where Response A occupies the beginning and Response B occupies the middle-to-end, Response A benefits from primacy attention and Response B benefits from recency attention, but the magnitudes are not symmetric across model families and sizes.
Instruction-tuned models reduce this asymmetry but do not eliminate it. When a model is fine-tuned on preference data where human annotators reviewed pairs in a fixed order, the model learns associations between positional patterns and quality assessments that were never intended to be part of the evaluation rubric. Those learned associations surface as position bias.
How Prompt Design Amplifies Bias
Position bias is not purely a product of model architecture. The structure of the evaluation prompt significantly shapes how much bias appears. Several design choices systematically amplify or dampen position effects:
Asymmetric framing occurs when the prompt refers to the two responses using unequal language, such as "here is the first response" versus "here is another response." The word "first" subtly signals that the ordering is meaningful, which may cause the model to treat position as a legitimate evaluation signal.
Distance between responses affects how coherently the model can compare them. When Response A and Response B are separated by a long rubric section or detailed evaluation criteria, the model's contextual representation of A has partially faded by the time it fully processes B. Shorter, more structured prompts with explicit separators reduce this effect.
The label convention matters more than it might seem. Using labels like "Response A" and "Response B" is better than "Response 1" and "Response 2" because ordinal numbers ("first," "second") are more strongly associated with ranked sequences in training data. Using neutral labels, or randomly assigning labels like arbitrary colors or codes, further weakens the positional signal.
The position of evaluation criteria affects which response the model mentally compares to the rubric most vividly. If evaluation criteria are stated after both responses, the most recent response (B) is compared more directly to the freshly active criteria. If criteria come before both responses, both are evaluated against the same contextually distant rubric.
Understanding these design factors is practically useful: careful prompt engineering can reduce position bias even without any algorithmic mitigation. But it cannot eliminate it entirely, which is why measurement and swapping remain necessary.
The Three Biases: Position, Verbosity, and Sycophancy
These three biases form a cluster that often co-occur and interact. Understanding each individually is necessary before examining how they compound. This section digs into the mechanics of each bias, how to define it precisely, how to quantify it, and what drives it.
Position Bias: Order Effects on Verdicts
Position bias is a structural problem. The prompt template for pairwise evaluation typically looks like this: the question, then Response A, then Response B, then the instruction to judge. This fixed structure means that whichever response occupies position A always appears first. If the judge model has even a slight preference for first-position responses, and most do to some degree, this creates a systematic advantage that accumulates across many evaluations. At scale, even a 5 to 10 percent bias rate can substantially distort a leaderboard.
The degree of bias varies by model family and size. Smaller models tend to show stronger and less consistent position biases because they have less reliable reasoning capacity: they rely more on shallow heuristics like "the first thing mentioned was more relevant" rather than careful comparison. Larger, instruction-tuned models show reduced but non-zero position effects. Models with longer context windows may exhibit different patterns because the distance between the two responses grows as prompt length increases, changing which response is more salient at decision time.
Measuring position bias requires running evaluations in both orders. For each question, you submit the pair as (Response A, Response B) and also as (Response B, Response A), then compare the verdicts. The key metrics are as follows.
Swap consistency rate: The fraction of comparisons where the judge gives a consistent verdict regardless of order. Formally, for a set of comparison pairs, let be the verdict when pair has response first, and let be the verdict when is first (with labels relabeled to match original A and B identity). The swap consistency rate is:
where:
- : the total number of comparison pairs evaluated
- : the verdict for pair when response appears first in the prompt
- : the verdict for pair when the same responses are swapped, so appears first (relabeled to preserve original A/B identity)
- : the indicator function, returning 1 if the condition inside is true and 0 otherwise
A perfectly consistent judge scores SC = 1.0. Values around 0.7 to 0.8 are commonly observed in practice, meaning 20 to 30 percent of judgments flip based solely on presentation order.
Position preference rate: Beyond consistency, you want to know which direction the bias flows. The position-A preference rate counts how often the judge favors whichever response appears in position A, regardless of that response's actual quality:
where the numerator counts every comparison in which the judge selected the response occupying the first (A) slot as the winner, and is the total number of comparisons evaluated.
In an unbiased judge, . Systematic deviations reveal directional position bias. If , the judge prefers the first response 65 percent of the time, a significant structural advantage that inflates the performance of whatever model happens to be placed in position A.
Inconsistency decomposition: When a judge is inconsistent across orderings, you can decompose the inconsistencies by direction. Let be the count of cases where the judge preferred position A in both orderings (i.e., preferred the first-presented response each time, regardless of which actual response that was). Let be the reverse, where the judge consistently preferred the last-presented response. The first-position bias rate among inconsistent pairs is:
where:
- : count of inconsistent pairs where the judge picked position A in both orderings, always choosing the first-presented response
- : count of inconsistent pairs where the judge picked position B in both orderings, always choosing the last-presented response
- : the total count of inconsistent pairs, the only pairs that reveal directional positional preference
A value above 0.5 indicates first-position preference. A value below 0.5 indicates last-position preference. A value near 0.5 means that when the judge is inconsistent, the inconsistency is bidirectional and not systematically directional. This last case represents a less dangerous form of position sensitivity, closer to stochastic noise than systematic bias.
Verbosity Bias: Longer is Not Better
Verbosity bias is distinct from position bias but often correlated with it. LLM judges tend to score longer responses higher, even when the additional length does not convey additional information or quality. This has been observed across multiple judge models and evaluation tasks, and it creates a dangerous incentive: model developers can game LLM evaluation by making their models verbose rather than accurate or helpful.
The mechanism behind verbosity bias is not fully understood, but several factors contribute. First, length may be a proxy for effort in the judge's training data: in human-written text, longer responses often signal more thought and attention. A student who writes three pages of analysis is generally doing more work than one who writes three sentences, and human annotators who created preference data for RLHF may have absorbed this correlation. Second, longer responses tend to cover more ground, hitting topics the judge considers relevant even if some coverage is shallow or padded. If the judge uses a checklist-like internal evaluation of "did the response cover X, Y, and Z," a longer response is more likely to touch all three items even if it does so superficially. Third, formatting effects compound with length: longer responses often use more headers, bullet points, and structured presentation, which correlates with higher quality in judge training data.
Measuring verbosity bias requires controlling for quality while varying length. One approach is to create pairs of semantically equivalent responses at different lengths: take a high-quality short answer and expand it with relevant but redundant content, then ask the judge which is better. A biased judge consistently prefers the longer version. Another approach is to run a regression across many evaluation pairs. You regress judge scores on response quality (as measured by a human panel) and response length simultaneously. If length has a significant positive coefficient after controlling for quality, verbosity bias is present.
Formally, let be the judge score for response , be the human-rated quality score, and be the log-length (log is used to handle the wide range of possible response lengths). Verbosity bias exists when the regression:
yields with statistical significance, where:
- : the judge score for response (the quantity we want driven only by quality)
- : the human-annotated quality rating for response (the ground truth benchmark)
- : the log-length of response , capturing the length signal on a compressed scale
- : the intercept, representing the baseline score for a zero-quality, zero-length response
- : the coefficient on quality, which should be large and positive for a well-calibrated judge
- : the coefficient on length, which should be near zero for an unbiased judge; a positive value indicates verbosity bias
- : the residual error capturing unexplained variation
The ratio quantifies how much length affects scores relative to actual quality. A perfectly calibrated judge would have , meaning length contributes nothing independent of content. Research has found values as high as 0.4 in some judge models, meaning length contributes nearly as much as quality to the final verdict. A model that doubles its response length while keeping content quality constant receives nearly the same score boost as a model that substantially improves quality.
This ratio has a direct practical implication for model development. If you are fine-tuning a model and using an LLM judge with to evaluate your checkpoints, you can artificially inflate your evaluation scores by training your model to be more verbose. This kind of evaluation gaming, sometimes called "judge hacking," is a real concern in competitive model development and one reason that LLM judge calibration audits are worth the effort.
Sycophancy: Agreement as a Bias Signal
Sycophancy is the tendency of a model to agree with or validate whatever position appears to be endorsed in the prompt. It is one of the most studied failure modes in LLM alignment research, and it shows up in evaluation as a systematic bias toward outputs that match what the judge model expects or prefers to be true. In the context of LLM-as-Judge, sycophancy manifests in several distinct ways.
The most studied form is model self-preference: a judge may favor responses generated by the same model family as itself. If GPT-4 is your judge, it may systematically overrate GPT-4-generated outputs compared to responses from other model families, because the internal representations and stylistic patterns are more familiar. The stylistic preferences baked into GPT-4 through RLHF are the same preferences it was trained to reward, so outputs that match those preferences score higher regardless of independent quality. This creates a circular evaluation loop: the judge validates outputs that resemble its own training, and model developers who use that judge receive inflated scores precisely when they replicate the judge's stylistic blind spots.
A second form is prompt-induced sycophancy: the judge's verdict is influenced by cues in the evaluation prompt that signal a preferred answer. If you include a statement like "the user found Response A more helpful" in the prompt, many judge models will shift their verdict toward Response A, even when they were explicitly instructed to ignore user preference signals and evaluate objectively. Research has demonstrated this effect in controlled conditions using what are called "authority injection" experiments, where a fake expert endorsement is added to one response before presenting it to the judge. Well-designed judges should be immune to such cues, but current models are not reliably so.
A third form is tone sycophancy: judges systematically prefer confident, authoritative-sounding responses. A response that states "the answer is X" is often rated higher than one that states "the evidence suggests X, though some uncertainty remains," even when the latter is more epistemically accurate and scientifically appropriate. This reflects a real bias in training data: human annotators, when creating preference data for RLHF, tend to reward confident responses because confidence is associated with competence in everyday language. Models absorb this pattern and apply it in evaluation contexts, penalizing appropriate uncertainty expression as if it were a sign of weakness.
The self-preference form is particularly dangerous in practice because it is invisible in any single-judge evaluation. You cannot detect it by analyzing the judge's outputs alone. You can only detect it by comparing multiple judge models from different families on the same set of examples and looking for divergences that correlate with model authorship. If all your judges come from the same model family, self-preference inflates scores uniformly across all evaluations, and you lose the ability to cross-validate or catch the inflation.
Measuring sycophancy requires careful experimental design. For self-preference measurement, you need blind evaluations where response authorship is hidden from the judge, plus comparisons across multiple judge models from different families. If Judge A consistently rates Model A's outputs higher than Judge B does, while both rate Model B's outputs similarly, that divergence provides evidence of self-preference sycophancy in Judge A. For prompt-induced sycophancy, you can inject misleading cues into evaluation prompts (such as "the user preferred Response A") and measure whether the judge updates its verdict accordingly, compared to a control evaluation with no injection. The sycophancy rate is:
where the numerator counts how many judgments flipped toward the injected preference compared to control evaluations with no injection. Well-calibrated judges show sycophancy rates below 5 percent. Many current models exceed 15 to 25 percent in controlled experiments, meaning a substantial fraction of their verdicts can be reversed simply by telling them a user or an expert preferred the other response.
How the Three Biases Interact
Position bias, verbosity bias, and sycophancy do not operate independently in real evaluations. They interact and compound in ways that make their combined effect more severe than any individual bias suggests.
Consider a typical evaluation scenario: Response A is placed first in the prompt, Response A is longer than Response B, and Response A is written in the same style as the judge model. All three biases now point in the same direction, amplifying each other. The judge that would give Response A a 5 percent positional advantage, a 10 percent verbosity advantage, and a 5 percent self-preference advantage is not giving it a 20 percent combined advantage. The three effects interact through the model's internal scoring mechanism in a nonlinear way. The judge may assign a 65 to 70 percent win probability to Response A based on the combination, even if Response A is objectively the lower-quality answer.
The reverse also holds. When designing experiments to measure one bias, you need to control for the others. If you measure position bias using pairs where one response happens to be systematically longer or happens to match the judge's style, your position bias estimate will be confounded. Good measurement methodology requires carefully constructing pairs where quality, length, and authorship are controlled, varying only the ordering. This is why bias measurement frameworks typically include separate datasets designed to isolate each bias and a combined dataset to measure interaction effects.
Worked Example: Observing and Measuring Bias
Let's walk through a concrete scenario to see how these biases manifest in practice. Consider a question from a code explanation benchmark: "Explain the difference between a list and a tuple in Python." Two responses have been generated.
Response A (short, accurate): "Lists are mutable sequences that can be changed after creation (append, remove, etc.). Tuples are immutable sequences that cannot be modified. Use tuples for fixed collections of items, lists when you need to add or remove elements."
Response B (long, padded): "Great question! Python offers several sequence types, each with unique characteristics. Lists are one of Python's most versatile data structures. They are mutable, meaning you can add, remove, and modify elements after the list is created. You can use methods like append(), extend(), remove(), and pop() to modify a list. Tuples, on the other hand, are immutable. Once created, you cannot change their contents. They are often used for fixed collections of data, like database records or coordinate pairs. In summary, use lists when mutability is needed, tuples for fixed data. Both support indexing and slicing."
A human expert rating these on quality would likely rate them roughly equal. Both are accurate and cover the core distinction. Response B adds padding (the opening "Great question!", repetitive phrasing, excess detail that does not add clarity), but it is not inaccurate. A well-calibrated judge should rate them similarly or perhaps slightly prefer the concise Response A for its directness and signal-to-noise ratio.
Now we run the experiment. We ask the same judge model the same question in both orders.
Evaluation 1 (A first): The judge sees Response A first, then Response B. It rates Response B as better: "Response B provides a more complete explanation with practical method examples, which is helpful for understanding how to apply these sequence types."
Evaluation 2 (B first): The judge sees Response B first, then Response A. It again rates Response B as better: "Response B provides a thorough, well-structured explanation that covers multiple aspects of the comparison."
The verdict is consistent but for the wrong reasons. The judge preferred B in both orderings, but this is verbosity bias, not honest quality assessment. The longer response won regardless of order. If we ran this experiment across 100 similar pairs where quality and length point in opposite directions, we would find that length predicts the judge's winner better than human quality ratings in a meaningful fraction of cases.
Now let's modify the scenario to isolate position bias. Suppose we have two responses of identical length and structure, and human raters assign them equal scores. We run the swap experiment:
Evaluation 1 (A first): Judge picks A. "Response A directly addresses the key differences in a clear, organized manner."
Evaluation 2 (B first): Judge picks B (which is now in position A). "Response B directly addresses the key differences in a clear, organized manner."
The rationale given is nearly identical, but the winner flipped. The judge always picked the first response. This is pure position bias, completely independent of content. The rationale it generates is post-hoc justification for a position-driven decision, and the near-identical language in both verdicts reveals that the judge did not engage differently with the two responses in the two orderings.
This worked example illustrates an important diagnostic: track which response wins and the language the judge uses to justify its verdict. When swap-inconsistent judges produce nearly identical rationales regardless of which response won, the judge's reasoning is following the bias rather than generating it.
Code Implementation
Let's build a complete measurement framework for position and verbosity bias. We will create a toy evaluation dataset, run simulated pairwise comparisons with a biased mock judge, and compute the key metrics. In a real pipeline, the mock judge would be replaced by LLM API calls, but the measurement structure is identical.
First, let's set up imports and fix random seeds for reproducibility:
import random
import numpy as np
random.seed(42)
np.random.seed(42)We define data structures for our evaluation framework. The Response class tracks both text content and the true quality score, which in a real evaluation would come from human annotation. The judge never sees the quality score directly:
from dataclasses import dataclass, field
@dataclass
class Response:
"""A model response with metadata."""
text: str
quality: float # True quality score, 0-1 (unknown to judge)
length: int = field(init=False)
def __post_init__(self):
self.length = len(self.text.split())
@dataclass
class ComparisonPair:
"""A pair of responses with an associated question."""
question: str
response_a: Response
response_b: Response
true_winner: str # 'A' or 'B' based on true qualityNow let's create a toy dataset of comparison pairs. We deliberately include cases where longer responses are lower quality, which lets us test whether the judge incorrectly rewards length over accuracy:
from typing import List
def make_toy_dataset(n_pairs: int = 80) -> List[ComparisonPair]:
"""Create evaluation pairs with controlled quality and length."""
questions = [
"What is gradient descent?",
"Explain regularization.",
"What is overfitting?",
"How does backpropagation work?",
"What is a transformer?",
]
pairs = []
for i in range(n_pairs):
q = questions[i % len(questions)]
# Assign quality randomly, independent of length
q_a = round(random.uniform(0.3, 1.0), 2)
q_b = round(random.uniform(0.3, 1.0), 2)
# Assign lengths with some correlation to quality but not perfect
# ~30% of pairs: longer response is lower quality (to detect verbosity bias)
len_a = random.randint(20, 50)
len_b = random.randint(20, 50)
if i % 3 == 0:
# Make B longer but lower quality than A
if q_a > q_b:
len_b = random.randint(80, 150)
else:
len_a = random.randint(80, 150)
text_a = " ".join([f"word{j}" for j in range(len_a)])
text_b = " ".join([f"word{j}" for j in range(len_b)])
resp_a = Response(text=text_a, quality=q_a)
resp_b = Response(text=text_b, quality=q_b)
true_winner = "A" if q_a > q_b else "B"
pairs.append(ComparisonPair(q, resp_a, resp_b, true_winner))
return pairs
dataset = make_toy_dataset(n_pairs=80)Now we implement a mock judge with tunable biases. The judge combines the true quality signal with a position bias term (which overrides quality with some probability) and a verbosity bias term (which adds a length bonus to each response's effective score):
def biased_judge(
question: str,
response_a: Response,
response_b: Response,
position_bias_strength: float = 0.3,
verbosity_bias_strength: float = 0.25,
noise: float = 0.1,
) -> str:
"""
Simulate a biased judge that combines quality, position, and verbosity signals.
Returns 'A' or 'B' as the winner.
position_bias_strength: probability of flipping to first-position regardless of quality.
verbosity_bias_strength: extra score weight given to longer responses.
noise: random noise added to simulate stochastic judge behavior.
"""
# True quality scores
score_a = response_a.quality
score_b = response_b.quality
# Add verbosity bias: longer responses get a bonus proportional to their log-length share
len_ratio_a = np.log1p(response_a.length) / (
np.log1p(response_a.length) + np.log1p(response_b.length)
)
len_ratio_b = 1 - len_ratio_a
score_a += verbosity_bias_strength * len_ratio_a
score_b += verbosity_bias_strength * len_ratio_b
# Add random noise
score_a += random.gauss(0, noise)
score_b += random.gauss(0, noise)
# Quality-based verdict
quality_verdict = "A" if score_a > score_b else "B"
# Apply position bias: with probability position_bias_strength, override to first position
if random.random() < position_bias_strength:
return "A" # A is always first in this call's perspective
return quality_verdictNow we run the swap experiment. For each pair, we evaluate in both orderings and compare verdicts. The relabeling step in the swapped ordering is important: when the judge returns "A" in the swapped call, "A" corresponds to the original Response B, so we translate the verdict back to original labels:
def run_swap_experiment(
pairs: List[ComparisonPair], judge_fn, **judge_kwargs
) -> dict:
"""
Run pairwise evaluation in both orders and compute bias metrics.
Returns a dict with all key metrics.
"""
verdicts_original = [] # verdict when A is first
verdicts_swapped = [] # verdict when B is first (relabeled to original A/B)
true_winners = []
for pair in pairs:
# Original order: A first
v1 = judge_fn(
pair.question, pair.response_a, pair.response_b, **judge_kwargs
)
# Swapped order: B first; judge returns 'A' or 'B' from its perspective,
# where the new 'A' is the original B
v2_raw = judge_fn(
pair.question, pair.response_b, pair.response_a, **judge_kwargs
)
# Re-label: if judge picked 'A' in swapped order, that means original B won
v2 = "B" if v2_raw == "A" else "A"
verdicts_original.append(v1)
verdicts_swapped.append(v2)
true_winners.append(pair.true_winner)
# Swap consistency: fraction where both orderings agree
consistent = [
v1 == v2 for v1, v2 in zip(verdicts_original, verdicts_swapped)
]
swap_consistency = np.mean(consistent)
# Position preference rate: how often position-A wins in original order
ppr_a = np.mean([v == "A" for v in verdicts_original])
# Accuracy against true winner (using original order verdict)
accuracy = np.mean(
[v == tw for v, tw in zip(verdicts_original, true_winners)]
)
# First-position bias among inconsistent pairs
inconsistent_mask = [not c for c in consistent]
inconsistent_original = [
v for v, m in zip(verdicts_original, inconsistent_mask) if m
]
first_position_bias = (
np.mean([v == "A" for v in inconsistent_original])
if inconsistent_original
else 0.5
)
return {
"swap_consistency": swap_consistency,
"position_preference_rate_A": ppr_a,
"accuracy_vs_true": accuracy,
"first_position_bias_rate": first_position_bias,
"n_inconsistent": sum(inconsistent_mask),
"n_pairs": len(pairs),
}
results_biased = run_swap_experiment(
dataset,
biased_judge,
position_bias_strength=0.3,
verbosity_bias_strength=0.25,
noise=0.1,
)
results_ideal = run_swap_experiment(
dataset,
biased_judge,
position_bias_strength=0.0,
verbosity_bias_strength=0.0,
noise=0.05,
)=== Biased Judge === swap_consistency: 0.550 position_preference_rate_A: 0.688 accuracy_vs_true: 0.762 first_position_bias_rate: 0.806 n_inconsistent: 36 n_pairs: 80 === Near-Ideal Judge (no position/verbosity bias) === swap_consistency: 0.825 position_preference_rate_A: 0.613 accuracy_vs_true: 0.887 first_position_bias_rate: 0.429 n_inconsistent: 14 n_pairs: 80
The biased judge shows swap consistency well below 1.0, confirming that verdict order affects outcomes. The ideal judge achieves near-perfect consistency, with accuracy tracking true quality closely. The first-position bias rate above 0.5 in the biased condition confirms the directional preference for first-presented responses.
Now let's measure verbosity bias by computing the correlation between length and judge score, separately from quality:
def measure_verbosity_bias(
pairs: List[ComparisonPair], judge_fn, **judge_kwargs
) -> dict:
"""
Measure how strongly response length predicts judge verdicts
after controlling for true quality.
"""
quality_advantages = [] # quality_a - quality_b
length_advantages = [] # log(len_a) - log(len_b)
judge_chose_a = []
for pair in pairs:
q_diff = pair.response_a.quality - pair.response_b.quality
l_diff = np.log1p(pair.response_a.length) - np.log1p(
pair.response_b.length
)
verdict = judge_fn(
pair.question, pair.response_a, pair.response_b, **judge_kwargs
)
chose_a = 1 if verdict == "A" else 0
quality_advantages.append(q_diff)
length_advantages.append(l_diff)
judge_chose_a.append(chose_a)
quality_advantages = np.array(quality_advantages)
length_advantages = np.array(length_advantages)
judge_chose_a = np.array(judge_chose_a)
# Compute correlation of each predictor with judge choice
corr_quality = np.corrcoef(quality_advantages, judge_chose_a)[0, 1]
corr_length = np.corrcoef(length_advantages, judge_chose_a)[0, 1]
# Compute accuracy of quality-based predictor alone
quality_predicted = (quality_advantages > 0).astype(int)
quality_accuracy = np.mean(quality_predicted == judge_chose_a)
# Verbosity bias: fraction of pairs where length and quality disagree on winner,
# and the judge picks the longer (lower-quality) response
disagree_mask = (quality_advantages > 0.05) & (length_advantages < 0)
if disagree_mask.sum() > 0:
verbosity_bias_rate = np.mean(judge_chose_a[disagree_mask] == 0)
else:
verbosity_bias_rate = 0.0
return {
"corr_quality_with_choice": corr_quality,
"corr_length_with_choice": corr_length,
"quality_prediction_accuracy": quality_accuracy,
"verbosity_bias_rate": verbosity_bias_rate,
"n_quality_length_conflict": int(disagree_mask.sum()),
}
verbosity_biased = measure_verbosity_bias(
dataset,
biased_judge,
position_bias_strength=0.0,
verbosity_bias_strength=0.4,
noise=0.05,
)
verbosity_unbiased = measure_verbosity_bias(
dataset,
biased_judge,
position_bias_strength=0.0,
verbosity_bias_strength=0.0,
noise=0.05,
)=== Verbosity-Biased Judge === corr_quality_with_choice: 0.690 corr_length_with_choice: -0.328 quality_prediction_accuracy: 0.887 verbosity_bias_rate: 0.103 n_quality_length_conflict: 29 === Unbiased Judge === corr_quality_with_choice: 0.721 corr_length_with_choice: -0.425 quality_prediction_accuracy: 0.838 verbosity_bias_rate: 0.034 n_quality_length_conflict: 29
In the verbosity-biased condition, the correlation between length advantage and judge choice is clearly positive, even in cases where the longer response is the lower-quality one. The verbosity bias rate quantifies this directly: when quality and length point in opposite directions, how often does the judge pick length over quality?
Bias Mitigation Strategies
The most widely used mitigation for position bias is answer swapping: run every comparison twice, once in each order, and aggregate the results. If the judge produces consistent verdicts across both orderings, that verdict is reliable. If the verdicts conflict, you have several options for handling the disagreement:
- Declare a tie: When verdicts differ across orderings, record no winner. This is conservative but avoids position-contaminated decisions.
- Average scores: If you use continuous scores rather than binary preferences, average the scores from both orderings. This is the most principled approach when scoring is available.
- Use structured output: Ask the judge to output an independent score for each response (such as 1 through 10) rather than picking a winner by comparison. This avoids the comparison framing that position bias exploits, though it introduces its own calibration challenges.
- Confidence-weighted aggregation: Weight the verdict from each ordering by the judge's expressed confidence. If the judge is highly confident in both orderings and they agree, use the shared verdict. If it is highly confident in one ordering but uncertain in the other, use the confident verdict.
For verbosity bias, mitigation is harder because it is embedded in the model's learned representations rather than just the prompt framing. Several approaches are used in practice:
- Normalized scoring: Ask the judge to score responses per unit of information rather than in absolute terms, or explicitly instruct it to penalize unnecessary length. This helps but requires the model to understand "unnecessary length," which is itself a judgment call.
- Length-controlled pairs: When comparing responses, truncate all responses to the same length before evaluation. This is extreme but effective for isolating quality from length effects.
- Post-hoc length penalty: After collecting judge scores, run the verbosity regression to estimate and subtract the estimated length contribution from each score. This requires a calibration dataset with human quality annotations.
- Chain-of-thought rubric evaluation: Ask the judge to evaluate each response against explicit criteria before deciding. If each criterion is independently scored, verbosity has fewer opportunities to inflate the aggregate.
For sycophancy, mitigation is the most technically challenging because it requires either using judge models maximally dissimilar from the models being evaluated, or using multiple judges and looking for consensus. Neither approach is fully reliable, but both reduce the problem. When using multiple judges, a disagreement between judges from different model families is itself informative: it suggests that the verdict may be driven by family-specific preferences rather than objective quality.
def mitigated_judge(
question: str,
response_a: Response,
response_b: Response,
judge_fn,
aggregation: str = "consistency",
**judge_kwargs,
) -> str:
"""
Run swapped evaluation and aggregate verdicts to reduce position bias.
aggregation options:
'consistency': return verdict only if both orderings agree, else 'TIE'
'majority': use first verdict when tie-broken by original order
"""
# Original order
v1 = judge_fn(question, response_a, response_b, **judge_kwargs)
# Swapped order (relabeled back to A/B)
v2_raw = judge_fn(question, response_b, response_a, **judge_kwargs)
v2 = "B" if v2_raw == "A" else "A"
if v1 == v2:
return v1
if aggregation == "consistency":
return "TIE"
else:
return v1 # default to original order on disagreement
def evaluate_with_mitigation(
pairs: List[ComparisonPair], judge_fn, **judge_kwargs
) -> dict:
"""Compare mitigated vs unmitigated evaluation accuracy."""
unmitigated_correct = 0
mitigated_correct = 0
mitigated_ties = 0
for pair in pairs:
# Unmitigated
v_unmitigated = judge_fn(
pair.question, pair.response_a, pair.response_b, **judge_kwargs
)
if v_unmitigated == pair.true_winner:
unmitigated_correct += 1
# Mitigated
v_mitigated = mitigated_judge(
pair.question,
pair.response_a,
pair.response_b,
judge_fn,
**judge_kwargs,
)
if v_mitigated == pair.true_winner:
mitigated_correct += 1
if v_mitigated == "TIE":
mitigated_ties += 1
n = len(pairs)
return {
"unmitigated_accuracy": unmitigated_correct / n,
"mitigated_accuracy": mitigated_correct / (n - mitigated_ties)
if (n - mitigated_ties) > 0
else 0,
"tie_rate": mitigated_ties / n,
"n_pairs": n,
}
mitigation_results = evaluate_with_mitigation(
dataset,
biased_judge,
position_bias_strength=0.25,
verbosity_bias_strength=0.15,
noise=0.1,
)=== Mitigation Results === unmitigated_accuracy: 0.800 mitigated_accuracy: 0.878 tie_rate: 0.487 n_pairs: 80 Interpretation: Accuracy gain from mitigation: 0.078 Cost: 48.8% of pairs declared ties (require human review)
Answer swapping improves accuracy among pairs where the judge can reach a consistent verdict. The cost is that a fraction of comparisons result in ties and require fallback handling, typically human annotation or a second judge from a different model family.
Key Parameters
The key parameters for bias measurement and mitigation are:
- position_bias_strength: The probability that a judge overrides quality with a position-based decision. Measured as 1 minus the swap consistency rate, adjusted for the direction of inconsistencies.
- verbosity_bias_strength: The regression coefficient on log-length in a quality-controlled model, quantifying how much length contributes to judge scores beyond actual content quality.
- swap_consistency: The fraction of comparison pairs where the judge produces the same verdict in both orderings. This is the primary diagnostic metric for position bias.
- first_position_bias_rate: Among inconsistent comparisons, the fraction where the judge chose the first-presented response. Values near 0.5 indicate undirected noise; values far from 0.5 indicate systematic positional preference.
Visualizations
The four visualizations below illuminate the bias patterns we have measured: how position and verbosity bias interact to reduce swap consistency, how bias degrades judge accuracy, how verbosity bias distorts individual comparisons, and how mitigation effectiveness changes as bias strength grows.
First, a heatmap showing swap consistency across different combinations of position bias and verbosity bias:

Second, a bar chart comparing accuracy and position preference rate across judge configurations, showing how each bias degrades reliability:

Third, a scatter plot showing the relationship between response length difference and quality difference, colored by judge verdict. Points in the quadrant where one response is longer but the other is higher quality reveal verbosity bias directly:

Fourth, a line chart showing how swap consistency and accuracy evolve as position bias strength increases, comparing unmitigated evaluation with swap-mitigated evaluation:

Limitations and Impact
When Mitigation Fails
Answer swapping is the most practical and widely adopted mitigation for position bias, but it is not a complete solution. The most basic limitation is that swapping doubles the cost of every evaluation. For high-stakes comparisons evaluated infrequently, doubling cost is acceptable. For routine quality monitoring at scale, running thousands of daily comparisons at twice the cost and latency may be prohibitive. This creates a common inconsistency in production systems: rigorous swap-mitigated evaluation for model selection decisions, and unmitigated evaluation for ongoing monitoring. The two pipelines may measure different things.
More subtly, swapping does not eliminate bias; it surfaces inconsistency and removes affected comparisons from the analysis. If a judge shows 30 percent position bias, about 30 percent of comparisons become unresolvable ties, which must be handled by human annotation, a second judge, or accepted as measurement noise. The underlying quality ranking becomes less sharp, not more accurate. There is also a selection effect: the pairs that survive as consistent verdicts may be systematically different from the pairs that produce ties. If position bias is correlated with response difficulty (the judge is more susceptible to positional cues when the responses are close in quality), then the consistent subset over-represents easy comparisons, and the swap-mitigated leaderboard may be a better measure of large-gap discrimination than close-comparison reliability.
Verbosity bias is harder to mitigate than position bias because it is embedded in both the model's learned representations and the prompt framing. Length normalization helps but introduces a different distortion: in practice, appropriate response length varies by question, and forcing all responses to the same length can destroy meaningful signals. A complete answer to a multi-part question is better than a terse one, and length normalization cannot reliably distinguish helpful completeness from padding. The post-hoc regression adjustment requires a calibration dataset with human quality annotations, which is exactly the resource that LLM judges were meant to replace. If you have enough high-quality human annotations to calibrate the verbosity regression, you arguably have enough annotations to run the evaluation directly.
Sycophancy mitigation is the most technically challenging of the three. Self-preference requires either using judge models that are maximally different from the models being evaluated, or using multiple judges and looking for cross-family consensus. Neither approach is fully reliable. Multi-judge consensus can catch obvious cases where one judge strongly disagrees with others, but if all available judges share training lineage (for example, all are fine-tuned from the same base model family), self-preference may be systematic across the entire pool of judges. You cannot detect this within the pool; you would need an entirely different evaluation methodology to expose it.
Remaining Biases After Mitigation
Even after applying standard mitigation strategies, LLM judges retain several residual biases that are difficult to remove without basic changes to how judge models are trained.
Formatting preference compounds with verbosity bias. Responses using markdown headers, bullet points, numbered lists, and code blocks are rated higher by many LLM judges, even when plain-prose responses convey the same information more efficiently. This is not purely a length effect. It reflects structural preferences learned from human feedback data where formatted responses were rewarded, perhaps because formatted text is easier for human annotators to quickly parse and grade. A well-calibrated judge should evaluate content independent of presentation format, but current judges often do not.
Confidence tone bias causes judges to prefer assertive, confident-sounding responses over tentative ones, even when the uncertain response is more epistemically accurate. A response stating "the answer is X" may outscore a response stating "the evidence suggests X, though some uncertainty remains," even if the latter is more scientifically honest. This reflects a documented bias in human preference data: annotators tend to reward confident responses because confidence is associated with competence in everyday language. Models absorb this and apply it in evaluation contexts, inadvertently penalizing calibrated uncertainty expression.
Cultural and stylistic preferences mean that judges may systematically favor responses written in a particular voice, formality level, or rhetorical style that matches their training data distribution. Responses from non-native English speakers, or from technical domains with specialized writing conventions, may receive systematically lower scores not because of quality deficits but because of style mismatch with the judge's implicit preferences. This is particularly concerning for multilingual evaluation, where a judge trained predominantly on English-language feedback data may apply English stylistic norms to responses in other languages.
Recency bias in multi-turn evaluation affects judge models when evaluating the quality of full conversations rather than single-turn responses. Later turns in a long conversation may be weighted more heavily simply because they are closer to the end of the context window, independent of their actual contribution to conversation quality. A model that starts well but deteriorates in later turns may receive the same score as a model that improves over the course of a conversation, because the judge's attention is drawn toward the end regardless of the trajectory.
These residual biases matter because they distort the incentives for model development. If you optimize a model against LLM judge scores, you are partly optimizing for these biases, rather than quality improvement alone. Models that learn to write confident-sounding, well-formatted, moderately verbose responses in the judge's preferred style will score well even if their factual accuracy or reasoning depth has not improved. This dynamic, sometimes called "Goodhart's Law" in the evaluation context, is a basic tension in automated evaluation: the measure becomes the target, and the target is no longer the thing you cared about.
Practical Guidance for Production Pipelines
The practical implication of all these limitations is that LLM judges should never be the sole quality signal in a production evaluation system. They are most reliable when used within a system that provides multiple layers of validation and cross-checking.
Concretely, a reliable production evaluation pipeline typically combines LLM judges with human evaluation panels for periodic calibration. Every few weeks, you run the same evaluation questions through both the LLM judge and a small human annotation panel, and you measure how well the judge's verdicts correlate with human preferences. If the correlation degrades over time, it may indicate that the judge model has drifted, that the distribution of production responses has shifted, or that biases are growing in the judge.
Using multiple judges from different model families reduces self-preference bias and provides natural variance estimates. When two judges agree, the verdict is reliable. When they disagree, the disagreement itself is informative: it flags the comparison for human review and prevents a single judge's idiosyncratic preferences from determining outcomes at scale.
Decomposing evaluation into independent dimensions, rather than a single overall quality score, reduces the surface area for biases to operate. Separate scores for factual accuracy, clarity, completeness, and appropriate length are harder to game with padding or confident tone alone. Each dimension requires a different form of bias to inflate, and the aggregate across dimensions is more reliable than any single overall score.
Finally, running regular adversarial probes to check for bias drift is worth the investment as judge models are updated. A judge that showed 5 percent position bias after its initial deployment may show 15 percent bias six months later if its model weights are updated, its prompt template changes, or the distribution of evaluation tasks shifts. Adversarial probes, which are pairs specifically designed to expose positional, verbosity, and sycophancy biases, catch these regressions before they silently corrupt weeks of production evaluation data.
Summary
Position bias, verbosity bias, and sycophancy are not edge cases in LLM evaluation. They are systematic properties of current judge models that affect every evaluation pipeline at some level, and they require active measurement and mitigation rather than passive acceptance.
The key takeaways from this chapter are:
- Position bias causes LLM judges to prefer responses in particular ordinal positions, independent of quality. Swap consistency is the primary diagnostic metric. Answer swapping is the standard mitigation, though it doubles evaluation cost and converts position-sensitive comparisons to ties.
- Verbosity bias causes judges to prefer longer responses, creating incentives for padding over precision. It can be measured via quality-controlled regression and partially mitigated by length normalization or post-hoc regression adjustment.
- Sycophancy causes judges to agree with perceived consensus, including stylistic self-agreement with outputs from the same model family. Multi-judge consensus across different model families and cross-family comparisons reduce but do not eliminate this bias.
- Mitigation is partial, not complete. Swapping doubles evaluation cost, verbosity normalization distorts legitimate length signals, and no current mitigation fully removes sycophancy without basic training changes.
- Biases compound. A judge exhibiting all three biases simultaneously is substantially less reliable than the individual bias rates suggest, because the three effects can reinforce each other when they happen to point in the same direction for a given comparison.
- LLM judges are best used within systems, not as standalone arbiters. Calibration against human annotations, multi-judge consensus, per-dimension scoring, and adversarial bias probes together make automated evaluation reliable enough for production use.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about position bias in LLM judges.
Position Bias in LLM Judges Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!