Bradley-Terry Model: Converting Preferences to Rankings

Michael BrenndoerferDecember 23, 202563 min read

Part of Language AI Handbook

Explains how the Bradley-Terry model converts pairwise preferences into consistent rankings. Foundation for reward modeling in RLHF and Elo rating systems.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Bradley-Terry Model: Converting Preferences to Rankings

In the previous chapter, we explored how human preference data is collected through pairwise comparisons: annotators see two model responses and indicate which one they prefer. But raw preference counts aren't directly useful for training. If response A was preferred over response B in 73% of comparisons, and response B was preferred over response C in 68% of comparisons, how much better is A than C? How do we convert these pairwise judgments into consistent quality scores that can guide model optimization? The answer requires a statistical model that can synthesize noisy, local comparisons into a coherent global ordering.

The Bradley-Terry model, developed by Ralph Bradley and Milton Terry in 1952, provides an elegant statistical framework for exactly this problem. Originally designed for ranking chess players and sports teams, it assumes each item has a latent "strength" parameter, and the probability of one item beating another depends only on the ratio of their strengths. This simple assumption leads to a powerful model that converts noisy pairwise comparisons into consistent, transitive rankings. Think of it as a hidden scoreboard: each competitor carries an invisible numeric quality score, and every observed comparison is a noisy measurement of the true gap between those scores.

The model's power comes from its probabilistic nature. Rather than asserting that the better item always wins, the Bradley-Terry model acknowledges that comparisons are inherently noisy. Two responses might both be reasonable answers to a question, and different annotators might prefer different ones depending on their background, values, or attention level on any given day. By modeling this uncertainty explicitly, the model can still recover a consistent underlying ranking even when individual comparisons contain apparent contradictions. The noise is not a nuisance to be eliminated; it is a basic part of human preference that the model embraces.

For LLM alignment, the Bradley-Terry model is foundational. As we'll see in the upcoming chapter on Reward Modeling, the reward model training objective is directly derived from the Bradley-Terry likelihood. Understanding this model deeply will clarify why reward models are trained the way they are, and why later techniques like Direct Preference Optimization (DPO) can eliminate the need for explicit reward modeling altogether. The Bradley-Terry model provides the statistical spine that connects raw human annotations to gradient updates in a neural network.

You might wonder why a model from 1952 would be relevant to modern language AI. The answer is that the basic problem it solves, turning pairwise comparisons into a global ranking, has not changed. What has changed is the scale and the nature of the items being compared. Instead of ranking eight chess players, we need to rank responses generated by language models across billions of possible prompts. The Bradley-Terry framework scales to this challenge because its parameters can be replaced by a neural network, preserving the mathematical structure while gaining the ability to generalize across inputs. This connection between a mid-century statistical model and 21st-century neural network training is one of the most elegant links in the history of machine learning.

Before diving into the mathematics, it helps to build strong intuition about what the model is doing at a high level. We start with a collection of items (model responses, chess players, restaurant dishes) and a set of pairwise comparisons where one item is declared the winner. Our goal is to assign each item a quality score such that these scores best explain the observed comparison outcomes. The Bradley-Terry model says: assume each item has a latent strength, and the probability of winning a comparison is a simple function of the two items' strengths. We then find the strength assignments that maximize the probability of having observed exactly the data we did. This is maximum likelihood estimation, and it produces the most statistically efficient ranking given our observations.

Historical Context

The Bradley-Terry model was introduced in the 1952 paper "Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons" by Ralph A. Bradley and Milton E. Terry, published in Biometrika. The core probabilistic formulation, however, has independent origins. Statistician Ernst Zermelo described an essentially equivalent model in 1929 for ranking chess tournaments, and psychologist Louis L. Thurstone proposed a closely related model called the "law of comparative judgment" in 1927. These parallel developments reflect how basic the pairwise comparison problem is across fields. The model became widely known in applied statistics through Bradley and Terry's systematic treatment, and gained further attention when Arpad Elo independently derived it (with a different scaling convention) for the chess rating system that now bears his name. Today the framework is central to preference learning, online ranking systems, and reinforcement learning from human feedback in large language model alignment.

The Pairwise Comparison Framework

Consider a setting where we have nn items (in our case, model responses) that we want to rank based on human preferences. We don't observe quality scores directly; instead, we observe the outcomes of pairwise comparisons. When comparing items ii and jj, a human annotator chooses one as better. This setup reflects how humans naturally make judgments: rather than assigning absolute quality scores on some numerical scale, people find it far easier to simply say which of two options they prefer. The challenge lies in synthesizing these individual binary decisions into a coherent global ranking. Think of this like reading a book of movie reviews where every review says only "I liked this film better than that one," and you must reconstruct a ranked list of every film ever made.

The core difficulty is that pairwise comparisons give us only local, relative information. Comparison A-vs-B tells us something about A relative to B, but nothing about where either item falls on an absolute scale. Comparison B-vs-C tells us something about B relative to C. If we want to know how A compares to C without having directly compared them, we need a model that can chain these local relationships into global conclusions. This is the inference problem that Bradley-Terry solves: it uses the entire structure of the comparison graph to find the parameter values that best explain all observed outcomes simultaneously.

Human preferences are also inherently noisy. The same two model responses shown to different annotators might produce different preference outcomes, because different annotators bring different expertise, different values, and different interpretations of the comparison criteria. Even the same annotator might decide differently on two separate occasions, especially for close pairs where neither response is clearly superior. A good preference model must account for this noise rather than pretending that comparisons are deterministic. The Bradley-Terry model handles noise by treating comparisons as Bernoulli trials, where the strength parameters determine the bias of the coin but cannot guarantee any particular outcome.

Pairwise Comparison

A pairwise comparison is a judgment where an evaluator observes two items and selects which one is superior according to some criterion. The comparison produces a binary outcome: either item ii wins or item jj wins.

The key insight of the Bradley-Terry model is that we can explain these observed preferences through latent strength parameters. Each item ii has an associated strength πi>0\pi_i > 0 that represents its underlying quality. We never observe these strengths directly. Instead, we must infer them from the pattern of wins and losses across many comparisons. Items with higher strength values are better, and when two items are compared, the one with higher strength is more likely to win, but not guaranteed to win, which accounts for noise in human judgments. This probabilistic view is needed: it acknowledges that human preferences contain randomness, whether from inconsistency, variation in attention, or differences of opinion among annotators.

The basic assumption is that the probability of item ii being preferred over item jj depends only on the ratio of their strengths. We write this as:

P(i≻j)=πiπi+πjP(i \succ j) = \frac{\pi_i}{\pi_i + \pi_j}

where:

  • P(i≻j)P(i \succ j): the probability that item ii is preferred over item jj
  • πi,πj\pi_i, \pi_j: the latent strength parameters of items ii and jj (must be positive)

This formula has an intuitive interpretation that makes it easy to reason about. Consider first the case where πi=πj\pi_i = \pi_j: substituting into the formula gives P(i≻j)=πi2πi=0.5P(i \succ j) = \frac{\pi_i}{2\pi_i} = 0.5. This confirms our intuition that equally strong items have equal probability of being chosen. Now consider a more extreme case: if πi=3πj\pi_i = 3\pi_j, then P(i≻j)=3πj3πj+πj=34=0.75P(i \succ j) = \frac{3\pi_j}{3\pi_j + \pi_j} = \frac{3}{4} = 0.75. An item three times as strong wins 75% of the time. Notice that the absolute magnitudes of the strengths do not matter, only their ratio. Whether we have strengths of 3 and 1, or 300 and 100, the predicted win probability remains the same. This ratio-based formulation captures the intuitive notion that what matters is relative quality, not absolute quality on some universal scale.

The formula also has a satisfying symmetry property: P(i≻j)+P(j≻i)=πiπi+πj+πjπi+πj=1P(i \succ j) + P(j \succ i) = \frac{\pi_i}{\pi_i + \pi_j} + \frac{\pi_j}{\pi_i + \pi_j} = 1. This means the probabilities of the two outcomes of a comparison sum to one, as they must for a valid probability model. The stronger the item relative to its opponent, the closer its win probability gets to 1, but it can never reach 1 (which would require infinite relative strength). Even the weakest item has some probability of winning a comparison against the strongest. This reflects the reality that human judgments always have some element of unpredictability.

Out[4]:
Visualization
Line chart showing Bradley-Terry win probability rising from 0 toward 1 as the strength ratio increases, with labeled red dots at ratios 1, 2, 3, and 5 showing specific win-probability values.
Win probability $P(i \succ j)$ as a function of the strength ratio $\pi_i / \pi_j$. The curve intersects 0.5 when strengths are equal (ratio = 1) and asymptotically approaches 1.0 as the ratio increases, which reflects the diminishing returns of strength superiority. The red dots mark specific reference ratios to aid interpretation.

The curve in this figure reveals an important practical implication: the Bradley-Terry model becomes less informative about exact strength differences when items are very far apart in quality. The difference between a strength ratio of 8 and a ratio of 10 translates to almost no change in win probability, since both are close to 1. This saturation effect means that if you only compare very strong items against very weak ones, you learn little about the precise relative ordering of the strong items among themselves, or the weak items among themselves. Good experimental design for Bradley-Terry estimation should pair items of similar strength to get the most informative comparisons.

The Log-Linear Form

While the ratio formulation is intuitive, a more convenient parameterization uses log-strengths. This transformation opens up powerful connections to logistic regression and makes optimization considerably easier. It also reveals the deep mathematical structure that links Bradley-Terry to the neural network training objectives we will use for reward models. Let βi=log⁡(πi)\beta_i = \log(\pi_i). Since strengths must be positive, the log-strength can be any real number, which is mathematically more convenient for optimization algorithms that operate over unconstrained real-valued spaces. We can now derive the equivalent expression in terms of these log-strengths:

P(i≻j)=πiπi+πj=eβieβi+eβj(substitute π=eβ)=eβi−βj1+eβi−βj(divide numerator and denominator by eβj)=σ(βi−βj)(definition of sigmoid)\begin{aligned} P(i \succ j) &= \frac{\pi_i}{\pi_i + \pi_j} \\ &= \frac{e^{\beta_i}}{e^{\beta_i} + e^{\beta_j}} && \text{(substitute } \pi = e^{\beta} \text{)} \\ &= \frac{e^{\beta_i - \beta_j}}{1 + e^{\beta_i - \beta_j}} && \text{(divide numerator and denominator by } e^{\beta_j} \text{)} \\ &= \sigma(\beta_i - \beta_j) && \text{(definition of sigmoid)} \end{aligned}

where:

  • βi,βj\beta_i, \beta_j: the log-strength parameters of items ii and jj
  • πi,πj\pi_i, \pi_j: the latent strength parameters (β=log⁡π\beta = \log \pi)
  • σ(x)\sigma(x): the sigmoid function 11+e−x\frac{1}{1 + e^{-x}}

The function σ(x)\sigma(x) is the standard sigmoid familiar from logistic regression and neural network activations (as covered in our chapter on activation functions). The key insight here is that the log transformation converts a ratio of two quantities into a difference of their logs. The win probability depends not on the absolute strengths but on how much larger one log-strength is than the other. A difference of βi−βj=1\beta_i - \beta_j = 1 means item ii is e≈2.72e \approx 2.72 times stronger than item jj, and this corresponds to a win probability of σ(1)≈0.73\sigma(1) \approx 0.73. A difference of 2 means item ii is e2≈7.39e^2 \approx 7.39 times stronger, corresponding to a win probability of σ(2)≈0.88\sigma(2) \approx 0.88.

The sigmoid function has several properties that make it ideal for representing probabilities. It maps any real number to the interval (0,1)(0, 1). This keeps valid probabilities. It is symmetric around zero, meaning σ(−x)=1−σ(x)\sigma(-x) = 1 - \sigma(x), which captures the natural symmetry of pairwise comparisons: the probability that jj beats ii equals one minus the probability that ii beats jj. The sigmoid is also smooth and differentiable everywhere, which enables gradient-based optimization. This differentiability is important when the parameters β\beta are outputs of a neural network: we can backpropagate gradients through the sigmoid to update the network weights.

Think of the log-strength parameterization as converting the strength problem into a distance problem. Each item is placed on a one-dimensional number line by its log-strength value. The probability that item ii beats item jj is a function of the signed distance βi−βj\beta_i - \beta_j: positive means ii is to the right of jj on the line and thus more likely to win, negative means jj is more likely to win. The sigmoid converts this signed distance into a probability. The further apart two items are on this line, the more predictable their comparison outcomes become. Items clustered near the same point on the line will have comparison outcomes close to 50/50.

Out[5]:
Visualization
S-shaped sigmoid curve plotted from -6 to 6, passing through 0.5 at x=0, with horizontal dashed reference lines at 0, 0.5, and 1.
Sigmoid function characteristics. The sigmoid maps real-valued inputs to the (0, 1) range, with maximum gradient at zero. The S-shape reflects the fact that moderate strength differences discriminate most clearly, while extreme differences produce nearly certain outcomes.
Two overlapping sigmoid curves, one for sigma(x) and one for sigma(-x), illustrating their mirror symmetry about the 0.5 probability level.
Symmetry property of the sigmoid in the Bradley-Terry context. The symmetry $\sigma(x) + \sigma(-x) = 1$ ensures that the probabilities of mutually exclusive comparison outcomes sum to unity, a basic consistency requirement.

This log-linear form reveals that the Bradley-Terry model is equivalent to logistic regression where the probability of ii beating jj depends on the difference in their log-strengths. The sigmoid function ensures probabilities stay between 0 and 1, with the midpoint at 0.5 when βi=βj\beta_i = \beta_j. This equivalence to logistic regression lets us use decades of research on efficient optimization algorithms, regularization techniques, and statistical theory about confidence intervals and hypothesis tests. Maximum likelihood estimation for Bradley-Terry is a well-studied convex optimization problem with guaranteed convergence to a global optimum.

Bradley-Terry Probability

The probability that item ii is preferred over item jj in the Bradley-Terry model is:

P(i≻j)=σ(βi−βj)=11+e−(βi−βj)P(i \succ j) = \sigma(\beta_i - \beta_j) = \frac{1}{1 + e^{-(\beta_i - \beta_j)}}

where:

  • βi,βj\beta_i, \beta_j: the log-strength parameters of the two items
  • σ(⋅)\sigma(\cdot): the sigmoid function mapping the difference to a probability

Preference Strength and Transitivity

A important property of the Bradley-Terry model is that it produces transitive preferences. Transitivity means that if A is better than B, and B is better than C, then A must be better than C. If we know how often A beats B, and how often B beats C, the model automatically determines how often A beats C. This follows directly from the parameterization, and understanding why reveals something deep about the model's structure. Think of transitivity as the model insisting that quality lives on a one-dimensional line: if A is to the right of B, and B is to the right of C, then A must be to the right of C. There is no room for circular orderings in a linear arrangement.

The transitivity arises because each item has a single strength parameter that governs all of its comparisons. When we estimate βA\beta_A, βB\beta_B, and βC\beta_C, these values must simultaneously explain the A-versus-B comparisons, the B-versus-C comparisons, and the A-versus-C comparisons. The model cannot assign different "strengths" to A depending on which item it is facing. This constraint forces global consistency. Every comparison provides information that constrains all the parameters jointly, and the maximum likelihood estimates balance all this information to find the single assignment of positions on the number line that best explains everything we observed.

This global consistency is what makes the model so powerful for synthesizing large amounts of comparison data. With nn items, there are (n2)\binom{n}{2} possible pairwise matchups, but only n−1n-1 free parameters (since one parameter is the anchor). The model over-constrains the problem: we use many more observations than parameters, and the maximum likelihood estimate finds the parameter values that best reconcile all these constraints simultaneously. This over-determination is a feature, not a bug. It means that comparisons between pairs we have observed well can inform predictions about pairs we have observed rarely, through the chain of transitivity.

Consider three items with strengths βA=2\beta_A = 2, βB=1\beta_B = 1, and βC=0\beta_C = 0. The pairwise probabilities are:

P(A≻B)=σ(2−1)=σ(1)≈0.731P(B≻C)=σ(1−0)=σ(1)≈0.731P(A≻C)=σ(2−0)=σ(2)≈0.881\begin{aligned} P(A \succ B) &= \sigma(2 - 1) = \sigma(1) \approx 0.731 \\ P(B \succ C) &= \sigma(1 - 0) = \sigma(1) \approx 0.731 \\ P(A \succ C) &= \sigma(2 - 0) = \sigma(2) \approx 0.881 \end{aligned}

where:

  • βA,βB,βC\beta_A, \beta_B, \beta_C: the example strengths (2,1,02, 1, 0)
  • P(A≻B)P(A \succ B): the probability A beats B
  • σ(⋅)\sigma(\cdot): the sigmoid function applied to strength differences

Notice that even though A and B have the same preference probability over their respective "next" item (both beat their weaker neighbor about 73% of the time), A has a much higher probability of beating C than B does of beating C. The strengths compose properly: the gap between A and C is the sum of the gaps between A and B and between B and C. Specifically, βA−βC=(βA−βB)+(βB−βC)=1+1=2\beta_A - \beta_C = (\beta_A - \beta_B) + (\beta_B - \beta_C) = 1 + 1 = 2. This additive property of log-strengths is what makes the model so elegant. Strength differences accumulate predictably across chains of comparisons. The model implicitly answers the question "If A usually beats B and B usually beats C, how often should A beat C?" without needing us to collect A-versus-C comparisons directly.

Out[6]:
Visualization
Color-coded 3x3 heatmap showing Bradley-Terry win probabilities for items A, B, and C, with annotated probability values in each cell and a green-to-red color scale.
Pairwise probability matrix for three items with log-strengths $\beta_A=2$, $\beta_B=1$, and $\beta_C=0$. The probability of A beating C (0.88) exceeds that of A beating B (0.73), visually confirming how strength differences accumulate to produce transitive rankings. Each row shows the probability that the row item beats each column item.
Out[7]:
Visualization
Bar chart showing log-strength parameters for items A (2.0), B (1.0), and C (0.0), with annotated bidirectional arrows showing the unit differences between adjacent items and the total gap of 2 between A and C.
Log-strength parameters for items A, B, and C illustrating the additive composition property of the Bradley-Terry model. The total strength difference between A and C equals the sum of the intermediate gaps (1+1=2), which directly determines all pairwise win probabilities without requiring direct A-vs-C comparisons.

This transitivity is a feature, not a bug. It forces the model to find a consistent ranking that best explains all observed comparisons, even when individual comparisons contain noise or contradictions. If humans sometimes prefer B over A, the model can still conclude A is better overall if A wins more often than it loses. The model essentially averages over the noise to find the underlying signal. When we observe an intransitivity in the data (say, A beats B 70% of the time, B beats C 70% of the time, but C beats A 55% of the time), the model will not try to represent this cycle explicitly. Instead, it will find the one-dimensional ordering that minimizes the total surprise across all observed outcomes. This averaging behavior is highly desirable for LLM evaluation: no single human comparison should be treated as ground truth, but the aggregate pattern across thousands of comparisons should be trusted.

However, as we will discuss in the Limitations section, this built-in transitivity can also be a constraint when human preferences violate the one-dimensional ordering assumption. Different annotators may prefer different attributes of responses, and these multi-dimensional preferences can produce systematic intransitivities that a Bradley-Terry model cannot capture. For now, it is enough to note that transitivity is both what makes the model tractable and what limits its expressiveness.

The Bradley-Terry Likelihood

To estimate the strength parameters β=(β1,β2,…,βn)\beta = (\beta_1, \beta_2, \ldots, \beta_n) from observed comparisons, we use maximum likelihood estimation. The goal is to find the parameter values that make the observed data most probable. This is a principled statistical approach: among all possible strength assignments, we seek the one that best explains the observed comparisons. Think of it as asking: "If I had to design a world where these strength values are true, which world would most plausibly produce the specific comparison outcomes I observed?"

Maximum likelihood estimation has excellent statistical properties for this problem. The estimates are consistent, meaning they converge to the true parameters as we collect more data. They are asymptotically efficient, meaning they extract the maximum possible information from the data. And the optimization problem is convex under mild conditions, meaning gradient-based methods will always find the global optimum rather than a local one. These properties make Bradley-Terry both a theoretically sound and computationally convenient choice for preference learning.

Suppose we have a dataset of mm pairwise comparisons, where comparison kk involves items iki_k and jkj_k, and item wkw_k was chosen as the winner. The key assumption is that each comparison is independent, which is reasonable if different annotators make different comparisons, or if the same annotator's judgments don't depend on what they saw before. Under this independence assumption, the probability of the entire dataset is the product of the probabilities of each individual comparison outcome.

The likelihood of observing this data under the Bradley-Terry model is:

L(β)=∏k=1mP(wk wins comparison k)=∏k=1mσ(βwk−βℓk)L(\beta) = \prod_{k=1}^{m} P(w_k \text{ wins comparison } k) = \prod_{k=1}^{m} \sigma(\beta_{w_k} - \beta_{\ell_k})

where:

  • L(β)L(\beta): the likelihood of the observed dataset given parameters β\beta
  • mm: the total number of pairwise comparisons
  • wk,ℓkw_k, \ell_k: indices of the winning and losing items in the kk-th comparison
  • σ(⋅)\sigma(\cdot): the sigmoid function mapping the strength difference to a probability

The product form arises from independence: the probability of seeing all the comparisons we saw is the product of the probabilities of each individual comparison. Each term in the product is the Bradley-Terry probability that the actual winner wins, given the current strength parameters. The goal of maximum likelihood estimation is to find the β\beta values that make this product as large as possible. When the parameters correctly capture the true underlying strengths, the outcomes we observe will be the likely ones, making each factor close to 1 and the product large. When parameters are wrong, many observed outcomes will be surprises under the model, making the product small.

Taking the log-likelihood converts the product into a sum, which is more convenient both analytically and computationally. Logarithms are monotone transformations, so the β\beta that maximizes the log-likelihood is identical to the one that maximizes the likelihood itself:

log⁡L(β)=∑k=1mlog⁡σ(βwk−βℓk)(log of product is sum of logs)=∑k=1mlog⁡(11+e−(βwk−βℓk))(substitute sigmoid definition)=−∑k=1mlog⁡(1+e−(βwk−βℓk))(using log⁡(1/x)=−log⁡x)\begin{aligned} \log L(\beta) &= \sum_{k=1}^{m} \log \sigma(\beta_{w_k} - \beta_{\ell_k}) && \text{(log of product is sum of logs)} \\ &= \sum_{k=1}^{m} \log \left( \frac{1}{1 + e^{-(\beta_{w_k} - \beta_{\ell_k})}} \right) && \text{(substitute sigmoid definition)} \\ &= -\sum_{k=1}^{m} \log(1 + e^{-(\beta_{w_k} - \beta_{\ell_k})}) && \text{(using } \log(1/x) = -\log x \text{)} \end{aligned}

where:

  • log⁡L(β)\log L(\beta): the log-likelihood function to be maximized
  • mm: the total number of comparisons
  • wk,ℓkw_k, \ell_k: indices of the winning and losing items in comparison kk
  • βwk,βℓk\beta_{w_k}, \beta_{\ell_k}: log-strength parameters for the winner and loser
  • −log⁡(1+e−x)-\log(1 + e^{-x}): the log-sigmoid, equivalent to the negative softplus function

This is the negative of the binary cross-entropy loss where we predict that the winner should win. The connection to binary cross-entropy is important: it means that training a Bradley-Terry model is computationally identical to training a logistic regression classifier. The maximum likelihood estimates are found by maximizing this expression with respect to β\beta, typically using gradient-based optimization methods like gradient descent or L-BFGS. The objective function is concave in β\beta, which guarantees that gradient ascent will converge to the global maximum.

To understand the gradient intuitively, consider a single comparison where item ii beat item jj, and suppose our current estimate assigns them equal strength (βi=βj\beta_i = \beta_j). The model predicts a 50% win probability for item ii, but we observed that item ii won. The gradient will push βi\beta_i up and βj\beta_j down, increasing the predicted probability that ii beats jj to better match the observation. The gradient magnitude is largest when the model is most surprised (predicted probability far from 1 for a win or far from 0 for a loss) and smallest when the model correctly assigns high probability to the observed outcome. This self-correcting behavior is what makes gradient-based optimization work so well here.

The Identifiability Issue

There's a subtlety that arises when we try to estimate Bradley-Terry parameters: the model has a location invariance. If we add a constant cc to all parameters, the probabilities don't change at all:

σ((βi+c)−(βj+c))=σ(βi−βj)\sigma((\beta_i + c) - (\beta_j + c)) = \sigma(\beta_i - \beta_j)

where:

  • cc: an arbitrary constant added to all log-strength parameters
  • βi,βj\beta_i, \beta_j: the log-strength parameters
  • σ(⋅)\sigma(\cdot): the sigmoid function

This invariance occurs because the sigmoid only depends on differences between parameters, not on their absolute values. The constant cc cancels out in the subtraction. Mathematically, this means the parameters are not uniquely identifiable; there are infinitely many solutions that produce the same likelihood. If (β1,β2,β3)=(2,1,0)(\beta_1, \beta_2, \beta_3) = (2, 1, 0) is a solution, then so is (102,101,100)(102, 101, 100), and so is (−5,−6,−7)(-5, -6, -7). All these solutions make identical predictions. A gradient descent optimizer will wander along this flat direction forever without converging to a unique point.

To resolve this ambiguity, we typically fix one parameter (e.g., β1=0\beta_1 = 0) or add a constraint like ∑iβi=0\sum_i \beta_i = 0. This anchoring doesn't affect the relative rankings or prediction probabilities. It simply pins down the arbitrary location of the parameter scale so that optimization algorithms converge to a unique solution. The choice of anchor is arbitrary: we could anchor any item at any value, and the relative differences would remain the same. In practice, the conventional choice for LLM evaluation is to normalize so that the average rating across all models equals some baseline value, such as 1000 in the original Elo system or 0 in a centered parameterization.

The identifiability issue has a practical implication for how we interpret Bradley-Terry parameters: we should never interpret the absolute value of any single strength parameter, only the differences between parameters. A strength of β=3\beta = 3 means nothing by itself. What matters is that this item has a log-strength 2 units higher than an item with β=1\beta = 1, which translates to a win probability of σ(2)≈0.88\sigma(2) \approx 0.88 in head-to-head comparisons.

Connection to Elo Ratings

If the Bradley-Terry model sounds familiar, it should. The Elo rating system, used to rank chess players and now common in competitive gaming and language model evaluation, is a Bradley-Terry model in disguise. Understanding this connection illuminates both systems and explains why Elo ratings have become standard for comparing language models on public leaderboards.

The Elo system was designed for practical use in tournament settings where ratings need to be updated incrementally after each game rather than recomputed from scratch with every new observation. Rather than running a global optimization over all historical comparisons, Elo uses a simple online update rule: after each game, the winner gains points and the loser loses points, with the amount proportional to how surprising the outcome was given the current ratings. If the expected winner wins, little rating change occurs. If the underdog wins, a large rating transfer happens. This online update rule is an approximation to the maximum likelihood gradient, making Elo an efficient algorithm for the same underlying Bradley-Terry model.

Elo ratings are typically scaled so that a 400-point difference corresponds to a 10:1 expected score ratio. This scaling was chosen for historical and practical reasons: it produces ratings that are easy to interpret and compare. A rating difference of 400 points means the stronger player is expected to score about 10 points for every 1 point the weaker player scores over many games.

In standard Elo, the expected score of player AA against player BB is:

EA=11+10(RB−RA)/400E_A = \frac{1}{1 + 10^{(R_B - R_A)/400}}

where:

  • EAE_A: the expected score (win probability) of player AA
  • RA,RBR_A, R_B: the Elo ratings of players AA and BB
  • 400400: the scaling factor where a difference of 400 points implies a 10:1 odds ratio

The use of base-10 exponentials rather than natural exponentials is purely conventional: Arpad Elo chose this scaling when he developed the system in the 1960s because it produced convenient numbers for chess ratings. The mathematical structure is identical to Bradley-Terry; only the scaling differs.

We can rewrite this in terms of the sigmoid function by converting the base-10 exponential to base-ee:

EA=11+10(RB−RA)/400=11+eln⁡(10)⋅(RB−RA)/400(change of base: 10x=exln⁡10)=11+e−ln⁡10400(RA−RB)(rearrange exponent)=σ(ln⁡10400(RA−RB))(definition of sigmoid)\begin{aligned} E_A &= \frac{1}{1 + 10^{(R_B - R_A)/400}} \\ &= \frac{1}{1 + e^{\ln(10) \cdot (R_B - R_A)/400}} && \text{(change of base: } 10^x = e^{x \ln 10} \text{)} \\ &= \frac{1}{1 + e^{-\frac{\ln 10}{400} (R_A - R_B)}} && \text{(rearrange exponent)} \\ &= \sigma\left( \frac{\ln 10}{400} (R_A - R_B) \right) && \text{(definition of sigmoid)} \end{aligned}

where:

  • EAE_A: the expected score (win probability) for player AA
  • ln⁡10≈2.30\ln 10 \approx 2.30: the natural logarithm of 10
  • σ(⋅)\sigma(\cdot): the standard sigmoid function

This derivation demonstrates that the Elo rating is just a scaled version of the Bradley-Terry log-strength parameter. Specifically, if β\beta is the Bradley-Terry parameter, then the corresponding Elo rating is R=400ln⁡10β≈173.7βR = \frac{400}{\ln 10} \beta \approx 173.7 \beta. The mathematical structure is identical; only the units differ. Every Elo system is a Bradley-Terry model, and every Bradley-Terry model can be expressed as an Elo system with a particular scaling convention.

Out[8]:
Visualization
Scatter plot with a straight blue line showing the proportional conversion from Bradley-Terry log-strength beta to Elo rating, with five labeled reference points and an annotation displaying the scale factor 400/ln(10).
Linear relationship between Bradley-Terry log-strength parameters and Elo ratings. The scaling factor of approximately 173.7 maps small log-strength differences to the broader Elo scale, where a single unit increase in $\beta$ translates to roughly 174 Elo points. The reference points illustrate how the two coordinate systems relate to each other.

The connection matters because Elo-style ratings are now used to evaluate LLMs. Platforms like LMSYS Chatbot Arena use Elo ratings computed from human preference comparisons to rank models. When you vote for which model response you prefer, these votes become pairwise comparisons that update the models' Elo ratings. The Bradley-Terry framework provides the statistical foundation for these rankings. This keeps the ratings are consistent and interpretable. A model with a 200-point Elo advantage over another is predicted to win head-to-head comparisons about 76% of the time: a concrete, actionable interpretation that pure benchmark scores on fixed datasets cannot provide.

The Elo system also demonstrates an important practical consideration: rating systems must handle the problem of new entrants. When a new language model joins the arena, we have no comparison data for it, so we initialize it at some default rating. As it accumulates comparisons against known models, its rating converges toward the true value. How quickly this convergence happens depends on the K-factor parameter in the Elo update rule, which controls the magnitude of rating changes per comparison. The Bradley-Terry maximum likelihood framework does not have this initialization problem: it estimates all parameters jointly from all available comparisons, so new models with few comparisons get wide uncertainty intervals rather than noisy point estimates.

Worked Example

Let's work through a concrete numerical example to solidify these concepts and see how the mathematics plays out with real numbers from start to finish. This will make the abstract framework concrete before we move to the code implementation. Suppose we have three model responses (A, B, C) and the following comparison outcomes from human annotators:

Pairwise comparison outcomes for three model responses. Each row represents a comparison matchup, showing how many times each participant won across repeated comparisons.
ComparisonWinnerOccurrences
A vs BA7
A vs BB3
B vs CB8
B vs CC2
A vs CA9
A vs CC1

We have 30 total comparisons across three matchups. Looking at the raw data, A beats B 70% of the time, B beats C 80% of the time, and A beats C 90% of the time. These observed frequencies give a rough sense of the ranking (A is best, C is worst), but they are not the Bradley-Terry parameters. The parameters must be inferred by maximum likelihood, and they will differ from the simple observed frequencies because they must be globally consistent.

Step 1: Set the anchor. We set βC=0\beta_C = 0 as our anchor. This choice is arbitrary, but it resolves the identifiability issue and allows us to interpret the other parameters as "how much better than C" each item is. With this convention, βA\beta_A and βB\beta_B become the only free parameters.

Step 2: Write the log-likelihood. The log-likelihood function aggregates all our comparison data:

log⁡L(βA,βB)=7log⁡σ(βA−βB)+3log⁡σ(βB−βA)+8log⁡σ(βB−0)+2log⁡σ(0−βB)+9log⁡σ(βA−0)+1log⁡σ(0−βA)\begin{aligned} \log L(\beta_A, \beta_B) &= 7\log\sigma(\beta_A - \beta_B) + 3\log\sigma(\beta_B - \beta_A) \\ &\quad + 8\log\sigma(\beta_B - 0) + 2\log\sigma(0 - \beta_B) \\ &\quad + 9\log\sigma(\beta_A - 0) + 1\log\sigma(0 - \beta_A) \end{aligned}

where:

  • log⁡L\log L: the log-likelihood function aggregating all comparisons
  • 7,3,…7, 3, \dots: the observed counts (e.g., A beats B 7 times)
  • σ(⋅)\sigma(\cdot): the sigmoid function mapping strength differences to probabilities
  • βA,βB\beta_A, \beta_B: the unknown log-strength parameters
  • 00: the fixed anchor value for item C (βC=0\beta_C = 0)

Each term in this sum corresponds to a set of comparisons. The first term, 7log⁡σ(βA−βB)7\log\sigma(\beta_A - \beta_B), contributes to the likelihood every time A beats B, which happened 7 times. The second term accounts for the 3 times B beat A. Notice how the symmetry of preferences emerges: when A beats B, we use σ(βA−βB)\sigma(\beta_A - \beta_B), and when B beats A, we use σ(βB−βA)=1−σ(βA−βB)\sigma(\beta_B - \beta_A) = 1 - \sigma(\beta_A - \beta_B).

Step 3: Take derivatives. To find the maximum, we take partial derivatives and set them to zero. We rely on the derivative property of the log-sigmoid function: ddxlog⁡σ(x)=1−σ(x)\frac{d}{dx} \log \sigma(x) = 1 - \sigma(x). This elegant result comes from the chain rule and the fact that σ′(x)=σ(x)(1−σ(x))\sigma'(x) = \sigma(x)(1-\sigma(x)), giving ddxlog⁡σ(x)=σ′(x)σ(x)=1−σ(x)\frac{d}{dx}\log\sigma(x) = \frac{\sigma'(x)}{\sigma(x)} = 1 - \sigma(x).

Applying the chain rule to the terms involving βA\beta_A:

  1. When A wins (e.g., against B): The term is log⁡σ(βA−βB)\log \sigma(\beta_A - \beta_B). The derivative with respect to βA\beta_A is 1−σ(βA−βB)1 - \sigma(\beta_A - \beta_B).
  2. When A loses (e.g., against B): The term is log⁡σ(βB−βA)\log \sigma(\beta_B - \beta_A). The derivative with respect to βA\beta_A is (1−σ(βB−βA))⋅(−1)(1 - \sigma(\beta_B - \beta_A)) \cdot (-1), which simplifies to −σ(βA−βB)-\sigma(\beta_A - \beta_B).

Substituting the counts into the gradient equation for βA\beta_A:

∂log⁡L∂βA=7(1−σ(βA−βB))−3σ(βA−βB)+9(1−σ(βA))−1⋅σ(βA)\begin{aligned} \frac{\partial \log L}{\partial \beta_A} &= 7(1 - \sigma(\beta_A - \beta_B)) - 3\sigma(\beta_A - \beta_B) \\ &\quad + 9(1 - \sigma(\beta_A)) - 1\cdot\sigma(\beta_A) \end{aligned}

where:

  • ∂log⁡L∂βA\frac{\partial \log L}{\partial \beta_A}: the gradient of the log-likelihood with respect to βA\beta_A
  • βA,βB\beta_A, \beta_B: the log-strength parameters
  • σ(⋅)\sigma(\cdot): the sigmoid function
  • 1−σ(⋅)1 - \sigma(\cdot): the derivative of the log⁡σ(⋅)\log \sigma(\cdot) term for wins
  • −σ(⋅)-\sigma(\cdot): the derivative of the log⁡σ(⋅)\log \sigma(\cdot) term for losses

Using the property that 1−σ(x)=σ(−x)1 - \sigma(x) = \sigma(-x), we can write this more compactly as:

∂log⁡L∂βA=7σ(βB−βA)−3σ(βA−βB)+9σ(−βA)−σ(βA)\frac{\partial \log L}{\partial \beta_A} = 7\sigma(\beta_B - \beta_A) - 3\sigma(\beta_A - \beta_B) + 9\sigma(-\beta_A) - \sigma(\beta_A)

where:

  • σ(βB−βA)\sigma(\beta_B - \beta_A): the probability that B beats A
  • σ(−βA)\sigma(-\beta_A): the probability that C beats A (equivalent to 1−σ(βA)1 - \sigma(\beta_A))
  • βA,βB\beta_A, \beta_B: the log-strength parameters
  • σ(⋅)\sigma(\cdot): the sigmoid function

There is a beautiful interpretation of this gradient. The first term, 7σ(βB−βA)7\sigma(\beta_B - \beta_A), is 7 times the probability that B beats A under the current model. The second term, 3σ(βA−βB)3\sigma(\beta_A - \beta_B), is 3 times the probability that A beats B. At the maximum likelihood solution, the gradient equals zero, which means: the expected number of times B should beat A under the model equals the observed number (3), and the expected number of times A should beat B equals the observed number (7). This is the method of moments interpretation of maximum likelihood for Bradley-Terry: at the optimum, observed win counts equal expected win counts for each item.

Step 4: Solve numerically. The gradient equations don't have a closed-form solution because the sigmoid function is nonlinear. Unlike linear regression, where we can solve for parameters directly, the nonlinearity of the sigmoid means we must use numerical methods. We run gradient ascent (or equivalently, minimize the negative log-likelihood) using an efficient optimizer like L-BFGS. The convexity of the objective guarantees we will find the global maximum.

The true maximum likelihood estimates for our three-item example work out to approximately βA≈2.05\beta_A \approx 2.05, βB≈0.82\beta_B \approx 0.82, and βC=0\beta_C = 0 (fixed). Let's verify these make sense. The gap βA−βB≈1.23\beta_A - \beta_B \approx 1.23 gives σ(1.23)≈0.77\sigma(1.23) \approx 0.77, close to the observed 70% but adjusted upward by the model's use of A-vs-C information. The gap βB−βC=0.82\beta_B - \beta_C = 0.82 gives σ(0.82)≈0.69\sigma(0.82) \approx 0.69, close to the observed 80% but adjusted downward by the indirect information flowing through A. The Bradley-Terry model reconciles all three matchups simultaneously to find the best global explanation.

Code Implementation

We'll implement Bradley-Terry estimation from scratch using gradient descent, then show how it's applied in practice. The implementation closely mirrors the mathematics we derived above, which should make the connection between theory and code transparent.

In[9]:
Code
# Define our comparison data
# Each tuple: (winner_idx, loser_idx, count)
comparisons = [
    (0, 1, 7),  # A beats B, 7 times
    (1, 0, 3),  # B beats A, 3 times
    (1, 2, 8),  # B beats C, 8 times
    (2, 1, 2),  # C beats B, 2 times
    (0, 2, 9),  # A beats C, 9 times
    (2, 0, 1),  # C beats A, 1 time
]

n_items = 3  # A, B, C
item_names = ["A", "B", "C"]

Now let's define the negative log-likelihood function that we want to minimize. We negate the log-likelihood because standard optimizers minimize rather than maximize:

In[10]:
Code
def neg_log_likelihood(beta, comparisons):
    """
    Compute negative log-likelihood for Bradley-Terry model.
    beta[0] corresponds to item A, beta[1] to B, etc.
    Item C (index 2) is fixed at 0 for identifiability.
    """
    # Extend beta with the fixed anchor (beta_C = 0)
    beta_full = np.array([beta[0], beta[1], 0.0])

    nll = 0.0
    for winner, loser, count in comparisons:
        # Log probability that winner beats loser
        log_prob = np.log(sigmoid(beta_full[winner] - beta_full[loser]) + 1e-10)
        nll -= count * log_prob

    return nll

We also need the gradient for efficient optimization. The gradient tells the optimizer which direction to move in parameter space to improve the log-likelihood:

In[11]:
Code
def neg_log_likelihood_gradient(beta, comparisons):
    """
    Compute gradient of negative log-likelihood.
    """
    beta_full = np.array([beta[0], beta[1], 0.0])
    grad = np.zeros(2)  # Only for beta_A and beta_B (beta_C is fixed)

    for winner, loser, count in comparisons:
        # P(winner beats loser)
        prob = sigmoid(beta_full[winner] - beta_full[loser])

        # Gradient contribution
        # d/d(beta_winner) of log(sigma(x)) = 1 - sigma(x)
        # d/d(beta_loser) of log(sigma(x)) = -(1 - sigma(x))
        if winner < 2:  # Only update if not the anchor
            grad[winner] -= count * (1 - prob)
        if loser < 2:
            grad[loser] += count * (1 - prob)

    return grad

Now let's optimize to find the maximum likelihood estimates. We use L-BFGS-B, a quasi-Newton method that approximates the Hessian matrix and converges much faster than simple gradient descent for smooth convex objectives like ours:

In[12]:
Code
# Initial guess
beta_init = np.array([0.0, 0.0])

# Optimize using L-BFGS-B
result = minimize(
    neg_log_likelihood,
    beta_init,
    args=(comparisons,),
    method="L-BFGS-B",
    jac=neg_log_likelihood_gradient,
)

beta_A, beta_B = result.x
beta_C = 0.0  # Fixed anchor
Out[13]:
Console
Maximum Likelihood Estimates:
  β_A = 2.2156
  β_B = 1.3761
  β_C = 0.0000 (fixed anchor)

Optimization converged: True
Final negative log-likelihood: 14.3638

The estimated coefficients indicate that Item A is the strongest, followed by Item B, with Item C as the baseline. The optimization successfully converged to these maximum likelihood estimates. Item A's log-strength of roughly 2 units above C corresponds to a predicted 88% win rate against C, which is close to (but not exactly) the observed 90%. The Bradley-Terry model smooths the observed frequencies slightly, because it must simultaneously satisfy constraints from all three matchups rather than fitting each one independently.

Let's verify these estimates by computing the predicted win probabilities:

In[14]:
Code
# Compute predicted probabilities
beta_full = np.array([beta_A, beta_B, beta_C])

prob_A_beats_B = sigmoid(beta_A - beta_B)
prob_B_beats_C = sigmoid(beta_B - beta_C)
prob_A_beats_C = sigmoid(beta_A - beta_C)

# Calculate observed win rates dynamically from data
# comparisons is list of (winner, loser, count)
counts = {(w, l): c for w, l, c in comparisons}

obs_A_beats_B = counts[(0, 1)] / (counts[(0, 1)] + counts[(1, 0)])
obs_B_beats_C = counts[(1, 2)] / (counts[(1, 2)] + counts[(2, 1)])
obs_A_beats_C = counts[(0, 2)] / (counts[(0, 2)] + counts[(2, 0)])
Out[15]:
Console
Predicted Win Probabilities:
  P(A > B) = 0.698  (observed: 0.700)
  P(B > C) = 0.798  (observed: 0.800)
  P(A > C) = 0.902  (observed: 0.900)

The model closely matches the observed frequencies. Now let's convert these to Elo-style ratings for interpretability. The Elo scale is more familiar to practitioners and gives an intuitive sense of the magnitude of the quality differences:

In[16]:
Code
# Convert to Elo scale (where 400 points = 10:1 odds)
elo_scale = 400 / np.log(10)
elo_ratings = beta_full * elo_scale

# Normalize so average is 1500 (conventional Elo baseline)
elo_ratings = elo_ratings - np.mean(elo_ratings) + 1500
Out[17]:
Console
Elo-Style Ratings:
  A: 1677
  B: 1531
  C: 1292

The ratings align with our expectations: Item A has the highest rating. This reflects its dominance in the pairwise comparisons. A difference of approximately 200 points between A and B suggests a strong win probability for A, corresponding to roughly 76% expected wins in head-to-head matchups.

Visualizing the Preference Surface

Let's visualize how the Bradley-Terry probability varies with the strength difference, and where our estimated matchups fall on this curve:

Out[18]:
Visualization
S-shaped curve showing preference probability increasing from 0 to 1 as strength difference increases, with colored scatter points marking estimated differences for the A vs B, B vs C, and A vs C matchups.
Bradley-Terry preference probability curve as a function of strength difference $\beta_i - \beta_j$. The sigmoid function produces a logistic response where probabilities change most rapidly near zero difference and saturate asymptotically at the extremes. Colored points mark the estimated strength differences for the three matchups in our worked example.

The curve shows the characteristic sigmoid shape. When strengths are equal (βi−βj=0\beta_i - \beta_j = 0), the win probability is exactly 0.5. Small strength differences near zero produce probability changes close to linear, while large differences saturate toward 0 or 1. Notice that the A-vs-C point sits further to the right than the A-vs-B point (larger strength difference) and therefore has a higher predicted win probability.

Out[19]:
Visualization
Filled contour plot of the Bradley-Terry log-likelihood surface over a grid of beta_A and beta_B values, with a red star marking the maximum likelihood estimate and a viridis color scale showing log-likelihood magnitude.
Log-likelihood surface for the three-item worked example with $\beta_C=0$. The contours show how model fit varies with parameters $\beta_A$ and $\beta_B$, with the maximum likelihood estimate marked by the red star. The elongated shape reflects the correlation between parameters arising from the relative nature of pairwise comparisons.
Out[20]:
Visualization
Grouped bar chart comparing observed win frequencies (blue) and Bradley-Terry predicted probabilities (red) for the A vs B, B vs C, and A vs C matchups, with annotated values above each bar.
Observed win frequencies versus Bradley-Terry model predictions for the three pairwise matchups. The close alignment between empirical win rates (blue) and model probabilities (red) confirms the model's ability to fit the underlying preference structure. Small discrepancies arise because the model must be globally consistent rather than fitting each matchup independently.

Handling Real Preference Data

In practice, preference data comes as individual comparison records rather than aggregated counts. Large-scale LLM evaluation pipelines might receive thousands of individual annotation records, each showing which of two responses a specific annotator preferred to a specific prompt. Let's implement a more realistic version that handles raw comparison logs and scales to many items:

In[21]:
Code
class BradleyTerryModel:
    """
    Bradley-Terry model for pairwise comparisons.
    Estimates latent strength parameters from preference data.
    The last item (index n_items-1) is used as the fixed anchor at beta=0.
    """

    def __init__(self, n_items):
        self.n_items = n_items
        self.beta = np.zeros(n_items)  # Log-strength parameters
        self.fitted = False

    def _neg_log_likelihood(self, beta, winners, losers):
        """Compute NLL for optimization."""
        # Fix last item at 0 for identifiability
        beta_full = np.concatenate([beta, [0.0]])

        diffs = beta_full[winners] - beta_full[losers]
        # Use log1p for numerical stability: log(sigmoid(x)) = -log(1 + exp(-x))
        log_probs = -np.log1p(np.exp(-diffs))

        return -np.sum(log_probs)

    def _gradient(self, beta, winners, losers):
        """Compute gradient of NLL."""
        beta_full = np.concatenate([beta, [0.0]])
        diffs = beta_full[winners] - beta_full[losers]
        probs = sigmoid(diffs)

        grad = np.zeros(self.n_items - 1)
        for i in range(self.n_items - 1):
            # Contribution from comparisons where item i won
            win_mask = winners == i
            grad[i] -= np.sum(1 - probs[win_mask])

            # Contribution from comparisons where item i lost
            lose_mask = losers == i
            grad[i] += np.sum(1 - probs[lose_mask])

        return grad

    def fit(self, winners, losers):
        """
        Fit model to comparison data.

        Args:
            winners: Array of winner indices for each comparison
            losers: Array of loser indices for each comparison
        """
        winners = np.asarray(winners)
        losers = np.asarray(losers)

        beta_init = np.zeros(self.n_items - 1)

        result = minimize(
            self._neg_log_likelihood,
            beta_init,
            args=(winners, losers),
            method="L-BFGS-B",
            jac=lambda b, w, l: self._gradient(b, w, l),
        )

        self.beta = np.concatenate([result.x, [0.0]])
        self.fitted = True
        return self

    def predict_proba(self, item_i, item_j):
        """Predict probability that item_i beats item_j."""
        if not self.fitted:
            raise ValueError("Model must be fitted before prediction")
        return sigmoid(self.beta[item_i] - self.beta[item_j])

    def get_rankings(self):
        """Return items sorted by strength (best first)."""
        return np.argsort(-self.beta)

This implementation uses vectorized operations for efficiency. The _neg_log_likelihood method uses np.log1p(np.exp(-diffs)) rather than np.log(1 + np.exp(-diffs)), which is numerically more stable for large negative differences where np.exp(-diffs) might overflow. This kind of numerical care matters at scale, where comparisons can span items with very different quality levels.

Let's test this implementation with simulated preference data where we know the true underlying parameters:

In[22]:
Code
# Simulate data: 5 items with true strengths
np.random.seed(42)
n_items = 5
true_beta = np.array([2.0, 1.0, 0.0, -0.5, -1.5])

# Generate 500 random comparisons
n_comparisons = 500
winners = []
losers = []

for _ in range(n_comparisons):
    # Pick two random items
    i, j = np.random.choice(n_items, 2, replace=False)

    # Determine winner based on true probabilities
    prob_i_wins = sigmoid(true_beta[i] - true_beta[j])
    if np.random.random() < prob_i_wins:
        winners.append(i)
        losers.append(j)
    else:
        winners.append(j)
        losers.append(i)

winners = np.array(winners)
losers = np.array(losers)
In[23]:
Code
# Fit the model
model = BradleyTerryModel(n_items)
model.fit(winners, losers)

# Compare estimated to true parameters (both anchored at last item)
true_beta_anchored = true_beta - true_beta[-1]
Out[24]:
Console
Parameter Recovery:
Item   True β     Estimated β  Error   
------------------------------------
0      3.500      3.517        +0.017  
1      2.500      2.239        -0.261  
2      1.500      1.560        +0.060  
3      1.000      0.560        -0.440  
4      0.000      0.000        +0.000

The estimated parameters closely recover the true values. This shows that the Bradley-Terry model can reliably learn latent strengths from noisy pairwise comparisons. The errors are small relative to the true parameter gaps, meaning the model correctly recovers the ranking and approximate win probabilities even with just 500 comparisons across 5 items.

Visualizing Parameter Recovery

Out[25]:
Visualization
Scatter plot comparing true and estimated Bradley-Terry parameters for five items, with points clustered tightly along a dashed diagonal reference line showing near-perfect parameter recovery.
True versus estimated Bradley-Terry log-strength parameters for simulated data with 500 comparisons across 5 items. The tight clustering along the diagonal confirms accurate parameter recovery. Small deviations from the diagonal reflect the statistical uncertainty inherent in any finite sample.
Out[26]:
Visualization
Log-log error plot showing mean absolute error of Bradley-Terry parameter estimates declining from roughly 0.65 at 50 comparisons to below 0.1 at 2000 comparisons, with error bars showing trial-to-trial variability.
Mean absolute error of parameter estimates as a function of sample size, showing how estimation accuracy improves with more comparison data. The downward trend on the log-log scale indicates power-law convergence typical of maximum likelihood estimators, with error approximately halving each time the sample size quadruples.

The sample size plot reveals an important practical consideration for LLM evaluation. At 50 comparisons, the mean absolute error in log-strength is above 0.3, which translates to meaningful uncertainty in predicted win probabilities. At 500 comparisons, the error drops below 0.1, giving much more reliable estimates. At 2000 comparisons, the model has high accuracy. For a real evaluation platform with dozens or hundreds of models, getting 200 or more comparisons per model pair requires substantial annotation effort, which is why platforms like Chatbot Arena aggregate crowd-sourced votes from many users over extended time periods.

Connection to Reward Modeling

The Bradley-Terry model is the foundation of reward model training in RLHF. Understanding this connection reveals why reward models are trained the way they are and illuminates the mathematical underpinnings of alignment techniques. This is perhaps the most important section of this chapter for practitioners building LLM alignment pipelines: the reward model training objective is not arbitrary; it is a direct instantiation of the Bradley-Terry likelihood.

Recall from the previous chapter that preference data consists of pairs (yw,yl)(y_w, y_l) where ywy_w is the preferred response and yly_l is the dispreferred response to the same prompt xx. A reward model rθ(x,y)r_\theta(x, y) assigns a scalar score to each prompt-response pair. This score represents the model's estimate of how good the response is for that particular prompt. The key insight is that we can interpret this reward as a Bradley-Terry strength parameter. Rather than having one fixed strength per item, we have a neural network that maps any (prompt, response) pair to a strength. This generalization is what allows the model to transfer from training prompts to novel prompts encountered at inference time.

Concretely, the probability that response ywy_w is preferred over yly_l should follow the Bradley-Terry formula:

P(yw≻yl∣x)=σ(rθ(x,yw)−rθ(x,yl))P(y_w \succ y_l \mid x) = \sigma(r_\theta(x, y_w) - r_\theta(x, y_l))

where:

  • yw,yly_w, y_l: the preferred (winning) and dispreferred (losing) responses
  • xx: the prompt or context
  • rθ(x,y)r_\theta(x, y): the reward model parameterized by weights θ\theta
  • σ(⋅)\sigma(\cdot): the sigmoid function converting the reward difference into a probability

This formulation says that the probability of preferring one response over another depends on the difference in their reward scores, passed through a sigmoid. Responses with higher rewards are more likely to be preferred, but the relationship is probabilistic, letting for noise in human judgments. Training the reward model to maximize this probability is equivalent to Bradley-Terry maximum likelihood estimation, except now the strength parameters are neural network outputs rather than fixed scalars. Instead of learning one number per item, we learn a function that can assign rewards to any prompt-response pair.

The loss function for reward model training is:

L(θ)=−E(x,yw,yl)∼D[log⁡σ(rθ(x,yw)−rθ(x,yl))]\mathcal{L}(\theta) = -\mathbb{E}_{(x, y_w, y_l) \sim \mathcal{D}} \left[ \log \sigma(r_\theta(x, y_w) - r_\theta(x, y_l)) \right]

where:

  • L(θ)\mathcal{L}(\theta): the loss function to minimize (negative log-likelihood)
  • D\mathcal{D}: the dataset of preference pairs
  • E\mathbb{E}: the expectation over samples drawn from the dataset
  • σ(⋅)\sigma(\cdot): the sigmoid function
  • rθ(x,y)r_\theta(x, y): the reward model output for input xx and response yy
  • yw,yly_w, y_l: the preferred and dispreferred responses

This is exactly the Bradley-Terry negative log-likelihood, with the expectation taken over the preference dataset. The connection is direct: minimizing this loss trains the neural network to produce reward scores that best explain the observed human preferences according to the Bradley-Terry model. We'll explore the full details of reward model architecture and training in the next chapter, but understanding Bradley-Terry makes the training objective transparent: we're fitting a neural network to predict comparison outcomes, using the same statistical framework that has been used to rank chess players since the 1950s.

The neural reward model generalizes the classical Bradley-Terry model in one important way: it can handle items that have never been directly compared. In the classical setting with nn fixed items, we could only make predictions about pairs that were directly represented in the parameter space. The neural reward model can evaluate any response to any prompt, because it learns a general function rather than a lookup table. This generalization is what makes reward models practical for training language models, where the "items" being compared are generated text sequences from an enormous combinatorial space.

Out[27]:
Visualization
Horizontal bar chart showing fixed log-strength parameters for items A (2.0), B (1.2), C (0.5), and D (-0.3), illustrating the classical Bradley-Terry one-parameter-per-item setup.
Classical Bradley-Terry parameterization with one fixed scalar parameter per item. Each item maps to a static log-strength value, letting consistent pairwise ranking but giving no mechanism to generalize to unseen inputs.
Out[28]:
Visualization
Schematic diagram of a small neural network with an input layer labeled (x, y), a hidden layer, and a single output node labeled r_theta, representing the reward model that replaces fixed Bradley-Terry strength parameters.
Neural reward model architecture as the generalized Bradley-Terry strength function. A neural network $r_\theta(x, y)$ maps any context-response pair to a scalar reward, preserving the Bradley-Terry probabilistic interpretation while letting generalization to novel inputs never seen during training.

Limitations and Practical Considerations

The Bradley-Terry model makes assumptions that do not always hold in practice, and being clear-eyed about these limitations is needed for using the model responsibly in LLM alignment. Understanding where the model can break down helps practitioners design better data collection procedures, interpret results with appropriate caution, and choose extensions or alternatives when the standard model is insufficient.

The most significant limitation is the independence from irrelevant alternatives (IIA) assumption. The probability that A beats B doesn't depend on what other items exist in the comparison set. This means if you're comparing responses A and B, introducing a third response C shouldn't change how humans compare A to B. In practice, human preferences show strong context effects. A response that seems excellent when compared against a poor one might seem only adequate when compared against an outstanding one. Annotators calibrate their judgments to the comparison set they have seen, and their implicit sense of "what constitutes a good response" shifts based on the distribution of examples they encounter. The Bradley-Terry model ignores these calibration effects, which can cause the estimated strengths to depend on which comparisons were collected and in what order, rather than which reflects stable underlying qualities.

The model also assumes that preferences are governed by a single latent dimension: each item has one "strength" value that determines all its comparisons. Human preferences for text are often multidimensional. A response might be more helpful but less safe, or more creative but less accurate, or more concise but less complete. Different annotators may weight these dimensions differently based on their background and priorities. A single scalar reward compresses these tradeoffs into one number, potentially losing important information about the structure of preferences. When the true preference space is multidimensional, the Bradley-Terry model will give inconsistent parameter estimates across different subpopulations of annotators, because what looks like "strength" in one dimension may look like "weakness" in another. Research into multi-objective RLHF attempts to address this limitation by learning separate reward models for attributes such as helpfulness and safety, with factual accuracy scored independently as well.

Another challenge is intransitivity in human preferences. The Bradley-Terry model enforces transitivity by construction: if A is stronger than B and B is stronger than C, then A must be stronger than C. But humans sometimes exhibit cycles: preferring A over B, B over C, yet C over A. These cycles can arise from different annotators having different preferences, or from the same annotator weighing different factors in different comparisons. They can also emerge from strategic annotation behavior, where annotators reason about which responses are "better" according to an implicit rubric that may itself be inconsistent. The Bradley-Terry model smooths over these inconsistencies by finding the best one-dimensional ordering, which is usually desirable when intransitivities reflect noise. But when intransitivities reflect disagreement among annotator subgroups about what makes a response good, smoothing them away loses information about the diversity of human preferences.

The model also requires careful attention to data collection design. Items that are never directly or indirectly compared cannot be placed on the same ranking scale. If models A and B are only compared against models from different eras, and never against each other, their relative ordering cannot be estimated. In practice, LLM evaluation platforms address this by ensuring that all models have some comparisons against a common set of reference models, creating a connected comparison graph. The statistical precision of estimates also depends heavily on the graph structure: items that are frequently compared against many others have well-estimated parameters, while items that appear in few comparisons have high uncertainty.

A subtler issue is position bias in pairwise comparison collection. When annotators see two responses, their choice may be influenced by which response was shown first or on which side of the screen. Research on human annotation has found systematic biases where annotators prefer the first or the longer response regardless of actual quality. The Bradley-Terry model treats all comparisons as unbiased measurements of underlying quality, so position biases in the annotation process will corrupt the estimated parameters. Careful data collection protocols randomize presentation order, and some research has proposed extended models that include a nuisance parameter for position bias.

Finally, the Bradley-Terry model produces point estimates of strength parameters without automatic uncertainty quantification. Two items with widely different amounts of comparison data will both produce single scalar estimates, but the uncertainty in the estimate with few comparisons is much larger. Platforms that use Elo-style ratings often acknowledge this by noting when a model has a small number of comparisons, but the rating itself does not encode this uncertainty. Bayesian extensions of the Bradley-Terry model, such as placing priors over the strength parameters, can produce posterior distributions that naturally represent uncertainty, allowing proper statistical comparisons between items with different amounts of data.

Despite these limitations, the Bradley-Terry model has had a large impact on LLM alignment. Its simplicity makes it computationally tractable to train reward models on millions of preference comparisons. The resulting reward scores provide the training signal for PPO and other policy optimization algorithms. More recently, the mathematical connection between reward models and the Bradley-Terry likelihood enabled Direct Preference Optimization, which we'll cover later in this part. DPO derives an implicit reward model from the policy itself, eliminating the need for separate reward model training while still optimizing the same Bradley-Terry objective. This connection lets us derive the family of RLHF and DPO training procedures from the same preference model.

Summary

The Bradley-Terry model provides a principled framework for converting pairwise comparison data into consistent quality scores. Its core assumptions are:

  • Each item has a latent strength parameter βi\beta_i
  • The probability that item ii beats item jj is σ(βi−βj)\sigma(\beta_i - \beta_j)
  • Parameters are learned by maximum likelihood estimation

Key insights from this chapter:

  • Log-linear form: The Bradley-Terry probability is a sigmoid of the strength difference, which makes it equivalent to logistic regression on pairwise comparisons. This connection to a well-studied model enables efficient optimization and rich statistical theory.
  • Transitivity by design: The model automatically produces transitive rankings by placing all items on a one-dimensional scale. This is both a feature (consistency) and a limitation (it cannot represent multidimensional preference structure or systematic intransitivities).
  • Identifiability: Only differences in strength matter, so one parameter must be anchored. Absolute values have no meaning; only relative differences are interpretable.
  • Connection to Elo: The Elo rating system is a scaled Bradley-Terry model with an online update rule approximating maximum likelihood. The 400-point scaling factor converts natural log-strengths to a conventional rating scale.
  • Foundation for reward modeling: The reward model training objective is exactly the Bradley-Terry likelihood with neural network-parameterized strengths. Minimizing the binary cross-entropy loss on preference pairs is mathematically identical to fitting a Bradley-Terry model where each item's strength is a neural network output.
  • Sample efficiency matters: Accurate parameter estimation requires on the order of hundreds of comparisons per item pair, which motivates the large-scale crowd-sourced annotation efforts used to build alignment datasets.

In the next chapter on Reward Modeling, we'll see how this framework scales to neural networks that can generalize across prompts and responses. This provides the dense training signal needed for RLHF. We'll also examine how the reward model architecture is designed, how to avoid reward hacking, and how the trained reward model interacts with the policy optimization step.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the Bradley-Terry model and its role in preference learning.

Bradley-Terry Model

Question 1 of 80 of 8 completed
In the Bradley-Terry model, if item A has strength πA=6\pi_A=6 and item B has strength πB=2\pi_B=2, what is P(A≻B)?P(A\succ B)?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025bradleyterry, author = {Michael Brenndoerfer}, title = {Bradley-Terry Model: Converting Preferences to Rankings}, year = {2025}, url = {https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Bradley-Terry Model: Converting Preferences to Rankings. Retrieved from https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings
MLAAcademic
Michael Brenndoerfer. "Bradley-Terry Model: Converting Preferences to Rankings." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings>.
CHICAGOAcademic
Michael Brenndoerfer. "Bradley-Terry Model: Converting Preferences to Rankings." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Bradley-Terry Model: Converting Preferences to Rankings'. Available at: https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Bradley-Terry Model: Converting Preferences to Rankings. https://mbrenndoerfer.com/writing/bradley-terry-model-pairwise-preferences-rankings

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.