Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation

Michael BrenndoerferMarch 8, 202647 min read

Part of Language AI Handbook

Covers chance-corrected agreement metrics for NLP annotation reliability. Calculate Cohen's kappa, Fleiss' kappa, and Krippendorff's alpha with Python examples.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Inter-Annotator Agreement

Every supervised learning system rests on labeled data, and labeled data rests on human judgment. But human judgment is neither infallible nor consistent. Two linguists reading the same sentence may disagree about whether it expresses sincere anger or mild frustration. Two medical text annotators may draw entity boundaries at different places. Two quality raters evaluating an LLM response may have completely different intuitions about what "helpful" means. When we train a model on these annotations or use them as benchmarks, we inherit whatever inconsistencies the annotation process contained.

This is not a theoretical concern. Datasets built without systematic agreement checking have been shown to contain systematic disagreements that bias model behavior in surprising ways. A model trained on sentiment data where one annotator was consistently stricter than others will learn to predict a blend of their standards rather than any coherent human judgment. An evaluation benchmark where raters frequently disagree will produce unstable leaderboard rankings that change more with the choice of raters than with the actual capability of the models being tested.

Inter-annotator agreement (IAA) provides the statistical machinery to quantify this reliability. It measures how consistently multiple annotators label the same data, correcting for the possibility that they might agree simply by chance. Without this correction, a dataset with 90% positive examples and only two labels might show 82% "agreement" even if annotators are guessing randomly. This chapter explores the mathematics of chance-corrected agreement, from Cohen's foundational kappa for pairs of annotators through Krippendorff's alpha, which generalizes to any number of raters, any measurement scale, and missing data. Along the way, we examine how to handle the disagreements that inevitably arise, when high agreement can mislead you, and how to choose the right metric for your annotation task.

As we discussed in Part VI: Sequence Labeling, named entity recognition and part-of-speech tagging require careful annotation protocols. The metrics we develop here tell us whether those protocols are sufficiently clear to produce reproducible labels. In Part XXXVII: Alignment and RLHF, we saw that human preference data drives reward modeling. Before using such data, we must verify that human raters agree on what constitutes a "better" response. We'll explore preference-specific evaluation in the next chapter on Preference Evaluation, but the foundations we build here apply universally across all annotation tasks.

The Problem of Chance Agreement

Consider a simple binary classification task: determining whether a movie review is positive or negative. Two annotators label 100 reviews. They agree on 75 reviews and disagree on 25. Their raw agreement is 75%, which seems respectable. But what if 90 of those reviews are positive, a highly imbalanced dataset?

If both annotators simply guessed "positive" every time without reading a single review, they would agree on 90 reviews by chance alone. Their 75% observed agreement would be below chance expectation, showing systematic disagreement masked by class imbalance. Raw agreement percentages, often called "percentage of agreement" or PoP_o, fail to tell us anything useful without accounting for how much agreement chance alone would produce.

This prevalence problem becomes even more acute in real-world NLP tasks. Named entity recognition datasets often have a high proportion of "O" (outside) tokens compared to entity tokens. If 95% of tokens have no entity label, two annotators who both copy-paste the majority class would agree 90% of the time. A chance-corrected coefficient would expose this false agreement immediately.

Chance-Corrected Agreement

A chance-corrected agreement coefficient compares observed agreement (PoP_o) against expected agreement by chance (PeP_e), normalized by the maximum possible improvement:

Coefficient=Po−Pe1−Pe\text{Coefficient} = \frac{P_o - P_e}{1 - P_e}

where:

  • PoP_o: the observed agreement proportion (the fraction of items on which raters agree)
  • PeP_e: the expected agreement by chance (the fraction of items on which raters would agree if guessing randomly based on marginal distributions)

When observed agreement equals chance agreement, the coefficient is zero. When observed agreement is perfect (1.0), the coefficient reaches 1.0. Negative values indicate agreement worse than chance.

The denominator 1−Pe1 - P_e represents the maximum possible improvement over chance. If chance agreement is already high (say 0.9 due to class imbalance), there is little room for improvement, making high coefficient values difficult to achieve even with careful annotators. This explains why kappa coefficients sometimes show paradoxical behavior that we examine later. Understanding this denominator is key to understanding both the power and the limitations of chance-corrected metrics.

The logic behind chance-corrected agreement can be thought of as asking: "Given how frequently each annotator uses each label, what agreement would we expect if their decisions were statistically independent?" If annotator A uses the "positive" label 70% of the time and annotator B uses it 60% of the time, and they are operating independently, we would expect them to both say "positive" on 0.70×0.60=0.420.70 \times 0.60 = 0.42 of items purely by coincidence. Any agreement beyond that reflects reliability in the annotation process.

Cohen's Kappa

Cohen's kappa (κ\kappa), introduced by Jacob Cohen in 1960, remains the most widely cited agreement statistic for exactly two raters working with categorical data. It assumes that the raters are distinct individuals with potentially different tendencies (one might be stricter or more lenient than the other), and it does not assume the categories are ordered. Virtually every annotation paper in NLP that involves two raters reports this coefficient.

Mathematical Formulation

The starting point for Cohen's kappa is the contingency table (often called the agreement matrix) that shows how often each rater assigned each category. For two raters and kk categories, this is a k×kk \times k matrix where cell (i,j)(i, j) shows how many items rater 1 assigned to category ii while rater 2 assigned them to category jj.

Let pijp_{ij} represent the proportion of items in cell (i,j)(i, j). The observed agreement PoP_o is the sum of diagonal elements, since diagonal cells represent cases where both raters agreed:

Po=∑i=1kpiiP_o = \sum_{i=1}^{k} p_{ii}

where:

  • piip_{ii}: the proportion of items that both raters assign to category ii (the diagonal elements of the agreement matrix)
  • kk: the number of categories
  • PoP_o: the total observed agreement (sum of diagonal proportions)

The expected agreement by chance PeP_e is computed under the assumption that the two raters' decisions are statistically independent. If rater 1 assigns fraction pi+p_{i+} to category ii, and rater 2 independently assigns fraction p+ip_{+i} to category ii, then the expected fraction of items where both agree on category ii is simply the product of these marginals:

Pe=∑i=1kpi+⋅p+iP_e = \sum_{i=1}^{k} p_{i+} \cdot p_{+i}

where:

  • pi+p_{i+}: the row marginal proportion (the proportion of items rater 1 assigns to category ii)
  • p+ip_{+i}: the column marginal proportion (the proportion of items rater 2 assigns to category ii)
  • PeP_e: the expected agreement by chance, calculated as the sum of products of marginal proportions

This is simply the probability that two independent draws from the respective marginal distributions would land on the same category, summed over all categories.

Cohen's kappa then combines these into the familiar chance-corrected formula:

κ=Po−Pe1−Pe=∑ipii−∑ipi+p+i1−∑ipi+p+i\begin{aligned} \kappa &= \frac{P_o - P_e}{1 - P_e} \\ &= \frac{\sum_{i} p_{ii} - \sum_{i} p_{i+} p_{+i}}{1 - \sum_{i} p_{i+} p_{+i}} \end{aligned}

where:

  • κ\kappa: Cohen's kappa coefficient (chance-corrected agreement ranging from −1-1 to 11)
  • PoP_o: the observed agreement proportion
  • PeP_e: the expected agreement by chance
  • piip_{ii}: the proportion of agreement on category ii
  • pi+p_{i+}, p+ip_{+i}: the marginal proportions for category ii

Properties and Assumptions

Cohen's kappa treats the raters as fixed entities. We are measuring agreement between these specific two people, not generalizing to a population of potential raters. This is the basic philosophical difference between Cohen's kappa and later metrics like Fleiss' kappa and Krippendorff's alpha: those metrics treat raters as interchangeable samples from some broader rater population.

The metric is symmetric: swapping rater 1 and rater 2 does not change the value. This makes it appropriate when there is no natural distinction between the two raters, such as two independent coders annotating the same corpus. When one rater is designated "gold standard" and the other is being evaluated, weighted agreement with the gold standard might be more informative.

Cohen's kappa also assumes that all disagreements are equally costly. If your categories are "negative," "neutral," and "positive," kappa penalizes a "negative" vs "positive" disagreement exactly as much as a "negative" vs "neutral" disagreement. This is appropriate for nominal scales but becomes problematic for ordinal data. Weighted kappa addresses this limitation by letting you to specify that some disagreements are worse than others.

Scott's Pi

Before Cohen's kappa, Scott (1955) proposed π\pi, which uses a single pooled marginal distribution rather than separate marginals for each rater. Scott's pi assumes both raters draw from the same underlying category distribution, which makes it appropriate when raters are interchangeable samples from a rater population rather than specific individuals. Cohen argued that raters often have different biases: one annotator might apply the "positive" label more liberally than another, making separate marginals more faithful to reality. In practice, the two metrics give similar results when rater biases are small.

Weighted Kappa for Ordinal Scales

When categories have a natural ordering, the gap between Cohen's standard kappa and what you care about can be substantial. Consider a 5-point toxicity scale ranging from 1 (completely benign) to 5 (highly toxic). Standard kappa treats a rater disagreement of 1 vs 2 the same as 1 vs 5, even though the latter represents a far more serious discrepancy for downstream use.

Weighted kappa introduces a penalty matrix wijw_{ij} that specifies how much credit to give for each type of agreement or near-agreement. The weighted observed and expected agreement become:

Pow=∑i=1k∑j=1kwij⋅pijP_o^w = \sum_{i=1}^{k} \sum_{j=1}^{k} w_{ij} \cdot p_{ij} Pew=∑i=1k∑j=1kwij⋅pi+⋅p+jP_e^w = \sum_{i=1}^{k} \sum_{j=1}^{k} w_{ij} \cdot p_{i+} \cdot p_{+j}

and weighted kappa is:

κw=Pow−Pew1−Pew\kappa_w = \frac{P_o^w - P_e^w}{1 - P_e^w}

where:

  • wijw_{ij}: the agreement weight for the pair of labels (i,j)(i, j), ranging from 1 (full credit) to 0 (no credit)
  • PowP_o^w: the weighted observed agreement
  • PewP_e^w: the weighted expected agreement
  • κw\kappa_w: the weighted kappa coefficient

Two common weighting schemes are linear weighting (wij=1−∣i−j∣k−1w_{ij} = 1 - \frac{|i-j|}{k-1}) and quadratic weighting (wij=1−(i−j)2(k−1)2w_{ij} = 1 - \frac{(i-j)^2}{(k-1)^2}). Quadratic weighting penalizes large disagreements disproportionately, which makes it the preferred choice when large errors are particularly harmful. Cohen's quadratic weighted kappa is identical to the intraclass correlation coefficient (ICC) under certain distributional assumptions, connecting it to the broader literature on reliability in psychology and medicine.

Interpretation Benchmarks

Landis and Koch (1977) proposed widely cited benchmarks for interpreting kappa values:

  • <0.00< 0.00: Poor agreement
  • 0.00−0.200.00 - 0.20: Slight agreement
  • 0.21−0.400.21 - 0.40: Fair agreement
  • 0.41−0.600.41 - 0.60: Moderate agreement
  • 0.61−0.800.61 - 0.80: Substantial agreement
  • 0.81−1.000.81 - 1.00: Almost perfect agreement

However, these benchmarks face significant criticism and should not be applied mechanically. Kappa values depend heavily on prevalence and the difficulty of the task. A κ\kappa of 0.6 might represent excellent agreement for subtle pragmatic phenomena like sarcasm detection, but poor agreement for clear-cut factual entity recognition where experienced annotators should achieve κ>0.9\kappa > 0.9. The benchmarks were derived empirically from medical studies and do not generalize automatically to NLP. Always interpret kappa values in the context of your task, your annotators' expertise, and the guidelines you provided.

A more principled approach is to set task-specific thresholds before annotation begins, based on your application requirements. If a model will be deployed in a high-stakes setting, you might require κ>0.8\kappa > 0.8 before accepting any annotation. If you are exploring a new task where guidelines are still being developed, κ>0.5\kappa > 0.5 with careful analysis of disagreement patterns might be acceptable.

Fleiss' Kappa

Cohen's kappa does not generalize to more than two raters. When you have a crowdsourcing setup with ten workers or a research project where five domain experts each annotate the full corpus, you need a different approach. Fleiss' kappa (1971) extends the concept to any fixed number of raters n≥2n \geq 2, though it assumes the raters are interchangeable rather than distinct individuals.

This assumption of interchangeability is what distinguishes Fleiss' kappa from Cohen's. When you use Fleiss' kappa, you are implicitly treating your raters as random samples from a population of potential annotators, not specific individuals whose particular biases you care about. This makes it appropriate for crowdsourcing platforms like Amazon Mechanical Turk, where annotators are indeed sampled from a large pool.

Mathematical Structure

Consider NN items and kk categories. For each item ii, let nijn_{ij} be the number of raters who assigned it to category jj, where ∑j=1knij=n\sum_{j=1}^{k} n_{ij} = n (each item gets exactly nn ratings). The proportion of raters assigning item ii to category jj is:

pij=nijnp_{ij} = \frac{n_{ij}}{n}

where:

  • nijn_{ij}: the number of raters who assigned item ii to category jj
  • nn: the total number of raters per item
  • pijp_{ij}: the proportion of raters assigning item ii to category jj

The observed agreement for item ii measures the extent to which raters agree on that specific item. Think of it as counting all pairs of raters who agreed and expressing this as a fraction of all possible rater pairs:

Pi=1n(n−1)∑j=1knij(nij−1)=1n(n−1)[(∑j=1knij2)−n]\begin{aligned} P_i &= \frac{1}{n(n-1)} \sum_{j=1}^{k} n_{ij}(n_{ij} - 1) \\ &= \frac{1}{n(n-1)} \left[ \left(\sum_{j=1}^{k} n_{ij}^2\right) - n \right] \end{aligned}

where:

  • PiP_i: the extent of agreement among raters for item ii (ranging from 0 to 1)
  • nijn_{ij}: the count of raters assigning item ii to category jj
  • nn: the total number of raters
  • kk: the number of categories
  • The first form counts agreeing pairs directly; the second simplifies computation using the sum of squares

To understand this formula, consider what happens at the extremes. If all nn raters choose the same category jj, then nij=nn_{ij} = n and nij(nij−1)=n(n−1)n_{ij}(n_{ij} - 1) = n(n-1). The sum equals n(n−1)n(n-1) and Pi=1P_i = 1. If raters split perfectly (each choosing a different category), then nij=1n_{ij} = 1 for each jj and nij(nij−1)=0n_{ij}(n_{ij} - 1) = 0 for all jj, giving Pi=0P_i = 0.

The mean observed agreement across all NN items is:

Po=1N∑i=1NPiP_o = \frac{1}{N} \sum_{i=1}^{N} P_i

where:

  • PoP_o: the mean observed agreement across all items
  • NN: the total number of items being rated
  • PiP_i: the agreement score for item ii

The chance agreement PeP_e uses the proportion of all assignments falling into each category. This is a single pooled distribution across all raters, which reflects the interchangeability assumption:

Pj=1N⋅n∑i=1NnijP_j = \frac{1}{N \cdot n} \sum_{i=1}^{N} n_{ij} Pe=∑j=1kPj2P_e = \sum_{j=1}^{k} P_j^2

where:

  • PjP_j: the overall proportion of assignments to category jj across all items and raters
  • PeP_e: the expected chance agreement (the sum of squared category proportions)

The intuition for Pe=∑jPj2P_e = \sum_j P_j^2 is clean: if we sampled two raters at random and both assigned labels independently from the pooled distribution, the probability they would pick the same category jj is Pj⋅Pj=Pj2P_j \cdot P_j = P_j^2. Summing over all categories gives the total chance agreement.

Fleiss' kappa is then:

κF=Po−Pe1−Pe\kappa_F = \frac{P_o - P_e}{1 - P_e}

where:

  • κF\kappa_F: Fleiss' kappa coefficient
  • PoP_o: the mean observed agreement across items
  • PeP_e: the expected chance agreement based on category proportions

Key Differences from Cohen's Kappa

The most important practical difference between Fleiss' and Cohen's kappa is the choice of marginal distributions. Cohen's kappa uses separate marginals for each rater (pi+p_{i+} for rater 1 and p+ip_{+i} for rater 2). This reflects the fact that rater 1 might use "positive" 70% of the time while rater 2 uses it only 50% of the time. Fleiss' kappa uses a single pooled marginal (PjP_j), treating all raters as if they draw from the same distribution. This pooled approach is identical to what Scott's pi uses for two raters.

This has a subtle but important consequence. If you have exactly two raters with different biases and you apply Fleiss' kappa, you get Scott's pi rather than Cohen's kappa. The two coefficients can give substantially different values when rater biases differ substantially. When choosing between them, consider whether rater-specific tendencies are meaningful information (use Cohen's kappa) or noise to be averaged away (use Fleiss' kappa or Krippendorff's alpha).

When you have many raters or view raters as interchangeable samples from a larger population, Fleiss' kappa is the natural choice. In crowdsourcing research, it is particularly common because the identity of individual workers is less important than the overall reliability of the pool.

Krippendorff's Alpha

Krippendorff's alpha (α\alpha) represents the most general agreement coefficient available. It accommodates:

  • Any number of raters (varying per item if needed)
  • Any number of categories
  • Any measurement scale (nominal, ordinal, interval, ratio)
  • Missing data, where some items receive fewer ratings than others

This flexibility makes it the preferred metric for complex annotation schemes in modern NLP, particularly when different items might have different numbers of ratings, when using ordinal scales like Likert items, or when the annotation project spans a long period where not all raters annotate all items. Klaus Krippendorff, who introduced the metric in 1970 and refined it extensively in subsequent decades, designed it explicitly for the messiness of real-world content analysis.

The General Form

Krippendorff's alpha uses a disagreement-based formulation rather than an agreement-based one, but the underlying logic is identical to the chance-correction framework we have seen throughout this chapter:

α=1−DoDe\alpha = 1 - \frac{D_o}{D_e}

where:

  • α\alpha: Krippendorff's alpha coefficient
  • DoD_o: the observed disagreement among raters
  • DeD_e: the expected disagreement by chance

When Do=0D_o = 0 (perfect agreement), α=1\alpha = 1. When Do=DeD_o = D_e (agreement equals chance), α=0\alpha = 0. When Do>DeD_o > D_e (worse than chance), α<0\alpha < 0.

For nominal data (categories without order), the observed disagreement for an item is the proportion of rater pairs that disagree:

Do=1n(n−1)∑cnc(n−nc)D_o = \frac{1}{n(n-1)} \sum_{c} n_c (n - n_c)

where:

  • DoD_o: the observed disagreement for an item
  • ncn_c: the count of raters choosing category cc
  • nn: the total number of raters for that item
  • The sum is taken over all categories cc

The expected disagreement assumes a multinomial distribution based on the overall category frequencies across all data. If the probability of any given annotation being category cc is πc\pi_c, then the probability that two independent annotations disagree is:

De=1−∑cπc2D_e = 1 - \sum_{c} \pi_c^2

where:

  • DeD_e: the expected disagreement by chance
  • πc\pi_c: the overall proportion of assignments to category cc across all data
  • The term ∑cπc2\sum_{c} \pi_c^2 represents the probability of chance agreement (the same formula as PeP_e in Fleiss' kappa)

Notice that for nominal data with complete observations, Krippendorff's alpha and Fleiss' kappa give identical results. The differences between them emerge when data is missing or when using non-nominal measurement scales.

Handling Missing Data

This is where Krippendorff's alpha truly distinguishes itself. Real annotation projects almost always have missing data. Annotators drop out, items are too difficult for some raters, or different items are assigned to different subsets of raters. Cohen's kappa and Fleiss' kappa require complete data for each item.

Krippendorff's alpha handles missing data through coincidence matrices rather than contingency tables. Instead of requiring every rater to label every item, we compute the observed disagreement from all pairs of observations that are present. For each item ii with mi≥2m_i \geq 2 valid ratings, we consider all ordered pairs of raters and record which categories they chose. Items with fewer than 2 valid ratings are simply skipped.

For mm items with varying numbers of raters, we construct a coincidence matrix where cell (c,k)(c, k) contains the number of times any rater pair for any item assigned one to category cc and the other to category kk. The diagonal contains agreements; off-diagonals contain disagreements. This matrix is symmetric, and each item with nin_i raters contributes ni(ni−1)n_i(n_i - 1) entries.

The resulting coincidence matrix is a sufficient statistic for computing both DoD_o and DeD_e, regardless of how many ratings each item has. This is the elegant mathematical property that lets alpha handle arbitrary missingness patterns without any special-casing.

Weighting Schemes

For ordinal, interval, or ratio scales, not all disagreements are equal. A 5-star quality rater who gives 3 stars while their colleague gives 4 stars is closer to agreement than one who gives 1 star. Krippendorff's alpha incorporates a difference function δck2\delta_{ck}^2 that weights disagreements by their distance:

α=1−∑c,kockδck2∑c,keckδck2\alpha = 1 - \frac{\sum_{c,k} o_{ck} \delta_{ck}^2}{\sum_{c,k} e_{ck} \delta_{ck}^2}

where:

  • ocko_{ck}: the observed coincidences between categories cc and kk
  • ecke_{ck}: the expected coincidences between categories cc and kk
  • δck2\delta_{ck}^2: the squared distance between categories cc and kk (weighting function)
  • The numerator sums observed disagreements weighted by distance; the denominator sums expected disagreements weighted by distance

Different measurement scales suggest different distance functions:

  • Nominal: δck2=0\delta_{ck}^2 = 0 if c=kc = k, else 11 (binary disagreement, all errors equal)
  • Ordinal: δck2=(∑g=min⁡(c,k)max⁡(c,k)ng−nc+nk2)2\delta_{ck}^2 = \left(\sum_{g=\min(c,k)}^{\max(c,k)} n_g - \frac{n_c + n_k}{2}\right)^2 (rank-based distances)
  • Interval: δck2=(c−k)2\delta_{ck}^2 = (c - k)^2 (squared difference in values)
  • Ratio: δck2=(c−kc+k)2\delta_{ck}^2 = \left(\frac{c - k}{c + k}\right)^2 (squared relative difference)

For Likert-scale ratings of LLM response quality, interval alpha is usually appropriate because the scale is treated as having equal spacing between levels. For judgments that have a natural zero point and where the ratio between values is meaningful (such as response latency in milliseconds), ratio alpha is more appropriate. The choice of distance function should be driven by the measurement theory underlying your scale, not by which choice makes the number look best.

Worked Example

Let's calculate all three metrics on a concrete example. Suppose three annotators (A, B, C) label 10 sentences for sentiment: Positive (P), Neutral (N), or Negative (G).

Sentiment annotations for 10 items by three annotators.
ItemABC
1PPP
2PPN
3NNN
4PPP
5NPN
6GGG
7PNP
8NNN
9PPP
10GPG

Six items (1, 3, 4, 6, 8, 9) are unanimously agreed upon. The four disagreement items (2, 5, 7, 10) each have a 2-1 split, with different categories causing trouble in each case.

Cohen's Kappa (A vs B)

First, construct the agreement matrix between A and B:

Agreement matrix between annotators A and B.
B:PB:NB:G
A:P410
A:N120
A:G101

Observed agreement:

Po=4+2+110=0.7\begin{aligned} P_o &= \frac{4 + 2 + 1}{10} \\ &= 0.7 \end{aligned}

Row marginals for A: pP+=0.5p_{P+} = 0.5, pN+=0.3p_{N+} = 0.3, pG+=0.2p_{G+} = 0.2 Column marginals for B: p+P=0.6p_{+P} = 0.6, p+N=0.3p_{+N} = 0.3, p+G=0.1p_{+G} = 0.1

Expected agreement:

Pe=(0.5×0.6)+(0.3×0.3)+(0.2×0.1)=0.30+0.09+0.02=0.41\begin{aligned} P_e &= (0.5 \times 0.6) + (0.3 \times 0.3) + (0.2 \times 0.1) \\ &= 0.30 + 0.09 + 0.02 \\ &= 0.41 \end{aligned}

Cohen's kappa:

κ=0.7−0.411−0.41=0.290.59≈0.492\begin{aligned} \kappa &= \frac{0.7 - 0.41}{1 - 0.41} \\ &= \frac{0.29}{0.59} \\ &\approx 0.492 \end{aligned}

This falls in the "moderate" range. Notice that A uses "Negative" 20% of the time while B uses it only 10% of the time. This rater-level difference in base rates is exactly what Cohen's kappa captures through separate marginals.

Fleiss' Kappa (All Three)

For Fleiss' kappa, we calculate PiP_i for each item:

  • Item 1 (P,P,P):
P1=13(2)(32−3)=66=1.0\begin{aligned} P_1 &= \frac{1}{3(2)}(3^2 - 3) \\ &= \frac{6}{6} \\ &= 1.0 \end{aligned}
  • Item 2 (P,P,N):
P2=16(22+12+0−3)=36=0.333\begin{aligned} P_2 &= \frac{1}{6}(2^2 + 1^2 + 0 - 3) \\ &= \frac{3}{6} \\ &= 0.333 \end{aligned}
  • Item 3 (N,N,N): 1.01.0

  • Item 4 (P,P,P): 1.01.0

  • Item 5 (N,P,N):

P5=16(12+22−3)=26=0.333\begin{aligned} P_5 &= \frac{1}{6}(1^2 + 2^2 - 3) \\ &= \frac{2}{6} \\ &= 0.333 \end{aligned}
  • Item 6 (G,G,G): 1.01.0

  • Item 7 (P,N,P):

P7=16(22+12−3)=0.333\begin{aligned} P_7 &= \frac{1}{6}(2^2 + 1^2 - 3) \\ &= 0.333 \end{aligned}
  • Item 8 (N,N,N): 1.01.0

  • Item 9 (P,P,P): 1.01.0

  • Item 10 (G,P,G):

P10=16(12+22−3)=0.333\begin{aligned} P_{10} &= \frac{1}{6}(1^2 + 2^2 - 3) \\ &= 0.333 \end{aligned}

Mean observed agreement:

Po=110(1+0.333+1+1+0.333+1+0.333+1+1+0.333)=7.33210=0.733\begin{aligned} P_o &= \frac{1}{10}(1 + 0.333 + 1 + 1 + 0.333 + 1 + 0.333 + 1 + 1 + 0.333) \\ &= \frac{7.332}{10} \\ &= 0.733 \end{aligned}

Category proportions (total assignments = 10 items ×\times 3 raters = 30):

  • PP=15/30=0.5P_P = 15/30 = 0.5
  • PN=10/30=0.333P_N = 10/30 = 0.333
  • PG=5/30=0.167P_G = 5/30 = 0.167

Expected agreement:

Pe=0.52+0.3332+0.1672=0.25+0.111+0.028=0.389\begin{aligned} P_e &= 0.5^2 + 0.333^2 + 0.167^2 \\ &= 0.25 + 0.111 + 0.028 \\ &= 0.389 \end{aligned}

Fleiss' kappa:

κF=0.733−0.3891−0.389=0.3440.611≈0.563\begin{aligned} \kappa_F &= \frac{0.733 - 0.389}{1 - 0.389} \\ &= \frac{0.344}{0.611} \\ &\approx 0.563 \end{aligned}

Krippendorff's Alpha

For nominal data with 3 raters per item, we build the coincidence matrix. Each item contributes ordered pairs of ratings, one per ordered pair of raters.

Item 5: A(N), B(P), C(N). Ordered pairs: (A,B):N(A,B):N-PP, (A,C):N(A,C):N-NN, (B,A):P(B,A):P-NN, (B,C):P(B,C):P-NN, (C,A):N(C,A):N-NN, (C,B):N(C,B):N-PP. This contributes oNP+=2o_{NP} += 2, oPN+=2o_{PN} += 2, oNN+=2o_{NN} += 2.

Counting all contributions:

  • oPPo_{PP}: Items 1(6), 2(2), 4(6), 7(2), 9(6) =22= 22
  • oNNo_{NN}: Items 3(6), 5(2), 8(6) =14= 14
  • oGGo_{GG}: Items 6(6), 10(2) =8= 8
  • oPN=oNPo_{PN} = o_{NP}: Items 2(2), 5(2), 7(2) =6= 6 each
  • oPG=oGPo_{PG} = o_{GP}: Item 10(2) =2= 2 each
  • oNG=oGNo_{NG} = o_{GN}: None =0= 0

Total: 22+14+8+6+6+2+2=6022 + 14 + 8 + 6 + 6 + 2 + 2 = 60 (which matches 10×3×2=6010 \times 3 \times 2 = 60 ordered pairs).

Observed disagreement:

Do=oPN+oNP+oPG+oGP60=6+6+2+260=1660=0.267\begin{aligned} D_o &= \frac{o_{PN} + o_{NP} + o_{PG} + o_{GP}}{60} \\ &= \frac{6 + 6 + 2 + 2}{60} \\ &= \frac{16}{60} \\ &= 0.267 \end{aligned}

Expected disagreement:

De=1−(PP2+PN2+PG2)=1−(0.52+0.3332+0.1672)=1−0.389=0.611\begin{aligned} D_e &= 1 - (P_P^2 + P_N^2 + P_G^2) \\ &= 1 - (0.5^2 + 0.333^2 + 0.167^2) \\ &= 1 - 0.389 \\ &= 0.611 \end{aligned}

Alpha:

α=1−0.2670.611=1−0.437=0.563\begin{aligned} \alpha &= 1 - \frac{0.267}{0.611} \\ &= 1 - 0.437 \\ &= 0.563 \end{aligned}

For this case with complete data and nominal categories, Fleiss' kappa and Krippendorff's alpha give the same value (0.563), as theory predicts. Cohen's kappa between specific rater pairs varies. The A vs B pair showed κ=0.492\kappa = 0.492; A vs C and B vs C would produce different values because each pair has different marginal distributions.

Code Implementation

Let's implement these calculations in Python, verifying our manual computation and showing library usage. The code also handles a missing-data scenario that illustrates Krippendorff's alpha's unique capabilities.

In[4]:
Code
import numpy as np

# Define our annotation data: 10 items, 3 annotators (A, B, C)
# Categories: 0=Negative(G), 1=Neutral(N), 2=Positive(P)
annotations = np.array(
    [
        [2, 2, 2],  # Item 1: P,P,P
        [2, 2, 1],  # Item 2: P,P,N
        [1, 1, 1],  # Item 3: N,N,N
        [2, 2, 2],  # Item 4: P,P,P
        [1, 2, 1],  # Item 5: N,P,N
        [0, 0, 0],  # Item 6: G,G,G
        [2, 1, 2],  # Item 7: P,N,P
        [1, 1, 1],  # Item 8: N,N,N
        [2, 2, 2],  # Item 9: P,P,P
        [0, 2, 0],  # Item 10: G,P,G
    ]
)

n_items, n_raters = annotations.shape
categories = np.unique(annotations)
k = len(categories)

Now we calculate Cohen's kappa between annotators A (column 0) and B (column 1).

In[5]:
Code
# Cohen's Kappa between A and B
rater_a = annotations[:, 0]
rater_b = annotations[:, 1]

# Using sklearn
kappa_sklearn = cohen_kappa_score(rater_a, rater_b)


# Manual calculation for verification
def cohens_kappa_manual(rater1, rater2, categories):
    n = len(rater1)
    # Agreement matrix
    agreement = np.zeros((len(categories), len(categories)))
    for i in range(n):
        idx1 = np.where(categories == rater1[i])[0][0]
        idx2 = np.where(categories == rater2[i])[0][0]
        agreement[idx1, idx2] += 1

    agreement = agreement / n  # Convert to proportions

    # Observed agreement
    p_o = np.trace(agreement)

    # Expected agreement (marginals)
    p_i_plus = np.sum(agreement, axis=1)
    p_plus_j = np.sum(agreement, axis=0)
    p_e = np.sum(p_i_plus * p_plus_j)

    kappa = (p_o - p_e) / (1 - p_e)
    return kappa, p_o, p_e


kappa_manual, p_o, p_e = cohens_kappa_manual(rater_a, rater_b, categories)
Out[6]:
Console
Cohen's Kappa (A vs B) using sklearn: 0.492
Manual calculation: kappa = (0.70 - 0.41) / (1 - 0.41) = 0.492

The manual calculation confirms our earlier arithmetic, showing moderate agreement between annotators A and B.

Out[7]:
Visualization
3x3 heatmap agreement matrix for annotators A and B on negative, neutral, and positive sentiment.
Agreement matrix between annotators A and B showing the count of label assignments in each category pair. Diagonal cells (upper-left to lower-right) represent items where both raters agreed. The three off-diagonal disagreements are spread evenly across category boundaries, with no single category pair accounting for all errors.

Next, we implement Fleiss' kappa for all three raters.

In[8]:
Code
def fleiss_kappa(annotations, categories):
    """
    Calculate Fleiss' kappa for multiple raters.
    annotations: n_items x n_raters array
    """
    n_items, n_raters = annotations.shape
    k = len(categories)

    # Count matrix: n_items x k
    count_matrix = np.zeros((n_items, k))
    for i in range(n_items):
        for j in range(n_raters):
            cat_idx = np.where(categories == annotations[i, j])[0][0]
            count_matrix[i, cat_idx] += 1

    # P_i for each item: proportion of agreeing ordered rater pairs
    P_i = (np.sum(count_matrix**2, axis=1) - n_raters) / (
        n_raters * (n_raters - 1)
    )
    P_bar = np.mean(P_i)

    # Overall proportion of assignments to each category
    P_j = np.sum(count_matrix, axis=0) / (n_items * n_raters)

    # Expected agreement
    P_e = np.sum(P_j**2)

    kappa = (P_bar - P_e) / (1 - P_e) if (1 - P_e) != 0 else 0

    return kappa, P_bar, P_e, P_j, P_i, count_matrix


kappa_fleiss, p_bar, p_e_fleiss, cat_props, P_i_values, count_mat = (
    fleiss_kappa(annotations, categories)
)
Out[9]:
Console
Fleiss' Kappa: 0.564
Observed agreement (P_bar): 0.733
Expected agreement (P_e): 0.389
Category proportions: {'G': '0.167', 'N': '0.333', 'P': '0.500'}

Fleiss' kappa shows moderate agreement across all three raters, slightly higher than the pairwise Cohen's kappa between A and B. Pooling all three raters raises the observed agreement while producing a slightly lower chance-agreement baseline for this dataset.

Out[10]:
Visualization
Bar chart of per-item Fleiss kappa agreement scores, with unanimous items in teal and disagreement items in blue.
Per-item agreement scores ($P_i$) for the Fleiss kappa calculation. Items 1, 3, 4, 6, 8, and 9 show perfect unanimous agreement ($P_i = 1.0$, shown in teal), while the four disagreement items (2, 5, 7, 10) each score 0.33, corresponding to a 2-vs-1 split among the three raters. The horizontal dashed line marks the mean observed agreement $P_o = 0.733$.

Now we implement Krippendorff's alpha, which requires building the coincidence matrix.

In[11]:
Code
def krippendorff_alpha(annotations, categories, metric="nominal"):
    """
    Calculate Krippendorff's alpha.
    Handles varying numbers of raters per item (missing data via NaN).
    """
    n_items, n_raters = annotations.shape
    k = len(categories)

    # Build coincidence matrix
    coincidence = np.zeros((k, k))

    for i in range(n_items):
        # Get valid annotations for this item (non-NaN if we had missing data)
        valid = annotations[i, :]
        n_valid = len(valid)

        # All ordered pairs of raters
        for r1, r2 in itertools.product(range(n_valid), range(n_valid)):
            if r1 != r2:
                c1 = np.where(categories == valid[r1])[0][0]
                c2 = np.where(categories == valid[r2])[0][0]
                coincidence[c1, c2] += 1

    total_coincidences = np.sum(coincidence)

    # Observed disagreement (for nominal: proportion of off-diagonal)
    if metric == "nominal":
        D_o = (total_coincidences - np.trace(coincidence)) / total_coincidences
    else:
        raise NotImplementedError("Only nominal metric implemented here")

    # Expected disagreement based on marginals
    category_counts = np.sum(coincidence, axis=1)
    n_total = np.sum(category_counts)
    probs = category_counts / n_total

    D_e = 1 - np.sum(probs**2)

    alpha = 1 - (D_o / D_e) if D_e != 0 else 0

    return alpha, D_o, D_e, coincidence


alpha, d_o, d_e, coin_mat = krippendorff_alpha(annotations, categories)
Out[12]:
Console
Krippendorff's Alpha: 0.564
Observed disagreement (D_o): 0.267
Expected disagreement (D_e): 0.611

Krippendorff's alpha matches Fleiss' kappa in this case because we have complete data, fixed numbers of raters, and nominal categories. The advantage of alpha becomes apparent with missing data or ordinal scales.

Out[13]:
Visualization
3x3 coincidence matrix heatmap for Krippendorff alpha showing diagonal agreements and off-diagonal disagreements in orange.
Coincidence matrix for Krippendorff's alpha, showing the frequency of ordered rater-pair label combinations across all items. Diagonal elements represent agreements between rater pairs. The dominant off-diagonal entries (P-N and N-P) indicate that positive-vs-neutral confusion accounts for most disagreements, while positive-negative confusion is rare.
Out[14]:
Visualization
Bar chart comparing Cohen kappa, Fleiss kappa, and Krippendorff alpha values with colored agreement-strength bands in the background.
Side-by-side comparison of the three inter-annotator agreement metrics on the worked example. Cohen's kappa for the A-B pair falls in the moderate range, which reflects the specific divergence in base rates between those two annotators. Fleiss kappa and Krippendorff alpha agree to three decimal places because the data is complete and nominal, placing the three-rater agreement in the moderate range according to Landis-Koch benchmarks. Background bands show the conventional interpretation regions.

Let's demonstrate handling missing data, a key strength of Krippendorff's alpha. We remove two ratings and observe that alpha remains computable and close to the complete-data value.

In[15]:
Code
# Simulate missing data: remove some annotations
annotations_missing = annotations.copy().astype(float)
annotations_missing[2, 1] = np.nan  # Item 3, rater B missing
annotations_missing[7, 2] = np.nan  # Item 8, rater C missing


def krippendorff_alpha_missing(annotations, categories):
    """Alpha with missing data handling."""
    n_items, n_raters = annotations.shape
    k = len(categories)

    coincidence = np.zeros((k, k))

    for i in range(n_items):
        # Filter out NaN values
        valid = annotations[i, ~np.isnan(annotations[i, :])]
        n_valid = len(valid)

        if n_valid < 2:
            continue

        for r1, r2 in itertools.product(range(n_valid), range(n_valid)):
            if r1 != r2:
                c1 = np.where(categories == valid[r1])[0][0]
                c2 = np.where(categories == valid[r2])[0][0]
                coincidence[c1, c2] += 1

    total = np.sum(coincidence)
    D_o = (total - np.trace(coincidence)) / total

    probs = np.sum(coincidence, axis=1) / total
    D_e = 1 - np.sum(probs**2)

    return 1 - (D_o / D_e)


alpha_missing = krippendorff_alpha_missing(annotations_missing, categories)
Out[16]:
Console
Annotations with missing data (NaN):
[[ 2.  2.  2.]
 [ 2.  2.  1.]
 [ 1. nan  1.]
 [ 2.  2.  2.]
 [ 1.  2.  1.]
 [ 0.  0.  0.]
 [ 2.  1.  2.]
 [ 1.  1. nan]
 [ 2.  2.  2.]
 [ 0.  2.  0.]]

Krippendorff's Alpha with missing data: 0.467
Original (complete data): 0.564

The alpha coefficient remains stable even with missing annotations. This makes it suitable for real-world annotation projects where not every rater labels every item. Cohen's kappa and Fleiss' kappa would require you to either drop items with missing data or impute values before computation.

Let's now compute pairwise Cohen's kappa for all three annotator pairs to show how agreement varies across pairs.

In[17]:
Code
# Compute pairwise Cohen's kappa for all annotator pairs
pair_labels = ["A vs B", "A vs C", "B vs C"]
pair_kappas = []
pair_p_o = []
pair_p_e = []

for r1, r2 in [(0, 1), (0, 2), (1, 2)]:
    kappa_val, po, pe = cohens_kappa_manual(
        annotations[:, r1], annotations[:, r2], categories
    )
    pair_kappas.append(kappa_val)
    pair_p_o.append(po)
    pair_p_e.append(pe)
Out[18]:
Console
A vs B: kappa=0.492  (P_o=0.700, P_e=0.410)
A vs C: kappa=0.844  (P_o=0.900, P_e=0.360)
B vs C: kappa=0.355  (P_o=0.600, P_e=0.380)
Out[19]:
Visualization
Bar chart of pairwise Cohen kappa values for annotator pairs A-B, A-C, and B-C with the Fleiss kappa reference line.
Pairwise Cohen's kappa values for all three annotator combinations. The A vs C pair shows the highest agreement among binary comparisons, while A vs B shows the lowest. The variation across pairs illustrates why aggregating over raters with Fleiss kappa or Krippendorff alpha provides a more stable summary when you have more than two annotators.

Handling Disagreement

High inter-annotator agreement validates our annotation scheme, but what happens when agreement is low? Disagreement is not always noise. Sometimes it signals ambiguous examples, underspecified guidelines, or subjective phenomena. Understanding and handling disagreement appropriately is as important as measuring it.

Adjudication and Majority Voting

The simplest approach to resolving disagreement is adjudication: a third expert reviews cases where raters disagree and determines the "correct" label. This works well for objective tasks with clear gold standards (such as syntactic parsing or coreference resolution) but risks imposing artificial certainty on subjective phenomena. When a quality rater says a response is "helpful" and another says it is "not helpful," adjudication forces a binary choice that erases real ambiguity.

For multi-rater scenarios, majority voting selects the most common label. With two raters, this requires a tie-breaker rule. Majority voting preserves the most likely interpretation but discards information about uncertainty. If two annotators rate sentiment as Positive and Neutral while a third says Negative, majority voting yields Positive, but the disagreement suggests this might be a borderline case that a model should approach with calibrated uncertainty rather than confident classification.

Soft Labels and Probabilistic Targets

Rather than forcing hard choices, we can preserve disagreement as soft labels. If three annotators label an item as [Positive, Positive, Neutral], the soft label becomes a probability distribution: P(Positive)=0.67P(\text{Positive}) = 0.67, P(Neutral)=0.33P(\text{Neutral}) = 0.33.

Training models on soft labels rather than hard majority votes captures the nuance of borderline cases. Research on label smoothing and distribution matching has shown that models trained with soft targets can produce better-calibrated predictions. As we discussed in Part XXXVII: Alignment and RLHF, reward models are often trained on preference probabilities derived from multiple human judgments rather than binary choices. This approach acknowledges that some comparisons are close calls and that a model should assign similar scores to the competing responses.

The degree of annotation disagreement can itself be used as a signal. Items where annotators consistently disagree might indicate that the underlying construct is inherently ambiguous, or that the item sits at a decision boundary in the feature space. Both interpretations are informative for model design: the first suggests we should collect more annotations or refine the task definition, the second suggests the model should be allowed to output uncertainty.

Modeling Uncertainty

Low agreement items might deserve different treatment during training. Items with high annotator disagreement can be:

  • Weighted down in the loss function to reduce their influence on learned representations
  • Excluded from training if disagreement exceeds a threshold, keeping only high-quality signal
  • Flagged for guideline refinement when systematic disagreement patterns emerge
  • Treated as belonging to a separate "ambiguous" class, giving the model an explicit way to express uncertainty

In evaluation, disagreement-aware metrics report both performance on high-agreement items (where labels are reliable) and performance on the full dataset. A model that performs well only on unambiguous items might be exploiting superficial features rather than understanding the task. Similarly, a benchmark where 30% of items have majority-vote accuracy below 60% among human raters should not be treated as having a clean gold standard.

Guideline Iteration

Persistent disagreement often indicates guideline deficiencies. When annotators systematically disagree on entity boundaries in NER or on whether a statement constitutes a "hallucination," the annotation guidelines need clarification. Iterative annotation, where initial rounds inform guideline refinement, typically improves kappa scores in subsequent rounds.

However, be cautious of over-fitting guidelines to specific annotators. If you refine guidelines until two particular annotators agree, you may have simply matched the idiosyncrasies of those two people, not improved generalizability. Ideal guideline development involves diverse annotators, multiple annotation rounds, and explicit documentation of edge cases so that future annotators can apply the same standards.

A practical workflow for high-quality annotation projects looks like this. First, annotate a small pilot batch (100-200 items) and measure kappa. Second, analyze disagreement patterns to identify which item types, which category distinctions, or which annotator pairs are most problematic. Third, update guidelines to address specific sources of confusion. Fourth, annotate a second pilot batch and measure kappa again. Only once kappa exceeds your threshold should you proceed with the full dataset. This iterative approach front-loads cost but dramatically reduces the risk of discovering systematic quality problems after annotating tens of thousands of items.

Limitations and Impact

Chance-corrected agreement coefficients, while needed tools in the NLP practitioner's toolkit, suffer from well-documented paradoxes and limitations. Understanding these issues is not just academic: they determine which coefficient you should choose for a given task and how you should report and interpret your results.

The Prevalence Problem

When class distributions are highly skewed, kappa coefficients can be surprisingly low even with high raw agreement. Consider a medical diagnosis task where 95% of cases are "healthy." Two doctors might agree on 90% of cases (both saying healthy) but disagree on the 5% who are sick. Raw agreement is 90%, but if they randomly guessed healthy 95% of the time, chance agreement would be:

Pe=0.952+0.052=0.9025+0.0025=0.905\begin{aligned} P_e &= 0.95^2 + 0.05^2 \\ &= 0.9025 + 0.0025 \\ &= 0.905 \end{aligned}

The kappa would be:

κ=0.90−0.9051−0.905=−0.0050.095≈−0.05\begin{aligned} \kappa &= \frac{0.90 - 0.905}{1 - 0.905} \\ &= \frac{-0.005}{0.095} \\ &\approx -0.05 \end{aligned}

This result suggests negative agreement despite 90% accuracy. The paradox arises because chance agreement is already 90.5% given the class imbalance, leaving essentially no room for improvement. The kappa denominator 1−Pe=0.0951 - P_e = 0.095 amplifies any noise in the numerator.

The implication for NLP practice is significant. For sequence labeling tasks like NER, where most tokens are "outside" entities, kappa values will routinely look poor even for high-quality annotators. This prevalence paradox means kappa is unreliable for heavily imbalanced datasets. Some researchers recommend reporting both raw agreement and prevalence-adjusted bias-adjusted kappa (PABAK), though this metric loses the chance-correction interpretation. Others recommend reporting agreement separately for each class (macro-averaged kappa per category) to identify which specific categories are problematic.

Out[20]:
Visualization
Dual-axis plot showing Cohen kappa falling and chance agreement rising as positive-class prevalence increases, with constant 90 percent raw agreement.
Illustration of the prevalence paradox: as positive-class prevalence increases from 0.5 to 0.95, Cohen''s kappa (red, left axis) declines sharply even when raw observed agreement is held constant at 90 percent. The blue dashed line (right axis) shows how chance agreement ($P_e$) rises simultaneously, squeezing the denominator of the kappa formula. When $P_e$ approaches $P_o$, kappa approaches zero or becomes negative despite high raw accuracy.

The Bias Problem

Cohen's kappa assumes fixed marginals for each rater. When raters have different base rates (one is more lenient than another), kappa paradoxically decreases compared to Scott's pi, which pools marginals. Byrt, Bishop, and Carlin (1993) formalized this as the "bias" component of kappa's limitations: if rater A uses "positive" 80% of the time but rater B uses it 40% of the time, the pooled marginal approach assumes an intermediate rate, while Cohen's kappa computes chance agreement using these divergent marginals separately, creating a lower expected agreement.

Fleiss' kappa and Krippendorff's alpha use pooled marginals, effectively treating rater differences as noise to be averaged out. This makes them more stable but less sensitive to systematic rater biases. The choice between coefficients depends on whether you view rater differences as meaningful signal (annotator A is systematically more strict) or measurement error (raters should be interchangeable). In practice, when two specific experts consistently differ in their labeling tendencies, Cohen's kappa better captures the challenge of reconciling their perspectives.

Limits of Categorical Agreement

Standard kappa statistics assume categorical data. Modern NLP increasingly uses continuous ratings (quality scores from 1-5) or rankings (preference pairs). While Krippendorff's alpha handles ordinal and interval scales through distance weighting, the interpretation becomes complex. A disagreement between ratings of 1 and 2 on a 5-point scale might mean something different than 4 vs 5, even if the numerical difference is identical: the bottom of the scale might be harder to distinguish perceptually than the middle.

For preference evaluation, which we will cover in the next chapter, we often use ranking agreement metrics like Kendall's τ\tau or the Bradley-Terry model discussed in Part XXXVII: Alignment and RLHF. These capture the relative nature of preferences better than categorical agreement. A rater who consistently ranks response A above response B agrees with another rater who does the same, even if they would assign different absolute quality scores to each response.

Agreement vs. Validity

High inter-annotator agreement does not guarantee validity. Annotators might agree consistently while being consistently wrong relative to some objective standard, or they might agree on superficial features while missing the deeper linguistic phenomenon we care about. Agreement validates the reproducibility of the measurement, not the correctness of the construct.

This distinction matters when using human judgments as training data for LLMs. In Part XXXVI: Instruction Tuning, we saw that instruction-following datasets rely on human judgments of response quality. High IAA ensures consistency in those judgments, but if the annotators share systematic biases (cultural, temporal, or ideological), the resulting model inherits those biases. A group of annotators who all come from the same demographic background might agree strongly with each other while being unrepresentative of the broader population the model serves. We will explore bias measurement and mitigation in Part LVIII: Bias and Fairness.

The appropriate framework is to view IAA as necessary but not sufficient for dataset quality. You need agreement to ensure the labels are consistent, and you need validity studies (correlation with external criteria, expert review, or task performance) to ensure the labels measure what you intend.

Practical Thresholds

Despite the Landis and Koch benchmarks, context determines what constitutes "good" agreement. For objective tasks like part-of-speech tagging on news text, κ>0.8\kappa > 0.8 is achievable and expected. For subjective tasks like sentiment analysis of tweets or assessing whether an LLM response "shows empathy," κ=0.6\kappa = 0.6 might represent excellent agreement given the inherent ambiguity.

Some practitioners use κ=0.67\kappa = 0.67 as a minimum threshold for reliable data, while κ>0.8\kappa > 0.8 indicates data suitable for algorithmic training without additional review. Values below 0.4 suggest the annotation scheme needs substantial revision before proceeding to large-scale labeling. These are starting points, not rules: always justify your threshold relative to your task requirements and your downstream use case.

When reporting IAA in a paper or technical report, always include the raw agreement PoP_o alongside the chance-corrected coefficient, note the number of raters and items, report how missing data was handled if applicable, and indicate which specific variant of each metric you computed (standard vs weighted kappa, nominal vs ordinal alpha). This transparency allows readers to assess the quality of the annotation even if they would have chosen different thresholds.

Summary

Inter-annotator agreement turns the vague notion of "label quality" into quantifiable, comparable statistics that can be tracked across annotation rounds, compared across datasets, and used to make principled decisions about when data is ready for use. We have examined three complementary approaches, each suited to different annotation scenarios:

  • Cohen's kappa measures agreement between two specific raters with potentially different biases, which makes it ideal for validating a primary annotator against an expert reviewer or for measuring consistency between two automated systems. It preserves rater-specific marginal distributions and is the standard choice when rater identity matters.
  • Fleiss' kappa generalizes to any number of raters treated as interchangeable samples from a population, suitable for crowdsourcing scenarios where you care about the pool as a whole rather than individual workers. It uses pooled marginals and equals Scott's pi in the two-rater case.
  • Krippendorff's alpha provides the most general framework, handling missing data, varying numbers of raters per item, and different measurement scales through configurable distance metrics. It is the preferred choice for complex annotation schemes with incomplete data or ordinal ratings.

All three coefficients share the chance-correction framework, comparing observed agreement against expected random agreement. This correction prevents inflated scores on imbalanced datasets but introduces sensitivity to prevalence and marginal distributions. The prevalence paradox and the bias problem are the most common pitfalls in practice, and both can lead to misleading conclusions if you report kappa values without understanding what drives them.

When agreement is low, we have options beyond simple majority voting. Soft labels preserve the uncertainty inherent in borderline cases. Iterative guideline refinement addresses systematic disagreement by clarifying ambiguous decision boundaries. Disagreement-weighted training reduces the influence of unreliable labels. The appropriate strategy depends on whether disagreement represents noise to be eliminated or signal to be preserved.

As we move toward evaluation methodologies that use LLMs as judges, the principles of chance-corrected agreement remain central. Validating an automated judge requires measuring its agreement with human judgments using the same statistical rigor we apply to human annotators. The next chapter on Preference Evaluation will explore how these agreement metrics apply specifically to the pairwise and ranking judgments that drive modern alignment techniques like RLHF and DPO, and how to design preference annotation pipelines that achieve both high IAA and valid coverage of the preference space.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about inter-annotator agreement and chance-corrected reliability coefficients.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026cohenfleiss, author = {Michael Brenndoerfer}, title = {Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation. Retrieved from https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability
MLAAcademic
Michael Brenndoerfer. "Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability>.
CHICAGOAcademic
Michael Brenndoerfer. "Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation'. Available at: https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Cohen, Fleiss & Krippendorff: IAA Metrics & Implementation. https://mbrenndoerfer.com/writing/inter-annotator-agreement-kappa-alpha-reliability

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.