BLEU Score: Evaluating Translation Quality with N-grams

Michael BrenndoerferFebruary 22, 202645 min read

Part of Language AI Handbook

Explains how the BLEU score evaluates machine translation quality using modified n-gram precision and brevity penalties.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

BLEU Score

How do you measure the quality of a machine translation when there are many correct ways to translate the same sentence? Unlike classification tasks with a single right answer, text generation produces open-ended outputs where multiple phrasings can be equally valid. This fundamental challenge plagued early research in statistical machine translation during the 1990s. Researchers found themselves trapped in a difficult methodological loop: to improve translation systems, they needed to evaluate them, but human evaluation required fluent bilingual speakers, cost significant money, and took days or weeks to complete. This bottleneck severely limited experimental iteration, as researchers could only test a handful of configurations per study.

The Bilingual Evaluation Understudy (BLEU) score, introduced by Papineni et al. in 2002, emerged as the solution that dominated the field for two decades. BLEU approaches evaluation through a deceptively simple intuition: good translations share words and phrases with professional human translations. By counting matching n-grams between a candidate translation and one or more reference translations, BLEU provides a numerical score that correlates reasonably well with human judgment while remaining computationally cheap. The brilliance of BLEU lies in its recognition that while we cannot easily quantify "good translation" abstractly, we can operationalize the concept through surface-level lexical overlap with high-quality reference translations.

BLEU combines two core components: modified n-gram precision to measure accuracy, and a brevity penalty to discourage overly short translations. As we discussed in Part II, Chapter 2: N-grams, n-grams capture local word order and collocation patterns. BLEU extends this idea by comparing n-grams between machine output and human references, checking unigram overlap along with bigrams, trigrams, and 4-grams to ensure fluency and grammatical coherence. This multi-scale approach acknowledges that translation quality operates at multiple levels: individual word choice matters, but so does the proper pairing of adjacent words and the maintenance of longer phrase structures.

Historical Context and Motivation

To appreciate why BLEU mattered so much when it appeared, you need to understand the state of machine translation evaluation in the late 1990s. Statistical machine translation (SMT) systems were rapidly improving, but the research community lacked a shared yardstick for measuring progress. Different research groups used different evaluation protocols, sometimes reporting results on proprietary test sets, making it nearly impossible to compare systems across institutions or determine which architectural choice drove improvements.

Human evaluation, when it did occur, was carried out in varying formats: adequacy ratings on a 5-point scale, fluency ratings on a separate scale, or preference judgments where annotators chose between two system outputs. These formats did not always agree with each other, and even the same format produced highly variable results depending on annotator selection, briefing procedures, and fatigue effects. A five-point adequacy scale used in one laboratory might bear little resemblance to what another laboratory considered a "4 out of 5" translation. Interlabeler agreement was poor, reproducibility was low, and the community had no way to determine whether a new method was better or simply happened to be evaluated more favorably.

Papineni and colleagues were working at IBM Research, at that time one of the centers of statistical machine translation research. Their goal was pragmatic: they needed to run dozens of experiments per day, comparing variants of word alignment models, language models, and decoding algorithms. Human evaluation at that scale was economically impossible. They set out to find an automatic metric that would correlate with human judgment well enough to serve as a reliable proxy during development, even if it was not perfect.

The key insight in the BLEU paper was methodological as much as technical. Rather than trying to build a perfect evaluation metric, Papineni et al. asked a more tractable question: can we build a metric that, when comparing two systems on the same test set, almost always agrees with human evaluators about which system is better? System-level ranking turns out to be far easier than absolute quality assessment. This design decision, targeting ranking agreement rather than absolute correlation, explains why BLEU works reasonably well for its intended purpose while failing badly in many situations where absolute quality matters.

The paper demonstrated that BLEU correlated highly with human judgments across multiple language pairs and test sets, reporting rank correlation coefficients above 0.99 for system-level comparisons. This remarkable number convinced the community to adopt the metric, and within a few years, reporting BLEU scores became essentially mandatory for machine translation publications. Shared evaluation tasks like NIST and later WMT (Workshop on Machine Translation) standardized on BLEU, creating the competitive benchmarking infrastructure that drove rapid progress through the mid-2000s.

Modified N-gram Precision

The foundation of BLEU lies in precision calculated over n-grams of varying lengths. However, naive precision counting fails catastrophically for translation evaluation. Consider a candidate translation "the the the the the" matched against a reference "The cat sat on the mat." A simple unigram precision calculation would yield 5/5 = 100% because every word in the candidate appears in the reference, despite the candidate being complete gibberish. This pathology demonstrates why standard precision from information retrieval is inadequate for generation tasks: it assumes that the presence of correct elements constitutes success, without considering whether the output is coherent or exploits the reference through repetition.

Modified Precision

BLEU uses modified n-gram precision, which clips the count of each n-gram in the candidate to its maximum count in any single reference translation. This prevents the exploitation of repeated words and ensures precision reflects meaningful matches.

The modification works as follows. For each n-gram type appearing in the candidate translation, we count how many times it appears in the candidate (the candidate count), then find the maximum number of times it appears in any single reference (the reference count). The clipped count is the minimum of these two values. We sum these clipped counts for all n-grams in the candidate and divide by the total number of n-grams in the candidate. This clipping mechanism is a critical safeguard against degenerate outputs. This keeps systems cannot game the metric by repeating high-frequency words like "the" or "and" to inflate their scores.

To see exactly how clipping works, consider the adversarial candidate "the the the the the" against the reference "The cat sat on the mat." The unigram "the" appears five times in the candidate (candidate count = 5), but only once in the reference (reference count = 1). After clipping, the count for "the" becomes min(5, 1) = 1. Since there are five unigrams in the candidate but only one receives credit, the clipped precision is 1/5 = 0.20. The metric correctly identifies this as a low-quality output, even though naive precision would have scored it perfectly.

Mathematically, for n-grams of order nn, the modified precision pnp_n aggregates clipped counts across all candidate sentences and divides by the total number of n-grams:

pn=∑C∈Candidates∑n-gram∈CCountclip(n-gram)∑C∈Candidates∑n-gram∈CCount(n-gram)p_n = \frac{\sum_{C \in \text{Candidates}} \sum_{n\text{-gram} \in C} \text{Count}_{\text{clip}}(n\text{-gram})}{\sum_{C \in \text{Candidates}} \sum_{n\text{-gram} \in C} \text{Count}(n\text{-gram})}

where:

  • pnp_n: the modified precision for n-grams of order nn (e.g., p1p_1 for unigrams, p2p_2 for bigrams)
  • CC: an individual candidate translation in the corpus
  • Candidates\text{Candidates}: the set of all candidate translations being evaluated
  • n-gramn\text{-gram}: a contiguous sequence of nn words from the candidate
  • Count(n-gram)\text{Count}(n\text{-gram}): the raw count of the n-gram appearing in candidate CC
  • Countclip(n-gram)\text{Count}_{\text{clip}}(n\text{-gram}): the clipped count, defined as min⁡(Count(n-gram),max reference count)\min(\text{Count}(n\text{-gram}), \text{max reference count})
  • max reference count\text{max reference count}: the maximum number of times the n-gram appears in any single reference translation

The clipping mechanism ensures that repeated words in the candidate cannot inflate the score beyond what is supported by the references. By taking the minimum between the candidate count and the maximum reference count, we count only as many matches as exist in at least one reference. This conservative approach reflects a fundamental principle of BLEU: a candidate should only receive credit for translation content that is attested in the reference set, preventing hallucination or over-generation from being rewarded.

This modified precision is computed separately for each n-gram order (typically unigrams through 4-grams), then combined geometrically with equal weighting. The choice of 4-grams as the upper bound reflects empirical findings that longer n-grams capture essential fluency patterns while remaining robust to the inevitable variations in phrase ordering that occur across valid translations. Unigrams capture adequacy (are the right concepts present?), while 4-grams capture fluency (do the words flow together naturally?).

Why Multiple N-gram Orders?

The decision to combine precision across four n-gram orders is not arbitrary; each order captures a qualitatively different aspect of translation quality. Understanding what each order measures helps you interpret BLEU scores and diagnose where a system falls short.

Unigram precision (p1p_1) measures vocabulary overlap: does the candidate use the right words? High unigram precision means the translation selected appropriate lexical items, even if they appear in the wrong order. A translation that scrambles every word but picks exactly the right vocabulary will score well on unigrams. Unigrams measure adequacy.

Bigram precision (p2p_2) measures the correctness of adjacent word pairs. This is where basic grammatical patterns emerge. Articles must precede their nouns ("the cat" vs. "cat the"), auxiliaries must precede their main verbs ("is running" vs. "running is"), and prepositions must follow the correct head words. A candidate can achieve high unigram precision but low bigram precision by selecting all the right words but arranging them incorrectly at the local level.

Trigram precision (p3p_3) captures short phrasal patterns: determiner-adjective-noun sequences ("the big cat"), verb-object-preposition structures ("sitting on the"), and similar three-word collocations that characterize fluent text. Languages have strong statistical preferences for particular three-word sequences; correct trigrams indicate the translation follows the syntactic and collocational patterns of the target language.

4-gram precision (p4p_4) captures longer phrasal consistency and is the most demanding requirement. Four-word matches require a phrase that is lexically correct and grammatically arranged in the specific order found in the references. Low 4-gram precision is a strong signal of non-fluency, even when the vocabulary selection is reasonable.

The geometric mean of these four values combines evidence across n-gram orders. A system that achieves perfect unigram precision but near-zero 4-gram precision is essentially a bag-of-words generator that ignores word order, and the geometric mean correctly assigns it a very low score. The exponential combination ensures that weakness at any n-gram order cannot be fully compensated by strength elsewhere.

Out[3]:
Visualization
Line chart with three lines showing modified precision (y-axis) vs n-gram order 1 through 4 (x-axis) for perfect, partial, and poor translations.
Precision values by n-gram order for three translation quality levels. Perfect translations maintain high precision across all orders, partially correct translations show declining precision for longer n-grams, and poor translations drop sharply after unigrams. The multiplicative effect of the geometric mean means that weakness at any order strongly penalizes the final score.

The Brevity Penalty

Modified n-gram precision alone creates a perverse incentive: shorter translations achieve higher precision scores. A candidate consisting of a single word that appears in the reference would achieve perfect precision, yet conveys almost no information. This problem, known as the "sparse translation" or "under-translation" issue, plagued early statistical systems that learned to minimize risk by producing short, safe outputs. To counteract this strategic shortening, BLEU introduces a brevity penalty (BP) that penalizes translations shorter than their references.

The brevity penalty compares the candidate length cc against the effective reference length rr. The effective reference length is determined by finding, for each candidate sentence, the reference length closest to the candidate length (in absolute difference). For corpus-level BLEU, these lengths are summed across all sentences before computing the ratio. This "closest match" approach prevents the penalty from being overly harsh when multiple references exist with varying lengths, allowing the candidate to align with the most comparable reference length.

BP={1if c>re(1−r/c)if c≤rBP = \begin{cases} 1 & \text{if } c > r \\ e^{(1 - r/c)} & \text{if } c \leq r \end{cases}

where:

  • BPBP: the brevity penalty factor (ranging from 0 to 1)
  • cc: the total length of the candidate translation (in words)
  • rr: the effective reference length, defined as the length of the reference closest to cc in absolute difference

When the candidate is longer than the reference (c>rc > r), no penalty is applied (BP=1BP = 1). This reflects the asymmetry in translation evaluation: while omitting content is a serious error, adding extra information (over-translation) is already penalized by the precision component, which will not match n-grams beyond what exists in the reference. When the candidate is shorter, the penalty grows exponentially as the ratio r/cr/c increases. A candidate half the length of the reference receives a penalty of e−1≈0.37e^{-1} \approx 0.37, while a candidate one-tenth the length receives e−9≈0.0001e^{-9} \approx 0.0001, effectively zeroing out the score. The exponential formulation ensures that the penalty accelerates rapidly as translations become dangerously short, protecting against systems that might attempt to cheat by producing single-word outputs.

Notice the deliberate asymmetry in the penalty design. Over-translation (producing a longer output) does not receive an explicit length penalty, only an implicit precision penalty from n-grams in the candidate that do not match anything in the reference. Under-translation (producing a shorter output) receives both the precision penalty for any missed content and the explicit brevity penalty for the length difference. This asymmetry reflects the original designers' view that under-translation is the more dangerous pathology for statistical systems. Neural systems exhibit the opposite tendency, often generating long, repetitive outputs, which is why modern neural MT evaluation sometimes augments BLEU with explicit penalties for over-generation.

Out[4]:
Visualization
Line chart showing the BLEU brevity penalty from approximately 0 to 1.1 against the candidate-to-reference length ratio from 0 to 2, with callouts at ratios 0.1 and 0.5.
Brevity penalty as a function of the candidate-to-reference length ratio. The penalty decays exponentially when the candidate is shorter than the reference (c/r < 1) and remains at 1.0 when the candidate is longer. The exponential form means that even mild under-translation at 90% of reference length receives a modest 10% penalty, but severe under-translation at 50% length cuts the score by 63%.

The Complete BLEU Formula

BLEU combines the brevity penalty with the geometric mean of modified n-gram precisions across different orders. The standard formulation uses 4-gram precision (weights for unigrams, bigrams, trigrams, and 4-grams), though variations exist for specific use cases. The geometric mean is chosen over the arithmetic mean because it enforces a stricter requirement: systems must perform well across all n-gram orders to achieve a high score, rather than compensating for poor performance in one order with excellence in another.

BLEU=BP⋅exp⁡(∑n=1Nwnlog⁡pn)\text{BLEU} = BP \cdot \exp\left(\sum_{n=1}^N w_n \log p_n\right)

where:

  • BLEU\text{BLEU}: the final BLEU score (between 0 and 1)
  • BPBP: the brevity penalty factor
  • NN: the maximum n-gram order (typically 4)
  • wnw_n: the weight for n-grams of order nn (typically uniform: wn=1/Nw_n = 1/N)
  • pnp_n: the modified precision for n-grams of order nn
  • exp⁡\exp: the exponential function, which converts the weighted log average back to a linear scale
  • ∑n=1N\sum_{n=1}^N: summation over all n-gram orders from 1 to NN

The formula combines the precisions using a geometric mean (via the logarithmic formulation), which ensures that all n-gram orders must perform well to achieve a high score. If any pnp_n is zero, the logarithm approaches negative infinity, making the overall score zero. In practice, smoothing is applied to avoid this edge case. The weights wnw_n allow adjusting the relative importance of different n-gram orders, though uniform weighting is standard. Some researchers adjust these weights to emphasize fluency (higher weights on 3-grams and 4-grams) or adequacy (higher weights on unigrams and bigrams) depending on their specific evaluation priorities.

The logarithmic formulation also provides a useful mathematical property: adding scores across n-gram orders in log space is equivalent to multiplying them in probability space. This means the final BLEU score is proportional to the product of all four precision values (scaled by the brevity penalty), not their sum. A translation achieving p1=p2=p3=p4=0.8p_1 = p_2 = p_3 = p_4 = 0.8 would yield a geometric mean of 0.8, while a translation achieving p1=1.0p_1 = 1.0, p2=1.0p_2 = 1.0, p3=1.0p_3 = 1.0, p4=0p_4 = 0 would score exactly 0 because the geometric mean collapses when any factor is zero. This strictness reflects the belief that good translations must demonstrate fluency at all n-gram levels, not just word matching.

BLEU-N Variants

The choice of maximum n-gram order NN defines different BLEU variants. The standard BLEU-4 (also written BLEU) uses N=4N=4 with uniform weights. Other variants occasionally appear in the literature:

BLEU-1 uses only unigram precision and is equivalent to modified unigram precision with a brevity penalty. It measures adequacy without any fluency constraint and is sometimes used to evaluate lexical coverage in tasks where word order flexibility is high, such as information extraction or certain forms of text simplification.

BLEU-2 adds bigram precision, capturing basic grammatical dependencies. Some summarization papers use BLEU-2 because summaries may legitimately reorder content, and 4-gram precision can be overly harsh when paraphrasing is encouraged.

NIST is a variant developed by the National Institute of Standards and Technology that weights n-gram matches by their information content. Common n-grams like "the cat" receive lower weight than rare n-grams like "photosynthetic apparatus". This weighting principle is more linguistically motivated than uniform weights, as correctly translating a rare technical term deserves more credit than matching a common function word. NIST also uses a different length penalty formulation.

For most machine translation benchmarks, BLEU-4 remains the default, and when researchers write "BLEU" without qualification they almost always mean BLEU-4 with corpus-level aggregation.

Corpus vs. Sentence-Level BLEU

BLEU operates at two distinct granularities with important practical differences that researchers must understand to avoid common evaluation pitfalls.

Sentence-level BLEU computes scores on individual sentence pairs. This provides fine-grained feedback useful for analyzing specific errors, but suffers from extreme volatility. A single sentence might contain no 4-grams matching the reference, yielding a zero score despite being partially correct. This "zero problem" occurs frequently because higher-order n-grams are sparse; a 10-word sentence contains only seven 4-grams, and if none happen to match the specific reference phrasing, the geometric mean collapses to zero. Sentence-level scores are also highly sensitive to reference diversity; with only one or two reference translations, legitimate variations not captured in the references receive unfairly low scores.

The zero problem is not a minor edge case. For a 10-word candidate translated into English, the probability of at least one 4-gram matching a single reference by chance is quite low, even for a good translation. This means that sentence-level BLEU-4 will report zero for many individually acceptable sentences. Researchers who aggregate sentence-level BLEU scores by averaging them often find the aggregate is dramatically lower than corpus-level BLEU on the same test set, and the discrepancy grows with shorter sentences.

Corpus-level BLEU aggregates statistics across entire test sets before computing precision. Counts of matching n-grams are summed across all sentences, as are total n-gram counts and lengths. This aggregation provides more stable, reliable scores because zeros in individual sentences are overwhelmed by matches elsewhere. Research consistently shows that corpus-level BLEU correlates significantly better with human judgment than sentence-level BLEU, with correlation coefficients often doubling or tripling when moving from sentence to corpus aggregation.

The distinction matters enormously in practice. Papineni et al.'s original paper validated BLEU at the corpus level, and the high correlation with human judgment that made the metric famous was measured at that level. When researchers later attempted to use sentence-level BLEU for tasks like minimum risk training or reward shaping in reinforcement learning, they encountered the zero-score problem and had to develop smoothing techniques to make sentence-level scoring viable.

The choice between them depends on your goal. Use corpus-level BLEU for comparing systems or tracking overall progress, as it provides the statistical stability necessary for reliable comparisons. Use sentence-level BLEU cautiously for error analysis, and consider smoothing techniques (adding small constants to zero counts) to avoid zero scores on short sentences. When using sentence-level BLEU for training objectives or reinforcement learning, specialized smoothing methods such as Lin smoothing or the addition of epsilon values become essential to maintain gradient flow.

Worked Example

Let us walk through a concrete example to see how these components interact. Consider translating a German sentence into English, where the German original might have been "Die Katze saß auf der Matte."

Candidate: "The cat sat on the mat" Reference 1: "The cat sat on the mat" Reference 2: "The cat was sitting on the mat"

First, we calculate modified n-gram precisions.

  • Unigrams (p1p_1): Every word in the candidate appears in at least one reference. "The" appears twice in the candidate and twice in Reference 1, so its clipped count is 2. All other unigrams appear once in both candidate and Reference 1. Total clipped unigrams: 6. Total unigrams: 6. p1=6/6=1.0p_1 = 6/6 = 1.0.

  • Bigrams (p2p_2): The candidate contains "The cat", "cat sat", "sat on", "on the", "the mat". All five appear in Reference 1 exactly as written. p2=5/5=1.0p_2 = 5/5 = 1.0.

  • Trigrams (p3p_3): "The cat sat", "cat sat on", "sat on the", "on the mat". All four appear in Reference 1. p3=4/4=1.0p_3 = 4/4 = 1.0.

  • 4-grams (p4p_4): "The cat sat on", "cat sat on the", "sat on the mat". All three appear in Reference 1. p4=3/3=1.0p_4 = 3/3 = 1.0.

Since the candidate exactly matches Reference 1, all precisions are perfect. The candidate length c=6c = 6 and the closest reference length is also 6, so r=6r = 6. The brevity penalty is 1 (since c=rc = r).

BLEU=BP⋅exp⁡(∑n=14wnlog⁡pn)=1.0⋅exp⁡(0.25⋅log⁡1+0.25⋅log⁡1+0.25⋅log⁡1+0.25⋅log⁡1)(substituting values)=1.0⋅exp⁡(0.25⋅(0+0+0+0))(since log⁡1=0)=1.0⋅exp⁡(0)=1.0⋅1=1.0\begin{aligned} \text{BLEU} &= BP \cdot \exp\left(\sum_{n=1}^4 w_n \log p_n\right) \\ &= 1.0 \cdot \exp(0.25 \cdot \log 1 + 0.25 \cdot \log 1 + 0.25 \cdot \log 1 + 0.25 \cdot \log 1) && \text{(substituting values)} \\ &= 1.0 \cdot \exp(0.25 \cdot (0 + 0 + 0 + 0)) && \text{(since } \log 1 = 0\text{)} \\ &= 1.0 \cdot \exp(0) \\ &= 1.0 \cdot 1 \\ &= 1.0 \end{aligned}

Now consider a flawed candidate: "The cat sat on mat" (missing "the" before "mat").

  • Unigrams: All five words appear in references. p1=5/5=1.0p_1 = 5/5 = 1.0.

  • Bigrams: Candidate bigrams are [("The", "cat"), ("cat", "sat"), ("sat", "on"), ("on", "mat")]. Reference 1 bigrams include ("on", "the") and ("the", "mat") but not ("on", "mat"). So 3 out of 4 bigrams match. p2=3/4=0.75p_2 = 3/4 = 0.75.

  • Trigrams: Candidate trigrams are [("The", "cat", "sat"), ("cat", "sat", "on"), ("sat", "on", "mat")]. Reference 1 contains the first two but not ("sat", "on", "mat"). p3=2/3≈0.667p_3 = 2/3 \approx 0.667.

  • 4-grams: Candidate 4-grams are [("The", "cat", "sat", "on"), ("cat", "sat", "on", "mat")]. Reference 1 contains the first but not the second. p4=1/2=0.5p_4 = 1/2 = 0.5.

Brevity penalty: candidate length 5, closest reference length 6. Since c<rc < r, BP=e(1−6/5)=e−0.2≈0.819BP = e^{(1-6/5)} = e^{-0.2} \approx 0.819.

BLEU=BP⋅exp⁡(∑n=14wnlog⁡pn)=0.819⋅exp⁡(0.25⋅log⁡1+0.25⋅log⁡0.75+0.25⋅log⁡0.667+0.25⋅log⁡0.5)(substituting precisions)≈0.819⋅exp⁡(0.25⋅(0−0.288−0.405−0.693))(evaluating logarithms)≈0.819⋅exp⁡(−0.346)(summing and multiplying)≈0.819⋅0.708(exponentiating)≈0.58\begin{aligned} \text{BLEU} &= BP \cdot \exp\left(\sum_{n=1}^4 w_n \log p_n\right) \\ &= 0.819 \cdot \exp(0.25 \cdot \log 1 + 0.25 \cdot \log 0.75 + 0.25 \cdot \log 0.667 + 0.25 \cdot \log 0.5) && \text{(substituting precisions)} \\ &\approx 0.819 \cdot \exp(0.25 \cdot (0 - 0.288 - 0.405 - 0.693)) && \text{(evaluating logarithms)} \\ &\approx 0.819 \cdot \exp(-0.346) && \text{(summing and multiplying)} \\ &\approx 0.819 \cdot 0.708 && \text{(exponentiating)} \\ &\approx 0.58 \end{aligned}

This demonstrates how missing a single word cascades through higher-order n-grams, dramatically reducing the score. The error in the bigram "on mat" propagates into the trigram "sat on mat" and the 4-gram "cat sat on mat", creating a compounding penalty that reflects the grammatical incoherence of the missing article. This sensitivity to local context explains why BLEU effectively distinguishes between fluent and disfluent translations, even when the underlying vocabulary overlap remains high.

Edge Cases in the Worked Example

The example above uses a clean, well-formed sentence, but real-world BLEU computation encounters several important edge cases worth understanding.

Multiple references with different lengths. When Reference 2 ("The cat was sitting on the mat", length 7) is also considered, the closest reference to candidate length 5 is Reference 1 (length 6, difference 1), not Reference 2 (length 7, difference 2). The effective reference length r=6r = 6 even though a shorter reference would reduce the brevity penalty. This "closest match" rule is important; an implementation that always uses the shortest reference would be more lenient, while one that always uses the longest would be more strict. The closest-match rule is the standard per the original paper.

The zero precision problem. If the candidate contained a 4-gram that appeared in neither reference, and no other 4-grams matched either, p4=0p_4 = 0 and the logarithm is undefined. The standard approach is to set the BLEU score to 0 when any precision is 0, which matches the logarithmic formulation's behavior as pn→0p_n \to 0. For sentence-level BLEU, this happens frequently for short sentences and is why smoothing was developed.

Case sensitivity. The original BLEU paper lowercases all tokens before comparison. This is the standard convention, as "The" and "the" should be considered the same word for evaluation purposes. Some implementations preserve case, which can lead to slightly different scores and makes comparisons across tools unreliable. Always verify that the tool you are using matches the case normalization expected by the benchmark you are comparing against.

Code Implementation

Let us implement BLEU from scratch to understand the mechanics, then compare with standard libraries. We will calculate both sentence-level and corpus-level scores. Implementing the metric manually illuminates the importance of edge cases such as empty strings, varying reference counts, and the numerical stability of logarithmic calculations.

First, we need utilities for extracting n-grams and computing clipped counts.

In[5]:
Code
from collections import Counter
from typing import List, Tuple


def get_ngrams(tokens: List[str], n: int) -> List[Tuple[str, ...]]:
    """Extract n-grams from a list of tokens."""
    return [tuple(tokens[i : i + n]) for i in range(len(tokens) - n + 1)]


def clipped_count(
    candidate_ngrams: Counter, reference_ngrams_list: List[Counter]
) -> int:
    """Sum of clipped counts for candidate n-grams against multiple references."""
    total = 0
    for ngram, count in candidate_ngrams.items():
        # Max count of this n-gram in any single reference
        max_ref_count = max(ref.get(ngram, 0) for ref in reference_ngrams_list)
        total += min(count, max_ref_count)
    return total


def modified_precision(
    candidate: List[str], references: List[List[str]], n: int
) -> float:
    """Calculate modified n-gram precision."""
    candidate_ngrams = Counter(get_ngrams(candidate, n))
    if not candidate_ngrams:
        return 0.0

    reference_counters = [Counter(get_ngrams(ref, n)) for ref in references]

    clipped = clipped_count(candidate_ngrams, reference_counters)
    total = sum(candidate_ngrams.values())

    return clipped / total if total > 0 else 0.0

Now we implement the brevity penalty and the full BLEU calculation.

In[6]:
Code
import math


def brevity_penalty(candidate_len: int, reference_len: int) -> float:
    """Calculate BP with effective reference length."""
    if candidate_len > reference_len:
        return 1.0
    else:
        return math.exp(1 - reference_len / candidate_len)


def closest_reference_length(
    candidate: List[str], references: List[List[str]]
) -> int:
    """Find the reference length closest to candidate length."""
    c_len = len(candidate)
    ref_lengths = [len(ref) for ref in references]

    # Find reference with length closest to candidate
    closest = min(ref_lengths, key=lambda r: abs(r - c_len))
    return closest


def sentence_bleu(
    candidate: List[str],
    references: List[List[str]],
    weights: List[float] = None,
    smoothing: float = 1e-10,
) -> float:
    """Calculate sentence-level BLEU."""
    if weights is None:
        weights = [0.25, 0.25, 0.25, 0.25]  # Uniform 1-4 gram weights

    max_n = len(weights)

    # Calculate modified precisions for each n-gram order
    precisions = []
    for n in range(1, max_n + 1):
        p = modified_precision(candidate, references, n)
        # Smoothing: avoid log(0)
        if p == 0:
            p = smoothing
        precisions.append(p)

    # Geometric mean of precisions
    log_sum = sum(w * math.log(p) for w, p in zip(weights, precisions))
    geo_mean = math.exp(log_sum)

    # Brevity penalty
    c_len = len(candidate)
    r_len = closest_reference_length(candidate, references)
    bp = brevity_penalty(c_len, r_len)

    return bp * geo_mean

Let us test this with our earlier examples.

In[7]:
Code
# Example 1: Perfect match
candidate1 = ["the", "cat", "sat", "on", "the", "mat"]
reference1 = ["the", "cat", "sat", "on", "the", "mat"]
reference2 = ["the", "cat", "was", "sitting", "on", "the", "mat"]

score1 = sentence_bleu(candidate1, [reference1, reference2])

# Example 2: Missing word
candidate2 = ["the", "cat", "sat", "on", "mat"]
score2 = sentence_bleu(candidate2, [reference1, reference2])

# Example 3: Word substitution (synonym)
candidate3 = ["the", "feline", "sat", "on", "the", "mat"]
score3 = sentence_bleu(candidate3, [reference1, reference2])

# Example 4: Adversarial repetition
candidate4 = ["the", "the", "the", "the", "the", "the"]
score4 = sentence_bleu(candidate4, [reference1, reference2])
Out[8]:
Console
Perfect match BLEU:          1.0000
Missing 'the' BLEU:          0.5789
Synonym 'feline' BLEU:       0.5373
Adversarial repetition BLEU: 0.0000

The perfect match achieves a score near 1.0, while the missing word and synonym substitutions cause significant drops. Notice that "feline" receives a harsh penalty despite being semantically equivalent to "cat". This highlights a fundamental limitation we will discuss shortly: BLEU operates on string matching, not semantic equivalence, making it blind to paraphrase and synonymy. The adversarial repetition example shows the clipping mechanism at work: despite repeating the most common word, the score is very low because the clipped count for "the" is capped at its maximum reference occurrence.

For corpus-level BLEU, we aggregate counts across all sentences before calculating precisions. This avoids the volatility of sentence-level scoring and provides the stability necessary for reliable system comparison.

In[9]:
Code
def corpus_bleu(
    candidates: List[List[str]],
    references_list: List[List[List[str]]],
    weights: List[float] = None,
) -> float:
    """Calculate corpus-level BLEU by aggregating counts."""
    if weights is None:
        weights = [0.25, 0.25, 0.25, 0.25]

    max_n = len(weights)
    total_candidate_len = 0
    total_reference_len = 0

    # Aggregate counts for each n-gram order
    clipped_counts = [0] * max_n
    total_counts = [0] * max_n

    for candidate, references in zip(candidates, references_list):
        total_candidate_len += len(candidate)
        total_reference_len += closest_reference_length(candidate, references)

        for n in range(1, max_n + 1):
            cand_ngrams = Counter(get_ngrams(candidate, n))
            ref_counters = [Counter(get_ngrams(ref, n)) for ref in references]

            clipped_counts[n - 1] += clipped_count(cand_ngrams, ref_counters)
            total_counts[n - 1] += sum(cand_ngrams.values())

    # Calculate precisions
    precisions = []
    for clipped, total in zip(clipped_counts, total_counts):
        if total == 0:
            precisions.append(0.0)
        else:
            precisions.append(clipped / total)

    # Geometric mean
    log_sum = sum(
        w * math.log(p) if p > 0 else float("-inf")
        for w, p in zip(weights, precisions)
    )
    if log_sum == float("-inf"):
        geo_mean = 0.0
    else:
        geo_mean = math.exp(log_sum)

    # Brevity penalty on corpus lengths
    bp = brevity_penalty(total_candidate_len, total_reference_len)

    return bp * geo_mean

Now we compare sentence-level versus corpus-level scoring on a mini test set.

In[10]:
Code
# Mini corpus with two sentences
candidates = [
    ["the", "cat", "sat", "on", "the", "mat"],
    ["a", "dog", "ran", "quickly"],
]

refs = [
    [
        ["the", "cat", "sat", "on", "the", "mat"],
        ["the", "cat", "was", "on", "the", "mat"],
    ],
    [["a", "dog", "ran", "fast"], ["the", "dog", "ran", "quickly"]],
]

# Sentence-level average
sent_scores = [sentence_bleu(c, r) for c, r in zip(candidates, refs)]
avg_sent = sum(sent_scores) / len(sent_scores)

# Corpus-level
corp_score = corpus_bleu(candidates, refs)
Out[11]:
Console
Sentence-level average: 0.5016
Corpus-level BLEU: 0.9306
Individual sentence scores: ['1.0000', '0.0032']
Out[12]:
Visualization
Bar chart comparing BLEU score values for Sentence 1, Sentence 2, the sentence-level average, and the corpus-level BLEU score.
Comparison of sentence-level BLEU scores versus corpus-level aggregation on a two-sentence test set. Corpus-level BLEU produces a more stable estimate by pooling n-gram statistics before computing precision, which avoids the zero-score problem that affects individual short sentences. The individual sentence scores show the variance inherent in sentence-level evaluation.

Corpus-level BLEU typically differs from the average of sentence-level scores because it aggregates statistics globally. Notice how the second sentence, with its synonym differences ("quickly" vs "fast"), drags down the average more than it might affect corpus-level scores where other matches compensate. This mathematical property makes corpus-level evaluation the gold standard for research publications and system comparisons.

In practice, you should use established libraries like sacrebleu or nltk rather than implementing BLEU yourself, as they handle edge cases (empty strings, smoothing, multiple references) robustly and follow standardized preprocessing conventions that ensure comparability across studies.

In[13]:
Code
import subprocess

# Install sacrebleu for comparison (using uv since pip is not available in venv)
subprocess.check_call(["uv", "pip", "install", "sacrebleu", "-q"])

import sacrebleu

# Ensure required imports are available for the code block

# Prepare strings for sacrebleu (expects detokenized strings)
candidate_strs = ["the cat sat on the mat", "a dog ran quickly"]
reference_strs = [
    ["the cat sat on the mat", "the cat was on the mat"],  # refs for sentence 1
    ["a dog ran fast", "the dog ran quickly"],  # refs for sentence 2
]

# sacrebleu expects references as list of lists (transposed)
refs_formatted = [
    [ref[i] for ref in reference_strs] for i in range(len(reference_strs[0]))
]

bleu = sacrebleu.corpus_bleu(candidate_strs, refs_formatted)
Out[14]:
Console
sacrebleu corpus BLEU: 93.06
Our implementation:    93.06

The sacrebleu score matches our manual corpus calculation when both use the same tokenization and case normalization. Note that sacrebleu reports scores on a 0-100 scale (multiply by 100) by convention, a tradition inherited from earlier machine translation competitions where integer scores were preferred for leaderboard display.

SacreBLEU and Standardization

One of the most persistent practical problems with BLEU is that its score depends heavily on text preprocessing: how text is tokenized, whether it is lowercased, how punctuation is handled, and whether scripts that do not use spaces as word boundaries receive special treatment. Two research groups could legitimately compute BLEU on the same system output using the same references and report scores differing by several points simply because they use different tokenization tools. This comparability problem undermined the very purpose of having a standardized metric.

The sacrebleu library, introduced by Post in 2018, addresses this problem by standardizing the entire BLEU pipeline. It enforces a specific tokenization scheme (based on the Moses tokenizer), provides canonical reference sets for standard benchmarks, and appends a BLEU signature to each score that encodes all the parameters used in computation. A score reported as BLEU+case.mixed+numrefs.1+smooth.exp+tok.13a+version.1.5.1 is fully reproducible by anyone with the same library, references, and candidate translations.

SacreBLEU Signature

SacreBLEU encodes all preprocessing choices in a "signature" appended to each score. When reporting BLEU in a paper, include the full signature so readers can reproduce your results exactly. For example: BLEU+case.mixed+numrefs.1+smooth.exp+tok.13a.

The sacrebleu library also provides access to canonical reference sets for many standard benchmarks (WMT, FLORES, etc.) directly from the library, so you do not need to manage reference files yourself. This dramatically reduces the risk of accidentally using different references than the benchmark intends.

Beyond standardization, sacrebleu introduced several smoothing methods for sentence-level evaluation. The smooth_method='exp' option (exponential smoothing) uses a geometric series to fill in zero counts for higher-order n-grams, preventing zero scores on short sentences while preserving the relative ordering of translations. Let us explore these smoothing methods:

In[15]:
Code
import sacrebleu

# Short sentence that would get zero 4-gram precision without smoothing
short_candidate = "The cat sat"
short_reference = "The cat sat on the mat"

# Without smoothing (BLEU collapses to 0 due to zero 4-gram precision)
# With smoothing (reasonable score preserved)
bleu_no_smooth = sacrebleu.sentence_bleu(
    short_candidate, [short_reference], smooth_method="none"
)
bleu_floor = sacrebleu.sentence_bleu(
    short_candidate, [short_reference], smooth_method="floor"
)
bleu_add_k = sacrebleu.sentence_bleu(
    short_candidate, [short_reference], smooth_method="add-k"
)
bleu_exp = sacrebleu.sentence_bleu(
    short_candidate, [short_reference], smooth_method="exp"
)
Out[16]:
Console
No smoothing:     36.79
Floor smoothing:  36.79
Add-k smoothing:  36.79
Exponential:      36.79

Floor smoothing adds a small constant (typically 0.01) to any zero n-gram precision before computing the logarithm. Add-k smoothing adds kk to both the numerator and denominator of each precision. Exponential smoothing (the sacrebleu default) applies a geometric series to gradually smooth out zero values. Each method makes different assumptions about how a zero n-gram count should be interpreted, and the choice can significantly affect sentence-level scores for short texts.

Key Parameters

The key parameters for BLEU calculation are:

  • weights: Weights for each n-gram order (typically unigrams through 4-grams). Uniform weighting (0.25 each) is standard, but adjusting these allows you to emphasize matching at specific granularities. Higher weights on larger n-grams reward fluency and word order accuracy, while higher weights on unigrams prioritize vocabulary coverage. Some implementations use non-uniform weights such as [0.1, 0.2, 0.3, 0.4] to emphasize longer matches.

  • smoothing: Small constant (default 1e-10) added to zero precision values to prevent log(0) errors, particularly important for sentence-level BLEU on short sentences where higher-order n-grams may not overlap with references. Different smoothing methods exist, including Laplace smoothing (adding 1 to all counts) and Lin smoothing, which becomes more aggressive for higher-order n-grams.

  • max_n: The maximum n-gram order to consider. Standard BLEU uses 4. Reducing to 2 or 3 is sometimes appropriate for tasks where local fluency is expected but longer phrase matches are unreliable due to legitimate translation variation.

  • case: Whether to lowercase all text before comparison. Case normalization is standard for most benchmarks but can matter for proper noun translation quality, where capitalizing the correct entity name might be meaningful.

  • tokenization: How to split text into tokens. Different tokenizers produce different token sequences, which changes n-gram counts. The sacrebleu library standardizes this with its tok parameter, and using a consistent tokenizer is essential for reproducible comparisons.

Limitations and Impact

BLEU revolutionized machine translation by providing a standardized, automatic evaluation metric that enabled rapid iteration. Before BLEU, researchers relied on expensive and slow human evaluations, limiting experimental throughput. BLEU allowed nightly regression testing and systematic optimization, directly contributing to the rapid improvement of statistical and later neural machine translation systems. The metric facilitated the development of shared tasks and benchmarks, creating a competitive research environment that accelerated progress from 2002 through the deep learning revolution.

However, BLEU suffers from significant limitations that have become more apparent as the field evolved. Understanding these limitations helps you avoid misuse of BLEU and understand what automatic evaluation can and cannot tell you about a generation system.

Semantic Blindness

The most glaring limitation is that BLEU is semantically blind. BLEU treats words as atomic tokens with no understanding of meaning. As we saw in our example, "cat" and "feline" produce no match despite being synonyms. Similarly, "quickly" and "fast" are treated as completely different words. This penalizes valid translation variation and fails to capture meaning preservation when word choice differs. The metric cannot distinguish between a translation that preserves meaning using different words and one that completely mistranslates the content, provided the surface forms do not match the reference.

This semantic blindness becomes more damaging for languages with rich morphology. In German, the same noun can take many different inflected forms depending on case and number. A translator who correctly identifies the right grammatical form but differs from the reference inflection receives no credit, even though the meaning is preserved. In agglutinative languages like Finnish or Turkish, the situation is even more extreme: a single word in these languages can correspond to an entire phrase in English, and matching that word exactly may be impossible even when the translation is semantically perfect.

The limitations are not merely theoretical. Studies have repeatedly shown cases where BLEU strongly disagrees with human judgment because it misses valid paraphrases. A translation system that learned to always echo the exact phrasing of its training references would score higher on BLEU than one that produced more natural, varied output, even though human evaluators would prefer the latter.

Word Order Sensitivity

BLEU also struggles with word order flexibility. While higher-order n-grams nominally check fluency, they penalize perfectly valid reorderings. In many languages, word order differs substantially from English; a translation might be fully fluent and accurate yet receive a low BLEU score because phrases appear in different positions than in the reference. German, for instance, often places verbs at the end of clauses, while Japanese follows a subject-object-verb order that requires extensive reordering when translating to English. BLEU unfairly penalizes translations that respect target language syntax when it differs from the specific reference translation provided.

Consider translating from a free-word-order language like Russian or Latin. The same sentence can appear with the subject before or after the verb, with adjectives in various positions relative to their nouns, and with prepositional phrases in multiple valid locations. A human translator would choose the most natural English order based on information structure and focus, which might differ from what the reference translator chose. BLEU would penalize both as equivalent errors even though one might be significantly more natural.

Weak Sentence-Level Correlation

The metric exhibits weak correlation with human judgment at the sentence level. While corpus-level BLEU roughly tracks system rankings, individual sentences can have terrible BLEU scores despite being good translations, or high BLEU scores while being ungrammatical. This makes BLEU unsuitable for applications requiring reliable sentence-level quality estimation, such as filtering training data or confidence estimation. A translation might be perfectly adequate but receive zero BLEU because it uses a completely valid paraphrase not present in the limited reference set.

The correlation problem is compounded by reference sparsity. Standard test sets typically include one to four reference translations per source sentence. Even with four references, a professional translator would still produce variations not captured in the reference set. The more creative or varied the correct translation space, the worse BLEU's correlation with human judgment becomes. For poetry, idiomatic expressions, or culturally embedded text, the space of valid translations can be enormous, and the handful of references in a test set cannot adequately represent it.

Domain and Language Pair Sensitivity

Domain sensitivity presents another challenge. BLEU scores are not comparable across different language pairs or domains. A BLEU of 30 might represent state-of-the-art for Chinese-English translation but poor quality for Spanish-English, where linguistic similarities allow for higher scores. Even within a language pair, scores vary by genre: news translation typically achieves higher BLEU than technical manuals or literary text, because news language is more formulaic and predictable, leading to more reference overlap.

This has led to the unfortunate practice of "BLEU score inflation" where papers report seemingly high numbers that are unimpressive when accounting for dataset difficulty. You must normalize scores against baseline systems to make meaningful comparisons. A new neural architecture claiming a BLEU of 45 on a news translation task is far less impressive if the previous best system already achieved 44, whereas the same absolute number on a difficult literary translation task might represent a substantial breakthrough.

Out[17]:
Visualization
Grouped bar chart showing simulated BLEU scores for three language pairs across three text domains.
Simulated BLEU scores across language pairs and text domains, illustrating how scores are not comparable across these dimensions. Spanish-English translation of news text reaches the highest BLEU scores due to linguistic similarity and domain regularity. Chinese-English and Japanese-English scores are substantially lower in all domains. Literary translation consistently shows lower BLEU across all language pairs because of the wide space of valid paraphrases. This variability means that BLEU scores must always be interpreted relative to baselines within the same language pair and domain.

The Goodhart's Law Problem

Perhaps the deepest limitation of BLEU is what happens when it becomes the primary optimization target. Goodhart's Law states that when a measure becomes a target, it ceases to be a good measure. Research groups began optimizing specifically for BLEU: developing smoothing methods tuned to maximize BLEU while producing less fluent text, selecting training data that maximized reference overlap, and designing architectures that were good at n-gram matching rather than semantic preservation. Some researchers noted that BLEU-optimized systems could produce translations that scored well on the metric but were judged as significantly worse by human evaluators who assessed overall meaning and naturalness.

The phenomenon of "BLEU hacking" manifested in several ways. Minimum Bayes risk (MBR) decoding, which selects the translation most similar to a set of candidates, was shown to produce outputs that scored well on BLEU but were bland and repetitive. Certain beam search strategies produced translations that successfully matched common n-grams by repeating safe phrases. The research community gradually recognized that maximizing BLEU was not the same as maximizing translation quality, leading to calls for supplementary metrics and more rigorous human evaluation protocols.

Beyond BLEU: Complementary Metrics

Despite these limitations, BLEU remains relevant as a diagnostic tool. When a system shows low BLEU, it often indicates serious problems with fluency or adequacy. However, modern research increasingly supplements BLEU with semantic metrics that address its core weaknesses.

METEOR (Metric for Evaluation of Translation with Explicit ORdering) was an early attempt to address semantic blindness. It incorporates synonym matching through WordNet, stemming to handle morphological variants, and a different recall-precision trade-off that emphasizes recall more heavily than BLEU. METEOR also uses a fragmentation penalty that rewards contiguous phrase matches over scattered word matches. This provides a finer-grained assessment of word order.

TER (Translation Edit Rate) measures the number of edits required to transform the candidate into any reference translation. Edits include insertions, deletions, substitutions, and phrase shifts (reordering entire phrases). TER penalizes word order differences through shifts and is less sensitive to exact surface form matches, making it complementary to BLEU in system comparisons.

BERTScore addresses semantic blindness directly by computing similarity between candidate and reference using contextual embeddings from BERT. Rather than requiring exact string matches, BERTScore matches tokens based on their embedding similarity, allowing semantically equivalent paraphrases to receive high scores. We will explore BERTScore in detail in the next chapter, where we examine how neural representations of meaning enable more robust automatic evaluation.

COMET (Crosslingual Optimized Metric for Evaluation of Translation) takes a different approach entirely, training a neural regressor to predict human quality assessments. Given a source sentence, candidate translation, and reference, COMET produces a single quality score calibrated to human direct assessment ratings. Because it is trained to match human judgment, COMET correlates significantly better with human evaluators than BLEU, particularly at the sentence level.

The practical recommendation is to report multiple metrics. BLEU provides a computationally cheap, widely understood baseline. COMET provides the most human-aligned quality estimate. BERTScore captures semantic equivalence. Together, these three metrics give a more complete picture of system quality than any single number can convey.

Interpreting BLEU Scores in Practice

Understanding what BLEU scores mean in practice requires context. The table below provides rough guidelines for interpreting corpus-level BLEU scores in standard news domain machine translation benchmarks, based on the field's accumulated experience:

Out[18]:
Visualization
Horizontal bar chart showing BLEU score ranges mapped to qualitative quality descriptions from 0-10 (mostly wrong) up to 60 plus (near human quality).
Rough interpretation guidelines for corpus-level BLEU scores in news domain machine translation, showing the quality ranges from unintelligible output to near-human translation quality. These thresholds are approximate and vary substantially by language pair and domain. The same score may represent excellent performance for a difficult language pair and mediocre performance for an easier one.

These ranges are approximate and highly domain-dependent. A BLEU of 35 is exceptional for Japanese-English literary translation but mediocre for Spanish-English news. Always compare against the best published baseline for your specific language pair and domain.

When evaluating your own system, the most useful interpretation of BLEU is relative, not absolute. A 2-point improvement over a competitive baseline on a well-studied benchmark is meaningful, while a 10-point improvement on a dataset of your own creation may reflect dataset construction choices more than real progress. The research community has learned, sometimes painfully, that BLEU improvements on held-out test sets do not always generalize to user satisfaction in production systems.

Summary

BLEU score evaluates text generation by measuring n-gram overlap between candidate and reference translations, combining modified precision across multiple n-gram orders with a brevity penalty to discourage short outputs. The modified precision clips n-gram counts to their maximum occurrence in any reference, preventing gaming through repetition and ensuring that candidates only receive credit for content present in the reference set.

Key implementation details distinguish corpus-level from sentence-level calculation. Corpus BLEU aggregates statistics across entire test sets before computing precision. This provides stable scores suitable for system comparison. Sentence-level BLEU computes scores independently, suffering from volatility and frequent zeros on short sentences due to the sparsity of higher-order n-grams.

The formula multiplies a brevity penalty by the geometric mean of n-gram precisions:

BLEU=BP⋅exp⁡(∑n=1Nwnlog⁡pn)\text{BLEU} = BP \cdot \exp\left(\sum_{n=1}^N w_n \log p_n\right)

where:

  • BLEU\text{BLEU}: the final BLEU score (between 0 and 1)
  • BPBP: the brevity penalty factor
  • NN: the maximum n-gram order (typically 4)
  • wnw_n: the weight for n-grams of order nn (typically uniform: wn=1/Nw_n = 1/N)
  • pnp_n: the modified precision for n-grams of order nn

The brevity penalty BPBP is defined as:

BP={1if c>re(1−r/c)if c≤rBP = \begin{cases} 1 & \text{if } c > r \\ e^{(1 - r/c)} & \text{if } c \leq r \end{cases}

where:

  • cc: the total length of the candidate translation (in words)
  • rr: the effective reference length, found by choosing the reference closest in length to the candidate

BLEU enabled the rapid progress of machine translation research through standardized, reproducible evaluation. Its limitations, including semantic blindness, word order sensitivity, and poor sentence-level correlation, have driven the development of complementary metrics like METEOR, BERTScore, and COMET. For modern research, the recommendation is to report BLEU alongside at least one semantically-aware metric and always compare against strong baselines within the same language pair and domain. The sacrebleu library provides standardized computation that ensures reproducibility across research groups.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the BLEU score evaluation metric.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026bleuscore, author = {Michael Brenndoerfer}, title = {BLEU Score: Evaluating Translation Quality with N-grams}, year = {2026}, url = {https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). BLEU Score: Evaluating Translation Quality with N-grams. Retrieved from https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp
MLAAcademic
Michael Brenndoerfer. "BLEU Score: Evaluating Translation Quality with N-grams." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp>.
CHICAGOAcademic
Michael Brenndoerfer. "BLEU Score: Evaluating Translation Quality with N-grams." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp.
HARVARDAcademic
Michael Brenndoerfer (2026) 'BLEU Score: Evaluating Translation Quality with N-grams'. Available at: https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). BLEU Score: Evaluating Translation Quality with N-grams. https://mbrenndoerfer.com/writing/bleu-score-machine-translation-evaluation-nlp

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.