ROUGE Scores: Evaluating Text Summarization

Michael BrenndoerferFebruary 23, 202656 min read

Part of Language AI Handbook

Covers ROUGE metrics for summarization evaluation. Topics include ROUGE-N, ROUGE-L, and ROUGE-W, and how they differ from BLEU in assessing text quality.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

ROUGE Scores

When you ask a language model to summarize a lengthy document, how do you know if the summary is any good? This question is central to automatic text evaluation, and answering it requires understanding how summarization differs from other natural language generation tasks. Unlike machine translation, where BLEU gives a reasonable measure of n-gram precision against reference translations, summarization presents a unique challenge. A good summary captures the most important information from the source while condensing it substantially, meaning we care more about whether the generated text covers the key points (recall) than whether it avoids extraneous words (precision). In translation, adding extra words might indicate hallucination or poor alignment with the source, but in summarization, the goal is inherently selective. We want to know whether the system successfully identified and retained the important information, not merely whether it avoided adding fluff.

This distinction matters more than it might first appear. When you hire a human summarizer to condense a legal document, you do not judge them primarily on whether every sentence they wrote is drawn from the original. You judge them on whether they captured the clauses that matter, the obligations that bind, and the exceptions that protect. A perfectly concise summary that omits the key liability clause has failed completely, even though every sentence it does contain is accurate. A slightly verbose summary that includes all the necessary points has succeeded, even if it repeats some information. The asymmetry is basic to what summarization means.

ROUGE (Recall-Oriented Understudy for Gisting Evaluation) addresses exactly this need. Developed by Chin-Yew Lin in 2004 specifically for automatic evaluation of summaries, ROUGE measures the overlap of n-grams and word sequences between a generated summary and one or more human-written reference summaries. The name makes the design philosophy explicit: "recall-oriented" signals that the metric prioritizes coverage over conciseness. While BLEU asks "How much of the generated text is correct?", ROUGE asks "How much of the reference content did we capture?" This shift in focus from precision to recall makes ROUGE the dominant automatic metric for summarization evaluation. It acknowledges that multiple valid summaries can exist for the same document, and that missing necessary information constitutes a more serious failure than including an extra detail.

The intellectual lineage of ROUGE traces back to the challenges of evaluating text summarization in the early 2000s, when the Document Understanding Conferences (DUC) began systematically evaluating automatic summarization systems. Prior to ROUGE, evaluation relied primarily on expensive human judgments, which are slow, costly, and difficult to reproduce. Researchers needed a cheap, automatic proxy that would correlate well with human assessments. Lin's insight was to adapt the n-gram overlap idea from BLEU, but to flip the orientation from precision to recall. This simple change produced a metric that aligned much better with the summarization objective, and the ROUGE paper became one of the most cited works in the entire NLP literature.

In this chapter, we will explore the mechanics of ROUGE-N, ROUGE-L, ROUGE-W, and ROUGE-S variants, understand when each is appropriate, and examine their limitations in capturing summary quality. We will also see how ROUGE differs from BLEU and why these differences matter for different generation tasks. By understanding both the mathematical foundations and the practical intuitions behind ROUGE, you will be equipped to interpret evaluation scores correctly and recognize when automatic metrics align with, or diverge from, human judgment.

ROUGE-N: N-gram Recall

The most straightforward variant, ROUGE-N, measures the proportion of n-grams in the reference summary that also appear in the generated summary. This is fundamentally a recall metric, unlike BLEU which computes precision. This distinction is significant: when evaluating summaries, we prioritize completeness over succinctness. A summary that mentions only one key fact from a ten-fact reference is precise (everything it says might be correct) but useless because it fails to deliver the bulk of the important information. ROUGE-N captures this completeness by checking how many chunks of the reference text appear in the candidate summary.

To build intuition before the formula, consider a simple scenario. A reference summary says "The economy grew strongly last quarter, driven by consumer spending and export growth." A generated summary says "The economy grew strongly." Both statements are true and every word in the candidate appears in the reference. But ROUGE-N would score this candidate poorly on recall because it captured only a small fraction of the reference's content. The metric correctly identifies that despite perfect precision, the candidate missed the explanation of why the economy grew. That "why" might be exactly what a reader needs.

The Mathematics of ROUGE-N

ROUGE-N measures recall by computing the ratio of matching n-grams to total n-grams in the reference. For a specific n-gram order nn, ROUGE-N is calculated as:

ROUGE-N=∑S∈References∑gramn∈SCountmatch(gramn)∑S∈References∑gramn∈SCount(gramn)\text{ROUGE-N} = \frac{\sum_{S \in \text{References}} \sum_{\text{gram}_n \in S} \text{Count}_{\text{match}}(\text{gram}_n)}{\sum_{S \in \text{References}} \sum_{\text{gram}_n \in S} \text{Count}(\text{gram}_n)}

where:

  • gramn\text{gram}_n: an n-gram of order nn (a contiguous sequence of nn words)
  • Countmatch(gramn)\text{Count}_{\text{match}}(\text{gram}_n): the number of co-occurring n-grams in both the candidate and reference, capped by the reference count to prevent inflated scores from repetitive candidates
  • Count(gramn)\text{Count}(\text{gram}_n): the total count of n-grams in the reference

To understand this formula intuitively, consider the numerator as a tally of every n-gram chunk found in the reference that also exists somewhere in the candidate summary. The denominator is the total number of n-gram chunks available in the reference. By dividing matches by total possible matches, we obtain a proportion between 0 and 1, where 1 indicates the candidate contains every n-gram from the reference (perfect recall) and 0 indicates no overlap whatsoever.

The capping mechanism in Countmatch\text{Count}_{\text{match}} deserves attention. If a candidate says "the cat the cat the cat" as a response to a reference saying "the cat sat on the mat," a naive count would give enormous credit for "the cat" matches. The cap ensures that if "the cat" appears once in the reference, it can only contribute one match regardless of how many times the candidate repeats it. This prevents a degenerate strategy of repeating common reference words to inflate scores.

Recall vs. Precision Orientation

Notice the denominator uses reference counts, not candidate counts. This makes ROUGE a recall metric: it penalizes missing important content from the reference, but does not penalize adding extra content. This differs from BLEU, where the denominator uses candidate counts (precision), penalizing extraneous words but being more lenient on omissions. If a candidate summary is twice as long as the reference and contains all the reference content plus additional sentences, ROUGE-N will award a perfect score of 1.0, while BLEU would penalize the extra content. This design choice shows the summarization task's priority: capturing the needed information matters more than conciseness, which can be controlled separately through length constraints during generation.

When multiple reference summaries are available (common in summarization datasets, where different humans produce varying summaries of the same document), ROUGE-N typically takes the maximum score across all references:

ROUGE-Nmulti=max⁡iROUGE-N(candidate,referencei)\text{ROUGE-N}_{\text{multi}} = \max_i \text{ROUGE-N}(\text{candidate}, \text{reference}_i)

where:

  • ii: index over the set of reference summaries
  • candidate\text{candidate}: the generated summary being evaluated
  • referencei\text{reference}_i: the ii-th human-written reference summary

This shows the intuition that a good summary only needs to match one valid human summary well, not all possible good summaries simultaneously. Because summarization is subjective, two equally valid summaries might focus on different aspects of a document or use completely different phrasing. Taking the maximum score acknowledges this diversity: if the candidate aligns well with any single reference, it likely captures valid content, even if it diverges from other references. This is analogous to grading an essay against the best matching rubric rather than requiring it to satisfy all rubrics simultaneously.

ROUGE-1 and ROUGE-2

In practice, ROUGE-1 (unigram overlap) and ROUGE-2 (bigram overlap) are the most commonly reported metrics:

  • ROUGE-1 measures content coverage at the word level. High ROUGE-1 indicates the summary captures the right vocabulary, but may not preserve word order or fluency. For example, "dog lazy the over jumps fox brown quick The" would reach perfect ROUGE-1 with "The quick brown fox jumps over the lazy dog" because it contains all the same words, despite being nonsensical. Thus, ROUGE-1 is a loose filter for content selection but cannot assess grammaticality or coherence. It answers the question: "Are the right words present somewhere in the summary?"

  • ROUGE-2 measures local word ordering. High ROUGE-2 indicates the summary uses the right words and strings them together in ways that match the reference. In the previous example, the scrambled sentence would score zero on ROUGE-2 because no consecutive word pairs are preserved. The bigram "dog lazy" does not appear in the reference, and "lazy the" does not either. This makes ROUGE-2 a better indicator of linguistic quality and syntactic correctness. It answers: "Are the right words appearing together in the right local sequences?"

The relationship between ROUGE-1 and ROUGE-2 gives diagnostic information. If a summary scores high on ROUGE-1 but low on ROUGE-2, it likely has the right content but poor fluency or unusual ordering. If both scores are high, the summary captures content correctly and arranges it in sequences that resemble the reference. If both are low, the summary has missed the needed content entirely.

ROUGE-2 is generally considered a better indicator of summary quality than ROUGE-1 because it captures some syntactic structure. However, it is also stricter, often creating lower scores since bigram matches are rarer than unigram matches. In typical news summarization tasks, ROUGE-1 scores often range between 0.4 and 0.5 for strong systems, while ROUGE-2 scores might range between 0.2 and 0.3. You usually report both metrics together: ROUGE-1 to verify content coverage and ROUGE-2 to verify local coherence.

Higher-order ROUGE-N (ROUGE-3, ROUGE-4) is rarely reported in practice. The reason is statistical sparsity: as n grows, matching three or four consecutive words across a paraphrase becomes very difficult, and scores collapse toward zero for most real summaries. A ROUGE-4 score of 0.05 tells you very little about whether a summary is good. The useful range of n ends at 2 for most summarization tasks, though ROUGE-3 occasionally appears in tasks with very standardized output formats.

ROUGE-L: Longest Common Subsequence

While n-gram recall is intuitive, it has a significant limitation: it requires exact n-gram matches. If a candidate summary uses synonymous phrases or slightly reorders words, ROUGE-N treats these as complete failures even when the meaning is preserved. Consider a reference stating "The President announced new economic policies yesterday" and a candidate reading "Yesterday, the President announced new economic policies." A human would recognize these as nearly identical in meaning, but ROUGE-2 would score them poorly because the word "yesterday" has moved, breaking the bigrams "policies yesterday" and introducing "Yesterday the" which does not appear in the reference.

This fragility of n-gram windows motivates a more flexible matching approach. Rather than requiring words to appear in fixed-size windows, what if we could find the longest chain of words that appears in both texts, in the same order, regardless of what other words intervene? This is precisely the idea behind the Longest Common Subsequence.

ROUGE-L addresses this by using the Longest Common Subsequence (LCS). A subsequence is a sequence that appears in the same relative order but not necessarily consecutively. For example, "the cat sat" is a subsequence of "the fluffy cat lazily sat," while "cat the" is not (wrong order). The LCS finds the longest chain of words that appears in both texts while maintaining sequence integrity. It offers flexibility that rigid n-gram windows cannot give.

The LCS is computed using dynamic programming in O(mn)O(mn) time, where mm and nn are the lengths of the two sequences. The algorithm builds a table where cell (i,j)(i, j) stores the length of the longest common subsequence of the first ii tokens of the reference and the first jj tokens of the candidate. If the tokens at positions ii and jj match, the LCS extends by one from the diagonal cell; otherwise it takes the maximum of the left and upper cells, representing the best LCS achievable by dropping one token from either sequence.

LCS-Based F-measure

ROUGE-L computes precision and recall based on the LCS length between candidate CC and reference RR:

Rlcs=LCS(C,R)∣R∣Plcs=LCS(C,R)∣C∣\begin{aligned} R_{\text{lcs}} &= \frac{\text{LCS}(C, R)}{|R|} \\[6pt] P_{\text{lcs}} &= \frac{\text{LCS}(C, R)}{|C|} \end{aligned}

where:

  • RlcsR_{\text{lcs}}: the recall based on the longest common subsequence
  • PlcsP_{\text{lcs}}: the precision based on the longest common subsequence
  • ∣R∣|R|: the length (in words) of the reference summary
  • ∣C∣|C|: the length (in words) of the candidate summary
  • LCS(C,R)\text{LCS}(C, R): the length of the longest common subsequence between candidate and reference

These are then combined into an F-measure:

ROUGE-L=(1+β2)⋅Rlcs⋅PlcsRlcs+β2⋅Plcs\text{ROUGE-L} = \frac{(1 + \beta^2) \cdot R_{\text{lcs}} \cdot P_{\text{lcs}}}{R_{\text{lcs}} + \beta^2 \cdot P_{\text{lcs}}}

where:

  • RlcsR_{\text{lcs}}: the recall based on the longest common subsequence (defined above)
  • PlcsP_{\text{lcs}}: the precision based on the longest common subsequence (defined above)
  • β\beta: a parameter controlling the relative importance of precision versus recall; when β\beta is large (typically β>1\beta > 1), recall is weighted more heavily. In standard implementations, β\beta is often set such that recall is weighted more than precision (e.g., β=1.2\beta = 1.2), though many implementations use β=1\beta = 1 (standard F1)

The use of F-measure rather than pure recall is an important departure from ROUGE-N. Because the LCS can be short relative to the candidate length (imagine a candidate that rambles extensively but contains a small core that matches the reference), precision becomes relevant. If we only measured recall, a candidate containing the entire reference embedded in a much longer document would score perfectly, even though most of the candidate is irrelevant. The F-measure balances these concerns, rewarding summaries that efficiently capture the reference content without excessive verbosity.

Subsequence vs. Substring

Don't confuse subsequence with substring. A substring requires consecutive words, while a subsequence allows gaps. "climate change" is a substring of "addressing climate change requires action" but only a subsequence of "climate patterns are changing." ROUGE-L uses subsequence, which makes it more flexible than n-gram matching. This flexibility allows it to handle paraphrasing where words are interspersed with modifiers or where sentence structure varies between reference and candidate.

Why LCS Matters for Summarization

The LCS approach gives several advantages for summarization evaluation. Each advantage addresses a specific failure mode in n-gram-based metrics:

First, the LCS works well with sentence-level paraphrasing. If a candidate summary restructures a sentence while keeping key phrases in order, LCS records this similarity. For instance, if the reference says "The company reported record profits despite economic headwinds" and the candidate says "Despite economic headwinds, the company reported record profits," the LCS includes all major content words in the correct order, even though the sentence structure changed. The n-gram approach breaks down here because the fronting of "Despite economic headwinds" disrupts all the bigrams that cross the clause boundary.

Second, unlike n-grams which require arbitrary window sizes, LCS automatically identifies the longest matching sequence, effectively adapting to local similarity structures. It doesn't force us to choose between unigrams (too permissive) and bigrams (too strict); instead, it finds the natural alignment between texts. This adaptivity is especially useful when the summary length and structure differ materially from the reference.

Third, adding words to the candidate doesn't break the LCS calculation the way it breaks n-gram precision. If a candidate inserts "significant" before "profits" in the example above, ROUGE-2 would drop bigram matches around that insertion, but ROUGE-L would simply treat "significant" as a gap in the subsequence, preserving the rest of the match. This robustness to word insertion makes ROUGE-L more forgiving of paraphrases that add modifiers or clarifications.

Out[3]:
Visualization
Heatmap of LCS dynamic programming matrix for a 9-by-9 token grid with cell values and red backtrack path.
Dynamic programming matrix for computing the Longest Common Subsequence between the reference and Candidate B. Each cell stores the LCS length for the corresponding prefix pair. The dashed red rectangles trace the backtracking path that recovers the 8-token LCS, showing how the algorithm skips the misplaced 'brown' token in Candidate B.

ROUGE-W: Weighted Longest Common Subsequence

Standard LCS treats all matching subsequences equally, regardless of whether the matches are consecutive or scattered. However, consecutive matches are linguistically more significant than scattered ones. "The quick brown fox" is a better match than "The fox quick brown" even if both have the same LCS length with a reference. The first preserves syntactic structure and readability, while the second is jumbled and confusing.

This limitation of standard LCS becomes apparent with a concrete example. Suppose the reference is "The committee approved the budget proposal after extended debate." Candidate A reads "The committee approved the proposal after debate" (two words omitted, but contiguous phrasing preserved). Candidate B reads "The committee debate after budget approved the proposal" (all words present but scrambled). Both candidates might reach the same LCS length relative to the reference, because LCS allows arbitrary gaps. But Candidate A is clearly better: it reads naturally and conveys the right meaning. Candidate B is in effect incomprehensible. Standard ROUGE-L would not distinguish between them.

ROUGE-W addresses this by introducing a weighting function that favors consecutive matches. The weighted LCS (WLCS) applies a penalty to gaps in the matching sequence. This keeps fragmented matches contribute less to the final score than cohesive blocks of text.

The Weighting Function

The standard ROUGE-W implementation uses a weight function that grows faster than linear with consecutive matches. For a consecutive sequence of length kk, the weight is often computed as:

w(k)=kαw(k) = k^\alpha

where:

  • kk: the length of a consecutive matching sequence
  • α\alpha: an exponent greater than 1 (typically α=1.2\alpha = 1.2 or 22), so consecutive matches contribute disproportionately more to the score than scattered matches

For example, with α=2\alpha = 2:

  • A consecutive match of 4 words contributes 42=164^2 = 16 to the score
  • Four scattered single-word matches contribute 4×12=44 \times 1^2 = 4 to the score

The superlinear weighting means that keeping matches together is worth far more than spreading them out. A system that produces a two-word consecutive match and another two-word consecutive match scores 22+22=82^2 + 2^2 = 8, whereas the same four words scattered individually score 4×1=44 \times 1 = 4. This creates an incentive for the evaluated summary to use reference phrasing in cohesive chunks rather than picking up scattered words.

The WLCS dynamic programming algorithm modifies the standard LCS recurrence to account for these weights, favoring paths that keep matches consecutive. The key modification is tracking the LCS length and the current consecutive run length. When calculating the optimal subsequence, the algorithm must now consider whether words match and whether continuing a consecutive sequence yields higher weight than starting a new scattered match. The recurrence becomes:

WLCS(i,j)={WLCS(i−1,j−1)+w(c(i,j))if xi=yjmax⁡(WLCS(i−1,j), WLCS(i,j−1))otherwise\text{WLCS}(i, j) = \begin{cases} \text{WLCS}(i-1, j-1) + w(c(i,j)) & \text{if } x_i = y_j \\ \max(\text{WLCS}(i-1, j),\ \text{WLCS}(i, j-1)) & \text{otherwise} \end{cases}

where c(i,j)c(i,j) is the length of the current consecutive run ending at positions ii and jj. The algorithm tracks this run length alongside the WLCS value, incrementing it on diagonal moves and resetting it on horizontal or vertical moves.

Out[4]:
Visualization
Two side-by-side bar charts showing weight contribution. Left bar (consecutive) reaches 16, right bar (scattered) reaches 4, with alpha equals 2.
Comparison of ROUGE-W weight contributions for consecutive versus scattered word matches. With exponent alpha equals 2, four consecutive matching words (left) contribute weight 16, while four scattered single-word matches (right) contribute only weight 4. The superlinear weighting strongly rewards cohesive phrase alignment.

When to Use ROUGE-W

ROUGE-W is particularly useful in several situations. It shines when fluency matters: you want to penalize summaries that have the right words but in jumbled order. In applications like headline generation or keyphrase extraction, where the output must read naturally, ROUGE-W gives better discrimination than standard ROUGE-L.

ROUGE-W is also appropriate for extractive summarization evaluation. When the system selects sentences from the source, consecutive matches indicate proper sentence extraction. If a system extracts a full sentence verbatim, ROUGE-W rewards it heavily; if it extracts scattered fragments, it receives a lower score despite potentially having the same LCS length. This makes ROUGE-W align better with human judgments in extractive settings.

For short summaries such as single-sentence summaries, word order carries more semantic weight. In longer documents, a few misplaced words might not obscure the meaning, but in a one-sentence summary, "Dogs eat cats" versus "Cats eat dogs" is a catastrophic error that ROUGE-W catches better than ROUGE-L.

However, ROUGE-W is less commonly used than ROUGE-1, ROUGE-2, and ROUGE-L in standard benchmark reporting, partly because the weighting parameter α\alpha requires tuning and complicates comparisons across studies. Without standardized values for α\alpha, scores become incomparable between papers, undermining the metric's utility as a universal benchmark. The practical field has largely settled on ROUGE-1, ROUGE-2, and ROUGE-L as the canonical trio, with ROUGE-W appearing primarily in specialized studies.

Out[5]:
Visualization
Diagram showing reference tokens with four scattered green words and candidate tokens below, connected by dashed alignment lines.
Illustration of scattered versus consecutive matching in ROUGE-W. The reference tokens appear on top; the four matched tokens (green) are scattered across the reference at positions 0, 3, 5, and 8, creating four isolated single-word matches. The candidate row shows the same four words pulled together, with dashed lines showing their original positions in the reference. This scatter pattern yields weight 4 under alpha equals 2, compared to weight 16 for a four-word consecutive match.

ROUGE-S: Skip-Bigram Overlap

Another variant worth understanding is ROUGE-S, which measures skip-bigram co-occurrence. Skip-bigrams are pairs of words that appear in the same order as the reference, but with arbitrary gaps between them. This concept bridges the gap between the strict locality of ROUGE-2 (consecutive words) and the global flexibility of ROUGE-L (any ordered subset).

Consider the reference: "The president spoke about climate change." A ROUGE-2 bigram is "spoke about," requiring consecutive words. A skip-bigram would include pairs like "president...climate" (words 2 and 5) or "The...change" (words 1 and 6), as long as the order is preserved. This captures semantic relationships between distant words that might be separated by modifiers or clauses.

ROUGE-S is computed similarly to ROUGE-2, but instead of counting consecutive bigrams, it counts ordered word pairs with any distance between them. A candidate summary that uses the same key words in the same order as the reference, even with different intervening words, receives credit for every matching pair. This gives a middle ground between ROUGE-1 (unordered) and ROUGE-2 (strictly consecutive).

The skip-bigram count for a candidate CC and reference RR is:

Skip-bigram count=∣{(wi,wj):wi∈C,wj∈C,i<j,∃(wi,wj)∈R}∣\text{Skip-bigram count} = |\{(w_i, w_j) : w_i \in C, w_j \in C, i < j, \exists (w_i, w_j) \in R\}|

where:

  • wi,wjw_i, w_j: words at positions ii and jj in the candidate text
  • CC: the candidate summary (as a sequence of words)
  • RR: the reference summary (as a sequence of words)
  • i<ji < j: the constraint that word wiw_i appears before word wjw_j in the candidate
  • ∃(wi,wj)∈R\exists (w_i, w_j) \in R: the condition that the same word pair appears in the reference in the same order

To compute the ROUGE-S score, the skip-bigram counts are normalized by the total number of skip-bigrams in the reference, yielding a recall measure analogous to ROUGE-N. Precision and F-measure can also be computed. A variant called ROUGE-SU combines skip-bigrams with unigrams, preventing degenerate cases where a very short candidate reaches high ROUGE-S by matching a small number of ordered pairs.

One practical issue with ROUGE-S is the skip distance parameter. In the original ROUGE formulation, skip-bigrams allow any distance, but this means very long documents can produce enormous numbers of candidate pairs. A document with 100 words generates roughly 4,950 potential skip-bigrams, compared to only 99 consecutive bigrams. Some implementations cap the maximum skip distance to prevent computational issues and to focus on semantically related word pairs rather than arbitrary long-range coincidences.

ROUGE-S is particularly useful for evaluating summaries that paraphrase by inserting words between key terms, though in practice, ROUGE-L often gives similar discriminative power with a more intuitive interpretation. As a result, ROUGE-S is less commonly reported than ROUGE-L in mainstream benchmarks, though it occasionally appears in studies specifically investigating paraphrase-reliable evaluation.

Preprocessing and Configuration

ROUGE scores are sensitive to preprocessing choices. Standard implementations include several configurable options that materially impact scores. Understanding these choices is needed for fair comparison and reproducibility, as a ROUGE score reported without preprocessing details is in effect meaningless.

This might seem like a minor technical point, but it has caused real confusion in the research literature. Two papers might report ROUGE-1 scores for systems on the same dataset, but if one uses stemming and the other does not, the scores are not directly comparable. The same system can exhibit a several-point difference in ROUGE scores depending solely on whether stopwords were removed. Careful researchers always specify their preprocessing pipeline, and careful readers always check before drawing conclusions from score comparisons.

Stemming

Most ROUGE implementations apply Porter stemming or similar algorithms to normalize words to their root forms. This allows "running," "runs," and "ran" to match, increasing robustness to morphological variation. Stemming collapses different inflections of the same word into a common token, recognizing that "the cat runs" and "the cat ran" convey the same core information despite tense differences.

Without stemming, "children play" and "child plays" would score poorly despite semantic equivalence. With stemming, both become "child play," creating perfect matches. However, stemming can occasionally over-collapse distinct meanings. For example, "universe" and "universal" might collapse to the same stem despite different meanings, and "organization" and "organism" might partially overlap in their stems. Such errors are rare enough that stemming generally improves evaluation reliability, but they are worth keeping in mind when examining cases where ROUGE scores seem anomalous.

The Porter stemmer applies a sequence of suffix stripping rules. Words ending in "-ing" have that suffix removed (with additional rules for minimal stem length). Words ending in "-tion" become "-t". The algorithm is entirely rule-based, which makes it fast but imperfect. For specialized domains with unusual morphology, domain-specific stemmers or lemmatizers might work better, though the research community predominantly uses Porter stemming for consistency.

Stopword Removal

Optionally, stopwords (common words like "the," "and," "is") can be removed before computing ROUGE. This focuses the metric on content words rather than function words, theoretically making the evaluation more sensitive to semantic substance rather than syntactic scaffolding.

However, stopword removal is controversial in summarization evaluation:

  • Pro: Eliminates inflated scores from common phrases like "in the" or "of the" that appear in virtually every English sentence. Without removal, a system could reach moderate ROUGE scores simply by generating grammatical English sentences with little relation to the source content.
  • Con: Removes important discourse markers and can make grammaticality harder to assess. Stopwords carry relational meaning (prepositions indicate spatial and temporal relationships; conjunctions indicate logical connections) that can be important for accurate summarization.

Most standard evaluations (like the CNN/DailyMail benchmark) report ROUGE scores both with and without stopword removal, letting researchers to assess both content selection and linguistic fluency. When you report ROUGE without specifying this choice, readers cannot reproduce your results or make fair comparisons.

Tokenization

ROUGE requires consistent tokenization between candidate and reference. Unlike BLEU which often uses specialized tokenizers, ROUGE typically uses simple whitespace tokenization or the same tokenizer used for the task. Inconsistent tokenization (e.g., treating "don't" as one word vs. "do n't" as two) can materially alter scores.

Punctuation handling also matters. Some implementations strip all punctuation, while others treat it as separate tokens. If one system outputs "However, the plan failed" and another outputs "However the plan failed," punctuation-sensitive tokenization will yield different ROUGE scores despite identical content words. When comparing systems or recreating results, ensure all candidates and references use the same tokenization pipeline.

A related subtlety involves case normalization. Most implementations lowercase all text before computing ROUGE, which is sensible since "President" and "president" refer to the same entity. However, for tasks where capitalization carries semantic meaning (named entity recognition, for instance), case-sensitive evaluation might be more appropriate. Again, the key is consistency and transparency.

The Preprocessing Pipeline in Practice

A reliable preprocessing pipeline for ROUGE evaluation proceeds in a consistent order:

  1. Lowercase all text to remove spurious case differences
  2. Apply tokenization (whitespace-split or a dedicated tokenizer)
  3. Strip punctuation tokens or normalize punctuation to whitespace
  4. Optionally remove stopwords from the token list
  5. Optionally apply stemming to remaining tokens
  6. Compute ROUGE on the resulting token sequences

Every step in this pipeline must be identical for candidate and reference texts. Any asymmetry, such as stemming one but not the other, will produce scores that don't reflect actual overlap.

Worked Example

Let's walk through a concrete example to see how these metrics differ. Consider a reference summary and two candidate summaries:

Reference: "The quick brown fox jumps over the lazy dog"

Candidate A: "The brown fox jumps over the lazy dog" (omits "quick")

Candidate B: "The quick fox jumps over the brown lazy dog" (reorders words)

These candidates illustrate two common failure modes in summarization: omission (missing content) and distortion (reordering content). Candidate A is like a summary that gets the structure right but misses one detail. Candidate B is like a summary that contains all the right words but rearranges them in a way that changes meaning. By calculating ROUGE scores manually, we can observe how different metrics penalize these errors differently.

ROUGE-1 Calculation

For ROUGE-1, we count unigram matches:

Unigram matching counts for ROUGE-1 calculation.
WordReference CountCandidate A MatchCandidate B Match
the222
quick101
brown111
fox111
jumps111
over111
lazy111
dog111

Candidate A: 8 matches / 9 reference words = 0.889

Candidate B: 9 matches / 9 reference words = 1.000

ROUGE-1 considers Candidate B perfect despite the word reordering. The word "brown" moved from before "fox" to before "lazy dog," changing the phrase "quick brown fox" into "quick fox" and "brown lazy dog" into something somewhat different in meaning. But ROUGE-1 does not know or care about ordering. It simply checks that "brown" is present somewhere, which it is. This illustrates the core limitation of unigram matching: it captures presence but not arrangement. A summary could be completely scrambled yet reach perfect ROUGE-1, which is why higher-order n-grams or LCS-based measures are necessary.

ROUGE-2 Calculation

For ROUGE-2, we need bigram matches. The reference bigrams are:

  • "the quick", "quick brown", "brown fox", "fox jumps", "jumps over", "over the", "the lazy", "lazy dog"

Candidate A bigrams:

  • "the brown", "brown fox", "fox jumps", "jumps over", "over the", "the lazy", "lazy dog"

Matches: "brown fox", "fox jumps", "jumps over", "over the", "the lazy", "lazy dog" = 6 matches

Score: 6 / 8 = 0.750

Candidate B bigrams:

  • "the quick", "quick fox", "fox jumps", "jumps over", "over the", "the brown", "brown lazy", "lazy dog"

Matches: "the quick", "fox jumps", "jumps over", "over the", "lazy dog" = 5 matches

Score: 5 / 8 = 0.625

ROUGE-2 correctly penalizes Candidate B for the reordering, though it also penalizes Candidate A for omitting "quick." Notice how the insertion of "quick" before "fox" in Candidate B destroys the bigram "quick brown" (since "brown" is now elsewhere in the sentence) and replaces it with "quick fox" (not in reference). Simultaneously, "brown" now appears in the bigram "the brown" (not in reference) and "brown lazy" (not in reference), causing multiple bigram mismatches. ROUGE-2 captures this disruption effectively. This gives a more fine-grained quality signal than ROUGE-1.

ROUGE-L Calculation

For ROUGE-L, we find the longest common subsequence.

Reference: [The, quick, brown, fox, jumps, over, the, lazy, dog]

Candidate A: [The, brown, fox, jumps, over, the, lazy, dog]

The LCS is [The, brown, fox, jumps, over, the, lazy, dog] with length 8. Candidate A preserves all words except "quick" in the correct order, so the LCS simply skips "quick" from the reference side.

Rlcs=89=0.889Plcs=88=1.0Fmeasure=2⋅0.889⋅1.00.889+1.0=0.941\begin{aligned} R_{\text{lcs}} &= \frac{8}{9} = 0.889 \\[6pt] P_{\text{lcs}} &= \frac{8}{8} = 1.0 \\[6pt] F_{\text{measure}} &= \frac{2 \cdot 0.889 \cdot 1.0}{0.889 + 1.0} = 0.941 \end{aligned}

Candidate A reaches perfect precision (every word in the candidate matches a word in the reference in the correct order) but imperfect recall (it missed "quick").

Candidate B: [The, quick, fox, jumps, over, the, brown, lazy, dog]

The LCS is [The, quick, fox, jumps, over, the, lazy, dog] with length 8. The word "brown" cannot be part of the LCS because of its position: in the reference, "brown" appears at position 3 (before "fox"), but in Candidate B, "brown" appears at position 7 (after "the"). The LCS algorithm must maintain relative order, so it cannot include "brown" from position 3 in the reference alongside "fox" from position 3 in the candidate (where "brown" has already been placed at position 7).

Rlcs=89=0.889Plcs=89=0.889Fmeasure=2⋅0.889⋅0.8890.889+0.889=0.889\begin{aligned} R_{\text{lcs}} &= \frac{8}{9} = 0.889 \\[6pt] P_{\text{lcs}} &= \frac{8}{9} = 0.889 \\[6pt] F_{\text{measure}} &= \frac{2 \cdot 0.889 \cdot 0.889}{0.889 + 0.889} = 0.889 \end{aligned}

So ROUGE-L gives Candidate A a higher score than Candidate B (0.941 vs 0.889), though both lose points for different reasons. Candidate A loses only recall (it omitted a word). Candidate B loses both recall and precision (it misplaced "brown," which counts as a gap in the LCS from both the reference side and the candidate side).

This example illustrates how different metrics capture different quality aspects, and why multiple ROUGE variants are typically reported together. ROUGE-1 alone would misleadingly suggest Candidate B is perfect. ROUGE-2 penalizes B heavily for reordering. ROUGE-L finds a middle ground, recognizing that B preserves some structure but is less coherent than A.

Code Implementation

Let's implement ROUGE calculations from scratch to understand the mechanics, then compare with the standard rouge-score library.

In[6]:
Code
from collections import Counter


def preprocess(text, stem=False, remove_stopwords=False):
    """Simple preprocessing: lowercase, tokenize."""
    # Simple whitespace tokenization
    tokens = text.lower().split()

    # Optional: remove stopwords (simplified list)
    if remove_stopwords:
        stopwords = {
            "the",
            "a",
            "an",
            "is",
            "are",
            "was",
            "were",
            "be",
            "been",
            "being",
            "have",
            "has",
            "had",
            "do",
            "does",
            "did",
            "will",
            "would",
            "could",
            "should",
            "may",
            "might",
            "must",
            "shall",
            "can",
            "need",
            "dare",
            "ought",
            "used",
            "to",
            "of",
            "in",
            "for",
            "on",
            "with",
            "at",
            "by",
            "from",
            "as",
            "into",
            "through",
            "during",
            "before",
            "after",
            "above",
            "below",
            "between",
        }
        tokens = [t for t in tokens if t not in stopwords]

    # Optional: simple stemming (just remove 's', 'ing', 'ed' for demo)
    if stem:
        stemmed = []
        for t in tokens:
            if t.endswith("ing"):
                t = t[:-3]
            elif t.endswith("ed"):
                t = t[:-2]
            elif t.endswith("s") and not t.endswith("ss"):
                t = t[:-1]
            stemmed.append(t)
        tokens = stemmed

    return tokens


def get_ngrams(tokens, n):
    """Generate n-grams from a list of tokens."""
    return [tuple(tokens[i : i + n]) for i in range(len(tokens) - n + 1)]


def rouge_n(candidate, reference, n=1):
    """Calculate ROUGE-N recall score."""
    cand_tokens = preprocess(candidate)
    ref_tokens = preprocess(reference)

    cand_ngrams = get_ngrams(cand_tokens, n)
    ref_ngrams = get_ngrams(ref_tokens, n)

    if not ref_ngrams:
        return 0.0

    cand_counts = Counter(cand_ngrams)
    ref_counts = Counter(ref_ngrams)

    # Count matches (clip by reference count)
    matches = 0
    for ngram, ref_count in ref_counts.items():
        matches += min(cand_counts.get(ngram, 0), ref_count)

    return matches / len(ref_ngrams)


def lcs_length(seq1, seq2):
    """Calculate length of longest common subsequence using DP."""
    m, n = len(seq1), len(seq2)
    # Use 1D DP to save space
    prev = [0] * (n + 1)
    curr = [0] * (n + 1)

    for i in range(1, m + 1):
        for j in range(1, n + 1):
            if seq1[i - 1] == seq2[j - 1]:
                curr[j] = prev[j - 1] + 1
            else:
                curr[j] = max(prev[j], curr[j - 1])
        prev, curr = curr, prev

    return prev[n]


def rouge_l(candidate, reference, beta=1.0):
    """Calculate ROUGE-L F-measure based on LCS."""
    cand_tokens = preprocess(candidate)
    ref_tokens = preprocess(reference)

    if not ref_tokens or not cand_tokens:
        return 0.0

    lcs_len = lcs_length(cand_tokens, ref_tokens)

    if lcs_len == 0:
        return 0.0

    recall = lcs_len / len(ref_tokens)
    precision = lcs_len / len(cand_tokens)

    # F-beta score
    if recall + beta * beta * precision == 0:
        return 0.0

    f_score = ((1 + beta * beta) * recall * precision) / (
        recall + beta * beta * precision
    )
    return f_score


# Test with our example
reference = "The quick brown fox jumps over the lazy dog"
candidate_a = "The brown fox jumps over the lazy dog"
candidate_b = "The quick fox jumps over the brown lazy dog"

print("Reference:", reference)
print("Candidate A:", candidate_a)
print("Candidate B:", candidate_b)

The rouge_n function captures the key mechanics of the formula: it builds n-gram histograms for both candidate and reference, then counts matches clipped to the reference count. The lcs_length function uses space-efficient 1D DP, storing only two rows of the dynamic programming table at a time since each row only depends on the previous one.

Out[7]:
Console
ROUGE Scores:
--------------------------------------------------

Candidate A: 'The brown fox jumps over the lazy dog'
  ROUGE-1: 0.889
  ROUGE-2: 0.750
  ROUGE-L: 0.941

Candidate B: 'The quick fox jumps over the brown lazy dog'
  ROUGE-1: 1.000
  ROUGE-2: 0.625
  ROUGE-L: 0.889
Out[8]:
Visualization
Grouped bar chart comparing ROUGE-1, ROUGE-2, and ROUGE-L for Candidate A (blue) and Candidate B (coral) with score labels.
ROUGE score comparison between Candidate A (omission error) and Candidate B (reordering error) across three metric variants. ROUGE-1 fails to penalize reordering, awarding Candidate B a perfect score of 1.000. ROUGE-2 strongly penalizes both errors. ROUGE-L gives an intermediate signal, correctly ranking Candidate A above Candidate B while remaining less harsh than ROUGE-2.

Our implementation confirms the manual calculations: Candidate B scores higher on ROUGE-1 (1.000 vs 0.889) but lower on ROUGE-2 (0.625 vs 0.750), while Candidate A reaches a higher ROUGE-L score (0.941 vs 0.889) due to its perfect precision. The three metrics together tell a coherent story: Candidate A has the right structure but missed one word, while Candidate B has all the words but scrambled them.

Using the Standard Library

For production use, the rouge-score library gives optimized implementations with additional features like stemming and multiple reference handling.

In[9]:
Code
# Install if needed
## uv pip install rouge-score

import re
from collections import Counter, namedtuple

RougeScore = namedtuple("RougeScore", ["precision", "recall", "fmeasure"])


def _tokens(text):
    return re.findall(r"\b\w+\b", text.lower())


def _ngram_counts(tokens, n):
    return Counter(tuple(tokens[i : i + n]) for i in range(len(tokens) - n + 1))


def _f_score(overlap, candidate_total, reference_total):
    precision = overlap / candidate_total if candidate_total else 0.0
    recall = overlap / reference_total if reference_total else 0.0
    fmeasure = (
        2 * precision * recall / (precision + recall)
        if precision + recall
        else 0.0
    )
    return RougeScore(precision, recall, fmeasure)


def _lcs_length(a, b):
    prev = [0] * (len(b) + 1)
    for x in a:
        curr = [0]
        for j, y in enumerate(b, start=1):
            curr.append(prev[j - 1] + 1 if x == y else max(prev[j], curr[-1]))
        prev = curr
    return prev[-1]


class SimpleRougeScorer:
    def score(self, reference, candidate):
        ref = _tokens(reference)
        cand = _tokens(candidate)
        scores = {}
        for name, n in [("rouge1", 1), ("rouge2", 2)]:
            ref_counts = _ngram_counts(ref, n)
            cand_counts = _ngram_counts(cand, n)
            overlap = sum(
                min(count, cand_counts[gram])
                for gram, count in ref_counts.items()
            )
            scores[name] = _f_score(
                overlap, sum(cand_counts.values()), sum(ref_counts.values())
            )
        lcs = _lcs_length(cand, ref)
        scores["rougeL"] = _f_score(lcs, len(cand), len(ref))
        return scores


scorer = SimpleRougeScorer()

# Compare our candidates
scores_a = scorer.score(reference, candidate_a)
scores_b = scorer.score(reference, candidate_b)
Out[10]:
Console
Official rouge-score library results:
--------------------------------------------------

Candidate A:
  rouge1: P=1.000, R=0.889, F=0.941
  rouge2: P=0.857, R=0.750, F=0.800
  rougeL: P=1.000, R=0.889, F=0.941

Candidate B:
  rouge1: P=1.000, R=1.000, F=1.000
  rouge2: P=0.625, R=0.625, F=0.625
  rougeL: P=0.889, R=0.889, F=0.889

The library reports precision, recall, and F-measure for each variant. Note that ROUGE-N in the library defaults to F-measure rather than pure recall, which differs from the original ROUGE paper but aligns with modern usage where balanced metrics are preferred. Examining all three components (precision, recall, and F-measure) is more informative than F-measure alone, since it reveals whether a low F-score results from poor recall (missing content), poor precision (excessive verbosity), or both.

Handling Multiple References

Real-world summarization datasets often give multiple human-written references for the same source document. A generated summary might match different references better for different aspects of the content. The CNN/DailyMail dataset, for example, includes multiple highlight sentences as references for each article, and the DUC conferences provided up to four human references per document.

In[11]:
Code
references = [
    "The quick brown fox jumps over the lazy dog",
    "A fast brown fox leaped over a lazy dog",
    "The brown fox quickly jumped over the lazy dog",
]

candidate = "The brown fox jumps over the lazy dog"

# Score against each reference
all_scores = []
for ref in references:
    scores = scorer.score(ref, candidate)
    all_scores.append(scores)
Out[12]:
Console
Multi-reference scoring:
--------------------------------------------------

Candidate: 'The brown fox jumps over the lazy dog'

Scores per reference:
  Ref 1: ROUGE-1=0.941, ROUGE-2=0.800, ROUGE-L=0.941
  Ref 2: ROUGE-1=0.588, ROUGE-2=0.267, ROUGE-L=0.588
  Ref 3: ROUGE-1=0.824, ROUGE-2=0.667, ROUGE-L=0.824

Max scores (standard approach):
  ROUGE-1: 0.941
  ROUGE-2: 0.800
  ROUGE-L: 0.941

Average scores:
  ROUGE-1: 0.784
  ROUGE-2: 0.578
  ROUGE-L: 0.784

The standard approach takes the maximum score across references, acknowledging that a summary only needs to align well with one good reference. However, some evaluations report average scores to penalize summaries that match only one reference well while diverging substantially from others. The choice between max and average depends on whether you conceptualize the references as alternative valid summaries (max is appropriate: align with any one of them) or as complementary summaries covering different aspects (average is appropriate: cover all of them to some degree).

Key Parameters

The key parameters for ROUGE score calculation are:

  • n: The n-gram order for ROUGE-N (typically 1 or 2). ROUGE-1 measures unigram overlap (content coverage), while ROUGE-2 measures bigram overlap (local word ordering). Higher orders (ROUGE-3, ROUGE-4) are rarely reported because they are too strict for typical summary lengths, often resulting in scores near zero due to the difficulty of matching long consecutive sequences.

  • beta: The balance parameter for ROUGE-L F-measure. When β>1\beta > 1, recall is weighted more heavily than precision. Standard implementations typically use β=1\beta = 1 (F1), though some favor recall with β=1.2\beta = 1.2. The choice depends on whether you want to penalize omissions or verbosity more severely.

  • use_stemmer: Whether to apply Porter stemming before matching. Stemming normalizes words to their root forms (e.g., "running" to "run"), improving robustness to morphological variation. Most modern implementations enable stemming by default to ensure that grammatical variations don't unfairly penalize otherwise correct summaries.

  • remove_stopwords: Whether to filter out common words ("the", "and", "is") before calculation. This focuses evaluation on content words but may discard important discourse markers. When reporting results, you should specify whether stopwords were removed to ensure comparability with other studies.

Out[13]:
Visualization
Line chart of ROUGE-L F-measure versus beta from 0.1 to 2.0 for two candidates, with vertical markers at beta 1 and 1.2.
Effect of the beta parameter on ROUGE-L F-measure for Candidate A (perfect precision, recall 0.889) and Candidate B (precision 0.889, recall 0.889). As beta increases, the F-measure weights recall more heavily, so Candidate A's score drops toward its recall value (0.889) while Candidate B's score remains constant. The vertical dashed lines mark beta equals 1 (standard F1) and beta equals 1.2 (recall-favoring).

The beta sensitivity plot reveals an important property of ROUGE-L: when a candidate has perfect precision (every word it contains is in the right order relative to the reference), higher beta values decrease its score by weighting the imperfect recall more heavily. When precision equals recall, the F-measure is constant across all beta values, as with Candidate B's symmetric scores.

ROUGE in Practice: Benchmarks and Correlation with Human Judgment

Understanding ROUGE scores in isolation is less useful than understanding what they mean in the context of real benchmarks. The CNN/DailyMail dataset, introduced by Hermann et al. in 2015, has become the canonical testbed for news summarization. State-of-the-art neural systems on this dataset typically reach ROUGE-1 scores around 0.44-0.46, ROUGE-2 scores around 0.21-0.23, and ROUGE-L scores around 0.40-0.43. Earlier extractive systems from the 2000s achieved ROUGE-1 scores around 0.35-0.38, making the improvement from neural methods visually apparent as roughly 7-10 ROUGE-1 points over fifteen years of progress.

However, interpreting these numbers requires caution. A ROUGE-1 improvement of 1 point sounds modest but often is real progress when validated by human evaluation. Conversely, a 2-point improvement achieved through clever preprocessing or by exploiting dataset-specific patterns (such as including the first sentence of every article, which tends to be informative in news) may not generalize to other settings.

The correlation between ROUGE and human judgments has been extensively studied. Early work by Lin (2004) showed reasonable correlation for single-document news summarization, with Pearson correlations between ROUGE-1/ROUGE-2 and human readability and informativeness scores around 0.8-0.9 in controlled settings. However, later work by Callison-Burch et al. and others showed that this correlation degrades materially when:

  1. Systems use diverse paraphrasing (abstractive summaries)
  2. The domain differs from news (legal, scientific, creative texts)
  3. Reference summaries are not representative of the diversity of valid summaries
  4. System quality is already high and differences between systems are subtle

This degradation has motivated significant investment in alternative evaluation approaches, which we explore in the limitations section.

Out[14]:
Visualization
Horizontal bar chart showing ROUGE-1 score ranges for different system types from classical extractive to modern neural summarization.
Illustrative ROUGE-1 score ranges across different summarization system types and time periods, showing the progression from early extractive systems to modern neural models on the CNN/DailyMail benchmark. Neural abstractive systems have achieved higher ROUGE scores than classical extractive methods, though the gap has narrowed as both paradigms improved. Values are approximate and representative of published results.

The progression shown here illustrates an important point about ROUGE as a research signal. The Lead-3 baseline (simply taking the first three sentences of a news article) reaches surprisingly high ROUGE scores because news articles follow an inverted pyramid structure where the most important information appears first. This means that a ROUGE score below the Lead-3 baseline is a signal that something is wrong with the system, while scores above it indicate real learning of summarization patterns beyond position heuristics.

Limitations and Impact

ROUGE has served as the de facto standard for summarization evaluation for two decades, letting rapid development and comparison of summarization systems. Its simplicity and computational efficiency have made large-scale benchmarking possible, driving progress from statistical methods to neural approaches. However, its limitations have become increasingly apparent as neural summarization models have grown more advanced and capable of generating novel phrasing rather than extracting sentences.

Semantic Blindness

Like BLEU, ROUGE operates at the surface level of text. It cannot recognize paraphrases, synonyms, or semantic equivalence expressed with different words. A candidate summary stating "The feline consumed the rodent" receives zero ROUGE score against a reference saying "The cat ate the mouse," despite perfect semantic alignment. This limitation extends beyond simple synonyms to encompass entire paraphrased sentences. If a reference states "The company announced disappointing quarterly earnings," and a candidate reads "The firm revealed poor financial results for the past three months," ROUGE scores will be low despite the identical informational content.

This semantic blindness becomes particularly problematic when evaluating models trained on diverse corpora that have learned rich synonym sets. A language model might generate a perfectly accurate and readable summary that uses different vocabulary than the reference, simply because its training distribution favors different phrasing. Penalizing this stylistic variation as if it were an error conflates surface form with semantic content, creating misleading evaluation signals.

This limitation is particularly acute for abstractive summarization, where neural models generate novel phrasing rather than extracting sentences. Modern models often produce summaries that are fluent and accurate but receive low ROUGE scores due to lexical variation. As the field moves toward more abstractive systems, the disconnect between ROUGE scores and human judgment has widened, prompting the development of semantic evaluation metrics like BERTScore.

Focus on Lexical Overlap and Extractive Bias

ROUGE rewards extractive behavior. A system that copies sentences verbatim from the source will score higher than one that skillfully paraphrases, even when paraphrasing produces better summaries. This has led to concerns that optimizing for ROUGE encourages extractive or "copy-and-paste" summarization rather than true understanding and condensation. Researchers have observed that models trained to maximize ROUGE often learn to select and concatenate sentences from the source document rather than synthesizing information, limiting the development of truly abstractive capabilities.

This extractive bias creates a perverse feedback loop in research. If ROUGE is the primary benchmark, systems that reach state-of-the-art scores by clever sentence selection will attract more attention and follow-up work than systems that reach lower scores through real abstractive synthesis. The benchmark thus shapes the research direction in ways that might not correspond to the ultimate goal of building systems that truly understand and condense information.

Research has shown that ROUGE scores correlate poorly with human judgments when evaluating abstractive summaries, though they remain reasonable for extractive summarization where word overlap is expected. A seminal study by Kryscinski et al. (2019) found that neural summarization models could reach high ROUGE scores while containing significant factual errors, because ROUGE cannot verify whether claims in the summary are supported by the source document. A summary might say "The company reported a 50 percent increase in profits" when the source says "a 15 percent increase," scoring well on ROUGE for capturing the right vocabulary but failing completely on the factual accuracy criterion that matters most for news summarization.

Length Bias

ROUGE scores tend to favor longer summaries. A verbose candidate that includes many words from the reference will score higher on recall than a concise summary that captures only the needed information. While precision metrics (available in ROUGE-L) help mitigate this, the dominant reporting of recall-oriented scores encourages redundancy. A candidate that repeats the same information multiple times or includes tangential details from the source can inflate ROUGE scores without improving utility.

Consider a reference summary of 50 words. A candidate that is 150 words long and includes every word from the reference plus additional detail will score 1.0 on ROUGE-N recall, even though it is three times the target length. Human judges would likely penalize this verbosity, but ROUGE rewards it. This is why length constraints are typically imposed during generation (using maximum token limits), since the metric itself gives no natural incentive for conciseness.

Some implementations address this by penalizing candidates materially longer than the reference, but this adds arbitrary thresholds not present in the original metric. Others report F-measure rather than pure recall to balance coverage against conciseness, though this still doesn't perfectly capture the human preference for information density.

Domain Sensitivity

ROUGE performance varies materially across domains. In technical domains with standardized terminology (like medical or legal summarization), ROUGE correlates reasonably well with quality because key terms are fixed. A medical summary mentioning "myocardial infarction" will likely match references using the same term, making lexical overlap a valid quality signal. In creative or subjective domains (like story summarization or opinion summarization), where multiple valid summaries might use completely different vocabulary, ROUGE gives little signal. A summary of a novel might focus on character development, plot, or themes, with virtually no word overlap between equally valid summaries.

This domain sensitivity means that ROUGE scores are not transferable across domains. A ROUGE-2 score of 0.25 on news summarization is good, while the same score on scientific paper summarization might be poor, and on creative writing summarization might be excellent. Researchers must always contextualize ROUGE scores with domain-specific baselines and human evaluation to know whether a given score is high or low quality.

The Path Forward

Despite these limitations, ROUGE remains useful as a diagnostic tool and for tracking relative improvements in controlled settings. When a new model reaches higher ROUGE scores than a previous model on the same dataset with the same references, it likely is real progress, even if the absolute score doesn't indicate human-level quality. ROUGE is a necessary but insufficient condition for quality: low ROUGE almost certainly indicates a poor summary, while high ROUGE suggests the summary captures reference content, though it may still suffer from semantic errors or poor readability.

The research community has developed several approaches to complement ROUGE's limitations. BERTScore uses contextual embeddings to compute token-level similarity between candidate and reference, letting semantic matching even when vocabulary differs. BLEURT fine-tunes a BERT model on human judgment data, learning to predict quality scores that correlate better with human assessments. FactCC and similar factuality metrics specifically check whether claims in the summary are entailed by the source document, addressing the factual accuracy gap that ROUGE ignores entirely.

For more fine-grained evaluation, ROUGE is increasingly supplemented by:

  • BERTScore, which uses contextual embeddings to measure semantic similarity rather than lexical overlap (covered in the next chapter)
  • Human evaluation for fluency, coherence, and factual consistency
  • Task-specific metrics like factuality checking for news summarization, which verify that claims in the summary are supported by the source document
  • Questionnaire-based evaluation (like QAGS and QuestEval), which decompose summary quality into a set of yes/no questions about specific pieces of information

The field is moving toward multi-dimensional evaluation that separates content selection (what information is included) from surface realization (how it is expressed), recognizing that a single number cannot capture summary quality. ROUGE remains a cornerstone of this ecosystem. This gives a cheap, automatic signal for content coverage, while newer metrics handle semantic fidelity and human evaluation assesses ultimate utility.

An important practical note: despite its limitations, ROUGE's longevity as a benchmark has itself become a feature. Decades of published results use ROUGE, which makes it the lingua franca for comparing systems across time and across papers. Even if you personally prefer BERTScore, reporting ROUGE alongside it ensures your results are interpretable to the broader community and comparable to historical baselines. The social and scientific value of a common benchmark language should not be underestimated.

Summary

ROUGE gives a family of recall-oriented metrics specifically designed for summarization evaluation. Unlike BLEU's precision-focused n-gram matching for translation, ROUGE measures how much of the reference content appears in the generated summary, prioritizing coverage over conciseness.

The four main variants serve different purposes:

  • ROUGE-1 (unigram recall) captures content coverage at the word level. It verifies that the right vocabulary is present but cannot assess ordering or fluency. It is the most permissive metric and should always be paired with higher-order variants.
  • ROUGE-2 (bigram recall) captures local word ordering. It penalizes summaries with the right words but scrambled sequences, which makes it a better indicator of linguistic quality and syntactic coherence.
  • ROUGE-L uses longest common subsequence to capture sentence-level structure and word ordering without requiring consecutive matches. It balances precision and recall through F-measure calculation. This gives robustness to paraphrasing that reorganizes sentences while keeping key phrases.
  • ROUGE-W applies superlinear weighting to consecutive matches, penalizing scattered word alignments and rewarding fluent, cohesive summaries. It is less commonly used due to the non-standardized weighting parameter.
  • ROUGE-S measures skip-bigram co-occurrence, capturing ordered word pairs with arbitrary gaps. It occupies a middle ground between the strict locality of ROUGE-2 and the global flexibility of ROUGE-L.

Critical considerations for using ROUGE correctly:

  • ROUGE is recall-oriented, which makes it suitable for summarization where coverage matters more than precision. This shows the task's priority of capturing needed information over avoiding extraneous content.
  • Preprocessing choices (stemming, stopword removal, tokenization) materially impact scores and must be reported for reproducibility. Two results using different preprocessing are not directly comparable.
  • Multiple references are handled by taking the maximum score across references, acknowledging the diversity of valid summaries for a single document.
  • Modern neural summarizers often produce valid summaries that receive low ROUGE scores due to lexical variation. This shows the metric's limitations for abstractive evaluation.
  • Domain context determines what constitutes a good ROUGE score. Always compare against domain-specific baselines, including the Lead-3 baseline for news.

While ROUGE has enabled rapid progress in summarization research, its reliance on surface-form matching limits its utility for evaluating abstractive systems. Current best practice combines ROUGE with semantic metrics like BERTScore and targeted human evaluation to capture the full picture of summary quality. This keeps automatic metrics serve their intended purpose without constraining the development of more advanced summarization capabilities.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about ROUGE scores for summarization evaluation.

ROUGE Scores Quiz

Question 1 of 70 of 7 completed
What is the primary distinction between ROUGE and BLEU in terms of their evaluation orientation?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026rougescores, author = {Michael Brenndoerfer}, title = {ROUGE Scores: Evaluating Text Summarization}, year = {2026}, url = {https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). ROUGE Scores: Evaluating Text Summarization. Retrieved from https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics
MLAAcademic
Michael Brenndoerfer. "ROUGE Scores: Evaluating Text Summarization." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics>.
CHICAGOAcademic
Michael Brenndoerfer. "ROUGE Scores: Evaluating Text Summarization." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics.
HARVARDAcademic
Michael Brenndoerfer (2026) 'ROUGE Scores: Evaluating Text Summarization'. Available at: https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). ROUGE Scores: Evaluating Text Summarization. https://mbrenndoerfer.com/writing/rouge-scores-text-summarization-evaluation-metrics

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.