Text Summarization: Extractive and Abstractive Methods

Michael BrenndoerferJanuary 19, 202667 min read

Part of Language AI Handbook

Covers extractive and abstractive summarization, from TextRank and MMR to BART and LLMs, with ROUGE and BERTScore evaluation techniques.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Summarization

Reading an entire book to answer a single question, or wading through a 50-page research paper to find a key finding, is a frustrating experience that most of us know well. Automatic summarization tackles this problem directly: given a long document or collection of documents, produce a shorter text that preserves the most important information. The goal is intelligent compression that keeps what matters and discards what does not.

Text summarization is one of the oldest problems in natural language processing, and one of the hardest. It requires a system to understand content at multiple levels simultaneously: individual word meanings, sentence relationships, document structure, and the pragmatic question of what a reader needs. A good summary distills the document instead of copying its first sentences or sampling at random that respects both what was said and why it matters.

The field divides naturally into two approaches: extractive and abstractive summarization. Extractive methods identify and copy the most important sentences from the original document. Abstractive methods generate new text that captures the meaning, potentially using words and phrasings that never appeared in the source. As we will see, these approaches have very different technical requirements, strengths, and failure modes. We have covered the foundations of language modeling and generation in earlier parts of this handbook; this chapter applies those foundations to the specific challenge of summarization, with attention to evaluation, practical applications, and the limitations that still make this a hard open problem.

The history of automatic summarization stretches back to the 1950s, when Hans Peter Luhn at IBM proposed selecting sentences based on word frequency statistics. The intuition was simple: words that occur more frequently in a document signal its topic, and sentences packed with such words are the most representative. For decades, progress was slow because the task ultimately requires general language understanding, something classical statistics alone could not provide. The rise of neural sequence-to-sequence models in 2015 and 2016 brought the first fully abstractive systems capable of paraphrasing rather than copying. Large pretrained models like BART and T5 then pushed quality to human-competitive levels on standard benchmarks, though faithfulness and domain generalization remain active research challenges. Understanding the full spectrum from classical extractive methods to modern neural generation reveals why both approaches remain relevant and valuable today.

Before diving into the technical details, it is worth asking what properties we want a good summary to have. Informativeness: the summary should convey the most important content from the source. Conciseness: the summary should be substantially shorter than the source. Coherence: the summary sentences should flow naturally and not feel like a random collection of fragments. Faithfulness: the summary should not introduce facts, claims, or relationships that were not present in the source. And relevance: for query-focused summarization, the summary should specifically address what the reader needs to know, not just what is most prominent in the document. These properties often conflict. An extractive system achieves high faithfulness by copying, but copying limits coherence when the selected sentences come from different parts of the document. An abstractive system can produce highly coherent, fluent output but at the cost of faithfulness risk. The right tradeoff depends entirely on the application.

Extractive Summarization

Extractive summarization frames the problem as a selection task: which sentences from the document are most important? This formulation avoids the harder generation problem entirely and guarantees that the output is grammatical (since the sentences come from a human-written source). The challenge shifts to scoring and ranking.

The simplest scoring approach treats sentences as bags of words and asks: which sentences contain the most representative vocabulary? This intuition leads to frequency-based methods, where words that appear often in the document are assumed to be important, and sentences that contain many such words score highly. While simple, this insight captures something real: documents are about their topic, and the topic is reflected in which words get used repeatedly. If a document about climate change uses "carbon dioxide," "emissions," and "temperature" throughout, then sentences containing those words are likely to be more central to the document's purpose than sentences about peripheral context.

The major limitation of frequency-based scoring is that it treats all individual sentences in isolation, without regard to how they relate to each other. A sentence that shares vocabulary with many other sentences is more central to the document's topic network than one that shares vocabulary with only a few. This observation motivates graph-based approaches that model sentence relationships explicitly.

TF-IDF Based Scoring

Building on the classical text representation techniques we explored in Chapter 2, TF-IDF scoring applies directly to sentence extraction. In a multi-document summarization context, TF-IDF helps identify terms that are important within a document but not too common across all documents in the collection. For single-document summarization, a simpler frequency weighting suffices.

Let f(w,s)f(w, s) denote the frequency of word ww in sentence ss, and F(w)F(w) denote the total frequency of word ww in the document. A basic sentence score is:

score(s)=∑w∈sf(w,s)F(w)\text{score}(s) = \sum_{w \in s} \frac{f(w, s)}{F(w)}

where:

  • ss: the sentence being scored
  • ww: a word in sentence ss
  • f(w,s)f(w, s): the count of word ww in sentence ss
  • F(w)F(w): the total count of word ww across the entire document

This formulation gives higher scores to sentences containing words that are locally concentrated (high f(w,s)f(w, s)) but not ubiquitous throughout the document (moderate F(w)F(w)). The intuition is that words repeated heavily in one part of the document signal a local topic focus. Stopwords like "the" or "is" would have high F(w)F(w) relative to f(w,s)f(w, s), pulling their contribution toward zero and effectively filtering them out without requiring an explicit stopword list.

In practice, you preprocess text by removing stopwords and applying stemming or lemmatization before computing frequencies. This prevents common function words from dominating the scores and ensures that "running," "ran," and "runs" are treated as the same underlying concept. The resulting score approximates how much each sentence is "about" the main topic of the document, at least as measured by shared vocabulary.

The limitation of pure frequency scoring is that it treats all frequent words as equally informative. In a document about machine learning, "model" and "data" appear constantly and score well, but they carry little discriminative information. The TF-IDF adjustment in multi-document settings addresses this by downweighting terms that are common across many documents. For single-document summarization, you often need additional heuristics to distinguish informative sentences from those that merely restate background context. This gap between what is frequent and what is important motivates more sophisticated approaches.

Graph-Based Methods: TextRank

A significant advancement over simple frequency counting is the graph-based approach, exemplified by TextRank (Mihalcea and Tarau, 2004). Rather than scoring sentences in isolation, TextRank treats the summarization problem as a graph ranking problem, similar in spirit to Google's PageRank algorithm for web pages.

The core idea is elegant: a sentence is important if it is similar to many other important sentences. Importance propagates through similarity relationships rather than being assigned independently to each sentence. This recursive definition captures something that frequency scoring misses: a sentence might not use many distinctive words, but if it is thematically central (similar to a large fraction of the document's other sentences), it is probably important. A sentence introducing a key concept that is then discussed throughout the rest of the document will share vocabulary with all those later sentences and therefore accumulate a high score even if it does not repeat distinctive terms internally.

The algorithm works in three steps. First, build a graph where each node is a sentence. Then add weighted edges between sentence pairs, where the edge weight captures their similarity. Finally, run a recursive ranking algorithm to assign importance scores, letting importance flow from sentence to sentence along similarity edges.

The similarity between two sentences sis_i and sjs_j is typically measured using a normalized word-overlap formula. Let WiW_i denote the set of unique words in sentence sis_i and WjW_j the set of unique words in sjs_j. The similarity is:

sim(si,sj)=∣Wi∩Wj∣log⁡∣Wi∣+log⁡∣Wj∣\text{sim}(s_i, s_j) = \frac{|W_i \cap W_j|}{\log|W_i| + \log|W_j|}

where:

  • Wi∩WjW_i \cap W_j: the set of words appearing in both sentences
  • ∣Wi∩Wj∣|W_i \cap W_j|: the number of shared words (the intersection size)
  • log⁡∣Wi∣\log|W_i| and log⁡∣Wj∣\log|W_j|: length normalization terms that prevent longer sentences from automatically scoring higher simply because they contain more words

Without the log normalization in the denominator, a 50-word sentence and a 10-word sentence sharing 5 words would produce the same raw overlap as a 10-word sentence and a 10-word sentence sharing 5 words, even though the first pair is less semantically aligned. The logarithm applies a soft penalty that grows with sentence length but does not grow as fast as the raw word count would. This normalization is important in practice because documents often contain sentences of very different lengths, and naive overlap would systematically favor the longer ones.

The ranking score is then computed iteratively. Let WS(si)\text{WS}(s_i) denote the score assigned to sentence sis_i. Starting from a uniform initialization, scores are updated according to:

WS(si)=(1−d)+d⋅∑sj∈In(si)wji∑sk∈Out(sj)wjk⋅WS(sj)\text{WS}(s_i) = (1 - d) + d \cdot \sum_{s_j \in \text{In}(s_i)} \frac{w_{ji}}{\sum_{s_k \in \text{Out}(s_j)} w_{jk}} \cdot \text{WS}(s_j)

where:

  • dd: a damping factor (typically d=0.85d = 0.85), representing the probability that a random traversal follows a similarity link rather than jumping uniformly to a random sentence
  • In(si)\text{In}(s_i): the set of sentences that link to sis_i (all other sentences in a fully connected graph)
  • wjiw_{ji}: the weight of the directed edge from sjs_j to sis_i, equal to sim(sj,si)\text{sim}(s_j, s_i)
  • ∑sk∈Out(sj)wjk\sum_{s_k \in \text{Out}(s_j)} w_{jk}: the sum of all edge weights leaving sjs_j, used to normalize so that outgoing weights sum to 1

The term (1−d)(1 - d) provides a uniform base score, preventing sentences with no incoming edges from scoring zero. This mirrors the "teleportation" term in PageRank. The iteration continues until scores converge, typically within 30 to 50 steps. Once scores stabilize, sentences are ranked in descending order and the top kk are selected.

The convergence property of this iteration is guaranteed by the Perron-Frobenius theorem: because the graph is fully connected (every sentence has an edge to every other sentence, even if the weight is very small) and the damping factor keeps a non-zero base probability, the transition matrix is ergodic and the power iteration converges to a unique stationary distribution. In practice, you can check convergence by measuring the maximum absolute difference between successive score vectors and stopping when this falls below a small threshold like 10−610^{-6}.

A key strength of TextRank over frequency-based methods is that it does not require any external knowledge or training data. It is fully unsupervised and works on any domain. This makes it useful for documents in specialized areas where training data is scarce: medical records, legal documents, or internal corporate reports. It is also fast, running in time proportional to the square of the number of sentences, which is manageable for documents of typical length.

PageRank Intuition

TextRank's connection to PageRank is deeper than analogy. Both algorithms solve the problem of finding the most "central" nodes in a graph, where centrality is defined recursively: a node is central if it is linked to by other central nodes. In web ranking, links from important pages make a page important. In sentence ranking, similarity to important sentences makes a sentence important. The mathematical formulation is identical; only the interpretation of nodes and edges changes.

Position-Based Heuristics

Not all extractive approaches rely on content analysis. A surprisingly effective heuristic is position: sentences near the beginning of a document are often more important than sentences in the middle, because good writing typically front-loads key information.

In news articles, the "inverted pyramid" structure ensures the most important information appears first: the first paragraph answers who, what, when, where, and why, and subsequent paragraphs add detail in decreasing order of importance. This structure evolved from practical necessity: when newspapers had to cut articles from the bottom to fit available space, the most important content needed to survive truncation. An automatic summarizer that knows this convention can simply take the first one or two sentences of a news article and perform surprisingly well.

In scientific papers, the abstract, introduction, and conclusion sentences carry disproportionate importance. In long-form essays, section opening sentences often summarize what follows. In structured documents like annual reports, the executive summary and key metrics paragraphs carry most of the actionable content. These structural patterns are not accidental. They reflect conventions that writers adopt to help readers move through long documents. An automatic summarizer that ignores these conventions is throwing away strong signal that comes for free.

A common approach combines position scores with content scores using a weighted sum. If content_score(s)\text{content\_score}(s) is a frequency or graph-based importance score and position_score(s)\text{position\_score}(s) captures where the sentence appears (typically 1/rank(s)1/\text{rank}(s) where the rank is 1 for the first sentence), the final score is:

final_score(s)=α⋅content_score(s)+(1−α)⋅position_score(s)\text{final\_score}(s) = \alpha \cdot \text{content\_score}(s) + (1 - \alpha) \cdot \text{position\_score}(s)

where:

  • α∈[0,1]\alpha \in [0, 1]: a mixing weight controlling the balance between content and position signals
  • content_score(s)\text{content\_score}(s): frequency or graph-based importance score
  • position_score(s)\text{position\_score}(s): typically 1/rank(s)1/\text{rank}(s) where rank(s)\text{rank}(s) is the sentence's position in the document (1-indexed)

The optimal α\alpha varies by document type. News articles benefit from heavy position weighting because their structure is highly regular. Research papers benefit from content weighting with position bonuses for certain structural positions (abstract sentences, section headings). For highly conversational or unstructured text, position provides almost no useful signal and α\alpha should be close to 1. Choosing α\alpha is itself a modeling decision that can be tuned if labeled data is available, or set heuristically based on domain knowledge.

Redundancy and Coverage

A naive sentence ranking approach has a critical flaw: it will often select multiple sentences that say essentially the same thing. The top-scored sentences in a document about climate change might all be variants of "global temperatures are rising," producing a summary that is accurate but uninformative as a whole. Readers benefit much more from a summary that covers different aspects of the topic than from one that repeats the most prominent theme three times.

This problem is especially acute for TextRank. Because the algorithm assigns high scores to sentences that are similar to many other sentences, it naturally tends to select sentences that are all similar to each other. The most "central" sentences in a document cluster around the main theme. Selecting the top three by score can give you three near-paraphrases of the same core idea. You end up with a summary that is highly faithful to the main topic but completely fails to convey the document's breadth.

This motivates Maximum Marginal Relevance (MMR), a selection criterion that balances relevance against redundancy. When selecting the next sentence to add to the summary, MMR scores each remaining candidate ss as:

MMR(s)=λ⋅Sim1(s,Q)−(1−λ)⋅max⁡sj∈SSim2(s,sj)\text{MMR}(s) = \lambda \cdot \text{Sim}_1(s, Q) - (1 - \lambda) \cdot \max_{s_j \in S} \text{Sim}_2(s, s_j)

where:

  • QQ: the query or topic of interest (the document as a whole in generic summarization)
  • SS: the set of sentences already selected for the summary
  • Sim1(s,Q)\text{Sim}_1(s, Q): similarity of candidate ss to the query, measuring relevance
  • Sim2(s,sj)\text{Sim}_2(s, s_j): similarity of candidate ss to an already-selected sentence sjs_j, measuring redundancy
  • λ∈[0,1]\lambda \in [0, 1]: a parameter controlling the relevance-redundancy tradeoff; λ=1\lambda = 1 gives pure relevance ranking, λ=0\lambda = 0 gives pure diversity

The greedy MMR algorithm selects the highest-scoring sentence under this criterion, adds it to SS, recomputes scores for remaining candidates (since the redundancy term now includes the newly added sentence), and repeats until the desired summary length is reached.

This algorithm is greedy. Because MMR selects one sentence at a time and updates the redundancy calculation after each selection, the chosen sentences depend on the order of selection. The first sentence chosen is the most relevant to the query (no penalty for redundancy yet). The second sentence is the one that best balances relevance with distinctiveness from the first. This sequential process generally produces better coverage than selecting all kk sentences simultaneously by pure relevance, because each new selection is made with full awareness of what has already been included.

Maximum Marginal Relevance

MMR was originally developed for query-based document retrieval, where the goal was to return a diverse set of relevant documents rather than kk copies of the most relevant one. Its application to summarization is a natural extension: both problems require balancing individual item quality against collection-level diversity. The same algorithm applies whenever you want a small representative set from a larger pool, whether you are selecting documents, sentences, or any other unit of content.

Neural Extractive Summarization

Classical extractive methods treat sentence scoring as an independent classification or ranking problem, but neural extractive models learn richer sentence representations that capture contextual meaning rather than just surface word overlap.

Neural extractive summarizers typically work in two stages. First, a pre-trained encoder (such as BERT or a similar transformer model) processes the entire document and produces contextual embeddings for each sentence. Unlike TF-IDF or word overlap features, these embeddings capture semantic meaning: "automobile" and "car" would produce similar embeddings even though they share no vocabulary. Second, a classification head predicts whether each sentence should be included in the summary, trained on labeled datasets where human annotators marked summary-worthy sentences.

The key advantage over classical methods is that the encoder understands language rather than merely counting words. A sentence might be identified as important because it introduces a concept (even if that concept uses unusual vocabulary), because it contains a causal claim that other sentences elaborate on, or because it appears to be a thesis statement even in mid-document position.

Models like BertSum (Liu and Lapata, 2019) demonstrated that fine-tuning BERT for extractive summarization substantially outperforms classical methods on standard benchmarks. The model adds a special [CLS] token at the beginning of each sentence, and the representation of this token after BERT encoding is used as the sentence-level feature for the classification decision. By training jointly over all sentences in the document, the model can also capture cross-sentence dependencies that sentence-by-sentence scoring cannot.

The main limitation of neural extractive approaches is that they require labeled training data: you need documents where human annotators have identified which sentences are summary-worthy. For common domains like news, such datasets exist. For specialized domains, creating this annotation data is expensive and time-consuming. Classical methods like TextRank remain valuable precisely because they require no labeled data at all.

Abstractive Summarization

Abstractive summarization is fundamentally different from extraction. Instead of copying sentences, an abstractive system reads the source and generates a new text from scratch. This allows the summary to paraphrase, combine ideas from multiple sentences, omit irrelevant details, and express concepts in cleaner or more concise language than the original.

The cost of this flexibility is substantial technical complexity. An abstractive summarizer must do what extractive systems do not: understand content deeply enough to paraphrase it, decide what to omit without losing meaning, and produce fluent natural language. This is a generation problem, requiring all the machinery of language modeling we have built up in earlier parts of this handbook.

The practical implications are significant. Extractive systems can be deployed with no training data and run in milliseconds on a laptop. Abstractive systems require gigabytes of training data, substantial compute for training, and inference that is orders of magnitude slower. But when they work well, abstractive summaries are qualitatively superior: more concise, more fluent, and more informative than any extractive alternative could produce. An abstractive system can fuse facts from three separate paragraphs into a single tight sentence, something extraction can never achieve.

Sequence-to-Sequence Architecture

The dominant framework for abstractive summarization is sequence-to-sequence (seq2seq) learning, which we covered in depth in Chapter 9. A seq2seq model consists of an encoder that reads the source document and produces a latent representation, and a decoder that generates the summary token by token conditioned on that representation.

For summarization, the encoder processes the full input document. Each input token xix_i is transformed into a hidden state hih_i as part of the full encoding:

(h1,h2,…,hn)=Encoder(x1,x2,…,xn)(h_1, h_2, \ldots, h_n) = \text{Encoder}(x_1, x_2, \ldots, x_n)

where:

  • x1,x2,…,xnx_1, x_2, \ldots, x_n: the sequence of input tokens (the source document)
  • hih_i: the hidden state corresponding to input token xix_i, encoding contextual information about xix_i in the context of all other tokens
  • nn: the number of input tokens

The decoder then generates summary tokens autoregressively. At each time step tt, it predicts the next token yty_t conditioned on all encoder hidden states and all previously generated tokens:

P(yt∣y1,…,yt−1,x1,…,xn)=softmax(Wo⋅dt)P(y_t \mid y_1, \ldots, y_{t-1}, x_1, \ldots, x_n) = \text{softmax}\bigl(W_o \cdot d_t\bigr)

where:

  • dt=Decoder(h1,…,hn,y1,…,yt−1)d_t = \text{Decoder}(h_1, \ldots, h_n, y_1, \ldots, y_{t-1}): the decoder hidden state at time step tt
  • WoW_o: the output projection matrix mapping decoder hidden states to vocabulary logits
  • softmax(⋅)\text{softmax}(\cdot): the softmax function normalizing logits into a probability distribution over the vocabulary VV

Training uses teacher forcing: the model is trained to predict each ground-truth token given the true previous tokens, using cross-entropy loss summed over all positions:

L=−∑t=1Tlog⁡P(yt∗∣y1∗,…,yt−1∗,x1,…,xn)\mathcal{L} = -\sum_{t=1}^{T} \log P(y_t^* \mid y_1^*, \ldots, y_{t-1}^*, x_1, \ldots, x_n)

where:

  • yt∗y_t^*: the ground-truth (reference) token at position tt
  • TT: the total length of the target (reference) summary
  • The summation accumulates cross-entropy loss across all target positions

Teacher forcing introduces a subtle training-inference mismatch: during training, the model always sees the correct previous token, but during inference, it sees its own (possibly wrong) previous predictions. This "exposure bias" can cause errors to compound during inference. Scheduled sampling and other techniques partially address this, but it remains a known limitation of the standard training procedure. In practice, the quality of large pretrained models has reduced the practical significance of exposure bias compared to earlier, smaller models.

Attention in Summarization

Vanilla seq2seq models compress the entire source document into a fixed-length vector before decoding. For long documents, this creates a severe bottleneck: the model must memorize everything it might need from thousands of tokens into a single vector. Empirically, the performance of vanilla RNN-based seq2seq models degrades sharply as source length increases beyond a few dozen tokens. The information content of a thousand-word document simply cannot be faithfully compressed into a vector of a few hundred dimensions.

The attention mechanism, which we introduced in Chapter 10, solves this problem by allowing the decoder to selectively attend to relevant parts of the source at each generation step. At decoder step tt, attention computes a context vector ctc_t as a weighted sum of encoder hidden states:

ct=∑i=1nαti⋅hic_t = \sum_{i=1}^{n} \alpha_{ti} \cdot h_i

where the attention weight αti\alpha_{ti} measures how relevant source token ii is when generating summary token tt. The weights are computed using a softmax over unnormalized alignment scores etie_{ti}:

αti=exp⁡(eti)∑j=1nexp⁡(etj)\alpha_{ti} = \frac{\exp(e_{ti})}{\sum_{j=1}^{n} \exp(e_{tj})}

The alignment score etie_{ti} is computed using a small learned network that receives the current decoder state and each encoder hidden state:

eti=va⊤tanh⁡(Wast−1+Uahi)e_{ti} = v_a^\top \tanh(W_a s_{t-1} + U_a h_i)

where:

  • st−1s_{t-1}: the decoder hidden state at the previous time step t−1t-1
  • hih_i: the encoder hidden state for source token ii
  • Wa∈Rd×dW_a \in \mathbb{R}^{d \times d}: a learned weight matrix applied to the decoder state
  • Ua∈Rd×dU_a \in \mathbb{R}^{d \times d}: a learned weight matrix applied to the encoder state
  • va∈Rdv_a \in \mathbb{R}^d: a learned vector projecting the combined representation to a scalar score

The context vector ctc_t is then concatenated with the decoder state sts_t to produce the final representation used for predicting yty_t. This allows the model to "look back" at the source with different weights at each generation step, focusing on the relevant parts of the document for each summary word.

For a document about a scientific study, when generating the word "participants," the attention mechanism would naturally weight encoder states corresponding to source tokens like "subjects," "volunteers," or "patients" more heavily than tokens describing the methodology. This selective reading is necessary for summarization: different parts of the source are relevant at different points in the generation. The attention weights effectively implement a soft alignment between source positions and generation steps.

Attention also provides a useful interpretability tool. By visualizing the attention weights αti\alpha_{ti} as a matrix with summary positions on one axis and source positions on the other, you can see which source passages the model "read" while generating each summary word. This is far more interpretable than the black-box vector encoding of vanilla seq2seq, and it allows practitioners to diagnose cases where the model attended to the wrong source passages.

Copy Mechanism

A critical limitation of vanilla seq2seq models for summarization is their inability to reproduce uncommon words, names, or technical terms from the source. If "CRISPR-Cas9" appears in the source document but not in the training vocabulary, the decoder cannot generate it. It might replace it with "the gene editing technique" or, worse, hallucinate an incorrect term entirely.

This problem is not rare. News articles are dense with proper names, place names, and numerical figures. Scientific papers use highly technical vocabulary. Legal documents reference specific statutes, parties, and clause numbers. Medical records contain drug names, dosages, and diagnostic codes. In all these domains, faithfully reproducing terms from the source is essential, yet a fixed vocabulary decoder cannot guarantee it. A model that paraphrases every proper noun it encounters will produce fluent but potentially misleading summaries.

The copy mechanism, introduced by See et al. (2017) in their Pointer-Generator Networks, addresses this by allowing the decoder to either generate a word from the vocabulary or copy a word directly from the source. A soft switch pgen∈[0,1]p_{\text{gen}} \in [0,1] balances these options:

P(w)=pgen⋅Pvocab(w)+(1−pgen)⋅∑i: xi=wαtiP(w) = p_{\text{gen}} \cdot P_{\text{vocab}}(w) + (1 - p_{\text{gen}}) \cdot \sum_{i:\, x_i = w} \alpha_{ti}

where:

  • pgenp_{\text{gen}}: the probability of generating from vocabulary, computed from the current context as σ(wh⊤h∗+ws⊤st+wx⊤xt+bgen)\sigma(w_h^\top h^* + w_s^\top s_t + w_x^\top x_t + b_{\text{gen}})
  • Pvocab(w)P_{\text{vocab}}(w): the softmax distribution over the fixed vocabulary from the decoder
  • ∑i: xi=wαti\sum_{i:\, x_i = w} \alpha_{ti}: the total attention weight across all source positions ii where the token equals ww, representing the copy probability for word ww

If pgenp_{\text{gen}} is low, the model largely copies the most attended source tokens. If it is high, the model generates from its learned vocabulary. The model learns to use this switch contextually: proper nouns and rare terms get low pgenp_{\text{gen}} while common words and paraphrases get high pgenp_{\text{gen}}.

The copy mechanism also provides a natural way to handle out-of-vocabulary tokens. Even if a word was never seen during training, the model can copy it from the source whenever the attention weights concentrate on the corresponding source positions. This is particularly important for numerical values, which would otherwise need to be reproduced character by character or memorized during training. A study in a clinical setting might mention a specific patient age (73 years), a specific blood pressure reading (142/91), and a specific medication dose (10 mg twice daily). A copy-enabled model can reproduce all three values accurately; a vocabulary-only model might substitute plausible but incorrect numbers.

Pointer-Generator Networks

Pointer-Generator Networks combine a standard seq2seq generator with a "pointer" that copies from the source. This hybrid architecture was a significant advance for summarization because many summaries require copying exact text (names, dates, technical terms) while still generating new connective text. The name reflects the dual nature: the model can "point" to source tokens to copy them, or "generate" tokens from its vocabulary.

Coverage Mechanism

Even with a copy mechanism, early attention-based summarizers suffered from a specific failure mode: they would repeat phrases from the source multiple times in the same summary. The model might generate "the Amazon rainforest" three times in a three-sentence summary because the most prominent source tokens kept attracting high attention weights.

See et al. (2017) also introduced a coverage mechanism to address this. A coverage vector ctc_t accumulates the attention weights from all previous decoder steps:

ct=∑t′=0t−1αt′c_t = \sum_{t'=0}^{t-1} \alpha_{t'}

This vector tracks which source tokens have already been attended to. A coverage loss then penalizes the model for attending to source positions that have already been heavily covered:

cov_losst=∑imin⁡(αti,cti)\text{cov\_loss}_t = \sum_i \min(\alpha_{ti}, c_{ti})

The min⁡\min function captures the amount by which current attention overlaps with past attention. If token ii has already accumulated coverage cti=0.8c_{ti} = 0.8 and receives attention αti=0.5\alpha_{ti} = 0.5 at step tt, the coverage loss contribution is min⁡(0.5,0.8)=0.5\min(0.5, 0.8) = 0.5, penalizing this overlap. The total coverage loss is added to the main cross-entropy loss during training, pushing the model to spread its attention more evenly across source positions rather than repeatedly attending to the same ones.

Coverage mechanisms address repetition at the generation level, but they do not fully solve the deeper problem of hallucination. A model can generate novel hallucinated content even when it is not repeating itself. The distinction between repetition and hallucination is important for diagnosis: repetition means attending to the same correct source content too many times; hallucination means generating content with no correspondence to any source content.

Transformer-Based Summarization

Modern abstractive summarization is dominated by large pretrained Transformer models, building on the architectures we studied in Chapters 12 through 19. Models like BART and T5 are trained with pretraining objectives specifically designed for text-to-text generation, then fine-tuned on summarization datasets.

BART (Lewis et al., 2020) uses a denoising objective: the model is trained to reconstruct original text from various corrupted versions, including sentence permutation, token deletion, and text infilling. This pretraining teaches the model to understand text structure at a deep level, making it particularly effective for summarization. The key insight is that a model that can reconstruct a document from a heavily corrupted version must have developed a strong understanding of what makes a document coherent and informative. When fine-tuned on CNN/DailyMail or XSum, BART produces fluent, accurate summaries that substantially outperform earlier seq2seq approaches.

T5 (Raffel et al., 2020) frames all NLP tasks as text-to-text problems. Summarization is simply a task where the input is "summarize: [document]" and the output is the summary. This unified framing allows T5 to transfer knowledge across tasks: understanding developed while learning translation helps with summarization, and vice versa. Its large-scale pretraining on the Colossal Clean Crawled Corpus (C4) gives it broad language knowledge that transfers well to specialized summarization domains.

More recently, instruction-tuned large language models like GPT-4, Claude, and Llama have become the de facto choice for summarization in production applications. These models can follow detailed instructions: "Summarize in three bullet points, stressing the financial implications" or "Write a one-paragraph summary suitable for a general audience." This instruction-following capability, which earlier seq2seq models lacked, substantially increases the practical value of abstractive summarization. You are no longer limited to a fixed output format trained during fine-tuning; you can specify exactly what you want and the model adapts.

The key advantage of pretrained Transformers over earlier seq2seq models comes from scale and from the quality of the learned representations. A BART encoder that has processed hundreds of billions of words has learned rich semantic and syntactic structure that allows it to identify the key information in a document far more reliably than a model trained only on a summarization dataset. Pretraining provides the general language understanding; fine-tuning on summarization data teaches the model the specific output format and style. This division of labor between pretraining and fine-tuning is a pattern we have seen throughout this handbook, and it is particularly consequential for summarization because the task requires such broad language understanding.

Worked Example: From Extraction to Abstraction

To make these approaches concrete, let us trace through a simple example. Consider this short passage:

"The Amazon rainforest covers 5.5 million square kilometers across nine countries. It produces 20% of the world's oxygen and houses 10% of all species on Earth. Deforestation has accelerated in recent decades, driven by agricultural expansion and logging. Scientists warn that losing 20-25% of the Amazon could trigger a 'dieback' event where the forest can no longer sustain itself."

An extractive summarizer might score each sentence and select the top two. Using a simple word frequency approach (after removing stopwords), the most distinctive terms are "Amazon," "rainforest," "deforestation," "species," "oxygen," and "dieback." Sentence 1 scores highly for "Amazon" and "rainforest." Sentence 4 scores highly for "Amazon," "forest," and introduces "dieback" (a term that appears nowhere else, making it locally distinctive). A good extractive summary might be:

"The Amazon rainforest covers 5.5 million square kilometers across nine countries. Scientists warn that losing 20-25% of the Amazon could trigger a 'dieback' event where the forest can no longer sustain itself."

An abstractive summarizer, in contrast, might generate:

"The Amazon rainforest, covering 5.5 million km² and hosting 10% of Earth's species, faces accelerating deforestation that could trigger catastrophic ecosystem collapse."

The abstractive summary is more concise, combines facts from multiple sentences (the size and the species statistic), and uses "catastrophic ecosystem collapse" to rephrase the scientific concept of "dieback" in more accessible language. It also omits the oxygen statistic, making an editorial judgment that the species diversity and dieback risk are more relevant to a general reader.

This comparison reveals both the power and the danger of abstractive summarization. The power: the abstractive summary is more informative, more informative per word than the extractive version. The danger: the abstractive model could just as easily generate "the Amazon rainforest hosts 10% of Earth's species and is growing rapidly" by confusing deforestation with growth, or "scientists warn of a 30% threshold" by misremembering the 20-25% figure. Extractive systems cannot hallucinate because they copy from the source; abstractive systems can generate anything that looks plausible given the source context.

Notice also what the extractive summary loses: the oxygen production statistic from sentence 2 does not appear. The extractive system had no way to combine it with the species statistic into a single more informative sentence. Each sentence had to stand alone. This illustrates why abstractive summarization can produce shorter, denser summaries that pack more information per word, but only when the model correctly understands and faithfully represents all the source facts it chose to include.

Code Implementation

Let us implement both approaches. We will start with a simple TextRank extractor from scratch, then use a pretrained BART model for abstractive summarization.

Setup

TextRank Extractor

We begin by implementing sentence similarity and the TextRank scoring algorithm.

In[4]:
Code
import re

import numpy as np


def preprocess(text):
    """Tokenize and clean a sentence."""
    words = re.findall(r"\b[a-z]+\b", text.lower())
    stopwords = {
        "the",
        "a",
        "an",
        "is",
        "are",
        "was",
        "were",
        "be",
        "been",
        "have",
        "has",
        "had",
        "do",
        "does",
        "did",
        "will",
        "would",
        "can",
        "could",
        "may",
        "might",
        "shall",
        "should",
        "must",
        "in",
        "on",
        "at",
        "to",
        "for",
        "of",
        "and",
        "or",
        "but",
        "that",
        "this",
        "it",
        "its",
        "with",
        "by",
        "from",
        "not",
    }
    return [w for w in words if w not in stopwords]


def sentence_similarity(s1_words, s2_words):
    """Compute overlap-based sentence similarity."""
    set1, set2 = set(s1_words), set(s2_words)
    intersection = set1 & set2
    if len(set1) == 0 or len(set2) == 0:
        return 0.0
    denominator = np.log(len(set1) + 1) + np.log(len(set2) + 1)
    return len(intersection) / denominator if denominator > 0 else 0.0


def textrank(sentences, top_k=3, damping=0.85, max_iter=50):
    """Run TextRank and return top-k sentences by score."""
    n = len(sentences)
    tokenized = [preprocess(s) for s in sentences]

    # Build similarity matrix
    sim_matrix = np.zeros((n, n))
    for i in range(n):
        for j in range(n):
            if i != j:
                sim_matrix[i][j] = sentence_similarity(
                    tokenized[i], tokenized[j]
                )

    # Row-normalize
    row_sums = sim_matrix.sum(axis=1, keepdims=True)
    row_sums[row_sums == 0] = 1  # Avoid division by zero
    sim_matrix = sim_matrix / row_sums

    # Power iteration
    scores = np.ones(n) / n
    for _ in range(max_iter):
        new_scores = (1 - damping) / n + damping * sim_matrix.T.dot(scores)
        if np.max(np.abs(new_scores - scores)) < 1e-6:
            break
        scores = new_scores

    # Select top-k sentences by original order
    ranked_idx = np.argsort(scores)[::-1][:top_k]
    selected_idx = sorted(ranked_idx)
    return [sentences[i] for i in selected_idx], scores

Now we apply TextRank to a sample document about climate change:

In[5]:
Code
document = """
Global average temperatures have risen by approximately 1.1 degrees Celsius since the pre-industrial era.
The primary driver is the increase in greenhouse gases, particularly carbon dioxide from burning fossil fuels.
Sea levels are rising at an accelerating rate, threatening coastal communities around the world.
The past decade has been the warmest on record, with 2023 ranking as the hottest year ever measured.
Extreme weather events including heatwaves, floods, and wildfires have become more frequent and severe.
International agreements like the Paris Accord aim to limit warming to 1.5 degrees Celsius above pre-industrial levels.
Scientists warn that exceeding 2 degrees of warming could trigger irreversible tipping points in Earth's climate system.
Renewable energy adoption has accelerated dramatically, with solar and wind power now cheaper than fossil fuels in most markets.
Carbon capture technologies are being developed, but currently operate at a fraction of the scale needed.
Adaptation measures, including improved infrastructure and early warning systems, are being deployed in vulnerable regions.
""".strip().split("\n")

document = [s.strip() for s in document if s.strip()]

summary_sentences, scores = textrank(document, top_k=3)
Out[6]:
Console
TextRank Summary:
------------------------------------------------------------
  Global average temperatures have risen by approximately 1.1 degrees Celsius since the pre-industrial era.
  The primary driver is the increase in greenhouse gases, particularly carbon dioxide from burning fossil fuels.
  International agreements like the Paris Accord aim to limit warming to 1.5 degrees Celsius above pre-industrial levels.

All Sentence Scores (sorted):
  [0.1689] International agreements like the Paris Accord aim to limit warming to...
  [0.1353] The primary driver is the increase in greenhouse gases, particularly c...
  [0.1195] Global average temperatures have risen by approximately 1.1 degrees Ce...
  [0.1107] Adaptation measures, including improved infrastructure and early warni...
  [0.1027] Carbon capture technologies are being developed, but currently operate...
  [0.0898] Renewable energy adoption has accelerated dramatically, with solar and...
  [0.0756] Scientists warn that exceeding 2 degrees of warming could trigger irre...
  [0.0616] Extreme weather events including heatwaves, floods, and wildfires have...
  [0.0360] Sea levels are rising at an accelerating rate, threatening coastal com...
  [0.0150] The past decade has been the warmest on record, with 2023 ranking as t...

The TextRank algorithm selected the three sentences that are most central to the document's topic network: the core temperature finding, the sea level rise threat (highly connected to the most common themes), and the tipping point warning. Sentences about specific policy agreements or adaptation measures score lower because they are less similar to the main body of content. Notice that the ranking reflects thematic centrality rather than position: sentences from the middle of the document can outrank earlier sentences if they are more representative of the document's overall vocabulary.

The bar chart below shows the full TextRank score distribution across all ten sentences. The gap between high-scoring and low-scoring sentences tells us how distinctly the graph-based algorithm separates central from peripheral content.

Out[7]:
Visualization
Horizontal bar chart of TextRank scores for 10 sentences with top 3 highlighted in orange.
TextRank sentence importance scores for a ten-sentence climate change document. The three highest-scoring sentences (shown in orange) are selected for the extractive summary. Scores reflect thematic centrality: sentences sharing vocabulary with many other sentences score higher regardless of their position. A dashed vertical line marks the score cutoff for the top-3 selection.

Implementing MMR for Redundancy Reduction

To see how MMR improves on naive sentence ranking, we implement a comparison. The key addition is the redundancy term that penalizes candidates similar to already-selected sentences.

In[8]:
Code
def word_overlap_similarity(s1, s2):
    """Simple cosine-style similarity for MMR."""
    w1 = set(preprocess(s1))
    w2 = set(preprocess(s2))
    if not w1 or not w2:
        return 0.0
    return len(w1 & w2) / (len(w1) * len(w2)) ** 0.5


def mmr_summary(sentences, scores, top_k=3, lambda_param=0.6):
    """
    Select sentences using Maximum Marginal Relevance.
    lambda_param controls relevance vs. diversity tradeoff.
    """
    selected = []
    remaining = list(range(len(sentences)))

    for _ in range(top_k):
        best_idx = None
        best_score = -np.inf

        for i in remaining:
            relevance = scores[i]
            if selected:
                redundancy = max(
                    word_overlap_similarity(sentences[i], sentences[j])
                    for j in selected
                )
            else:
                redundancy = 0.0

            mmr_score = (
                lambda_param * relevance - (1 - lambda_param) * redundancy
            )

            if mmr_score > best_score:
                best_score = mmr_score
                best_idx = i

        selected.append(best_idx)
        remaining.remove(best_idx)

    selected_sorted = sorted(selected)
    return [sentences[i] for i in selected_sorted]
Out[9]:
Console
MMR Summary (lambda=0.6, balanced relevance-diversity):
------------------------------------------------------------
  The primary driver is the increase in greenhouse gases, particularly carbon dioxide from burning fossil fuels.
  International agreements like the Paris Accord aim to limit warming to 1.5 degrees Celsius above pre-industrial levels.
  Adaptation measures, including improved infrastructure and early warning systems, are being deployed in vulnerable regions.

Most Redundant Sentence Pair (before MMR):
  Similarity: 0.322
  Sent A: Global average temperatures have risen by approximately 1.1 degrees Ce...
  Sent B: International agreements like the Paris Accord aim to limit warming to...

MMR tends to pick a more diverse set of sentences compared to pure TextRank ranking, especially when the top-ranked sentences are thematically similar to each other. The lambda parameter gives you direct control over this tradeoff: setting lambda closer to 1.0 produces a summary very similar to the pure relevance ranking, while values closer to 0.5 push the algorithm to explore less central but more novel sentences. In practice, values between 0.5 and 0.7 tend to produce the best balance between coverage and relevance.

The following visualization shows how different lambda values affect which sentences MMR selects, revealing the tradeoff between thematic centrality and content diversity.

Out[10]:
Visualization
Grid heatmap showing MMR sentence selection across 5 lambda values and 10 sentences.
Effect of the lambda parameter on MMR sentence selection for the climate change document. Each row shows which 3 sentences (marked with X) are selected at a given lambda value, ranging from high diversity (lambda=0.2) to pure relevance (lambda=1.0). At high lambda, MMR converges to the same selection as pure TextRank. As lambda decreases, the selection shifts toward peripheral sentences, trading relevance for broader topic coverage.

Abstractive Summarization with BART

Now we use a pretrained BART model to generate abstractive summaries. BART-large-cnn is specifically fine-tuned on the CNN/DailyMail summarization dataset, which contains over 300,000 news article-summary pairs. This fine-tuning teaches the model the specific output style expected for news summarization: neutral tone, factual language, and the inverted-pyramid structure where the most important facts come first.

In[11]:
Code
# Uncomment to install if needed:
# !uv pip install transformers torch

from transformers import BartForConditionalGeneration, BartTokenizer

# Load BART fine-tuned on CNN/DailyMail
bart_model_name = "facebook/bart-large-cnn"
bart_tokenizer = BartTokenizer.from_pretrained(bart_model_name)
bart_model = BartForConditionalGeneration.from_pretrained(bart_model_name)
In[12]:
Code
import torch

# Summarize the climate change document
article = " ".join(document)
inputs = bart_tokenizer(
    article, return_tensors="pt", max_length=1024, truncation=True
)
with torch.no_grad():
    summary_ids = bart_model.generate(
        inputs["input_ids"],
        max_length=80,
        min_length=30,
        num_beams=4,
        early_stopping=True,
    )
bart_summary = bart_tokenizer.decode(summary_ids[0], skip_special_tokens=True)
Out[13]:
Console
BART Abstractive Summary:
------------------------------------------------------------
Global average temperatures have risen by approximately 1.1 degrees Celsius since the pre-industrial era. Sea levels are rising at an accelerating rate, threatening coastal communities around the world. Extreme weather events including heatwaves, floods, and wildfires have become more frequent and severe.

Original length: 160 words
Summary length: 42 words
Compression ratio: 0.26

The BART summary illustrates the key advantages of abstractive over extractive approaches: it paraphrases rather than copies, combines facts naturally, and produces a fluent narrative sentence rather than a collection of extracted fragments. The compression ratio gives us a sense of how aggressively BART condenses the content: a ratio of 0.25 means the summary is one quarter the length of the original. Notice that BART makes editorial decisions about what to include and how to frame it, decisions that no extractive system can make because it cannot rearrange or rephrase.

Key Parameters

The key parameters for summarization are:

  • top_k: The number of sentences to include in an extractive summary. Larger values improve coverage but reduce conciseness. The right value depends on the acceptable summary length for your application.
  • damping (TextRank): The fraction of score contributed by the graph structure versus the uniform base. Values around 0.85 follow the original PageRank parameter. Lower values make all sentences more equal; higher values amplify the scores of the most central nodes.
  • lambda_param (MMR): Balances relevance against diversity. Values above 0.5 favor relevance; values below 0.5 favor diversity. The right value depends on how much redundancy you expect in the source document.
  • max_length / min_length (BART): Control the length of the generated abstractive summary. The model will not exceed max_length tokens or go below min_length, regardless of content length. Setting these bounds correctly is important: too tight and the model truncates useful content; too loose and the model may pad with low-quality additions.
  • num_beams: The number of beams in beam search decoding. Higher values explore more of the generation space and often improve quality but at proportional computational cost. For summarization, 4 to 8 beams is typical.
  • do_sample: When False, uses greedy or beam search decoding; when True, enables stochastic sampling. For summarization, deterministic decoding is usually preferred because the goal is accurate representation of the source, not creative variation.

Summarization Evaluation

Evaluating summaries is difficult. Unlike translation or question answering, there is rarely a single correct answer. A good summary of a 10,000-word article could take many valid forms, using different sentences, different phrasings, and covering different subsets of the content. This ambiguity makes automatic evaluation challenging and human evaluation expensive.

The challenge is compounded by the multiplicity of valid summaries. Two equally skilled human summarizers given the same document will produce summaries that differ substantially in vocabulary and structure, yet both may be excellent. An automatic metric that rewards similarity to one reference summary will penalize the other, even if both are equally good. This multi-reference problem is a fundamental tension in summarization evaluation that no current metric fully resolves.

A further complication is that different evaluation criteria matter differently in different applications. A medical summary used to brief a clinician needs extreme faithfulness: a single hallucinated drug dosage could harm a patient. A news summary for a casual reader can afford some paraphrasing that slightly stretches the original claims. An internal corporate summary that will be read by domain experts can use technical vocabulary freely; one targeted at executives may need to simplify. Any single automatic metric that treats these scenarios identically will be misleading in at least one of them.

ROUGE Metrics

The dominant automatic evaluation metric for summarization is ROUGE (Recall-Oriented Understudy for Gisting Evaluation), introduced by Lin (2004). ROUGE measures the overlap between a generated summary and one or more reference summaries written by humans.

ROUGE-N measures n-gram overlap. The recall formulation counts how many of the reference n-grams appear in the candidate summary:

ROUGE-NR=∑S∈Refs∑gramn∈SCount_match(gramn)∑S∈Refs∑gramn∈SCount(gramn)\text{ROUGE-N}_{R} = \frac{\sum_{S \in \text{Refs}} \sum_{\text{gram}_n \in S} \text{Count\_match}(\text{gram}_n)}{\sum_{S \in \text{Refs}} \sum_{\text{gram}_n \in S} \text{Count}(\text{gram}_n)}

where:

  • Refs\text{Refs}: the set of reference summaries
  • gramn\text{gram}_n: each n-gram (sequence of nn consecutive tokens) in a reference summary
  • Count_match(gramn)\text{Count\_match}(\text{gram}_n): the count of gramn\text{gram}_n in the candidate summary, clipped to the count in the reference
  • Count(gramn)\text{Count}(\text{gram}_n): the count of gramn\text{gram}_n in the reference summary

This is a recall-oriented measure: it rewards summaries that cover the reference content. Precision measures the complementary direction (how much of the candidate is in the reference), and the F1 harmonic mean balances both:

ROUGE-NF1=2⋅P⋅RP+R\text{ROUGE-N}_{F_1} = \frac{2 \cdot P \cdot R}{P + R}

where PP is precision (fraction of candidate n-grams found in the reference) and RR is recall (fraction of reference n-grams found in the candidate).

ROUGE-1 uses unigram overlap and measures vocabulary coverage broadly. ROUGE-2 uses bigram overlap and is more sensitive to phrasing: two summaries covering the same facts in different words will score well on ROUGE-1 but potentially poorly on ROUGE-2 if they use different collocations. This is why ROUGE-2 is often a better discriminator between system quality levels, though it also penalizes valid paraphrases more harshly. In general, ROUGE-1 tells you whether the candidate covers the right topics; ROUGE-2 tells you whether it also uses similar phrasing to the reference.

ROUGE-L measures the longest common subsequence (LCS) between the candidate and reference, capturing structural similarity beyond exact n-gram matches. The recall-based formulation is:

ROUGE-LR=LCS(C,R)∣R∣\text{ROUGE-L}_{R} = \frac{\text{LCS}(C, R)}{|R|}

where:

  • CC: the candidate summary (as a token sequence)
  • RR: the reference summary (as a token sequence)
  • LCS(C,R)\text{LCS}(C, R): the length (in tokens) of the longest common subsequence of CC and RR
  • ∣R∣|R|: the number of tokens in the reference summary

Unlike ROUGE-N, ROUGE-L does not require the matching tokens to be contiguous. A candidate that captures the same sequence of ideas as the reference but skips connecting words can still score well under ROUGE-L. This makes it somewhat more robust to paraphrasing than ROUGE-2, though less sensitive to local phrasing choices. ROUGE-L is often used as a single composite metric when you want something in between the n-gram granularity of ROUGE-1 and ROUGE-2.

ROUGE Limitations

ROUGE is a proxy, not a true quality measure. A summary that scores well on ROUGE may be grammatically awkward, factually wrong, or miss key nuances. Conversely, a fluent and informative paraphrase may score poorly if it uses different vocabulary than the reference. Research has consistently shown that ROUGE correlates only moderately with human judgments of summary quality, with Spearman correlations often in the 0.2 to 0.5 range depending on the dimension being evaluated. Despite these limitations, ROUGE remains the standard benchmark metric because it is cheap to compute and provides a reproducible comparison point across systems.

BERTScore

BERTScore (Zhang et al., 2019) addresses ROUGE's vocabulary limitation by using contextual embeddings instead of exact token matches. For each token in the candidate summary, BERTScore finds the most similar token in the reference according to their BERT embeddings, and vice versa.

The precision component measures the average similarity of each candidate token to its closest reference match:

PBERT=1∣C∣∑ci∈Cmax⁡rj∈Rcos⁡(eci,erj)P_{\text{BERT}} = \frac{1}{|C|} \sum_{c_i \in C} \max_{r_j \in R} \cos(\mathbf{e}_{c_i}, \mathbf{e}_{r_j})

The recall component measures the average similarity of each reference token to its closest candidate match:

RBERT=1∣R∣∑rj∈Rmax⁡ci∈Ccos⁡(erj,eci)R_{\text{BERT}} = \frac{1}{|R|} \sum_{r_j \in R} \max_{c_i \in C} \cos(\mathbf{e}_{r_j}, \mathbf{e}_{c_i})

where:

  • CC: the set of candidate summary tokens
  • RR: the set of reference summary tokens
  • eci\mathbf{e}_{c_i}: the contextual embedding of candidate token cic_i from BERT (a vector in Rd\mathbb{R}^d)
  • erj\mathbf{e}_{r_j}: the contextual embedding of reference token rjr_j from BERT
  • cos⁡(u,v)=u⋅v∥u∥∥v∥\cos(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\| \|\mathbf{v}\|}: cosine similarity between two embedding vectors

The BERTScore F1 combines precision and recall in the standard way. BERTScore better captures semantic equivalence: "vehicle" and "car" would contribute positive similarity even though they never match under ROUGE-N. BERTScore correlates better with human judgments than ROUGE in most evaluations, but it comes with higher computational cost (computing BERT embeddings for every token in every summary and reference) and the additional complexity of choosing which BERT layer and model to use. The choice of BERT model matters because different models produce embeddings with different geometric properties, and the optimal choice varies by domain.

Human Evaluation Criteria

When human evaluation is conducted (typically for research publications or high-stakes applications), evaluators assess multiple dimensions of summary quality:

  • Faithfulness: Does the summary contain only information that can be supported by the source? Unfaithful summaries contain hallucinated facts. This is often the most critical dimension for production applications because users who discover a single false claim may lose trust in the entire system.
  • Relevance: Does the summary cover the most important aspects of the source? A highly faithful but irrelevant summary copies trivial details while omitting the key findings. Relevance is subjective and depends on what the reader needs.
  • Coherence: Is the summary well-organized and logical? Extracted sentences may be individually accurate but collectively incoherent if they come from different parts of the document with no narrative thread connecting them.
  • Fluency: Is the language grammatically correct and natural? Abstractive models sometimes produce disfluent text near their generation length limits, or when copying rare vocabulary from the source.

These dimensions often trade off against each other. Extractive systems score very high on faithfulness (since they copy) but sometimes lower on relevance and coherence. Abstractive systems can score higher on fluency and relevance but are more likely to introduce factual errors. Instruction-tuned LLMs generally score well on fluency and can be guided to prioritize faithfulness with appropriate prompting, but they require careful evaluation in high-stakes settings.

Human evaluation typically uses rating scales (1-5 or 1-7) on each dimension, administered through crowdsourcing platforms or expert panels. The cost ranges from a few cents per judgment on crowdsourcing platforms to hundreds of dollars per page for expert legal or medical review. This cost is why automatic metrics remain dominant in research, even though they are imperfect. For production systems, a practical approach is to use automatic metrics for rapid iteration during development, then conduct targeted human evaluation on a representative sample before deployment.

Implementing ROUGE Evaluation

In[14]:
Code
# Uncomment to install if needed:
# !uv pip install rouge-score

import re
from collections import Counter, namedtuple

RougeScore = namedtuple("RougeScore", ["precision", "recall", "fmeasure"])


def _tokens(text):
    return re.findall(r"\b\w+\b", text.lower())


def _ngram_counts(tokens, n):
    return Counter(tuple(tokens[i : i + n]) for i in range(len(tokens) - n + 1))


def _f_score(overlap, candidate_total, reference_total):
    precision = overlap / candidate_total if candidate_total else 0.0
    recall = overlap / reference_total if reference_total else 0.0
    fmeasure = (
        2 * precision * recall / (precision + recall)
        if precision + recall
        else 0.0
    )
    return RougeScore(precision, recall, fmeasure)


def _lcs_length(a, b):
    prev = [0] * (len(b) + 1)
    for x in a:
        curr = [0]
        for j, y in enumerate(b, start=1):
            curr.append(prev[j - 1] + 1 if x == y else max(prev[j], curr[-1]))
        prev = curr
    return prev[-1]


class SimpleRougeScorer:
    def score(self, reference, candidate):
        ref = _tokens(reference)
        cand = _tokens(candidate)
        scores = {}
        for name, n in [("rouge1", 1), ("rouge2", 2)]:
            ref_counts = _ngram_counts(ref, n)
            cand_counts = _ngram_counts(cand, n)
            overlap = sum(
                min(count, cand_counts[gram])
                for gram, count in ref_counts.items()
            )
            scores[name] = _f_score(
                overlap, sum(cand_counts.values()), sum(ref_counts.values())
            )
        lcs = _lcs_length(cand, ref)
        scores["rougeL"] = _f_score(lcs, len(cand), len(ref))
        return scores


scorer_rouge = SimpleRougeScorer()

# Reference summary (hypothetical human-written)
reference = (
    "Global temperatures have risen 1.1 degrees Celsius since pre-industrial times, "
    "driven by greenhouse gas emissions. Sea levels are rising and extreme "
    "weather events are increasing. Scientists warn that exceeding 2 degrees could "
    "trigger irreversible climate tipping points."
)

extractive_summary = " ".join(summary_sentences)
Out[15]:
Console
ROUGE Scores: Extractive vs Abstractive
=======================================================
Metric                 Extractive                 BART
-------------------------------------------------------
rouge1                     0.3333               0.4819
rouge2                     0.1591               0.2963
rougeL                     0.2889               0.4337

Note: ROUGE scores depend heavily on reference choice.
These scores reflect overlap with one specific reference summary.

ROUGE scores vary substantially based on the reference summary. In this particular single-reference comparison, BART scores higher than TextRank on all three metrics because its generated summary happens to align more closely with the reference wording. Both methods still score lower on ROUGE-2 than on ROUGE-1 or ROUGE-L because exact bigram matches are harder to preserve. A system that produces a summary using nearly identical phrasing to the reference will score very high on ROUGE-2, while a system producing an equally informative but differently worded summary will score lower, even if human judges prefer the latter. This reference-dependency problem is not a bug in ROUGE but an inherent consequence of using a single reference in a task where many valid summaries exist.

The bar chart below visualizes the ROUGE comparison directly, making the extractive-abstractive difference visible.

Out[16]:
Visualization
Grouped bar chart comparing ROUGE-1, ROUGE-2, and ROUGE-L scores for extractive and BART summaries.
ROUGE F1 scores comparing extractive (TextRank) and abstractive (BART) summaries against a single human reference summary. BART scores higher on ROUGE-1, ROUGE-2, and ROUGE-L in this example because its wording aligns more closely with this reference. Both methods score lower on ROUGE-2 than on ROUGE-1 or ROUGE-L, reflecting the stricter requirement for matching contiguous bigrams.

Summarization Applications

Summarization is not a single task but a family of related problems, each with distinct requirements and challenges. The choice of summarization approach depends heavily on the application domain, the acceptable error rate for hallucinations, the available compute budget, and whether training data for fine-tuning exists. Understanding this diversity is essential for choosing the right tool for a given situation.

Single-Document Summarization

The classic formulation: given one document, produce one summary. This is the most extensively studied variant, with benchmark datasets like CNN/DailyMail (news articles with human-written highlights), XSum (BBC articles with single-sentence summaries), and ArXiv/PubMed (scientific papers with abstracts as reference summaries).

The key challenge in single-document summarization varies by domain. News summarization benefits from the inverted pyramid structure of journalism, which front-loads the most important information. A simple extractive approach that takes the first two sentences performs remarkably well on news precisely because journalists are trained to make those sentences maximally informative. Scientific summarization requires understanding technical terminology and the logic of empirical claims: a summary of a clinical trial that omits the confidence intervals or effect sizes may be technically correct about the conclusion but misleading about the certainty. Legal document summarization demands extreme precision, since omitting a qualifier like "not" can fundamentally change meaning.

The length of the source document also varies dramatically across applications. A tweet thread and a 50-page contract are both single-document summarization tasks, but they require very different approaches. Very short sources (a few hundred words) may not need summarization at all, or may benefit from a single-sentence abstractive summary. Very long sources (tens of thousands of words) require special handling because most Transformer models have fixed context windows, and simply truncating the input discards potentially critical information.

Multi-Document Summarization

Multi-document summarization (MDS) takes a set of related documents and produces a summary covering the key information across all of them. This task appears in news aggregation (summarizing 20 articles about the same event), literature review (synthesizing dozens of papers on a research topic), and enterprise search (answering a query from a large document corpus).

The unique challenge of MDS is redundancy management at scale. If ten news articles all report that "the earthquake struck at 2:17 AM," the summary should mention this fact once, not ten times. But the articles may disagree on details (some say 2:15, others 2:20), requiring the system to handle conflicting information. Should the summary report the most common value? Average them? Report the uncertainty? These editorial decisions require judgment that purely statistical methods struggle to make correctly.

Additionally, different documents may cover different aspects of the same event. An MDS system must identify which facts appear in many documents (high consensus, important to include), which facts appear in only one document (low coverage, may be unique or may be noise), and how to fuse complementary information into a coherent narrative. The Cross-Document Structure Theory (CST) framework provides a linguistic vocabulary for these relationships (elaboration, contrast, identity, refinement), but automatically recognizing these relationships across documents remains challenging.

Query-Focused Summarization

Standard summarization asks "what is this document about?" Query-focused summarization asks "what does this document say about X?" Given a user query, the system extracts or generates text specifically addressing that query rather than producing a generic summary.

This task supports information retrieval pipelines. In a RAG (Retrieval-Augmented Generation) system, as we studied in Chapter 29, retrieved documents may be long and only partially relevant. Query-focused summarization can extract the relevant portion, reducing noise and improving downstream answer quality. Instead of passing an entire 5,000-word report to an LLM, query-focused summarization produces a 200-word excerpt focused on exactly the aspects relevant to the user's question.

Query-focused extractive methods use the query as the reference in MMR similarity calculations: sentences are selected for their relevance to the query, with redundancy relative to already-selected content penalized. The Sim1_1 term in the MMR formula measures similarity to the query rather than to the full document, biasing selection toward query-relevant content even if that content is not thematically central to the document as a whole. Abstractive query-focused methods condition the decoder on both the source document and the query, typically by prepending the query to the input: "question: [query] context: [document]."

Dialogue Summarization

Dialogue summarization condenses conversations into concise notes or minutes. This appears in customer service (summarizing support chat transcripts to update a CRM record), meetings (generating meeting minutes from transcripts), and medical consultations (summarizing doctor-patient conversations for clinical records).

Dialogue has fundamentally different structure from documents. It is interactive: understanding what one participant says often requires knowing what the other said previously. It is often incomplete: speakers rely on shared context and skip information they assume the other knows, using pronouns and demonstratives ("that thing we discussed last week") that only make sense with access to context outside the transcript. It contains noise: corrections, repairs, false starts, off-topic tangents, and filler language that adds length without adding content.

A key challenge is speaker attribution: whose views are being summarized? A customer complaining about a product and an agent explaining policy may be summarized differently depending on the use case (the agent's notes versus the customer's report). The factual content may be identical, but the framing and emphasis differ. Datasets like AMI (meeting summarization) and SAMSum (dialogue summarization) provide training data for this specialized setting, though they cover a relatively narrow range of dialogue types and may not generalize to specialized professional conversations.

Headline Generation

Headline generation is an extreme form of abstractive summarization: compress an entire article into a single phrase of 5-10 words. This task requires aggressive information selection and creative compression. The XSum dataset was built specifically for this use case, pairing BBC articles with their one-sentence introductory summaries.

Headline generation is particularly challenging because the summary must be both maximally compressed and highly engaging. A technically accurate but boring headline ("Company Releases Product Update") is worse than a vivid but slightly imprecise one for many publishers. This tension between informativeness and engagement is not captured by ROUGE scores. A headline optimized for ROUGE might faithfully reproduce the article's main claim in dry language. A headline optimized for human engagement might sacrifice some precision for impact.

The compression required for headlines also exposes a fundamental tension in abstractive summarization: shorter summaries require more inference. To write "Tech Giant's AI System Surpasses Doctors in Cancer Detection," you must infer that the key finding is the performance comparison, that "AI system" is the appropriate level of specificity, and that "surpasses" is the right verb given the reported statistics. A model that cannot make these inferences consistently will produce headlines that are either too vague ("New Research Published") or too specific ("Model Achieves 94.2% AUC on Lung Cancer Detection Dataset, Up from 91.4%").

Summarization in Production Pipelines

In modern production systems, summarization rarely operates in isolation. It functions as a component in larger pipelines. A document ingestion pipeline might summarize each incoming document before storing it in a retrieval index, making retrieval more efficient. A customer service platform might summarize each conversation before passing it to an LLM that generates a follow-up email, reducing the context the LLM needs to process. A financial analyst tool might summarize quarterly earnings reports into a standardized format before extracting specific metrics.

The integration of summarization into pipelines introduces new failure modes. An error in the summary (a hallucinated number, a missed key fact) can propagate through the pipeline and affect downstream outputs in ways that are hard to trace. A financial model that makes decisions based on a summarized earnings report might trade incorrectly because the summary misrepresented a quarterly revenue figure. This propagation problem makes faithfulness even more critical in pipeline contexts than in standalone summarization applications. When summaries are consumed by humans who can exercise judgment about questionable claims, a few errors are tolerable. When they are consumed by automated systems that trust their inputs, errors can compound.

Limitations and Challenges

Summarization is one of the more mature NLP applications, yet it remains far from solved. Several fundamental limitations affect current systems and deserve careful attention when deploying summarization in practice.

The radar chart below illustrates the quality tradeoffs between extractive and abstractive approaches across the main evaluation dimensions. No single approach dominates across all criteria, which is why the choice of method depends heavily on which properties matter most for your application.

Out[17]:
Visualization
Radar chart comparing extractive and abstractive summarization on faithfulness, fluency, relevance, coherence, and speed.
Quality tradeoffs between extractive and abstractive summarization across five evaluation dimensions. Extractive methods (blue) excel at faithfulness and inference speed but often score lower on fluency and relevance. Abstractive methods (orange) score higher on fluency and relevance but sacrifice faithfulness due to hallucination risk and are much slower to run. The radar shows why neither approach dominates: the optimal choice depends on whether faithfulness or readability is the primary concern for the application.

Faithfulness and Hallucination

The most pressing limitation of abstractive summarization is faithfulness. Large language models frequently generate summaries containing facts not present in, or contradicted by, the source document. A study of BART-generated summaries found that roughly 25% contained at least one factual error, ranging from incorrect numbers to wrong entity relationships.

This data quality issue follows directly from the generation mechanism. During decoding, the model balances many competing signals: the source content, the fluency of the text being generated, and the prior probability of token sequences learned from pretraining. When these signals conflict, the model does not always prioritize faithfulness. It may smooth over an awkward phrase from the source with a fluent but incorrect alternative. A model that has seen thousands of news articles stating that inflation "rose to 4.2%" will have a strong prior for that sentence pattern, and may reproduce a plausible-looking percentage even when the source document says 3.8%.

Hallucination in summarization takes several distinct forms. Intrinsic hallucination introduces information that contradicts the source: changing a date, inverting a causal relationship, or swapping two entities. Extrinsic hallucination introduces information that is not contradicted by the source but is also not present in it: adding a detail that seems plausible given the topic but was never stated. Both types are harmful, but extrinsic hallucination is harder to detect because the hallucinated content may be factually true (just not in this specific document) and therefore not obviously wrong to a reader.

Faithfulness problems are harder to detect than grammatical errors because a hallucinated fact can be indistinguishable from a correct fact to a reader who does not independently verify it. Automatic faithfulness metrics like FactCC and QAGS attempt to measure this by checking whether claims in the summary can be inferred from the source, but they are imperfect and expensive to compute. For high-stakes applications (medical, legal, financial), post-generation fact verification against the source is often necessary. One practical approach is to use an LLM to check each claim in the generated summary against the source document, producing a faithfulness score alongside the summary itself.

Length Bias and Truncation

Summarization models exhibit length biases in both directions. Very short documents are sometimes padded with repetition to meet minimum length constraints. Very long documents are often truncated before being summarized, because most Transformer models have context windows of 512 to 4,096 tokens. A research paper may be 10,000 words; summarizing only the first 1,000 words can produce misleading summaries that miss the conclusions entirely. This truncation problem is particularly insidious because the model produces a confident-looking summary without any indication that it never read the most important part of the document.

Long-context models (as we discussed in Chapter 15) partially address this, but even with longer windows, the model's ability to attend to distant parts of the input tends to degrade. The "lost in the middle" phenomenon describes a consistent finding in long-context LLM evaluations: information in the middle of a long document is recalled less reliably than information near the beginning or end. For summarization, this means that key findings buried in the middle sections of a long document may be systematically underrepresented.

Hierarchical summarization approaches, which first summarize sections and then summarize the section summaries, are a practical workaround that avoids the direct length constraint. But they introduce their own failure modes: information lost in the first stage of summarization cannot be recovered in the second stage. A critical caveat mentioned in the discussion section of a paper will not appear in a section summary that focuses on the main results, and therefore will not appear in the final summary either.

Evaluation Validity

ROUGE scores are widely reported but poorly correlated with human judgments of quality in many settings. Studies have found correlations between ROUGE-2 and human quality judgments as low as 0.2 to 0.3 depending on the evaluation dimension being measured. Models fine-tuned to maximize ROUGE can learn to produce verbose repetitions of common words that inflate recall without improving quality. This optimization pressure creates a metric gaming problem: the winning model under ROUGE may not be the model that produces the most useful summaries for readers.

BERTScore and newer learned metrics (SummEval, BLANC) show better correlation with human judgments but are more computationally expensive and themselves require validation against human data. The SummEval benchmark (Fabbri et al., 2021) systematically compared many automatic metrics against expert human judgments and found that no single metric consistently outperformed others across all quality dimensions. Coherence, in particular, is almost impossible to capture automatically because it requires understanding narrative flow across sentences, something that token-overlap metrics cannot access.

The field currently lacks a single trusted automatic metric that generalizes across domains and tasks. This creates a practical problem: when you compare two summarization systems, ROUGE might prefer one while BERTScore prefers the other and human judges prefer a third. Understanding which metric to trust requires domain-specific calibration against human judgments, which most practitioners do not have the resources to perform.

Domain Adaptation

A model fine-tuned on CNN/DailyMail news articles may perform poorly on scientific papers, legal documents, or medical records. News articles follow predictable structures, use general vocabulary, and have available training data in the millions of examples. Technical documents do not follow predictable structures, use specialized vocabulary that may not appear in general pretraining corpora, and often have no publicly available summarization training data.

The domain mismatch manifests in several ways. The model may produce summaries at the wrong level of technical detail (too simplified for expert readers, or too technical for the intended audience). It may hallucinate domain-specific facts by drawing on its general pretraining knowledge rather than the source document. It may miss domain-specific conventions: a medical summary that omits statistical significance values is incomplete by clinical standards, even if those values seem like minor technical details to a general reader.

This is especially problematic for organizations that want to summarize proprietary internal documents that differ substantially from public training data. Meeting transcripts, internal policy documents, and technical specifications have very different structure and vocabulary from news or academic literature. Fine-tuning on domain-specific data helps significantly but requires annotated examples, which are expensive to create. For many specialized domains, the cost of creating enough labeled data for fine-tuning makes pre-trained general models the only practical option, despite their limitations.

Abstraction Level Control

Current models offer limited control over the level of abstraction in the output. A user might want a one-sentence headline, a three-sentence executive summary, or a structured bullet-point breakdown of the same document. While length parameters partially control this, the style and structure of the output are harder to specify. Instruction-tuned models like GPT-4 and Claude show significantly better adherence to specific summarization instructions, but this capability is not universal and depends heavily on how the instruction is phrased and whether the model was trained on similar instruction patterns.

The mismatch between user intent and model output is often subtle. A user who asks for "a brief summary focusing on the key risks" may receive a summary that is brief and covers risks, but weights financial risks more heavily than operational or reputational risks based on the model's training distribution rather than the user's actual priorities. Making these preferences explicit in natural language requires users to anticipate and articulate every dimension of their intent, which is cognitively demanding and error-prone. Structured output formats (bullet points with specific headings) help constrain the output format but do not fully solve the content selection problem.

Compression Ratio and Information Loss

Every summary involves information loss, and the acceptable loss depends entirely on the use case. A 10:1 compression ratio (reducing a 1,000-word article to 100 words) necessarily discards 90% of the content. For a casual reader seeking a quick overview, this may be appropriate. For a researcher who needs complete understanding of the methodology, it is almost certainly too aggressive.

Current models lack principled mechanisms for deciding which 90% to discard. They use proxies (word frequency, position, attention weights) that correlate with importance on average but fail in specific cases. The supporting detail that appears only once in a document and never attracts high attention might be precisely the piece of information the reader needs. A summary system that consistently discards unique details in favor of repeated assertions will produce summaries that reflect popular content rather than critical content.

This problem is intrinsic to the task formulation. If the user's information need is not specified, the model must guess at it. The more aligned the model's guess is with typical readers of typical documents in its training distribution, the better it will perform on average. The more atypical the user's need or the document's topic, the more likely the model is to discard important information. Understanding this limitation helps practitioners set appropriate expectations: summarization systems work best when the user's information need closely matches the system's implicit assumptions about what is important.

Summary

Text summarization spans a spectrum from simple sentence selection to sophisticated language generation. Extractive methods like TextRank identify the most central sentences through graph-based ranking, modeling the insight that importance is relational rather than intrinsic. MMR improves on pure ranking by ensuring the selected sentences cover diverse aspects of the content, balancing relevance against redundancy through the lambda parameter. Neural extractive methods extend this further by using contextual embeddings to capture semantic meaning rather than surface word overlap.

Abstractive methods, built on seq2seq architectures with attention and copy mechanisms, can paraphrase and compress information more effectively than any extractive approach. The attention mechanism allows the decoder to selectively read relevant parts of the source at each generation step, the copy mechanism enables faithful reproduction of rare terms and proper names, and the coverage mechanism prevents repetition. Large pretrained models like BART and T5 bring broad language understanding to the summarization task, enabling higher quality output at the cost of substantial computational resources.

Evaluation remains a central challenge. ROUGE measures n-gram overlap with human references and is cheap but imperfectly correlated with human quality judgments. BERTScore uses semantic similarity from contextual embeddings to capture paraphrase equivalence. Human evaluation across dimensions of faithfulness, relevance, coherence, and fluency provides the most reliable quality signal but is expensive to collect at scale. The current lack of a single trusted automatic metric makes rigorous system comparison difficult and creates incentives for metric gaming.

The applications of summarization span single documents, multi-document sets, query-focused retrieval, dialogue, and headline generation, each bringing distinct structural challenges. As you build summarization systems, the most important design decisions are: which faithfulness-fluency tradeoff to accept (extractive is safer, abstractive is more powerful), how to handle domain mismatch between training data and deployment context, what evaluation criteria matter most for your use case, and how to verify that summaries are faithful before they enter production pipelines where errors can propagate.

In the next chapter, we turn to question answering, where many of the same foundations apply but the task shifts from condensing information to precisely locating and extracting specific answers.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about text summarization.

Text Summarization Quiz

Question 1 of 80 of 8 completed
What is the key difference between extractive and abstractive summarization?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026textsummarization, author = {Michael Brenndoerfer}, title = {Text Summarization: Extractive and Abstractive Methods}, year = {2026}, url = {https://mbrenndoerfer.com/writing/summarization}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Text Summarization: Extractive and Abstractive Methods. Retrieved from https://mbrenndoerfer.com/writing/summarization
MLAAcademic
Michael Brenndoerfer. "Text Summarization: Extractive and Abstractive Methods." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/summarization>.
CHICAGOAcademic
Michael Brenndoerfer. "Text Summarization: Extractive and Abstractive Methods." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/summarization.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Text Summarization: Extractive and Abstractive Methods'. Available at: https://mbrenndoerfer.com/writing/summarization (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Text Summarization: Extractive and Abstractive Methods. https://mbrenndoerfer.com/writing/summarization

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.