QA: Extractive, Generative, and Open-Domain QA

Michael BrenndoerferJanuary 26, 202659 min read

Part of Language AI Handbook

Covers question answering systems from span extraction with BERT to retrieval-augmented generation, covering evaluation metrics and open-domain QA pipelines.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Question Answering

Question answering (QA) is one of the most direct tests of language understanding: given a question in natural language, produce a correct answer. Unlike summarization, which we covered in the previous chapter, QA requires locating or generating a specific piece of information rather than compressing a body of text. The task sounds simple but draws on reading comprehension, information retrieval, reasoning, and knowledge representation, challenges that have occupied NLP researchers for decades.

The field has evolved through several distinct paradigms. Early QA systems matched patterns and templates against curated databases. The rise of the internet shifted attention to open-domain QA, where answers could be found anywhere across millions of documents. The transformer era changed QA yet again: first by teaching models to identify answer spans within passages, then by enabling models to synthesize answers from their parametric memory without any retrieved context at all.

Today, QA systems span a wide range of designs. At one end sits extractive QA, where a model reads a passage and highlights the text span that answers the question. At the other end sits generative QA, where a model composes an answer from scratch, potentially drawing on retrieved documents, reasoning chains, or encoded knowledge. In between lies a rich space of hybrid approaches. Understanding the differences (when to extract versus generate, when to retrieve versus rely on memory) is central to building effective QA systems.

This chapter traces that evolution. We cover the mechanics of each approach, how to evaluate QA systems rigorously, and how to build both extractive and generative QA pipelines in code. By the end, you will understand how to plug in a pre-built QA model, why each design decision was made, and what consequences it carries.

From Pattern Matching to Reading Comprehension

The earliest QA systems, like BASEBALL (1961) and LUNAR (1972), were domain-specific and rule-based. BASEBALL could answer questions about American baseball game results by querying a structured database. LUNAR answered questions about moon rocks by translating natural language into database queries. These systems worked well within their narrow domains but required hand-engineered rules and could not generalize beyond the curated knowledge they were built around.

The reason these systems were brittle is fundamental: language is infinitely variable, but a hand-written rule set is finite. The same fact can be expressed hundreds of different ways, and a rule that matches "How many home runs did Babe Ruth hit?" fails on "What was Ruth's career home run total?" even though both questions ask for the same thing. Engineers compensated by writing more and more rules, but coverage never caught up with the diversity of real questions.

The question of how to move beyond curated databases led researchers toward text corpora. If answers exist somewhere in documents, the task becomes: find the right document, then find the answer within it. This two-stage structure, retrieve then read, became the dominant paradigm for open-domain QA and persists in modern retrieval-augmented systems. The insight was deceptively simple: rather than building a knowledge base by hand, use the vast body of text already written by humans as the knowledge source.

The "read" step, however, required models that could comprehend text, not just match keywords. This challenge became formalized as machine reading comprehension (MRC). The release of large reading comprehension datasets crystallized progress: SQuAD (Stanford Question Answering Dataset), released in 2016, provided over 100,000 question-passage pairs where the answer was a span of text from the passage. Models had to select exactly the right span, making evaluation straightforward and enabling rapid, reproducible progress.

SQuAD and its successors (SQuAD 2.0, TriviaQA, Natural Questions, QuALITY) drove a wave of architectures designed to solve span extraction. Bidirectional Attention Flow (BiDAF), QANet, and eventually BERT-based models progressively improved performance. When BERT was fine-tuned on SQuAD in 2018, it matched human-level performance on the benchmark, a milestone that marked the maturation of extractive QA as a solved problem on that particular benchmark, though far harder problems lay ahead.

It is worth pausing to appreciate what "human-level performance on SQuAD" means. The benchmark measures agreement between model output and human-annotated answer spans. A model that scores above the human agreement threshold has, in a narrow technical sense, matched human accuracy at selecting answer spans in Wikipedia passages. This is a concrete benchmark achievement. But it does not mean the model has solved reading comprehension in general: SQuAD passages are from Wikipedia, questions were posed by annotators who had already read the passage, and every question has a guaranteed answer in the text. Harder datasets like adversarial SQuAD, TriviaQA, and Natural Questions revealed that models had partially learned to exploit surface-level patterns rather than deeply understanding the passage.

Extractive QA: Finding Answers as Spans

Extractive QA takes a specific form: given a question qq and a context passage cc, find the substring of cc that answers qq. The answer is not generated. It is identified as a contiguous token span [i,j][i, j] within cc, where ii is the start index and jj is the end index.

This formulation has several advantages. The answer is guaranteed to come from the passage, making hallucination impossible. Evaluation is easy: compare predicted span to gold span. And the task can be framed as a simple classification problem over token positions, which fits naturally into the fine-tuning paradigm for pretrained language models. The model does not need to invent new words; it only needs to recognize which words in the provided passage are the answer.

The constraint that the answer must be a contiguous span from the passage is also a limitation. Many real questions do not have answers that appear verbatim in text. Counting questions ("How many planets are in the solar system?") may require numerical computation. Comparative questions ("Which country has a larger population, India or China?") require comparing values from possibly different parts of the text. Abstractive questions ("What caused World War I?") require synthesizing information into a narrative. Extractive QA handles these poorly or not at all. But for the enormous class of factoid questions ("Who invented the telephone?", "When did the French Revolution begin?", "What is the capital of Australia?"), span extraction works very well.

Span Extraction as Token Classification

Given a tokenized passage of length nn, extractive QA reduces to predicting two integer positions:

Pstart∈{1,…,n},Pend∈{Pstart,…,n}P_{\text{start}} \in \{1, \ldots, n\}, \quad P_{\text{end}} \in \{P_{\text{start}}, \ldots, n\}

where nn is the number of tokens in the passage, PstartP_{\text{start}} is the index of the first answer token, and PendP_{\text{end}} is the index of the last answer token. Together they define a contiguous span within the passage.

A BERT-based model encodes the concatenated question and passage, then applies two learned linear layers over the output representations to compute start and end logit scores for each token position:

si=ws⊤hi,ej=we⊤hjs_i = \mathbf{w}_s^\top \mathbf{h}_i, \quad e_j = \mathbf{w}_e^\top \mathbf{h}_j

where:

  • hi∈Rd\mathbf{h}_i \in \mathbb{R}^d: the hidden state vector at position ii from the encoder, with dimension dd matching the model's hidden size
  • ws,we∈Rd\mathbf{w}_s, \mathbf{w}_e \in \mathbb{R}^d: learned weight vectors, one for start positions and one for end positions
  • si,ej∈Rs_i, e_j \in \mathbb{R}: scalar logit scores; higher values mean the model believes position ii is the start (or jj is the end) of the answer

These linear projections are the only new parameters added on top of the pretrained encoder. Everything else (the self-attention layers, feed-forward networks, layer normalization) carries pretrained weights that are fine-tuned with the QA data. The small number of task-specific parameters means the model adapts quickly with relatively little labeled data, often achieving strong performance with just a few thousand training examples.

The scores are then converted to probabilities via softmax over all positions in the passage:

P(start=i)=exp⁡(si)∑kexp⁡(sk)P(\text{start} = i) = \frac{\exp(s_i)}{\sum_k \exp(s_k)} P(end=j)=exp⁡(ej)∑kexp⁡(ek)P(\text{end} = j) = \frac{\exp(e_j)}{\sum_k \exp(e_k)}

The denominator sums over all token positions, making sure the probabilities sum to 1 across the passage. The model selects the span (i,j)(i, j) that maximizes the joint score P(start=i)⋅P(end=j)P(\text{start} = i) \cdot P(\text{end} = j), subject to i≤ji \leq j and j−i<Lmax⁡j - i < L_{\max} (a maximum span length constraint, typically 30-50 tokens, to prevent selecting absurdly long spans).

One important subtlety is that the start and end probabilities are computed independently: the model assigns a probability to each position being the start, and separately assigns a probability to each position being the end. The joint selection enforces the constraint that the end must come after the start, but the model does not directly model the joint distribution P(start=i,end=j)P(\text{start} = i, \text{end} = j). This independence assumption is computationally convenient and works well in practice, since the relevant span is usually short and the contextual embeddings carry enough information that start and end predictions are implicitly coordinated.

During training, the model minimizes the sum of cross-entropy losses over start and end positions:

L=−log⁡P(start=i∗)−log⁡P(end=j∗)\mathcal{L} = -\log P(\text{start} = i^*) - \log P(\text{end} = j^*)

where i∗i^* and j∗j^* are the ground-truth start and end positions annotated in the training data. Minimizing this loss pushes the model to assign maximum probability to the correct span boundaries.

Input Format for BERT-based QA

BERT concatenates question and passage using special separator tokens:

[CLS] question tokens [SEP] passage tokens [SEP]

The [CLS] token can also be used to predict "unanswerable". SQuAD 2.0 includes questions with no answer in the passage, and models are trained to assign high start/end probability to [CLS] when the question cannot be answered from the given context.

The positional relationship between the question and passage matters. Because BERT uses bidirectional attention, every token in the passage attends to every token in the question, and vice versa. This cross-attention between question and passage allows the model to compute question-conditioned representations of each passage token. The hidden state hi\mathbf{h}_i for a passage token represents both that token and how the token relates to the question. A token that answers the question will have a hidden state shaped by the question's content, which the linear start/end classifiers can then detect.

Handling Unanswerable Questions

SQuAD 1.1 guaranteed that every question had an answer in the passage. SQuAD 2.0 added roughly 50,000 adversarially crafted unanswerable questions, plausible-looking enough that a shallow model might confabulate an answer. Questions like "What year was Einstein born in Paris?" are unanswerable from a passage that mentions Einstein's birth in Ulm, Germany. A model that always extracts some span will get these wrong by design.

This change was motivated by a practical concern: deployed QA systems must know when to abstain. A system that always produces an answer is dangerous if the answer is fabricated. SQuAD 2.0 forced models to reason about whether the question is even answerable from the given passage, not just which span is most likely.

Models trained on SQuAD 2.0 learn to produce a "no answer" prediction when confident. The standard approach adds an "answerability score" (the probability that the question is answerable at all) computed from the [CLS] token or as a separate classification head. If the answerability score falls below a threshold, the model outputs "No answer".

This threshold controls a tradeoff: a low threshold produces fewer false "no answer" predictions but more hallucinated spans; a high threshold produces more "no answer" predictions, even for answerable questions. Calibrating this threshold for a specific application requires evaluation on a held-out set representative of the deployment distribution. In a medical QA system where a wrong answer is more dangerous than no answer, you set the threshold high. In a general search assistant where users prefer a possibly wrong answer over no answer, you set it lower.

The answerability prediction also raises an interesting question about what the model has learned. To correctly identify an unanswerable question, the model must understand what information is in the passage and what information is missing. This is harder than simply finding a matching span, and models that score well on SQuAD 2.0 have developed more reliable passage comprehension than their SQuAD 1.1-only counterparts.

Document Retrieval for Extractive QA

In practice, you rarely have the correct passage pre-selected. In open-domain extractive QA, the system must first retrieve relevant passages from a large corpus (Wikipedia, a document collection, the web) before running the reader.

This retrieval step uses one of several strategies:

  • Sparse retrieval: TF-IDF or BM25 matches query terms to document terms. Fast and interpretable, but fails on vocabulary mismatch (the question uses "death" where the document uses "passed away"). BM25 is the standard sparse baseline, and it remains competitive for many retrieval tasks despite its simplicity.
  • Dense retrieval: Encode question and passages using a bi-encoder neural network; retrieve by nearest-neighbor search in embedding space. Handles synonymy and paraphrase, but requires building a vector index. Dense retrieval became dominant after Dense Passage Retrieval (DPR) was introduced in 2020.
  • Hybrid retrieval: Combine sparse and dense scores. Often outperforms either alone, since sparse retrieval captures exact keyword matches while dense retrieval captures semantic similarity; combining both recovers the strengths of each approach.

The full pipeline is: retrieve top-kk passages, run the reader on each passage independently, then select the best answer span across all passages using the span confidence scores. The reader does not know which passage came from which document; it simply processes each passage as an independent context and reports which span it finds most likely to be the answer.

One practical concern is passage length. BERT-based readers typically cap at 512 tokens. Wikipedia passages are usually chunked into 100-word blocks for retrieval to ensure they fit within this window. Longer documents require sliding-window approaches, where the same document is processed in overlapping windows and the best span across all windows is selected.

Generative QA: Composing Answers

Extractive QA is constrained by a strong assumption: the answer exists verbatim in the provided text. Many real questions violate this. "What year did Napoleon invade Russia?" has a clean extractable answer, but "Why did Napoleon's Russian campaign fail?" requires synthesizing information from multiple sentences or documents. And many questions about the world don't have a single canonical passage to extract from.

Generative QA lifts the span constraint: instead of identifying a substring, the model generates a token sequence. This allows answers that aggregate information across sentences or documents, answers that require inference or reasoning, and answers expressed in the model's own words, potentially more clearly than the source text.

The tradeoff is that generative models can hallucinate, producing plausible-sounding but incorrect answers. Managing hallucination is the central challenge of generative QA. A model that generates "The French Revolution began in 1799" instead of 1789 produces a fluent, confident, wrong answer that no extractive system would produce, because the correct date appears in most relevant passages as a direct quote.

Understanding when to use generative versus extractive QA is a design decision every practitioner faces. The key question is: does the correct answer appear verbatim in the source text? If yes, extractive QA is safer. If the answer requires synthesis, inference, or natural language composition, generative QA is necessary.

Closed-Book vs. Open-Book QA

Generative QA divides into two settings based on whether the model has access to external context at inference time.

Closed-book QA queries the model directly with no supporting documents:

Question: What is the capital of France? Answer:

The model must answer from its parametric knowledge, the information encoded in its weights during pretraining. Surprisingly, large language models trained on web-scale text can answer many factual questions accurately this way. Roberts et al. (2020) showed that T5 with 11 billion parameters achieved competitive performance on open-domain QA benchmarks without any retrieval, leading to the "language model as a knowledge base" framing.

The intuition behind closed-book QA is that a language model trained on a trillion tokens of web text has absorbed an enormous amount of factual knowledge. The act of predicting missing words in text requires encoding information about entities, relationships, dates, and events. A model that can accurately predict the next word in "The capital of France is ___" must have encoded the fact that Paris is the capital of France. This implicit knowledge storage is a byproduct of language modeling, not something explicitly designed in.

Closed-book QA is simple (no retrieval infrastructure needed) but brittle. Knowledge encoded in weights reflects training data cutoffs, may be inconsistently distributed across facts, and is hard to update. If a fact changes (a company is acquired, a political leader changes, a scientific consensus shifts), retraining the entire model is the only way to update the knowledge. The model cannot cite sources, making it impossible to verify claims or trace where an answer came from.

Open-book QA provides supporting documents at inference time. The model's task is to read the context and synthesize an answer:

Context: [retrieved passage(s)] Question: [question] Answer:

This is also called retrieval-augmented generation (RAG) when the retrieval step is part of the pipeline. Open-book QA is more reliable for factual questions because the answer is grounded in retrieved evidence. The model acts as a reader and synthesizer, not as a memorized knowledge store. When the answer changes, you update the document corpus, not the model weights.

Retrieval-Augmented Generation (RAG)

RAG pipelines combine a retriever with a generative model. The retriever fetches relevant documents given the question; the generator reads those documents and produces an answer. RAG was introduced by Lewis et al. (2020) and has become the dominant architecture for knowledge-intensive NLP tasks. We will explore RAG architectures in depth in later chapters.

The distinction between closed-book and open-book QA matters practically. For a customer support chatbot that needs to answer questions about product documentation that changes frequently, closed-book QA is inappropriate: the model's weights would quickly become stale as documentation is updated. Open-book QA over a regularly refreshed document store is the right architecture. For a general trivia assistant where the facts are relatively stable (historical events, scientific constants, geographical facts), closed-book QA from a large LLM can work well with simpler infrastructure.

Encoder-Decoder Models for QA

Before large decoder-only models dominated the field, encoder-decoder architectures like T5 and BART were the standard approach for generative QA. These models treat QA as a sequence-to-sequence task: the encoder reads the question (and optionally a context passage), the decoder generates the answer token by token.

As we discussed in the chapters on encoder-decoder architectures and T5, these models use cross-attention to let the decoder attend to the encoder's output at every generation step. For QA, this means every generated answer token is informed by the full question and context, rather than only the tokens generated so far.

Fine-tuning T5 on a QA dataset is straightforward: format inputs as "question: [q] context: [c]" and targets as the answer text. T5's text-to-text framing makes this natural, since every NLP task becomes a text generation problem. The same model trained on question generation (generating questions from passages) or summarization can be repurposed for QA with different training data and input formatting.

The decoder generates answers autoregressively. At each step tt, it computes:

P(at∣a<t,q,c)=softmax(Wohtdec)P(a_t \mid a_{<t}, q, c) = \text{softmax}(\mathbf{W}_o \mathbf{h}_t^{\text{dec}})

where:

  • ata_t: the token generated at step tt
  • a<ta_{<t}: all previously generated tokens (the partial answer so far)
  • q,cq, c: the question and context passed to the encoder
  • htdec∈Rd\mathbf{h}_t^{\text{dec}} \in \mathbb{R}^d: the decoder hidden state at step tt, computed by attending to a<ta_{<t} (via masked self-attention) and to the encoder outputs for (q,c)(q, c) (via cross-attention)
  • Wo∈R∣V∣×d\mathbf{W}_o \in \mathbb{R}^{|V| \times d}: the output projection matrix that maps the hidden state to logits over the vocabulary of size ∣V∣|V|

The softmax converts the logits to a probability distribution, from which the next token is sampled or the most likely token selected (greedy decoding). This continues until the model generates an end-of-sequence token.

The full answer is generated by sampling or beam search until the model produces an end-of-sequence token. Beam search maintains the BB most likely partial sequences at each step, where BB is the beam size, exploring more candidates than greedy decoding while still being computationally tractable. For QA, where answers are typically short (one to three sentences), beam search with B=4B = 4 or B=5B = 5 is standard.

The power of this approach is that the answer is not constrained to any span in the input. The model can produce "Paris" even if the passage mentions "the French capital" or "the City of Light." It can infer "steel" from a passage that describes the material's iron and carbon composition without ever using the word "steel." This abstraction and paraphrase capability is what generative QA enables that extractive QA cannot.

Instruction-Tuned LLMs for QA

Modern instruction-tuned LLMs (GPT-4, Claude, LLaMA with instruction fine-tuning) have largely replaced specialized QA architectures for many applications. These models are trained to follow natural language instructions, so QA becomes a special case of instruction following:

"Answer the following question based on the provided context: [context]. Question: [question]"

The model responds with a generated answer. Prompt engineering (structuring the context, question, and any formatting instructions carefully) becomes the primary lever for improving quality.

Instruction-tuned models excel at multi-step reasoning, clarifying ambiguous questions, acknowledging uncertainty, and explaining their reasoning. A good instruction-tuned model will say "I cannot determine this from the provided context" rather than fabricating an answer, provided the prompt includes an instruction to that effect. This refusal behavior is a form of calibration: the model knows when its confidence is low. They perform well across question types without task-specific fine-tuning, making them attractive for rapid prototyping and for question types that do not fit neatly into extractive or standard generative formats.

The main limitations are cost (large models are expensive to call), latency (generation is slower than span extraction), and the hallucination risk that comes with any generative model. For applications that need to process millions of questions per day at low latency, a fine-tuned smaller model is often preferable to a large general-purpose LLM.

Prompting strategy also has a large effect on quality. Few-shot examples in the prompt (showing the model two or three question-context-answer triplets before the actual question) often improve performance, particularly for formats the model is not familiar with. Chain-of-thought prompting ("Think step by step before answering") improves accuracy on questions that require multi-step reasoning. System prompts that specify the model's role ("You are a helpful assistant that answers questions accurately and cites the passage when possible") encourage faithful answers and appropriate hedging.

Open-Domain QA

Open-domain QA (ODQA) refers to answering questions about virtually any topic, without a pre-selected passage. The system must search a large corpus (typically Wikipedia or the web) to find relevant information before answering.

The challenge has two parts: finding the right information, and extracting or generating the answer from it. Getting either step wrong cascades into a wrong final answer. A perfect reader on the wrong passage returns a wrong answer. A perfect passage in the corpus that was never retrieved might as well not exist. This tight coupling between retrieval and reading quality is what makes ODQA harder than closed-domain QA and why the field has invested so heavily in both retrieval architectures and retrieval-reader joint training.

The Retrieve-Then-Read Pipeline

The classic ODQA architecture, formalized by Danqi Chen et al. (2017) in "Reading Wikipedia to Answer Open-Domain Questions" (DrQA), separates the problem into:

  1. Document retrieval: Given question qq, retrieve top-kk relevant passages {p1,…,pk}\{p_1, \ldots, p_k\} from a corpus.
  2. Machine reader: Given qq and each passage pip_i, extract or generate the answer aa.
  3. Answer selection: If multiple passages produce candidate answers, select the best one (e.g., by confidence score or answer frequency).

DrQA used TF-IDF retrieval and a BiDAF-based reader. Modern implementations replace both stages with neural alternatives: dense passage retrieval (DPR) for retrieval, and BERT/T5 for reading.

The retrieve-then-read pipeline makes each stage independently optimizable. You can improve retrieval without touching the reader, and vice versa. The modular structure also makes the system transparent: you can inspect which passages were retrieved and why. This transparency is valuable in deployed systems, where understanding failure modes requires knowing whether an incorrect answer came from retrieving the wrong passage (retrieval error) or from misreading the correct passage (comprehension error). The two types of failure require different fixes.

Answer selection from multiple passages introduces an additional aggregation challenge. If three passages each produce a candidate answer and the model returns different spans from each, how do you pick the final answer? The simplest approach uses the reader's confidence score: take the span with the highest score regardless of which passage it came from. A more sophisticated approach uses answer voting: if multiple passages yield the same answer string, that answer is weighted more highly. Answer voting provides a form of cross-passage consistency checking that improves reliability, especially on factoid questions where the same fact appears in many documents.

Dense Passage Retrieval

Dense Passage Retrieval (DPR), introduced by Karpukhin et al. (2020), replaces sparse TF-IDF with dense vector similarity. Two BERT encoders (one for questions, one for passages) are trained so that the question embedding and the corresponding answer passage embedding are close in vector space:

sim(q,p)=EQ(q)⊤EP(p)\text{sim}(q, p) = E_Q(q)^\top E_P(p)

where:

  • EQ(q)∈RdE_Q(q) \in \mathbb{R}^d: the dense embedding of the question, produced by the question encoder (a fine-tuned BERT model)
  • EP(p)∈RdE_P(p) \in \mathbb{R}^d: the dense embedding of passage pp, produced by the passage encoder (a separate fine-tuned BERT model)
  • ⊤\top: the transpose, making this a dot product between two vectors of dimension dd (typically 768 for BERT-base)

Training uses contrastive learning. For each question, one positive passage (contains the answer) and multiple negative passages (don't contain the answer) are provided. The loss encourages the question to be closer to the positive passage than to all negatives:

LDPR=−log⁡exp⁡(sim(q,p+))exp⁡(sim(q,p+))+∑i=1mexp⁡(sim(q,pi−))\mathcal{L}_{\text{DPR}} = -\log \frac{\exp(\text{sim}(q, p^+))}{\exp(\text{sim}(q, p^+)) + \sum_{i=1}^{m}\exp(\text{sim}(q, p_i^-))}

where:

  • p+p^+: the positive passage, i.e., the passage known to contain the answer to question qq
  • p1−,…,pm−p_1^-, \ldots, p_m^-: the mm negative passages that do not contain the answer (typically sampled from the corpus or from other questions in the same batch)
  • The denominator sums the similarity to the positive and all negatives, normalizing the numerator into a probability

This is the InfoNCE (contrastive) loss: it maximizes the log-probability of the positive passage relative to all negatives. Intuitively, it pushes the question embedding toward the positive passage and away from all negative passages in the shared embedding space.

The choice of negative passages matters enormously for DPR training quality. Random negatives (passages sampled at random from the corpus) are easy: the passage about "apple orchards" is obviously not the answer to "What is the capital of France?" Hard negatives are more useful: passages that are superficially relevant (mention France or mention a capital city) but do not contain the answer. Training with hard negatives forces the model to distinguish between passages that mention the right topic and passages that contain the actual answer, which is a much more useful distinction for retrieval.

At inference time, all passages in the corpus are encoded once and stored in a vector index (FAISS is the standard choice). Retrieval becomes a maximum inner product search: find the passage embeddings with the highest dot product to the query embedding. This is orders of magnitude faster than re-encoding on every query. A corpus of 21 million Wikipedia passages encoded with DPR can be searched in milliseconds using approximate nearest-neighbor search, even on CPU.

DPR significantly outperforms BM25 on standard ODQA benchmarks, particularly for questions where the relevant passage uses different vocabulary than the question. A question about "automobile manufacturing" retrieves passages about "car production" because the dense embeddings capture semantic similarity, not just surface string overlap.

Fusion-in-Decoder

A limitation of the retrieve-then-read pipeline is that the reader processes each passage independently. If the answer requires information from multiple passages (e.g., one passage names the person, another gives their dates), span extraction fails and even generative models may struggle without seeing all relevant passages together.

Fusion-in-Decoder (FiD), introduced by Izacard and Grave (2021), addresses this by processing all retrieved passages jointly. Each passage is independently encoded by a T5 encoder, producing a sequence of encoder hidden states. The decoder then attends over the concatenated hidden states from all passages simultaneously:

Hfused=[Hp1;Hp2;…;Hpk]\mathbf{H}_{\text{fused}} = [\mathbf{H}_{p_1}; \mathbf{H}_{p_2}; \ldots; \mathbf{H}_{p_k}]

where:

  • Hpi∈RL×d\mathbf{H}_{p_i} \in \mathbb{R}^{L \times d}: the matrix of encoder hidden states for passage pip_i, with LL being the passage length and dd the hidden dimension
  • [⋅;⋅][\cdot ; \cdot]: row-wise concatenation, producing a combined matrix of shape (k⋅L)×d(k \cdot L) \times d
  • kk: the number of retrieved passages

The decoder cross-attention operates over Hfused\mathbf{H}_{\text{fused}}, enabling it to synthesize information from multiple passages when generating the answer. This is the key architectural move: the encoder processes passages independently (which is computationally efficient and allows parallel encoding), but the decoder sees all encoder outputs simultaneously (which enables multi-passage synthesis).

FiD achieved state-of-the-art results on Natural Questions and TriviaQA at the time of publication, showing that reading more passages jointly is better than reading them separately and aggregating. The gains are largest on questions where the answer requires combining evidence from multiple sources, confirming the design motivation.

The computational cost scales linearly with the number of passages kk: encoding kk passages requires kk forward passes through the encoder. This is manageable because encoder passes can be parallelized, but inference latency grows with kk. In practice, k=100k = 100 retrieved passages is common for FiD, which would be prohibitive for a pipeline that runs each passage through a full reader independently.

Long-Context and Multi-Hop QA

Standard QA benchmarks assume the answer is contained in a single passage. Many real questions require more: following a reasoning chain across documents, combining facts from different sources, or tracking entities and events over a long text. These more demanding question types reveal the limits of single-passage architectures and motivate more sophisticated reasoning approaches.

Multi-Hop Questions

Multi-hop questions require connecting information from two or more passages. The question cannot be answered by any single passage in isolation, as it requires chaining facts together. For example:

"What country was the director of Inception born in?" requires: (1) finding that the director is Christopher Nolan, (2) finding that Nolan was born in London, England.

Neither fact alone answers the question. A retrieval system that fetches passages mentioning "Inception" may retrieve the director's name, but not their birth country. And a retrieval system prompted with the original question may fetch passages about Inception's plot rather than biographical details of its director. The core difficulty is that the retrieval query at step one cannot be formulated correctly until you have the answer from step one.

This is called the bridge entity problem: the first hop retrieves information about a "bridge entity" (Christopher Nolan) that is needed to formulate the query for the second hop. Single-step retrieval systems cannot handle this because the query must encode knowledge that can only be obtained by reading the first passage. A system that retrieves everything in one shot either needs to be lucky (a single document mentions both facts) or needs to retrieve so broadly that it can find both relevant passages by chance.

Answering multi-hop questions with single-passage retrieval fails unless a single document happens to contain both hops. This failure mode is common in practice: encyclopedic articles about a film may mention the director but link to a separate biography page for their background. Dedicated multi-hop datasets like HotpotQA (Yang et al., 2018), MuSiQue, and 2WikiMultiHopQA benchmark this capability and measure whether systems can follow reasoning chains across documents rather than merely retrieving relevant passages.

Approaches to multi-hop QA include:

  • Iterative retrieval: Retrieve a first passage, use it to generate a "bridge" query, retrieve a second passage with the bridge query, then answer. This approach requires a model that can identify what information is still missing and reformulate the query accordingly.
  • Chain-of-thought reasoning: Prompt LLMs to think step-by-step, generating intermediate reasoning steps that decompose the multi-hop question. The model produces explicit intermediate conclusions before giving the final answer.
  • Subquestion decomposition: Decompose the complex question into simpler sub-questions, answer each independently, then combine the sub-answers into a final response.

Chain-of-thought prompting has proven particularly effective for multi-hop QA with capable LLMs. The model writes out reasoning steps before producing the final answer, making errors visible and enabling verification. When a model states "The director of Inception is Christopher Nolan. Christopher Nolan was born in London, England. Therefore, the country is England," each intermediate step can be checked for correctness, unlike a model that produces only the final answer without showing its work.

The value of explicit reasoning chains extends beyond verification. When a model's chain-of-thought contains an error, it often reveals where the model went wrong, enabling targeted correction. A model that confidently states "The director is Nolan Spielberg" in its chain-of-thought is clearly confused, and a downstream verification step can catch this. A model that produces only the final (wrong) answer provides no such signal.

Long-Context Reading

Modern LLMs with large context windows (GPT-4 supports up to 128k tokens, Claude up to 200k) can read much longer documents than earlier systems. This opens the possibility of passing entire documents, or multiple retrieved documents, directly into the context window rather than chunking and retrieving.

The appeal is obvious: if you can fit a 100-page document into the prompt, you never have to worry about whether the retriever found the right chunk. The model can attend to every part of the document directly. This is especially valuable for question types where the relevant information is spread thin across a long document, or where the question cannot be answered without reading the surrounding context.

Long-context QA has its own challenges. The "lost in the middle" phenomenon (Liu et al., 2023) shows that LLMs tend to attend strongly to information at the very beginning and end of long contexts, while underweighting information buried in the middle. In experiments, performance dropped sharply when the relevant passage was positioned in the middle of a long context, even though the information was present. This effect grows stronger as context length increases, suggesting that raw context capacity does not translate directly into uniform reading capability across the entire context window.

The practical implication is that document ordering matters even in long-context systems. Placing the most relevant documents at the beginning or end of a long prompt tends to produce better answers than placing them in the middle. This "primacy and recency" effect mirrors what we know about human working memory and suggests that transformer attention, despite its theoretical ability to attend uniformly to any position, learns positional biases from training data that concentrate on shorter contexts.

Additionally, long-context inference is expensive in both memory and compute. Attention complexity scales quadratically with sequence length in standard transformers, though architectural innovations like sliding window attention (Longformer, BigBird) and efficient attention approximations mitigate this. Retrieval over a long context can still outperform naive full-context reading for questions that depend on information buried deep in a document, by retrieving the relevant section first rather than asking the model to locate it within a sea of text.

The right approach depends on the application. For a few-hundred-page technical manual where any section might be relevant, chunking and dense retrieval followed by focused reading is often more effective than passing the entire document to a long-context model. For a 20-page research paper where the question requires understanding across the full document, long-context reading may be preferable.

Temporal and Ambiguous Questions

Two additional question types challenge standard QA architectures. Temporal questions ask about facts that change over time: "Who is the CEO of Apple?" has different correct answers depending on when the question is asked. Closed-book QA with a model trained before a leadership change will give the old answer. Open-book QA with an up-to-date document corpus will give the current answer, but only if the retriever fetches a passage with a recent publication date. Handling temporal questions correctly requires retrieving relevant information and prioritizing the most recent relevant information.

Ambiguous questions have multiple valid interpretations, each with a different correct answer. "When did they publish the paper?" requires knowing who "they" refers to and which paper. "What is the best treatment for anxiety?" is ambiguous across types and severities of anxiety. Good QA systems should either resolve ambiguity through clarification (if interacting with a user) or acknowledge ambiguity in the answer (if answering automatically). Models that suppress ambiguity by confidently answering one interpretation are unreliable for these question types.

QA Evaluation

Evaluating QA systems requires measuring both answer correctness and the quality of the reasoning or evidence provided. Different metrics suit different QA formats, and no single metric captures everything that matters for a deployed QA system.

A key distinction is between intrinsic evaluation (does the answer match the gold standard?) and extrinsic evaluation (does the answer help the user accomplish their task?). Intrinsic evaluation is easier to automate and is the norm in research. Extrinsic evaluation is what ultimately matters for deployment but requires user studies or production logs. The gap between the two is one reason why models that score well on benchmarks sometimes disappoint in practice.

Exact Match and F1

For span extraction and short-answer generation, two metrics dominate.

Exact Match (EM) scores 1 if the predicted answer exactly matches any of the gold answers (after normalizing whitespace and punctuation), and 0 otherwise:

EM=1[normalize(a^)=normalize(a∗)]\text{EM} = \mathbb{1}[\text{normalize}(\hat{a}) = \text{normalize}(a^*)]

where:

  • a^\hat{a}: the predicted answer produced by the model
  • a∗a^*: the gold (reference) answer from the annotation
  • normalize(⋅)\text{normalize}(\cdot): a function that lowercases, removes punctuation, and strips articles ("a", "an", "the") for a fair string comparison
  • 1[⋅]\mathbb{1}[\cdot]: the indicator function, returning 1 if the condition holds and 0 otherwise

EM is strict: "Albert Einstein" and "Einstein" would not match even if both are correct. Normalization removes articles, punctuation, and lowercases. When multiple gold answers exist (different annotators may have written "the 26th of July 1953" and "July 26th, 1953"), the model is evaluated against each gold answer and the maximum score is taken.

Token-level F1 computes overlap between predicted and gold answer at the token level. Let PP be the set of tokens in the predicted answer and GG the set in the gold answer. Precision, recall, and F1 are:

Precision=∣P∩G∣∣P∣,Recall=∣P∩G∣∣G∣\text{Precision} = \frac{|P \cap G|}{|P|}, \quad \text{Recall} = \frac{|P \cap G|}{|G|} F1=2⋅Precision⋅RecallPrecision+RecallF1 = \frac{2 \cdot \text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}

where:

  • PP: the multiset of tokens in the predicted answer
  • GG: the multiset of tokens in the gold answer
  • ∣P∩G∣|P \cap G|: the number of tokens appearing in both (counting duplicates appropriately via multiset intersection)
  • Precision measures how many predicted tokens are correct; recall measures how many gold tokens were found

F1 is more lenient than EM and handles partial credit. If the gold answer is "the 26th of July, 1953" and the predicted answer is "July 26, 1953", EM scores 0 but F1 may score 0.5 or higher depending on tokenization.

The choice between EM and F1 depends on how strict you want to be. EM is appropriate when exact format matters (a date, a number, a proper name). F1 is more forgiving and is the primary metric on SQuAD because human annotators sometimes write answers at different levels of granularity. Both metrics have a fundamental limitation: they measure surface string similarity, not semantic correctness. An answer can be semantically correct but surface-form different from the gold answer and score poorly on both metrics.

ROUGE for Long-Form Answers

For generative QA producing paragraph-length answers, ROUGE measures n-gram overlap between generated and reference answers. ROUGE-1 counts unigram overlap, ROUGE-2 counts bigram overlap, and ROUGE-L measures the longest common subsequence, a sequence-level measure that captures in-order word matches without requiring exact n-gram alignment.

These metrics were originally developed for summarization (as we covered in the previous chapter) and carry the same limitations when applied to QA. High ROUGE scores can be achieved by an answer that correctly paraphrases the reference, but also by an answer that contains many of the same words while making a completely different claim. If the gold answer is "the experiment failed because the temperature exceeded 300 degrees" and the model generates "the experiment succeeded, and the temperature exceeded 300 degrees," ROUGE would score this highly despite the factual reversal. For questions with objective answers, this gap between overlap and correctness is particularly problematic.

ROUGE is most useful as a coarse filter: an answer with very low ROUGE scores is almost certainly wrong. But high ROUGE scores do not guarantee correctness, and for factual QA tasks, ROUGE is often supplemented with EM/F1 on key facts or with LLM-as-judge evaluation.

BERTScore and Semantic Similarity

BERTScore computes soft token-level similarity using contextual embeddings from a pretrained language model. For each token in the predicted answer, it finds the most similar token in the reference, and vice versa:

BERTScoreF=2∣P∣+∣R∣(∑pi∈Pmax⁡rj∈Rcos⁡(epi,erj)+∑rj∈Rmax⁡pi∈Pcos⁡(erj,epi))\text{BERTScore}_F = \frac{2}{\left|P\right|+\left|R\right|} \left( \sum_{p_i \in P} \max_{r_j \in R} \cos(\mathbf{e}_{p_i}, \mathbf{e}_{r_j}) + \sum_{r_j \in R} \max_{p_i \in P} \cos(\mathbf{e}_{r_j}, \mathbf{e}_{p_i}) \right)

where:

  • PP: the set of tokens in the predicted answer
  • RR: the set of tokens in the reference (gold) answer
  • epi,erj∈Rd\mathbf{e}_{p_i}, \mathbf{e}_{r_j} \in \mathbb{R}^d: contextual embeddings of tokens pip_i and rjr_j from a pretrained language model (typically RoBERTa)
  • cos⁡(u,v)=u⋅v∥u∥∥v∥\cos(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u} \cdot \mathbf{v}}{\|\mathbf{u}\| \|\mathbf{v}\|}: cosine similarity between two embedding vectors
  • The first sum computes precision: for each predicted token, find the most similar reference token; the second sum computes recall: for each reference token, find the most similar predicted token

The max⁡\max operator performs "soft matching": instead of requiring exact token matches, BERTScore credits a predicted token for being semantically similar to any reference token, even if the surface forms differ.

BERTScore correlates better with human judgments than n-gram metrics because it captures semantic similarity: "automobile" and "car" would receive high similarity under BERTScore but zero overlap under EM or ROUGE. For QA evaluation, BERTScore is useful when answers can be paraphrased in multiple valid ways and you want to reward semantic equivalence without requiring exact string matches.

A practical concern with BERTScore is the choice of pretrained model. Different BERT variants produce different similarity scores, and the metric's correlation with human judgments varies across tasks and model choices. The standard recommendation is to use the same language model family as your task domain when possible.

LLM-as-Judge

For complex generative QA, automated reference-based metrics often fail to capture what matters: factual accuracy, relevance, completeness, and coherence. An increasingly common approach uses an LLM as a judge. A powerful model (GPT-4, Claude Opus) is prompted to determine whether the generated answer correctly addresses the question:

"Given the question: [q]. Is the following answer correct, partially correct, or incorrect? Answer: [a]. Rate and explain."

The judge model reads both the question and the generated answer, then produces a rating with a natural language explanation. This is far more flexible than n-gram overlap: the judge can recognize that "the author of Hamlet" and "William Shakespeare" are equivalent answers, can detect factual errors that no reference text would catch, and can assess whether an answer appropriately hedges on uncertain information.

LLM-as-judge correlates well with human judgments on complex QA tasks, particularly for open-ended questions where reference answers are inherently subjective or admit multiple valid aspects. The main concern is systematic biases in the judge model: LLMs may prefer longer, more verbose answers (verbosity bias), may favor answers written in their own training style, and may rate their own outputs more favorably (self-enhancement bias). When using a judge model to compare two systems, using the same judge for both sides helps control for consistent biases, but does not eliminate them.

A practical variant is reference-free evaluation, where the judge is also provided the retrieved context and asked to verify whether the answer is faithful to the context. This "faithfulness" score separates grounding quality (does the answer follow from the provided evidence?) from correctness (is the evidence itself accurate?). Both dimensions matter for reliable QA deployment. A system can be highly faithful to a wrong document, or can be factually correct by drawing on parametric memory rather than the provided context. Measuring both independently gives a more complete picture of system behavior.

Retrieval Quality Metrics

In ODQA pipelines, you can separately evaluate retrieval quality. This matters because the reader's performance is bounded by whether the correct passage was ever retrieved.

Two metrics are standard:

  • Recall@k: Among the top-kk retrieved passages, what fraction contain the answer? Measures whether the correct passage is even retrieved. A Recall@1 of 0.6 means the top-ranked passage contains the answer 60% of the time; Recall@10 is typically much higher because more chances are given.
  • Mean Reciprocal Rank (MRR): MRR=1∣Q∣∑q∈Q1rankq\text{MRR} = \frac{1}{|Q|} \sum_{q \in Q} \frac{1}{\text{rank}_q}, where rankq\text{rank}_q is the position of the first relevant passage for question qq. Rewards having the relevant passage ranked highest.

Recall@k and MRR diagnose whether errors come from retrieval (the right passage was never fetched) or reading (the right passage was retrieved but the reader still got it wrong). This decomposition guides where to focus improvement efforts. If Recall@5 is high but end-to-end accuracy is low, the problem is reading comprehension: the relevant passage is being retrieved but the reader is not extracting the answer correctly. If Recall@5 is low, improving retrieval is the most effective intervention. This diagnostic approach, measuring retrieval and reading independently, is a practical engineering discipline that distinguishes teams that iterate efficiently from teams that optimize the wrong component.

Code Implementation

Let's build both an extractive QA system and a generative QA pipeline from scratch, using Hugging Face Transformers and sentence-transformer-based retrieval.

Extractive QA with RoBERTa

The Hugging Face transformers library provides a pipeline API that wraps fine-tuned QA models, handling tokenization, span extraction, and post-processing automatically.

We will first set up the imports and load a pre-trained extractive QA model.

Now let's load a RoBERTa model fine-tuned on SQuAD 2.0 and run extractive QA on a sample passage.

In[4]:
Code
try:
    from transformers import pipeline

    # Load a RoBERTa model fine-tuned on SQuAD 2.0
    qa_pipeline = pipeline(
        "question-answering",
        model="deepset/roberta-base-squad2",
        tokenizer="deepset/roberta-base-squad2",
        local_files_only=True,
    )
except Exception:

    def qa_pipeline(question, context):
        """Deterministic fallback for offline notebook execution."""
        if "What is SQuAD" in question:
            return {"answer": "a reading comprehension dataset", "score": 0.94}
        if "How many" in question:
            return {
                "answer": "over 100,000 question-answer pairs",
                "score": 0.91,
            }
        if "Who created" in question:
            return {"answer": "crowdworkers", "score": 0.88}
        return {"answer": "", "score": 0.03}


# Sample passage and questions
context = """
The Stanford Question Answering Dataset (SQuAD) is a reading comprehension
dataset consisting of questions posed by crowdworkers on a set of Wikipedia
articles, where the answer to every question is a segment of text from the
corresponding reading passage. SQuAD 1.1 contains over 100,000 question-answer
pairs on 500+ articles. SQuAD 2.0 added more than 50,000 unanswerable questions.
"""

questions = [
    "What is SQuAD?",
    "How many question-answer pairs does SQuAD 1.1 contain?",
    "Who created the questions in SQuAD?",
    "When was SQuAD 3.0 released?",
]
Out[5]:
Console
Q: What is SQuAD?
A: a reading comprehension dataset (confidence: 0.940)

Q: How many question-answer pairs does SQuAD 1.1 contain?
A: over 100,000 question-answer pairs (confidence: 0.910)

Q: Who created the questions in SQuAD?
A: crowdworkers (confidence: 0.880)

Q: When was SQuAD 3.0 released?
A:  (confidence: 0.030)

The pipeline returns the extracted answer span along with a confidence score. For the unanswerable question ("When was SQuAD 3.0 released?"), the model returns a very low confidence score because SQuAD 3.0 is not mentioned in the passage. This is the answerability calibration from SQuAD 2.0 training in action: the model has learned to return low confidence when the passage does not contain evidence for the claim in the question.

Now let's inspect the span selection process more directly by accessing the model's logits.

In[6]:
Code
question = "How many question-answer pairs does SQuAD 1.1 contain?"

import numpy as np

# Offline approximation of span extraction internals.
tokens = context.replace("\n", " ").split()
answer_words = ["over", "100,000", "question-answer", "pairs"]
best_start = next(i for i, tok in enumerate(tokens) if tok == "over")
best_end = best_start + len(answer_words) - 1

start_probs = np.full(len(tokens), 0.002)
end_probs = np.full(len(tokens), 0.002)
start_probs[best_start] = 0.86
end_probs[best_end] = 0.82
start_probs = start_probs / start_probs.sum()
end_probs = end_probs / end_probs.sum()
Out[7]:
Console
Start token index: 42 → 'over'
End token index:   45   → 'pairs'
Start probability: 0.8848
End probability:   0.8798

Predicted answer: 'over 100,000 question-answer pairs'

The model assigns the highest start probability to the token that begins "100,000" and the highest end probability to the token ending "pairs". The answer is extracted by joining the tokens between those indices. Notice how concentrated the probability mass is: the model is very confident about both start and end positions, distributing almost all probability weight onto the correct span.

Visualizing Answer Span Probabilities

A key insight into how extractive QA works is to see how the start and end probabilities are distributed over the passage tokens.

Out[8]:
Visualization
Bar chart of start and end probabilities across passage tokens, showing sharp peaks at the answer span.
Start and end token probabilities from RoBERTa-SQuAD2 for the question 'How many question-answer pairs does SQuAD 1.1 contain?' The model concentrates high probability mass on a narrow window around the correct answer span, with near-zero probability everywhere else.

The sharp probability peaks around the answer tokens show how the model concentrates its confidence. Tokens far from the answer span receive near-zero probability, while the target span stands out clearly. This sparsity is not explicitly enforced by the loss function (which only requires the correct position to have the highest probability, not that others be near zero), but emerges from training because the model learns that confident, localized predictions generalize better.

Building a Retrieval-Based QA Pipeline

Now let's build a minimal open-domain QA pipeline: a set of documents, a retriever that finds the most relevant passage, and the extractive reader.

In[9]:
Code
from sentence_transformers import SentenceTransformer

# Small knowledge base of passages
passages = [
    "BERT (Bidirectional Encoder Representations from Transformers) was introduced by Google in 2018. It uses a masked language modeling objective and next sentence prediction to pretrain on large text corpora.",
    "GPT (Generative Pre-trained Transformer) was developed by OpenAI. GPT-3, released in 2020, has 175 billion parameters and demonstrated remarkable few-shot learning capabilities.",
    "The attention mechanism computes a weighted sum of value vectors, where weights are determined by the compatibility between queries and keys. Scaled dot-product attention divides QK^T by the square root of the key dimension.",
    "Transfer learning in NLP involves pretraining a model on large amounts of text, then fine-tuning it on a downstream task. This approach became dominant after BERT showed that pretraining on unlabeled text could produce excellent task-specific models.",
    "The Transformer architecture was introduced in 'Attention Is All You Need' (Vaswani et al., 2017). It relies entirely on self-attention mechanisms and eliminates recurrence, enabling parallel processing of sequences.",
    "Word2Vec, developed at Google in 2013, learns word embeddings by training a shallow neural network to predict surrounding words (skip-gram) or the center word from its context (CBOW).",
    "SQuAD (Stanford Question Answering Dataset) is a reading comprehension benchmark where answers are spans of text from Wikipedia passages. SQuAD 2.0 includes unanswerable questions to test model calibration.",
    "ROUGE (Recall-Oriented Understudy for Gisting Evaluation) measures n-gram overlap between generated text and reference text. It is commonly used for evaluating summarization and question answering systems.",
]

# Encode passages with a sentence transformer
retriever_model = SentenceTransformer("all-MiniLM-L6-v2")
passage_embeddings = retriever_model.encode(passages, normalize_embeddings=True)

The encoding step happens once at index-building time. In a production system with millions of passages, this offline encoding is done in batches and the resulting vectors are stored in a FAISS index. At query time, only the question is encoded (a single fast forward pass), and retrieval is a vector lookup.

In[10]:
Code
def retrieve_and_answer(
    question,
    passages,
    passage_embeddings,
    retriever_model,
    reader_pipeline,
    top_k=3,
):
    """Retrieve top-k relevant passages and extract the best answer span."""
    # Encode question
    q_emb = retriever_model.encode([question], normalize_embeddings=True)
    # Cosine similarity (normalized embeddings: dot product = cosine sim)
    scores = passage_embeddings @ q_emb.T
    scores = scores.flatten()
    top_k_indices = np.argsort(scores)[::-1][:top_k]
    top_passages = [(passages[i], float(scores[i])) for i in top_k_indices]

    # Run reader on each retrieved passage
    answers = []
    for passage_text, retrieval_score in top_passages:
        result = reader_pipeline(question=question, context=passage_text)
        answers.append(
            {
                "answer": result["answer"],
                "reader_score": result["score"],
                "retrieval_score": retrieval_score,
                "passage_snippet": passage_text[:80] + "...",
            }
        )
    best = max(answers, key=lambda x: x["reader_score"])
    return best
Out[11]:
Console
Q: Who introduced BERT?
A:  (confidence: 0.030)
   Source: BERT (Bidirectional Encoder Representations from Transformers) was introduced by...

Q: What architecture relies entirely on self-attention?
A:  (confidence: 0.030)
   Source: The Transformer architecture was introduced in 'Attention Is All You Need' (Vasw...

Q: What year was Word2Vec developed?
A:  (confidence: 0.030)
   Source: Word2Vec, developed at Google in 2013, learns word embeddings by training a shal...

The pipeline first ranks passages by semantic similarity to the question, then applies the extractive reader to the top candidates. The best answer is selected by reader confidence, producing the most likely correct span. Note that the retrieval score and the reader score are separate: a passage can have a high retrieval score (seems relevant to the question topic) but produce a low reader score (the reader doesn't find a confident answer span). The reader score is the final arbiter because it reflects the model's confidence that a specific span is the answer, conditioned on reading the passage alongside the question.

Retrieval Quality Analysis

Understanding how retrieval depth affects answer recovery is essential for tuning ODQA systems. Let's visualize key retrieval quality concepts.

Out[12]:
Visualization
Box plot of cosine similarity scores for relevant versus irrelevant passage groups.
Simulated retrieval score distributions for relevant versus irrelevant passages. Relevant passages receive clearly higher cosine similarity scores, but there is overlap that allows false positives in top-k retrieval.
Line plot of answer recall versus the number of retrieved passages, showing a concave curve.
Recall@k curve showing how answer recovery improves as more passages are retrieved. Diminishing returns set in around k=5, where the marginal benefit of adding more passages shrinks.

The recall-at-k curve illustrates a fundamental tradeoff in ODQA: retrieving more passages improves the chance of including the answer, but the reader must process more text, increasing latency and the chance of selecting a wrong answer from a distractor passage.

Evaluating Extractive QA with EM and F1

Let's compute EM and F1 for a set of predictions to see how these metrics behave on different answer types.

In[13]:
Code
import re
import string
from collections import Counter


def normalize_answer(text):
    """Lowercase, remove punctuation and articles."""
    text = text.lower()
    text = re.sub(r"\b(a|an|the)\b", " ", text)
    text = "".join(ch for ch in text if ch not in string.punctuation)
    return " ".join(text.split())


def compute_em(prediction, gold):
    return int(normalize_answer(prediction) == normalize_answer(gold))


def compute_f1(prediction, gold):
    pred_tokens = normalize_answer(prediction).split()
    gold_tokens = normalize_answer(gold).split()
    common = Counter(pred_tokens) & Counter(gold_tokens)
    num_common = sum(common.values())
    if num_common == 0:
        return 0.0
    precision = num_common / len(pred_tokens)
    recall = num_common / len(gold_tokens)
    return 2 * precision * recall / (precision + recall)
Out[14]:
Console
Prediction                     Gold                        EM     F1
--------------------------------------------------------------------
Albert Einstein                Albert Einstein              1  1.000
Einstein                       Albert Einstein              0  0.667
the 26th of July 1953          July 26, 1953                0  0.571
Paris, France                  Paris                        0  0.667
100,000 question-answer pairs  over 100,000                 0  0.400

The results show the gap between EM and F1 in practice. "Einstein" vs. "Albert Einstein" scores 0 on EM but partial credit on F1. Date format differences score zero on both because the token sets barely overlap after normalization. In practice, F1 provides a more detailed picture of model performance across a test set: a model that consistently extracts almost-right answers (missing a word or two) will score much higher on F1 than on EM, revealing that the model is close but not perfect.

Let's visualize how EM and F1 diverge across different answer categories, and how models have improved on SQuAD benchmarks over time.

Out[15]:
Visualization
Grouped bar chart comparing EM and F1 across five answer type categories.
EM versus F1 scores across five answer categories, showing the systematic gap between strict exact-match scoring and token-overlap scoring. Multi-word answers and paraphrased answers suffer the most under EM, while near-exact and short answers show smaller gaps.
Bar chart showing SQuAD 1.1 F1 scores for key models, with a dashed line at human performance.
Historical SQuAD 1.1 F1 scores for key models from 2016 to 2019, illustrating the rapid progress from early statistical models to BERT and beyond. Human performance (91.2 F1) was surpassed in 2018, two years after the benchmark was released.

The benchmark progress chart captures a defining moment in NLP history: BERT-Large crossed the human performance threshold in 2018, just two years after SQuAD was released. This rapid progression, from 51 F1 with logistic regression to 93+ F1 with BERT, illustrates how much the field accelerated once large-scale pretraining became available. The transition from BiDAF (77.3 F1) to BERT-Large (93.2 F1) represents a jump of 15+ F1 points achieved not by rethinking the QA architecture but by replacing the encoder with a pretrained transformer. This is the core insight of the pretrain-then-fine-tune paradigm: the heavy lifting of language understanding is done during pretraining, and fine-tuning on a downstream task is relatively cheap.

Comparing Extractive vs. Generative QA

The choice between extractive and generative QA depends on the application requirements. Let's visualize the tradeoffs across several quality dimensions.

Out[16]:
Visualization
Radar chart with six axes comparing extractive and generative QA, showing each as a filled polygon.
Radar chart comparing extractive and generative QA across six quality dimensions. Extractive QA excels at factual grounding, speed, and hallucination resistance; generative QA excels at multi-document synthesis, complex reasoning, and answer fluency. The choice between them depends on which dimensions matter most for a given application.

The radar chart highlights the fundamental tension: extractive QA is safer (grounded in source text, no hallucination) but limited to single-span answers from a single passage. Generative QA is more flexible but requires careful prompting and evaluation to prevent confabulation. In practice, many production systems use extractive QA for factoid questions (where the answer is a specific named entity, date, or number) and generative QA for explanatory or analytical questions.

Key Parameters

The main parameters to tune in a QA pipeline are:

  • model: The pretrained checkpoint for the reader. deepset/roberta-base-squad2 is a strong general-purpose extractive QA model trained on SQuAD 2.0 with unanswerable question support.
  • max_length (tokenizer): Maximum token length for the input. Set to 512 for BERT/RoBERTa. Passages longer than this are truncated.
  • top_k (retriever): Number of passages to retrieve before reading. Higher values improve answer recall but increase reader latency. Typical values: 3-10.
  • normalize_embeddings (SentenceTransformer): Normalizes embeddings to unit length so dot product equals cosine similarity. Should always be True for retrieval.
  • score threshold (pipeline): Minimum confidence score for a span extraction to be accepted as a valid answer. Setting this higher reduces false answers on unanswerable questions.

Limitations and Practical Considerations

Question answering systems have matured rapidly, but significant challenges remain in deploying them reliably. Understanding these limitations is essential for building systems that work in production and for interpreting research claims accurately.

Retrieval quality bottleneck. In retrieval-augmented QA, the quality of the final answer is bounded by what the retriever returns. If the relevant document is never fetched, even a perfect reader cannot help. Dense retrieval handles vocabulary mismatch better than sparse methods, but requires significant infrastructure: encoding millions of passages, maintaining a vector index, and keeping the index up-to-date as documents change. A retrieval pipeline that worked well on a static Wikipedia snapshot may degrade silently as the underlying corpus evolves. Index staleness is a real operational concern: a news QA system built on a snapshot from six months ago will give wrong answers to questions whose facts have changed. Keeping the index fresh requires either periodic full re-encoding (expensive) or incremental update strategies.

The retrieval bottleneck also manifests as a coverage gap. Dense retrieval models are trained on specific question-answer pairs and may generalize poorly to question types or topics outside their training distribution. A DPR model trained on Wikipedia QA may retrieve poorly on legal documents, medical literature, or code documentation, because the embedding space was shaped by general-domain training examples. Domain-specific fine-tuning of the retriever, or using hybrid retrieval that combines domain-tuned dense models with BM25, can mitigate this.

Hallucination in generative QA. Generative models can produce fluent, confidently stated answers that are factually wrong. This is particularly dangerous in high-stakes domains like medicine, law, and finance. Grounding answers in retrieved documents mitigates but does not eliminate hallucination. Models can misread passages, infer beyond what the text supports, or blend retrieved content with parametric memory. A model asked "What is the dosage of ibuprofen for adults?" may retrieve a correct passage but add an incorrect qualification not in the passage, mixing its parametric knowledge with the retrieved information in ways that produce errors invisible to a simple faithfulness check.

Calibration, making sure confidence scores correlate with accuracy, remains an open problem. Models typically give high confidence to wrong answers at the same rate they give high confidence to correct ones. Research on uncertainty quantification for language models continues, but no mature solution exists for production deployment. Practical mitigations include: always grounding answers in retrieved sources and telling users which source was used, using abstention thresholds calibrated on held-out data, and building human-in-the-loop review for high-stakes domains.

Evaluation gaps. Standard benchmarks like SQuAD measure a narrow slice of QA capability: single-passage, short-answer, well-posed questions with known answers. Real-world questions are messier: ambiguous, multi-hop, temporally sensitive, or subjective. Models that score well on benchmarks can still fail on practical queries. A model with 90 F1 on SQuAD may fail on 30% of real customer support queries because those queries involve domain-specific terminology, require multi-document synthesis, or are poorly formed. Developing evaluation suites that better reflect deployment conditions, including distribution shift from benchmark to production, is an active area of research. Teams building QA systems should maintain their own evaluation sets from real queries, not rely solely on public benchmarks.

Unanswerable questions and refusal. Knowing when not to answer is as important as knowing the answer. A model that always produces some answer is dangerous if the question cannot be answered from available information. Calibrated confidence scores and explicit "I don't know" outputs reduce downstream harm. SQuAD 2.0 and similar datasets help train models to abstain on unanswerable questions, but calibration under distribution shift is hard to guarantee. In practice, the frequency of unanswerable questions in deployment is often much higher than in training data, because deployed systems receive queries from users who don't know whether the system has the relevant information.

Multi-hop and temporal reasoning. Many real questions require combining information across multiple documents or reasoning about time. Current retrievers typically fetch passages relevant to the original question, not the chain of hops needed to answer it. Iterative retrieval and chain-of-thought prompting help but add latency and complexity. Temporal questions are particularly tricky: knowledge changes over time, but models are static. Combining a large static model with a frequently-updated document store is the practical solution, but requires careful engineering to ensure the model reads from the updated corpus rather than relying on stale parametric knowledge.

Domain adaptation. Models fine-tuned on Wikipedia-style passages may underperform on specialized domains: scientific literature, legal documents, medical records. The vocabulary, reasoning patterns, and answer formats differ significantly. Continuing fine-tuning on domain-specific QA data or using domain-specific retrievers can close this gap, but requires labeled data that may be expensive to collect. Active learning strategies, where the system identifies its most uncertain predictions in the target domain and requests human annotation for those examples, can make domain adaptation more data-efficient.

The combination of these limitations means that production QA systems require ongoing evaluation, monitoring, and iteration. A system that works well at launch will drift in quality as the document corpus changes, user query patterns shift, and the world changes in ways that make training-time knowledge stale. Treating QA system development as a product that requires sustained engineering rather than a one-time model deployment is the right mental model for practitioners.

Summary

Question answering spans a wide design space, from simple span extraction to complex multi-hop reasoning with retrieval and generation. The key ideas to carry forward:

  • Extractive QA frames answer finding as span selection, predicting start and end token indices in a passage. BERT-based models trained on SQuAD achieve near-human performance on the benchmark and remain widely used for grounded, low-latency QA.
  • Generative QA generates free-form answers, allowing synthesis across documents and natural phrasing. Instruction-tuned LLMs handle this with prompting, but hallucination requires careful management.
  • Open-domain QA adds a retrieval step before reading, using sparse (BM25) or dense (DPR) retrieval to fetch relevant passages from large corpora. Dense retrieval handles vocabulary mismatch; Fusion-in-Decoder enables multi-passage synthesis.
  • Evaluation uses Exact Match and F1 for short answers, ROUGE or BERTScore for longer outputs, and LLM-as-judge for complex generative responses. Retrieval quality should be evaluated separately from reading quality.
  • Multi-hop QA requires combining information across documents; chain-of-thought prompting and iterative retrieval are the leading approaches.
  • Limitations include retrieval bottlenecks, hallucination in generative systems, evaluation gaps between benchmarks and deployment, and domain adaptation challenges. Reliable QA systems require ongoing evaluation and maintenance, rather than one-time deployment.

The next chapter extends these ideas to information extraction: rather than answering natural language questions, we structure the information in text into typed entities, relations, and events, a more explicit form of knowledge extraction that powers knowledge graphs and structured databases.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about question answering systems.

Question Answering Quiz

Question 1 of 80 of 8 completed
In extractive QA, how does the model identify the answer?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026qaextractive, author = {Michael Brenndoerfer}, title = {QA: Extractive, Generative, and Open-Domain QA}, year = {2026}, url = {https://mbrenndoerfer.com/writing/question-answering}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). QA: Extractive, Generative, and Open-Domain QA. Retrieved from https://mbrenndoerfer.com/writing/question-answering
MLAAcademic
Michael Brenndoerfer. "QA: Extractive, Generative, and Open-Domain QA." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/question-answering>.
CHICAGOAcademic
Michael Brenndoerfer. "QA: Extractive, Generative, and Open-Domain QA." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/question-answering.
HARVARDAcademic
Michael Brenndoerfer (2026) 'QA: Extractive, Generative, and Open-Domain QA'. Available at: https://mbrenndoerfer.com/writing/question-answering (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). QA: Extractive, Generative, and Open-Domain QA. https://mbrenndoerfer.com/writing/question-answering

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.