Part of Language AI Handbook
Covers RAG system design by exploring retriever-generator interactions, timing strategies like iterative retrieval, and architectural variations like RETRO.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
RAG Architecture and Design Patterns
The previous chapter established why retrieval-augmented generation matters: it grounds language models in external knowledge, reducing hallucinations and letting access to information beyond the training cutoff. But understanding the motivation is only half the story. To build effective RAG systems, you need to understand how the components fit together and the design decisions that shape system behavior. A RAG system that retrieves poorly will produce wrong answers even if the generator is perfect. Conversely, a weak generator will fail to synthesize correct answers even when the retriever finds exactly the right documents. The interaction between these components, and the way you orchestrate that interaction, determines whether your system succeeds or fails in practice.
Think of a RAG system as a research assistant who has access to a library. When you ask a question, the assistant first searches the library's catalog and retrieves the most relevant books and articles. Then, with those materials spread on the desk, the assistant reads through them and composes an answer. The quality of the assistant's response depends on two things: how good the librarian is at finding relevant materials (retrieval), and how skilled the assistant is at reading comprehension and synthesis (generation). A brilliant researcher who retrieves the wrong books will give you a wrong answer. An excellent librarian who supplies perfect sources to an incompetent reader will also fail. Both components must work, and they must work together.
A RAG system consists of two primary components: a retriever that finds relevant documents from a knowledge base, and a generator (typically a large language model) that synthesizes an answer from those documents. These components interact through a carefully designed interface, and the timing and manner of their interaction significantly affects system performance. The retriever acts as a search engine, filtering millions of documents to identify the few most relevant ones. The generator then reads these documents and synthesizes a response that addresses your question. But this description understates the complexity: every decision about how to encode documents, what index structure to use, how many documents to retrieve, how to format the context, and when retrieval occurs can measurably shift final system quality.
This chapter examines each component in detail, explores when retrieval occurs relative to generation, and surveys the major architectural variations that have emerged since the original RAG paper. We also build a complete working pipeline so you can see these concepts in code. By the end, you'll understand the design space well enough to make informed decisions when building your own systems, and you'll know which trade-offs matter most for different use cases.
The idea of augmenting language models with retrieved context predates the transformer era. Early open-domain question answering systems like DrQA (2017) retrieved Wikipedia paragraphs and fed them to a reading comprehension model. The key insight was that retrieval could replace some of the knowledge stored in model parameters. The REALM paper from Google (2020) made retrieval differentiable and learned the retriever end-to-end. The original RAG paper from Facebook AI Research (2020) by Lewis et al. popularized the term "retrieval-augmented generation" and demonstrated that marginalizing over retrieved documents during training could significantly improve performance on knowledge-intensive tasks. Since then, the field has exploded: systems like Atlas, Toolformer, RETRO, and Self-RAG have each pushed different aspects of the retrieval-generation interface, and the architecture has become a standard component in production AI systems worldwide.
The Retriever Component
The retriever's job is deceptively simple: given a query, return the documents most likely to contain relevant information. However, this process is complex. The retriever must somehow measure the relevance between your question, expressed in natural language with all its ambiguity and variation, and documents that may discuss the same concepts using entirely different vocabulary. Someone asking about "automobile reliability" and a document discussing "car dependability" are clearly related, but a keyword matcher would miss the connection entirely. In practice, achieving effective retrieval involves three sub-systems working together: document encoding, indexing, and query-time retrieval. Each sub-system has its own design decisions, and getting them right requires understanding both the algorithmic options and the practical constraints of your deployment environment.
The retriever is often the weakest link in RAG systems. Language model generators have improved dramatically over the past few years, but retrieval remains a hard problem. The challenge is that relevance is inherently query-dependent: the same document might be highly relevant for one question and irrelevant for another. A document explaining how attention mechanisms work in transformers is invaluable for someone asking "how does self-attention work?" but useless for someone asking "what is the capital of France?" The retriever must make relevance judgments without knowing exactly what the generator needs, and it must make those judgments quickly enough to support interactive applications.
Think of the retriever as a sommelier at a restaurant. When you describe what you'd like to drink, the sommelier doesn't pour every bottle in the cellar, they draw on their knowledge of the wine list to select a few bottles they believe will satisfy you. Their recommendations depend on how well they understood your description, how well they know the cellar, and how good their model of your preferences is. A great sommelier finds exactly what you wanted. A mediocre one gives you something in the right category but misses the details. The retriever faces an analogous challenge: it must interpret your query and search a knowledge base to find the items most likely to satisfy the generator.
Document Encoding
Before retrieval can happen, documents must be transformed into a searchable representation. Raw text, while meaningful to humans, cannot be efficiently compared or searched by algorithms. The encoding process bridges this gap, converting human-readable text into mathematical objects that computers can manipulate. This process typically involves two stages that work together to prepare documents for rapid, accurate retrieval.
First, documents are split into smaller chunks. A 50-page technical report cannot be fed to a language model in its entirety, so we divide it into passages of a few hundred tokens each. The chunking strategy significantly affects retrieval quality: too small and you lose context, too large and you dilute relevant information with noise. Consider a chunk that contains only half a sentence: the retriever might match it to a query based on keyword overlap, but the generator would receive incomplete, potentially misleading information. Conversely, a chunk spanning multiple pages might contain one relevant paragraph buried among irrelevant content, which makes it harder for the generator to identify what matters. A common heuristic is 256 to 512 tokens per chunk with 10 to 20 percent overlap between adjacent chunks, but the right size depends on your documents and queries. Technical documentation with dense, self-contained paragraphs might warrant smaller chunks, while narrative text with ideas that span multiple paragraphs might need larger ones. We'll explore chunking strategies in depth in a later chapter, examining techniques like overlapping windows, semantic boundary detection, and hierarchical chunking.
Second, each chunk is converted into a numerical representation suitable for similarity search. This transformation is the heart of the retrieval process, determining how the system measures whether one piece of text relates to another. Two broad approaches dominate the field, each with distinct characteristics that make them suited to different scenarios.
Sparse representations like TF-IDF and BM25 represent documents as high-dimensional vectors where each dimension corresponds to a vocabulary term. As we covered in Part II, BM25 scores documents based on term frequency, document length normalization, and inverse document frequency. The intuition behind sparse representations is straightforward: documents are characterized by the words they contain. A document about neural networks will have non-zero values for terms like "neuron," "layer," and "activation," while a document about cooking will have non-zero values for entirely different terms. These representations are sparse because most terms don't appear in any given document, leaving most dimensions at zero. A vocabulary might contain 100,000 terms, but any single document uses only a few hundred, resulting in vectors that are more than 99 percent zeros. The sparsity is not a weakness, it enables the highly efficient inverted index data structure that makes BM25 retrieval lightning fast. Sparse methods are interpretable (you can see exactly which terms caused a document to rank highly), require no training data, and can be deployed with minimal infrastructure. They remain competitive baselines and are often used in hybrid systems.
Dense representations use neural networks to encode chunks into continuous vectors, typically with 384 to 4096 dimensions. Unlike sparse vectors where dimensions correspond to specific words, dense embedding dimensions capture abstract semantic features that emerge during training. These features don't have human-interpretable names; instead, they represent learned patterns that help distinguish relevant from irrelevant content. Two documents discussing the same concept with different vocabulary will have similar dense embeddings even if their sparse representations share few non-zero dimensions. For example, a query about "car maintenance" would find documents discussing "vehicle repair" or "automobile servicing" because the neural network has learned that these phrases occupy similar regions of the embedding space. This semantic understanding comes at a cost: dense retrieval requires training a neural encoder (or fine-tuning a pre-trained one), storing and searching high-dimensional vectors, and GPU infrastructure for encoding at scale. But the quality gains are substantial, especially for queries that involve paraphrase or conceptual relationships rather than exact keyword matching.
The key insight is that sparse and dense representations capture complementary signals. Sparse methods are precise for exact terminology, dense methods are flexible for semantic content. This observation motivates hybrid retrieval, which we'll discuss in the architecture variations section.


The choice between sparse and dense retrieval involves trade-offs we'll examine in the next chapter. Sparse methods run quickly without training data, and their results are easy to interpret, making them excellent baselines. Dense methods capture semantic similarity that keyword matching misses but require substantial training data and computational resources. For now, recognize that both approaches solve the same basic problem: converting text into vectors that can be compared for similarity.
The Index
Encoded documents are stored in an index optimized for fast similarity search. Without an index, finding the most similar documents to a query would require computing similarity with every document in the collection. For a knowledge base with millions of documents, this brute-force approach would take seconds or minutes per query, far too slow for interactive applications. The index structure depends on the retrieval approach, with each type of representation requiring specialized data structures that exploit the mathematical properties of the representation.
For sparse retrieval, inverted indexes map each vocabulary term to the list of documents containing that term. The name "inverted" reflects that instead of mapping documents to their terms (the natural way we think about a document), we map terms to their documents. When a query arrives, the system looks up query terms, retrieves candidate documents containing those terms, and scores them using BM25 or similar formulas. This is the same technology powering web search engines, refined over decades to handle billions of documents with sub-second query times. The inverted index dramatically reduces computation because most queries contain only a handful of terms, and each term appears in only a fraction of documents. If your query contains three terms and each term appears in 0.1 percent of a 1 million document corpus, you need to score at most 3,000 documents rather than 1 million. That is a 99.7 percent reduction in computation from a simple data structure.
For dense retrieval, vector indexes organize embeddings to enable efficient nearest-neighbor search. Comparing a query embedding against millions of document embeddings naively requires millions of distance calculations, each involving hundreds or thousands of floating-point operations. Specialized data structures like Hierarchical Navigable Small World graphs (HNSW) and Inverted File indexes (IVF) reduce this to thousands or hundreds of comparisons through clever approximations. HNSW builds a multi-layer graph where each node connects to its nearest neighbors, letting search move quickly toward the most relevant region of the embedding space. Think of it as a transit network: at the top layer you have a few high-speed connections that let you travel large distances quickly, while at the bottom layer you have local connections for fine-grained navigation. IVF partitions the embedding space into clusters (using k-means) and searches only the clusters closest to the query vector, reducing the number of direct comparisons dramatically. These approximations trade a small amount of accuracy for dramatic speed improvements, and in practice the quality loss is negligible while the speedup can be 100x or more. We'll cover these algorithms in detail in upcoming chapters on vector similarity search.
The index improves performance while encoding a particular view of the data. An inverted index assumes that matching documents to queries is primarily about term overlap. A vector index assumes that semantic similarity is the primary relevance signal. Hybrid indexes can combine both, maintaining both an inverted index for exact match and a vector index for semantic match. When you choose an index, you're making a statement about what "relevant" means in your application.
Retrieval at Query Time
When your query arrives, the retriever executes a specific sequence of operations. First, it encodes the query using the same method applied to documents. For sparse retrieval, this means computing BM25 term weights for the query terms. For dense retrieval, this means passing the query through the same neural encoder that embedded the documents. This symmetry is needed: queries and documents must inhabit the same representation space for similarity comparisons to be meaningful. If you encode documents with one model and queries with a different model, the resulting vectors occupy different spaces and comparisons will be meaningless.
Second, the retriever searches the index for the most similar documents. For sparse retrieval, this involves looking up query terms in the inverted index and combining the results using BM25 scoring. For dense retrieval, this involves traversing the vector index to find the nearest neighbors of the query embedding. In both cases, the index structure turns what would be an exhaustive search into a targeted lookup.
Third, the retriever returns those documents along with their similarity scores. These scores serve multiple purposes: they determine the ranking order, they can filter out low-quality matches, and they can inform the generator about the relative confidence of different sources. A document with a BM25 score of 8.2 and a document with a score of 0.3 should not be weighted equally when constructing the generator's prompt.
The parameter (typically 3 to 10 documents) balances coverage against noise. Too few documents risk missing relevant information; too many dilute the generator's context with marginally relevant passages. Research has consistently found that retrieving more documents helps up to a point, after which quality plateaus or even degrades as irrelevant content confuses the generator. The optimal value of depends on your use case: simple factual questions might need only one or two relevant passages, while complex analytical queries might benefit from synthesizing information across many sources.
# Minimal retriever demonstration using BM25
!uv pip install rank_bm25
from rank_bm25 import BM25Okapi
import numpy as np
# Sample knowledge base (in practice, thousands to millions of chunks)
documents = [
"The transformer architecture was introduced in the paper Attention Is All You Need in 2017.",
"BERT uses bidirectional self-attention to create contextualized word representations.",
"GPT models are decoder-only transformers trained with causal language modeling.",
"Retrieval-augmented generation combines neural retrievers with language model generators.",
"The attention mechanism allows models to focus on relevant parts of the input sequence.",
"Large language models are typically trained on trillions of tokens from web text.",
]
# Tokenize documents for BM25
tokenized_docs = [doc.lower().split() for doc in documents]
bm25 = BM25Okapi(tokenized_docs)
# Retrieve for a query
query = "How does the transformer architecture work?"
tokenized_query = query.lower().split()
scores = bm25.get_scores(tokenized_query)
# Get top-k documents
k = 3
top_indices = np.argsort(scores)[::-1][:k]Query: How does the transformer architecture work? Retrieved documents (ranked by BM25 score): 1. [Score: 3.072] The transformer architecture was introduced in the paper Attention Is All You Need in 2017. 2. [Score: 0.789] The attention mechanism allows models to focus on relevant parts of the input sequence. 3. [Score: 0.000] Large language models are typically trained on trillions of tokens from web text.

BM25 retrieves documents sharing query terms like "transformer" and "architecture." The scores reflect term overlap, frequency, and document length normalization as defined by the BM25 formula we covered in Part II. Notice that the highest-scoring document explicitly mentions both "transformer" and "architecture," while lower-ranked documents may match only one of these key terms or match through related vocabulary.
Query Encoding vs. Document Encoding
One subtlety that trips up many developers is that queries and documents are not always encoded identically, even when using the same encoder model. In asymmetric search, queries tend to be short, conversational questions while documents are longer, more formal passages. Some dense retrieval models are specifically trained with different prompting strategies for queries versus documents to bridge this distributional gap.
For example, the E5 family of embedding models prepends "query:" to queries and "passage:" to documents before encoding. The model has been trained to understand these prefixes as signals about the type of text being encoded, letting it to represent the query in a way that aligns well with relevant passages rather than with similar-looking queries. Models like BGE use instruction-tuning approaches where you can provide task-specific instructions like "Represent this sentence for searching relevant passages:". These seemingly minor formatting differences can substantially affect retrieval quality on asymmetric tasks, which is exactly the setup typical RAG systems use. When you use a pre-trained retrieval model, always read its documentation carefully to understand how it expects queries and documents to be formatted.
The Generator Component
The generator takes the retrieved documents and the original query, then produces a natural language response. In modern RAG systems, this is almost always a large language model, either an API-based model like GPT-4 or an open-weight model like LLaMA. The generator's role extends beyond simple information extraction: it must understand the query's intent, identify relevant portions of the retrieved documents, resolve any conflicts between sources, and synthesize a coherent, well-structured response. This is fundamentally a reading comprehension and synthesis task, and the generator's effectiveness depends heavily on both its underlying capabilities and how you present the context.
Think of the generator as a lawyer who has been handed a stack of case files and asked to write a brief. The lawyer doesn't just quote the files verbatim; they read critically, synthesize the relevant points, identify the strongest evidence, and construct a coherent argument. When sources conflict, they weigh the evidence. When sources are silent on a key point, a good lawyer acknowledges the gap rather than fabricating. The generator faces the same challenge: read the retrieved documents, identify what's relevant to the query, synthesize a response, and know when to say "I don't have enough information."
The quality of the generator's output depends on three interacting factors: the quality of the retrieved context (garbage in, garbage out), the quality of the prompt (which shapes what the model attends to and how it responds), and the intrinsic capability of the underlying model (which determines what complex synthesis and reasoning are possible). We can control all three, and improving any one of them generally improves the final answer.
Context Integration
The most straightforward approach to context integration is prompt concatenation: retrieved documents are inserted into the prompt before the query, giving the model access to relevant information through its standard attention mechanism. This approach requires no architectural modifications to the language model; we simply provide additional context that the model can reference during generation.
A typical RAG prompt structure looks like:
Context:
[Document 1]
[Document 2]
[Document 3]
Question: [User query]
Answer based on the context above:
The generator reads this concatenated input and generates a response, attending to both the query and the retrieved documents. This uses the in-context learning capabilities we discussed in Part XXVIII: the model learns to extract and synthesize information from the provided context without any weight updates. The model has been trained on countless examples of reading passages and answering questions, so it naturally applies these learned behaviors to the RAG setting. The point is that prompt concatenation turns the generation task into something the model already knows how to do well. You are not asking the model to learn a new skill; you are giving the raw material and asking it to apply an existing skill.
The context window is both the enabler and the constraint here. With a 128k or 200k token context window, you can retrieve many documents and present them all simultaneously. The model can attend to all of them in a single forward pass, drawing connections across documents that iterative approaches would miss. But larger contexts come with costs: latency increases because the model must process more tokens, and attention costs scale quadratically with sequence length (unless you use efficient attention variants). There is also evidence that models do not use all positions in a long context equally well, with positions in the middle of a very long context receiving systematically less attention than positions at the beginning and end.
Attention Over Retrieved Context
From the model's perspective, retrieved documents are simply additional tokens in the input sequence. There is no special mechanism distinguishing context from query; both are processed uniformly by the same transformer layers. The self-attention mechanism, which we covered extensively in Part XIII, allows every generated token to attend to every token in the context. This means the model can identify which parts of which documents are relevant to the query by computing high attention weights for informative passages and low weights for irrelevant ones. It can synthesize information across multiple documents, combining facts from different sources into a unified answer. When sources conflict or provide complementary information, the model can implicitly judge which source is more authoritative or relevant based on patterns learned during training.
The attention patterns often reveal which documents influenced the response. Tokens in the generated answer attend strongly to the specific passages they're drawing from. This creates an implicit citation mechanism. Researchers have exploited this property to build attribution systems that show which retrieved passages contributed to each part of the generated response. While not a formal guarantee of faithfulness, these attention patterns provide valuable interpretability. Tools like LLM-Grader and various attribution frameworks attempt to use these signals to produce document-level or span-level citations, helping users verify that claims in the generated answer are grounded in the retrieved context rather than hallucinated.
The generator's ability to use retrieved context depends on how well it was trained to do so. Base language models trained only on next-token prediction may not reliably distinguish context-based answers from parametric answers. Models fine-tuned specifically on reading comprehension tasks or instruction-following tend to ground their responses in the provided context more reliably. This is one reason why using a model fine-tuned for question answering or instruction following often improves RAG quality over using a raw language model, even if the base models have similar overall quality.
Prompt Engineering for RAG
The exact prompt template significantly affects generation quality. Small changes in wording can substantially alter the model's behavior, determining whether it faithfully uses the context, hallucinates confidently, or appropriately expresses uncertainty. Prompt engineering for RAG is a non-trivial skill, and the right template often depends on the specific model you're using.
Several considerations shape effective RAG prompts:
- Instruction clarity: Explicitly telling the model to base its answer on the provided context reduces hallucination. Phrases like "Answer based only on the documents above" or "If the information isn't in the context, say you don't know" provide clear behavioral guidance. Without such instructions, many models will blend retrieved context with their parametric knowledge, which can introduce outdated or incorrect information.
- Document ordering: Models sometimes exhibit position bias, attending more to documents at the beginning or end of the context. Important documents might be placed in these privileged positions, or the order might be randomized to reduce systematic bias. Some research suggests ordering by decreasing relevance score (most relevant first) helps, while other research suggests the opposite. The safest approach is to experiment on your specific domain and model.
- Attribution requests: Asking the model to cite which documents it used can improve traceability. Requests like "Reference the document numbers in your answer" encourage the model to explicitly connect claims to sources, which makes it easier to verify generated content.
- Uncertainty expression: Instructing the model to say "I don't know" when context is insufficient prevents confident hallucinations. Without this instruction, models often generate plausible-sounding answers even when they lack the information to do so correctly. An explicit instruction to express uncertainty shifts the model's behavior from confabulation toward appropriate epistemic humility.
- Format specification: Telling the model how to structure its response (bullet points, numbered steps, prose paragraphs) improves usability and consistency. Consistent output formats also make downstream parsing more reliable if your application processes the generated text programmatically.
def format_rag_prompt(query: str, documents: list[str]) -> str:
"""Format query and documents into a RAG prompt."""
context_parts = []
for i, doc in enumerate(documents, 1):
context_parts.append(f"[Document {i}]: {doc}")
context = "\n\n".join(context_parts)
prompt = f"""Use the following documents to answer the question. If the documents don't contain enough information to answer, say "I cannot answer this based on the provided documents."
{context}
Question: {query}
Answer:"""
return prompt
# Create RAG prompt with retrieved documents
retrieved_docs = [documents[idx] for idx in top_indices]
rag_prompt = format_rag_prompt(query, retrieved_docs)Use the following documents to answer the question. If the documents don't contain enough information to answer, say "I cannot answer this based on the provided documents." [Document 1]: The transformer architecture was introduced in the paper Attention Is All You Need in 2017. [Document 2]: The attention mechanism allows models to focus on relevant parts of the input sequence. [Document 3]: Large language models are typically trained on trillions of tokens from web text. Question: How does the transformer architecture work? Answer:
This prompt template makes the task explicit: use the documents, cite uncertainty when appropriate. The numbered document format enables the model to reference specific sources in its response, improving traceability and helping you verify the generated information. Notice how the instruction at the beginning explicitly constrains the model to the provided context, which is important for faithfulness.
System Prompt vs. User Prompt Placement
When using chat-based APIs (GPT-4, Claude, etc.), you have additional choices about where to place context. Retrieved documents can go in the system prompt (which persists across turns in a multi-turn conversation) or in the user message (which is specific to each turn). System prompt placement is useful when you have stable context documents that should inform all responses, like a company's product documentation. User message placement is better for dynamic context that changes per query, which is the typical RAG case. Some practitioners place general instructions in the system prompt ("You are a helpful assistant that answers questions based on provided context") and the actual retrieved documents in the user message alongside the query, striking a balance between reusability and freshness.
The placement decision also affects caching behavior. Some inference providers cache system prompts, so placing frequently-used context there can reduce latency and cost. Prefix caching allows the provider to reuse the KV-cache computation from a shared prefix, avoiding redundant processing. If your application has a fixed document set that many users query, system-prompt placement with prefix caching can significantly reduce latency.
Retrieval Timing
A important architectural decision is when retrieval occurs relative to generation. This timing choice affects latency, complexity, and the types of queries the system can handle effectively. Different timing strategies suit different use cases, and the choice determines the basic capabilities of the system rather than merely optimizing performance.
Think of retrieval timing as the difference between a student who reads all relevant materials before writing an essay (single-shot), one who writes sections and periodically stops to look up needed facts (iterative), and one who has a photographic memory aid that supplies relevant passages for each word they write (token-level). Each approach has different strengths: the first is simpler and faster, the second handles complex multi-part questions better, and the third maintains maximum relevance at maximum cost.
Single-Shot Retrieval
The simplest approach retrieves documents once, before generation begins. The query goes to the retriever, documents come back, and generation proceeds with that fixed context. The flow is linear: your query reaches the retriever, which searches the index and returns relevant documents, then those documents and the original query combine into a prompt that the generator processes to produce the final response.
Single-shot retrieval works well when:
- The query clearly specifies what information is needed
- Retrieved documents are likely to contain complete answers
- Low latency is necessary (one retrieval round-trip)
- The query does not require multi-hop reasoning across different topics
Most production RAG systems use single-shot retrieval because it handles straightforward queries quickly with a simple architecture. The design is easy to reason about and debug, which also makes optimization more direct. When something goes wrong, you can examine the retrieved documents to determine whether the problem was in retrieval (wrong documents retrieved) or in generation (right documents, wrong answer). This diagnostic clarity is a significant practical advantage.
The limitation of single-shot retrieval becomes apparent with complex queries. Consider "How did the policy changes in 2019 affect the 2020 research outcomes described in the Smith et al. study?" This query requires knowing what the 2019 policy changes were, what research outcomes the study described, and how the former affected the latter. A single retrieval might return some of the relevant documents, but might miss the policy context entirely, or return the study abstract without the relevant methodology section. Single-shot retrieval cannot adapt to what it finds.
Iterative Retrieval
Complex queries sometimes require multiple retrieval steps. The model might need to first retrieve background information, then use that to formulate more specific sub-queries. In iterative retrieval, the process begins with an initial retrieval based on your query. Partial generation may then reveal a gap in the available information, prompting the system to formulate a new, more specific query and retrieve again. The final generation step synthesizes all retrieved content into a complete response.
Consider the query: "Compare the economic impact of the 2008 and 2020 recessions." A single retrieval might not return sufficiently comparable information about both events. The query contains two distinct information needs, and the knowledge base might organize information about each recession separately. Iterative retrieval allows the system to retrieve documents about the 2008 recession first, then retrieve documents about the 2020 recession, and finally synthesize a comparison using both sets of documents.
The trade-off is latency and complexity. Each retrieval adds round-trip time and requires logic to determine when additional retrieval is needed. The system must decide how to decompose complex queries and when it has gathered sufficient information to generate a final answer. In production systems, iterative retrieval often involves a "planning" step where the model first identifies the sub-questions that need answering, then retrieves for each one, and finally synthesizes the collected information. This is closely related to agent-based systems where a language model orchestrates multiple tool calls, with the retriever being one of the available tools.
Iterative retrieval can also implement multi-hop reasoning, where answering one question provides context needed to answer the next. For example, to answer "Who founded the company that makes the most popular Python web framework?", a system might first retrieve "What is the most popular Python web framework?" (answer: Django), then retrieve "Who founded Django?", letting a two-hop answer that neither single query alone could provide. Multi-hop reasoning is particularly important for knowledge-intensive tasks where the answer requires combining information from multiple distinct sources.
Token-Level Retrieval
At the opposite extreme from single-shot, some architectures retrieve new documents for each token generated. The RETRO (Retrieval-Enhanced Transformer) architecture from DeepMind interleaves retrieval throughout generation, attending to freshly retrieved passages at regular intervals. This approach ensures that the retrieved context remains relevant even as generation shifts to new topics or sub-questions.
RETRO divides the input into chunks of a fixed size (typically 64 tokens) and retrieves nearest-neighbor passages from a massive database for each chunk. These retrieved passages are processed by a separate encoder and integrated into the main model through cross-attention layers. The effect is that every portion of the generation process has access to relevant passages from the retrieval database, dynamically updated as the generation proceeds. A passage relevant to the first part of a long answer might not be relevant to the third part, and RETRO can adapt by retrieving different passages for each segment.
This approach can maintain relevance as generation shifts topics, but the computational overhead is substantial. Each retrieval operation adds latency, and performing thousands of retrievals (one per chunk of generation) quickly becomes impractical for interactive applications. Token-level retrieval is primarily a research technique rather than a production pattern, though its insights have influenced more practical architectures. The key insight that retrieval should remain dynamically relevant throughout generation, rather than being fixed at query time, has influenced designs like kNN-LM and more recent memory-augmented approaches.

Query-Time vs. Index-Time Retrieval
An orthogonal timing dimension is when document processing occurs. This choice affects the trade-off between preparation time (when documents are added) and response time (when queries arrive).
Index-time processing pre-computes everything possible: chunking, embedding, and indexing happen once when documents are added to the knowledge base. Query-time only computes query encoding and index lookup. This approach minimizes latency for users but requires re-processing documents whenever the chunking strategy or embedding model changes. If you want to upgrade from a 384-dimensional embedding model to a better 1536-dimensional model, you need to re-embed every document in your knowledge base, which can take hours or days for large corpora.
Query-time processing delays some computation until a query arrives. For example, the system might store raw documents and compute embeddings using a query-dependent prompt that incorporates information about the specific query or user context. This enables more sophisticated relevance scoring at the cost of latency. Some hybrid approaches pre-compute base embeddings at index time but compute additional query-specific features at query time, combining the efficiency of pre-computation with the flexibility of dynamic scoring.
Most systems favor aggressive index-time processing to minimize query latency, accepting the overhead of re-indexing when configurations change. This matches the typical read/write pattern: documents are added infrequently, but queries arrive continuously. Optimizing for query latency by accepting higher index-building time is usually the right trade-off.
Architecture Variations
Since the original RAG paper, researchers and practitioners have developed numerous architectural variations. Understanding these helps you choose or design the right architecture for your use case. Simpler designs may sacrifice performance or capabilities, while more capable systems cost more to build and operate. The variations are not mutually exclusive; production systems often combine elements from multiple approaches.
Retrieve-then-Read (Sequential RAG)
The standard architecture we've been describing is sometimes called "retrieve-then-read": retrieve first, then read (generate). Documents pass through a clean interface: the retriever's output becomes the generator's input. This separation creates a clear contract between components: the retriever promises to return relevant documents, and the generator promises to synthesize them into an answer.
class SequentialRAG:
"""Basic retrieve-then-read RAG architecture."""
def __init__(self, retriever, generator, k=3):
self.retriever = retriever
self.generator = generator
self.k = k
def answer(self, query: str) -> str:
# Step 1: Retrieve
documents = self.retriever.search(query, k=self.k)
# Step 2: Generate
prompt = self.format_prompt(query, documents)
response = self.generator.generate(prompt)
return response
def format_prompt(self, query: str, documents: list[str]) -> str:
context = "\n\n".join(documents)
return f"Context:\n{context}\n\nQuestion: {query}\n\nAnswer:"This separation enables independent optimization of each component. You can upgrade the retriever without touching generation logic, or swap in a different LLM without modifying retrieval. This modularity also simplifies debugging: if answers are wrong, you can examine retrieved documents to determine whether the problem is in retrieval (wrong documents) or generation (right documents, wrong answer). The sequential architecture is the right starting point for most projects. Build it first, measure where quality falls short, and then consider more complex architectures only if the sequential version cannot meet your requirements.
Fusion-in-Decoder
The Fusion-in-Decoder (FiD) architecture processes each retrieved document independently through the encoder, then fuses their representations in the decoder. This approach addresses a scalability challenge with the standard concatenation approach.
Instead of concatenating all documents into a single input, which creates one long sequence that grows linearly with the number of documents, FiD keeps encoding costs constant per document by encoding each document-query pair separately. Each encoding produces a fixed-size representation, and these representations are concatenated for the decoder. The decoder then attends over all document representations simultaneously when generating the answer.
This approach scales better with the number of retrieved documents. Standard concatenation creates a sequence of length where is average document length and is the number of documents, and attention cost grows quadratically with sequence length. FiD keeps each encoder pass at length , with only the decoder needing to handle the combined representations. The decoder's cross-attention mechanism attends over the concatenated encoder outputs, effectively fusing information from all documents. This allows FiD to process many more documents than would be practical with simple concatenation, often improving quality on tasks that require synthesizing information from many sources.
The original FiD paper demonstrated state-of-the-art performance on open-domain question answering benchmarks by retrieving and fusing up to 100 documents per query. At that scale, concatenation would create sequences of tens of thousands of tokens, making generation prohibitively expensive, while FiD kept costs manageable by encoding documents independently. The trade-off is that FiD requires an encoder-decoder architecture (like T5), which makes it less applicable to pure decoder models like LLaMA or GPT, which are increasingly common in modern RAG systems.
Retrieval-Augmented Language Models (REALM and RAG)
The original RAG paper from Facebook AI and the REALM paper from Google both proposed end-to-end trainable retrieval. Rather than using a frozen retriever, these architectures backpropagate gradients through the retrieval step. This means the retriever can learn which documents are similar to the query and which documents are useful for generating correct answers.
The key insight is treating retrieval as a latent variable. In this formulation, we don't commit to using a single retrieved document. Instead, we consider multiple retrieved documents and weight their contributions by how relevant the retriever believes them to be. Given query and a set of retrieved documents , the probability of generating answer is computed by marginalizing over the documents:
where:
- : the generated answer we want to produce
- : the input query from the user
- : the -th retrieved document from the knowledge base
- : the total number of retrieved documents under consideration
- : the probability of the generator creating answer given query and document , measuring how well this specific document helps generate the correct answer
- : the probability that the retriever assigns to document being relevant to query , derived from the similarity score between the query embedding and the document embedding
This formula computes a weighted mixture of generation probabilities, where each document's contribution to the final probability is proportional to how relevant the retriever believes it to be. Think of it as holding a committee vote: each document votes on the answer, and its vote carries weight proportional to its relevance score. If document is highly relevant ( is large) and strongly supports the correct answer ( is large), it dominates the sum. Documents that are irrelevant or that point toward wrong answers receive low weights and contribute little.
The core advantage of this formulation is differentiability. Because both terms in the product are differentiable functions of model parameters, gradients flow from the generation loss back through the generation probability, through the retriever weights, to the document embeddings. If the generator consistently produces wrong answers when using certain documents, the retriever learns to down-weight those documents for similar queries. The retriever and generator co-adapt: the retriever learns to retrieve what helps the generator, and the generator learns to use whatever the retriever provides.
Training end-to-end is expensive because updating document embeddings requires re-indexing, which can take hours for large knowledge bases. This creates a practical challenge: if you update the document encoder at every gradient step, you need to re-embed and re-index all documents continuously, which is impractical. The original RAG paper uses asynchronous index updates, refreshing document embeddings periodically (e.g., every few thousand gradient steps) rather than after every update. However, the resulting retrievers are specifically optimized for the generation task, often outperforming generic retrievers by substantial margins on knowledge-intensive benchmarks. The trade-off is infrastructure complexity and training cost versus retrieval quality.
RETRO: Chunked Cross-Attention
The RETRO (Retrieval-Enhanced Transformer) architecture from DeepMind takes a different approach, building retrieval into the transformer architecture itself. Rather than treating retrieval as a preprocessing step, RETRO makes retrieval an integral part of the model's forward pass.
RETRO divides the input into chunks and retrieves neighbors for each chunk. These neighbors are documents from the knowledge base that are similar to the current chunk being processed. Retrieved neighbors are processed by a separate encoder, and their representations are integrated through chunked cross-attention layers interleaved with standard self-attention. This means the model alternates between attending to its own representations (self-attention over the input and generated text) and attending to retrieved content (cross-attention over the encoder outputs for retrieved passages). Every chunk of generated text has access to the most relevant passages from the retrieval database, dynamically updated as generation proceeds.
This tight integration allows RETRO to scale retrieval to trillions of tokens while keeping the main model relatively small. The model can effectively "look up" information during generation rather than storing all knowledge in its parameters. A 7B parameter RETRO model can match the performance of a 25B model without retrieval by using a massive retrieval database. This represents a major change in how we think about model capacity: some knowledge lives in parameters (where it's always available but may be outdated), while other knowledge lives in an external database (where it can be updated without retraining) accessed through retrieval. The implication is that the right strategy for scaling might be to scale the retrieval database rather than the model parameters, at least for knowledge-intensive tasks.
Self-RAG: Retrieval as a Learned Decision
Most RAG systems retrieve for every query, but retrieval isn't always necessary or beneficial. Simple queries like "What is 2+2?" don't benefit from retrieval, and retrieving irrelevant documents can hurt performance by introducing noise and confusing the generator. Self-RAG trains models to decide when to retrieve, what to retrieve, and how to use retrieved content, adapting retrieval behavior to the specific query at hand.
The Self-RAG model generates special control tokens during its forward pass, inserting them at appropriate points to indicate its decisions about retrieval and document relevance:
- [Retrieve]: Signals whether retrieval is needed for this query or generation segment
- [IsREL]: Assesses whether a retrieved passage is relevant to the current generation context
- [IsSUP]: Evaluates whether a passage supports the claim being generated
- [IsUSE]: Determines whether to use a passage in the final generated response
These reflection tokens enable the model to critique its own retrieval usage in real time, filtering irrelevant documents and identifying when it should rely on parametric knowledge instead of retrieved context. The model has been trained through a combination of supervised learning (to generate appropriate reflection tokens) and reinforcement learning (to optimize the quality of final answers), so its retrieval decisions are grounded in observed answer quality.
Self-RAG is more adaptive than systems that blindly retrieve for every query. It achieves competitive quality with selective retrieval, reducing both latency and cost for queries that don't need external information. The training approach also produces models that express more calibrated uncertainty: they retrieve more when they're less sure about the answer and rely on parametric knowledge when they're confident, rather than always deferring to the retriever. This behavior is closer to how a knowledgeable human expert would work: drawing on memory for familiar facts while looking things up for less familiar topics.
Hybrid Retrieval
Rather than choosing between sparse and dense retrieval, many production systems use both and combine their results. The intuition is simple: sparse retrieval is excellent at finding documents with exact term matches (useful when queries use precise technical terminology), while dense retrieval is excellent at finding semantically related documents (useful when queries use everyday language about technical topics). Combining both captures the strengths of each.
The most common combination strategy is Reciprocal Rank Fusion (RRF). Given two ranked lists of documents, one from sparse retrieval and one from dense retrieval, RRF assigns each document a score based on its rank in each list:
where:
- : the document being scored
- : the set of retrieval methods (e.g., sparse and dense)
- : the rank of document in the results from retrieval method (lower rank means more relevant)
- : a constant (typically 60) that dampens the influence of very high-ranked documents and prevents any single result from dominating
The key insight behind RRF is that rank is more reliable than score. BM25 scores and cosine similarities live on different scales and cannot be directly compared, but if a document ranks first in both the sparse and dense results, that convergent signal strongly suggests relevance. The constant was found empirically to work well across many retrieval tasks by reducing the impact of randomly high rankings. Documents that appear in the top results of both systems receive the highest RRF scores, with the score declining as rank falls.
RRF is simple to implement, requires no model training, and consistently outperforms either retriever alone on most benchmarks. More sophisticated fusion approaches learn weights for each retrieval signal, but RRF is an excellent default that works well across diverse domains.
Worked Example: Tracing a Query Through the Full Pipeline
Let's trace a specific query through every stage of a sequential RAG system to make the abstract components concrete. We'll use the query: "What is the key innovation that allows transformers to train faster than RNNs?"
Stage 1: Query Encoding (Sparse). The query is tokenized and lowercased: ["what", "is", "the", "key", "innovation", "that", "allows", "transformers", "to", "train", "faster", "than", "rnns"]. Stop words ("what", "is", "the", "that", "to", "than") are typically removed for BM25: ["key", "innovation", "allows", "transformers", "train", "faster", "rnns"]. BM25 term weights are computed for each remaining term based on their IDF values across the corpus.
Stage 2: Index Lookup. The retriever looks up each query term in the inverted index. "transformers" might appear in 1,247 documents, "train" in 3,821 documents, "faster" in 892 documents, and "innovation" in 654 documents. The inverted index returns candidate documents for each term, and BM25 scores are computed for each candidate based on the formula we studied in Part II. For our knowledge base, the highest-scoring document contains the sentence "The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training."
Stage 3: Top-k Selection. The retriever ranks all candidate documents by BM25 score and returns the top . The returned documents discuss transformers, self-attention, and the comparison with RNNs, all directly relevant to the query.
Stage 4: Context Formatting. The three retrieved documents are formatted into numbered context blocks with the query appended, following the template shown earlier. The full prompt is approximately 350 tokens.
Stage 5: Generation. The language model receives the 350-token prompt and generates a response attending to all tokens in the context. The model identifies the passage "replacing recurrence with self-attention" and "process all positions in parallel during training" as the key information. The generated response synthesizes this into a coherent answer: "The key innovation letting transformers to train faster than RNNs is the replacement of sequential recurrence with parallel self-attention. RNNs must process tokens one at a time, preventing parallelization across sequence positions. Transformers compute attention over all positions simultaneously, letting full GPU utilization during training."
Stage 6: Verification. In a production system, you would compare the generated claims against the retrieved documents to verify grounding. The claim about "parallel self-attention" is directly supported by Document 6 in the knowledge base. The claim about "sequential recurrence" preventing parallelization is a reasonable inference from the same passage.
This trace reveals why each component matters. If the retriever had returned documents about cooking instead of transformers, the generator's answer would have been wrong no matter how capable it was. If the generator had ignored the context and hallucinated, it might have said something plausible but incorrect. The final answer quality depends on both components working correctly in sequence.
Building a Complete RAG Pipeline
Let's assemble the components into a working RAG system. We'll use a sparse retriever for simplicity; the next chapter covers dense retrieval in depth. This implementation demonstrates the core patterns you'll use in production systems, even though real systems would connect to actual language model APIs.
from dataclasses import dataclass
import numpy as np
from rank_bm25 import BM25Okapi
@dataclass
class RetrievedDocument:
"""A document with its retrieval score."""
content: str
score: float
doc_id: int
class BM25Retriever:
"""BM25-based sparse retriever."""
def __init__(self, documents: list[str]):
self.documents = documents
self.tokenized_docs = [doc.lower().split() for doc in documents]
self.bm25 = BM25Okapi(self.tokenized_docs)
def retrieve(self, query: str, k: int = 3) -> list[RetrievedDocument]:
"""Retrieve top-k documents for a query."""
tokenized_query = query.lower().split()
scores = self.bm25.get_scores(tokenized_query)
top_indices = np.argsort(scores)[::-1][:k]
return [
RetrievedDocument(
content=self.documents[idx], score=scores[idx], doc_id=idx
)
for idx in top_indices
]class SimpleGenerator:
"""Placeholder generator (in practice, use an LLM API or local model)."""
def __init__(self):
pass
def generate(self, prompt: str, max_tokens: int = 256) -> str:
"""Generate a response given a prompt.
This is a placeholder. In practice, you would:
- Call an API (OpenAI, Anthropic, etc.)
- Run a local model (llama.cpp, vLLM, etc.)
"""
return "[Generated response based on retrieved context]"class RAGPipeline:
"""Complete RAG pipeline combining retriever and generator."""
def __init__(
self,
retriever: BM25Retriever,
generator: SimpleGenerator,
k: int = 3,
include_scores: bool = False,
):
self.retriever = retriever
self.generator = generator
self.k = k
self.include_scores = include_scores
def format_context(self, docs: list[RetrievedDocument]) -> str:
"""Format retrieved documents into context string."""
parts = []
for i, doc in enumerate(docs, 1):
if self.include_scores:
parts.append(f"[{i}] (score: {doc.score:.2f}) {doc.content}")
else:
parts.append(f"[{i}] {doc.content}")
return "\n\n".join(parts)
def build_prompt(self, query: str, context: str) -> str:
"""Construct the full prompt for the generator."""
return f"""Answer the question based only on the following context. If the context doesn't contain the answer, say "I don't have enough information to answer this question."
Context:
{context}
Question: {query}
Answer:"""
def answer(self, query: str) -> dict:
"""Run the full RAG pipeline."""
retrieved_docs = self.retriever.retrieve(query, k=self.k)
context = self.format_context(retrieved_docs)
prompt = self.build_prompt(query, context)
response = self.generator.generate(prompt)
return {
"query": query,
"retrieved_documents": retrieved_docs,
"prompt": prompt,
"response": response,
}# Build a knowledge base about ML concepts
knowledge_base = [
"Transformers use self-attention mechanisms that allow each token to attend to all other tokens in the sequence. This enables capturing long-range dependencies without the sequential processing limitations of RNNs.",
"The attention mechanism computes compatibility scores between queries and keys, then uses these scores to create weighted sums of values. This can be expressed as Attention(Q,K,V) = softmax(QK^T / sqrt(d_k))V.",
"BERT (Bidirectional Encoder Representations from Transformers) uses masked language modeling for pre-training, randomly masking 15% of input tokens and training the model to predict them.",
"GPT models are autoregressive, generating one token at a time while attending only to previously generated tokens. This causal attention prevents information leakage from future tokens.",
"Retrieval-augmented generation grounds language models in external knowledge by retrieving relevant documents before generation. This reduces hallucination and enables access to up-to-date information.",
"The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training.",
"Multi-head attention runs several attention operations in parallel, letting the model to capture different types of relationships at different positions.",
"Position encodings are necessary in transformers because self-attention is permutation-invariant. Sinusoidal encodings and learned position embeddings are common approaches.",
]
retriever = BM25Retriever(knowledge_base)
generator = SimpleGenerator()
rag = RAGPipeline(retriever, generator, k=3, include_scores=True)
result = rag.answer("What is the key innovation of transformers?")Query: What is the key innovation of transformers? ============================================================ Retrieved Documents: ============================================================ [Doc 5] Score: 5.025 The key innovation of transformers is replacing recurrence with self-attention. While RNNs process s... [Doc 7] Score: 1.061 Position encodings are necessary in transformers because self-attention is permutation-invariant. Si... [Doc 0] Score: 0.820 Transformers use self-attention mechanisms that allow each token to attend to all other tokens in th... ============================================================ Constructed Prompt: ============================================================ Answer the question based only on the following context. If the context doesn't contain the answer, say "I don't have enough information to answer this question." Context: [1] (score: 5.02) The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training. [2] (score: 1.06) Position encodings are necessary in transformers because self-attention is permutation-invariant. Sinusoidal encodings and learned position embeddings are common approaches. [3] (score: 0.82) Transformers use self-attention mechanisms that allow each token to attend to all other tokens in the sequence. This enables capturing long-range dependencies without the sequential processing limitations of RNNs. Question: What is the key innovation of transformers? Answer:
The pipeline retrieves the three most relevant documents (those mentioning "transformers," "attention," and related concepts) and then constructs a prompt instructing the model to answer based on this context. Notice how the BM25 scores reflect the degree of term overlap between the query and each document, with the highest-scoring document containing the exact phrase "key innovation of transformers."
Handling Edge Cases
Production RAG systems need to handle several edge cases gracefully. Not every query will have good matches in the knowledge base, and the system should behave appropriately in these situations rather than hallucinating or returning nonsensical results.
def check_retrieval_quality(
docs: list[RetrievedDocument], score_threshold: float = 1.0
) -> dict:
"""Assess whether retrieved documents are likely relevant."""
if not docs:
return {"status": "no_documents", "action": "fallback_to_parametric"}
max_score = max(doc.score for doc in docs)
avg_score = sum(doc.score for doc in docs) / len(docs)
if max_score < score_threshold:
return {
"status": "low_confidence",
"max_score": max_score,
"action": "warn_user_or_expand_search",
}
return {
"status": "good",
"max_score": max_score,
"avg_score": avg_score,
"action": "proceed_with_generation",
}
# Test with a query that has good matches
good_docs = retriever.retrieve("How does attention work in transformers?", k=3)
quality_good = check_retrieval_quality(good_docs)
# Test with a query outside the knowledge base
poor_docs = retriever.retrieve("What is the capital of France?", k=3)
quality_poor = check_retrieval_quality(poor_docs)Query with good matches: Max score: 1.078 Status: good Action: proceed_with_generation Query outside knowledge base: Max score: 1.722 Status: good Action: proceed_with_generation
When retrieval scores are low, the system might fall back to the model's parametric knowledge, return "I don't know" rather than hallucinating, or trigger a different retrieval strategy (expand query, search different index). The appropriate action depends on your use case. A customer support bot might need to always provide some answer, while a medical information system must be conservative and admit uncertainty. The threshold used in the quality check function should be tuned empirically on representative queries from your target domain, since BM25 scores are not standardized across corpora.
Key Parameters
Understanding the key parameters in a RAG pipeline allows you to tune the system for your specific requirements. These parameters interact with each other, and changing one often requires re-tuning others.
The most important parameters are:
- k: The number of documents to retrieve. Balances context coverage against noise. Typical values range from 3 to 10, with lower values for focused queries and higher values for broad research questions. Larger increases latency both from the retrieval step (more results to rank) and the generation step (longer context to process). Experiment with values from 1 to 20 and measure answer quality to find the sweet spot for your domain.
- score_threshold: The minimum similarity score required to consider a document relevant. This threshold depends heavily on the retrieval method and corpus, and should be tuned on representative queries. A threshold calibrated on one corpus may not transfer to another.
- max_tokens: The maximum number of tokens to generate in the response. Longer limits allow more complete answers but increase latency and cost. Consider the typical answer length for your domain when setting this parameter.
- chunk_size: The size of document chunks in tokens. Smaller chunks (128-256 tokens) are better for precise factual retrieval; larger chunks (512-1024 tokens) preserve more context and work better for questions requiring understanding of extended passages.
- chunk_overlap: The number of tokens that adjacent chunks share. Overlap of 10-20 percent reduces the chance of splitting relevant content across chunk boundaries at the cost of some index redundancy.
Data Flow Visualization
The following diagram illustrates how data flows through a RAG system:

The numbers indicate the sequence: (1) the query goes to the retriever, (2) relevant documents return from the knowledge base, (3) the query and retrieved context combine into a prompt, and (4) the generator produces the final response.
Comparing Architecture Variants
Different RAG architectures balance response quality against latency, and their implementation costs also vary. Understanding these trade-offs helps you choose the right approach for your application's requirements.

Sequential RAG dominates in production because it's fast and simple while delivering good quality. Fusion-in-Decoder trades some speed for better multi-document reasoning. Iterative retrieval handles complex queries but adds latency. End-to-end training maximizes quality but requires significant infrastructure investment. The right choice for your application depends on your latency budget, quality requirements, and engineering resources. Start with sequential RAG, measure its performance on your specific queries, and consider more complex architectures only if sequential RAG demonstrably falls short.
Limitations and Considerations
RAG architecture decisions involve inherent trade-offs that you should understand before building systems. These limitations don't make RAG ineffective; rather, they define the boundaries within which RAG excels and help you set appropriate expectations.
Context window constraints remain a basic limitation. Even with modern models supporting 100K or more token contexts, there's a practical limit to how many documents you can retrieve. As we discussed in Part XVII on efficient attention and Part XVIII on context extension, attention costs grow with sequence length, and models can struggle to effectively use information buried deep in long contexts. Research has shown that models tend to focus on information at the beginning and end of the context, with middle sections receiving less attention. This "lost in the middle" phenomenon means retrieval quality matters enormously: retrieving the right 3 documents beats retrieving 30 marginally relevant ones. You can mitigate this by placing the most relevant document first or last in the context, by reranking retrieved documents to prioritize the highest-quality ones, or by using architectures like FiD that avoid concatenation entirely.
The retriever-generator mismatch problem arises because these components are typically trained separately. A retriever optimized for traditional information retrieval metrics may not retrieve what is most useful for generation. A passage that scores highest on BM25 might lack the specific details the generator needs, while a lower-ranked passage contains exactly the right information. The retriever knows about word overlap and semantic similarity, but it doesn't know what the generator needs to answer the specific question it's facing. End-to-end training addresses this mismatch directly, but requires substantial infrastructure. Reranking approaches use a cross-encoder (which jointly encodes query and document) to re-score retrieved documents after initial retrieval, capturing relevance signals that the bi-encoder retriever misses. We'll cover cross-encoders and reranking in later chapters.
Latency accumulates across the pipeline in compound ways. Retrieval adds round-trip time to the knowledge base (potentially remote), and longer contexts increase generation time because the attention mechanism must process more tokens. For real-time applications, this motivates careful optimization: caching frequent queries, using smaller models, or pre-computing responses for common question patterns. Each additional retrieved document adds both retrieval time (more results to rank and return) and generation time (more tokens for the model to attend over), creating compound latency effects. The relationship between and total latency is often superlinear in practice, because vector index search time grows with and generation cost grows with total context length.
Attribution and faithfulness remain active research areas that no current approach fully solves. When a RAG system generates an answer, how do you verify it came from the retrieved documents rather than from hallucinated parametric knowledge? The model might confidently state a fact that appears nowhere in the retrieved context, drawing instead from knowledge encoded in its weights during pretraining. This is particularly concerning for knowledge that changes over time: a model trained before a major news event might ignore retrieved context about that event and generate answers based on outdated weights. Current approaches include asking models to cite sources, comparing responses with and without context, or using specialized fact-verification models. None are fully reliable, making RAG a tool that reduces rather than eliminates hallucination risk. You should understand that retrieved documents provide evidence for the answer, but the system may still make errors in how it interprets or synthesizes that evidence.
Knowledge base maintenance introduces operational challenges that compound over time. Documents in your knowledge base become outdated as the world changes. Adding new documents requires re-indexing. Removing documents to comply with deletion requests requires removing them from the document store and removing their embeddings from the vector index, which some index types handle poorly. The knowledge base is a living system that requires ongoing maintenance, not a static artifact. In production, you need processes for adding new content, removing stale or incorrect content, and periodically re-evaluating whether your chunking and encoding strategies still serve your users well.
Finally, evaluation is harder than it appears. Measuring RAG quality requires both component-level evaluation (is the retriever finding the right documents?) and end-to-end evaluation (is the final answer correct?). Component metrics like recall at and mean reciprocal rank tell you about retrieval quality, but a high recall doesn't guarantee good final answers if the generator fails to synthesize correctly. End-to-end metrics like exact match, F1, or human evaluation tell you about overall quality but don't pinpoint where the system is failing. Building a complete evaluation harness that covers both components, with test cases representing the diversity of queries your users will send, is one of the most important investments you can make in a production RAG system.
Summary
This chapter examined the architecture of retrieval-augmented generation systems, revealing how the seemingly simple "retrieve then generate" pipeline conceals important design decisions at every stage.
The retriever component turns documents into searchable representations (sparse or dense), stores them in specialized indexes, and returns the most relevant passages at query time. The choice of retrieval method significantly affects what information reaches the generator, with sparse methods excelling at exact term matching and dense methods excelling at semantic similarity. Hybrid retrieval combining both approaches often outperforms either alone.
The generator component synthesizes answers from retrieved context using in-context learning. Prompt design matters: clear instructions about grounding responses in context, appropriate document formatting, and explicit uncertainty handling all improve output quality. The generator sees retrieved documents as additional tokens in the prompt and attends to them through the same self-attention mechanism used everywhere in the transformer.
Retrieval timing ranges from single-shot (fast, simple, suitable for most queries) to iterative (handles complex multi-step reasoning) to token-level as in RETRO (maximum relevance at maximum cost). Most production systems use single-shot retrieval with optional iteration for hard queries.
Architecture variations like Fusion-in-Decoder, end-to-end RAG/REALM, RETRO, and Self-RAG demonstrate that the basic sequential pattern is just one point in a large design space. FiD processes documents independently to scale retrieval count. End-to-end training aligns the retriever directly with generation quality. Self-RAG makes retrieval an adaptive, query-dependent decision rather than an unconditional operation.
The limitations are real: context window constraints, retriever-generator mismatch, latency accumulation, and attribution challenges each bound what RAG can reliably achieve. Understanding these boundaries helps you build systems with appropriate safeguards and realistic expectations.
The next chapter dives deep into dense retrieval: how neural networks learn to embed queries and documents into the same vector space, letting semantic matching that goes far beyond keyword overlap. This is the technology that makes modern RAG systems substantially more capable than their sparse retrieval predecessors, and understanding it is needed for building production-quality systems.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about RAG architecture.
RAG Architecture Fundamentals
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!