RAG Architecture: Components, Timing & Design Patterns

Michael BrenndoerferJanuary 21, 202660 min read

Part of Language AI Handbook

Covers RAG system design by exploring retriever-generator interactions, timing strategies like iterative retrieval, and architectural variations like RETRO.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

RAG Architecture and Design Patterns

The previous chapter established why retrieval-augmented generation matters: it grounds language models in external knowledge, reducing hallucinations and letting access to information beyond the training cutoff. But understanding the motivation is only half the story. To build effective RAG systems, you need to understand how the components fit together and the design decisions that shape system behavior. A RAG system that retrieves poorly will produce wrong answers even if the generator is perfect. Conversely, a weak generator will fail to synthesize correct answers even when the retriever finds exactly the right documents. The interaction between these components, and the way you orchestrate that interaction, determines whether your system succeeds or fails in practice.

Think of a RAG system as a research assistant who has access to a library. When you ask a question, the assistant first searches the library's catalog and retrieves the most relevant books and articles. Then, with those materials spread on the desk, the assistant reads through them and composes an answer. The quality of the assistant's response depends on two things: how good the librarian is at finding relevant materials (retrieval), and how skilled the assistant is at reading comprehension and synthesis (generation). A brilliant researcher who retrieves the wrong books will give you a wrong answer. An excellent librarian who supplies perfect sources to an incompetent reader will also fail. Both components must work, and they must work together.

A RAG system consists of two primary components: a retriever that finds relevant documents from a knowledge base, and a generator (typically a large language model) that synthesizes an answer from those documents. These components interact through a carefully designed interface, and the timing and manner of their interaction significantly affects system performance. The retriever acts as a search engine, filtering millions of documents to identify the few most relevant ones. The generator then reads these documents and synthesizes a response that addresses your question. But this description understates the complexity: every decision about how to encode documents, what index structure to use, how many documents to retrieve, how to format the context, and when retrieval occurs can measurably shift final system quality.

This chapter examines each component in detail, explores when retrieval occurs relative to generation, and surveys the major architectural variations that have emerged since the original RAG paper. We also build a complete working pipeline so you can see these concepts in code. By the end, you'll understand the design space well enough to make informed decisions when building your own systems, and you'll know which trade-offs matter most for different use cases.

Historical Context

The idea of augmenting language models with retrieved context predates the transformer era. Early open-domain question answering systems like DrQA (2017) retrieved Wikipedia paragraphs and fed them to a reading comprehension model. The key insight was that retrieval could replace some of the knowledge stored in model parameters. The REALM paper from Google (2020) made retrieval differentiable and learned the retriever end-to-end. The original RAG paper from Facebook AI Research (2020) by Lewis et al. popularized the term "retrieval-augmented generation" and demonstrated that marginalizing over retrieved documents during training could significantly improve performance on knowledge-intensive tasks. Since then, the field has exploded: systems like Atlas, Toolformer, RETRO, and Self-RAG have each pushed different aspects of the retrieval-generation interface, and the architecture has become a standard component in production AI systems worldwide.

The Retriever Component

The retriever's job is deceptively simple: given a query, return the documents most likely to contain relevant information. However, this process is complex. The retriever must somehow measure the relevance between your question, expressed in natural language with all its ambiguity and variation, and documents that may discuss the same concepts using entirely different vocabulary. Someone asking about "automobile reliability" and a document discussing "car dependability" are clearly related, but a keyword matcher would miss the connection entirely. In practice, achieving effective retrieval involves three sub-systems working together: document encoding, indexing, and query-time retrieval. Each sub-system has its own design decisions, and getting them right requires understanding both the algorithmic options and the practical constraints of your deployment environment.

The retriever is often the weakest link in RAG systems. Language model generators have improved dramatically over the past few years, but retrieval remains a hard problem. The challenge is that relevance is inherently query-dependent: the same document might be highly relevant for one question and irrelevant for another. A document explaining how attention mechanisms work in transformers is invaluable for someone asking "how does self-attention work?" but useless for someone asking "what is the capital of France?" The retriever must make relevance judgments without knowing exactly what the generator needs, and it must make those judgments quickly enough to support interactive applications.

Think of the retriever as a sommelier at a restaurant. When you describe what you'd like to drink, the sommelier doesn't pour every bottle in the cellar, they draw on their knowledge of the wine list to select a few bottles they believe will satisfy you. Their recommendations depend on how well they understood your description, how well they know the cellar, and how good their model of your preferences is. A great sommelier finds exactly what you wanted. A mediocre one gives you something in the right category but misses the details. The retriever faces an analogous challenge: it must interpret your query and search a knowledge base to find the items most likely to satisfy the generator.

Document Encoding

Before retrieval can happen, documents must be transformed into a searchable representation. Raw text, while meaningful to humans, cannot be efficiently compared or searched by algorithms. The encoding process bridges this gap, converting human-readable text into mathematical objects that computers can manipulate. This process typically involves two stages that work together to prepare documents for rapid, accurate retrieval.

First, documents are split into smaller chunks. A 50-page technical report cannot be fed to a language model in its entirety, so we divide it into passages of a few hundred tokens each. The chunking strategy significantly affects retrieval quality: too small and you lose context, too large and you dilute relevant information with noise. Consider a chunk that contains only half a sentence: the retriever might match it to a query based on keyword overlap, but the generator would receive incomplete, potentially misleading information. Conversely, a chunk spanning multiple pages might contain one relevant paragraph buried among irrelevant content, which makes it harder for the generator to identify what matters. A common heuristic is 256 to 512 tokens per chunk with 10 to 20 percent overlap between adjacent chunks, but the right size depends on your documents and queries. Technical documentation with dense, self-contained paragraphs might warrant smaller chunks, while narrative text with ideas that span multiple paragraphs might need larger ones. We'll explore chunking strategies in depth in a later chapter, examining techniques like overlapping windows, semantic boundary detection, and hierarchical chunking.

Second, each chunk is converted into a numerical representation suitable for similarity search. This transformation is the heart of the retrieval process, determining how the system measures whether one piece of text relates to another. Two broad approaches dominate the field, each with distinct characteristics that make them suited to different scenarios.

Sparse representations like TF-IDF and BM25 represent documents as high-dimensional vectors where each dimension corresponds to a vocabulary term. As we covered in Part II, BM25 scores documents based on term frequency, document length normalization, and inverse document frequency. The intuition behind sparse representations is straightforward: documents are characterized by the words they contain. A document about neural networks will have non-zero values for terms like "neuron," "layer," and "activation," while a document about cooking will have non-zero values for entirely different terms. These representations are sparse because most terms don't appear in any given document, leaving most dimensions at zero. A vocabulary might contain 100,000 terms, but any single document uses only a few hundred, resulting in vectors that are more than 99 percent zeros. The sparsity is not a weakness, it enables the highly efficient inverted index data structure that makes BM25 retrieval lightning fast. Sparse methods are interpretable (you can see exactly which terms caused a document to rank highly), require no training data, and can be deployed with minimal infrastructure. They remain competitive baselines and are often used in hybrid systems.

Dense representations use neural networks to encode chunks into continuous vectors, typically with 384 to 4096 dimensions. Unlike sparse vectors where dimensions correspond to specific words, dense embedding dimensions capture abstract semantic features that emerge during training. These features don't have human-interpretable names; instead, they represent learned patterns that help distinguish relevant from irrelevant content. Two documents discussing the same concept with different vocabulary will have similar dense embeddings even if their sparse representations share few non-zero dimensions. For example, a query about "car maintenance" would find documents discussing "vehicle repair" or "automobile servicing" because the neural network has learned that these phrases occupy similar regions of the embedding space. This semantic understanding comes at a cost: dense retrieval requires training a neural encoder (or fine-tuning a pre-trained one), storing and searching high-dimensional vectors, and GPU infrastructure for encoding at scale. But the quality gains are substantial, especially for queries that involve paraphrase or conceptual relationships rather than exact keyword matching.

The key insight is that sparse and dense representations capture complementary signals. Sparse methods are precise for exact terminology, dense methods are flexible for semantic content. This observation motivates hybrid retrieval, which we'll discuss in the architecture variations section.

Out[3]:
Visualization
Bar chart showing a sparse vector where most values are zero and only four bins have non-zero heights, illustrating BM25 or TF-IDF sparse representation.
Sparse vector representation for a simulated vocabulary. The bar chart shows how most dimensions remain at zero, with active weights assigned only to the specific indices corresponding to terms present in the document, as used in BM25 and TF-IDF retrieval methods.
Out[4]:
Visualization
Horizontal heatmap strip with 16 cells colored from blue to red representing continuous embedding values, showing that all dimensions have non-zero weights.
Dense vector representation using a continuous embedding space. The heatmap strip shows non-zero values across every dimension, illustrating how dense models capture abstract semantic features rather than relying on exact keyword matches.

The choice between sparse and dense retrieval involves trade-offs we'll examine in the next chapter. Sparse methods run quickly without training data, and their results are easy to interpret, making them excellent baselines. Dense methods capture semantic similarity that keyword matching misses but require substantial training data and computational resources. For now, recognize that both approaches solve the same basic problem: converting text into vectors that can be compared for similarity.

The Index

Encoded documents are stored in an index optimized for fast similarity search. Without an index, finding the most similar documents to a query would require computing similarity with every document in the collection. For a knowledge base with millions of documents, this brute-force approach would take seconds or minutes per query, far too slow for interactive applications. The index structure depends on the retrieval approach, with each type of representation requiring specialized data structures that exploit the mathematical properties of the representation.

For sparse retrieval, inverted indexes map each vocabulary term to the list of documents containing that term. The name "inverted" reflects that instead of mapping documents to their terms (the natural way we think about a document), we map terms to their documents. When a query arrives, the system looks up query terms, retrieves candidate documents containing those terms, and scores them using BM25 or similar formulas. This is the same technology powering web search engines, refined over decades to handle billions of documents with sub-second query times. The inverted index dramatically reduces computation because most queries contain only a handful of terms, and each term appears in only a fraction of documents. If your query contains three terms and each term appears in 0.1 percent of a 1 million document corpus, you need to score at most 3,000 documents rather than 1 million. That is a 99.7 percent reduction in computation from a simple data structure.

For dense retrieval, vector indexes organize embeddings to enable efficient nearest-neighbor search. Comparing a query embedding against millions of document embeddings naively requires millions of distance calculations, each involving hundreds or thousands of floating-point operations. Specialized data structures like Hierarchical Navigable Small World graphs (HNSW) and Inverted File indexes (IVF) reduce this to thousands or hundreds of comparisons through clever approximations. HNSW builds a multi-layer graph where each node connects to its nearest neighbors, letting search move quickly toward the most relevant region of the embedding space. Think of it as a transit network: at the top layer you have a few high-speed connections that let you travel large distances quickly, while at the bottom layer you have local connections for fine-grained navigation. IVF partitions the embedding space into clusters (using k-means) and searches only the clusters closest to the query vector, reducing the number of direct comparisons dramatically. These approximations trade a small amount of accuracy for dramatic speed improvements, and in practice the quality loss is negligible while the speedup can be 100x or more. We'll cover these algorithms in detail in upcoming chapters on vector similarity search.

The index improves performance while encoding a particular view of the data. An inverted index assumes that matching documents to queries is primarily about term overlap. A vector index assumes that semantic similarity is the primary relevance signal. Hybrid indexes can combine both, maintaining both an inverted index for exact match and a vector index for semantic match. When you choose an index, you're making a statement about what "relevant" means in your application.

Retrieval at Query Time

When your query arrives, the retriever executes a specific sequence of operations. First, it encodes the query using the same method applied to documents. For sparse retrieval, this means computing BM25 term weights for the query terms. For dense retrieval, this means passing the query through the same neural encoder that embedded the documents. This symmetry is needed: queries and documents must inhabit the same representation space for similarity comparisons to be meaningful. If you encode documents with one model and queries with a different model, the resulting vectors occupy different spaces and comparisons will be meaningless.

Second, the retriever searches the index for the kk most similar documents. For sparse retrieval, this involves looking up query terms in the inverted index and combining the results using BM25 scoring. For dense retrieval, this involves traversing the vector index to find the nearest neighbors of the query embedding. In both cases, the index structure turns what would be an exhaustive search into a targeted lookup.

Third, the retriever returns those documents along with their similarity scores. These scores serve multiple purposes: they determine the ranking order, they can filter out low-quality matches, and they can inform the generator about the relative confidence of different sources. A document with a BM25 score of 8.2 and a document with a score of 0.3 should not be weighted equally when constructing the generator's prompt.

The parameter kk (typically 3 to 10 documents) balances coverage against noise. Too few documents risk missing relevant information; too many dilute the generator's context with marginally relevant passages. Research has consistently found that retrieving more documents helps up to a point, after which quality plateaus or even degrades as irrelevant content confuses the generator. The optimal value of kk depends on your use case: simple factual questions might need only one or two relevant passages, while complex analytical queries might benefit from synthesizing information across many sources.

In[5]:
Code
# Minimal retriever demonstration using BM25
!uv pip install rank_bm25
from rank_bm25 import BM25Okapi
import numpy as np

# Sample knowledge base (in practice, thousands to millions of chunks)
documents = [
    "The transformer architecture was introduced in the paper Attention Is All You Need in 2017.",
    "BERT uses bidirectional self-attention to create contextualized word representations.",
    "GPT models are decoder-only transformers trained with causal language modeling.",
    "Retrieval-augmented generation combines neural retrievers with language model generators.",
    "The attention mechanism allows models to focus on relevant parts of the input sequence.",
    "Large language models are typically trained on trillions of tokens from web text.",
]

# Tokenize documents for BM25
tokenized_docs = [doc.lower().split() for doc in documents]
bm25 = BM25Okapi(tokenized_docs)

# Retrieve for a query
query = "How does the transformer architecture work?"
tokenized_query = query.lower().split()
scores = bm25.get_scores(tokenized_query)

# Get top-k documents
k = 3
top_indices = np.argsort(scores)[::-1][:k]
Out[6]:
Console
Query: How does the transformer architecture work?

Retrieved documents (ranked by BM25 score):

1. [Score: 3.072]
   The transformer architecture was introduced in the paper Attention Is All You Need in 2017.

2. [Score: 0.789]
   The attention mechanism allows models to focus on relevant parts of the input sequence.

3. [Score: 0.000]
   Large language models are typically trained on trillions of tokens from web text.
Out[7]:
Visualization
Horizontal bar chart listing six documents ranked by BM25 score, with the top three bars colored green and the remaining three colored gray, showing a score gap between relevant and irrelevant documents.
Document relevance ranking based on BM25 scores for a transformer architecture query. A clear score gap separates the top three retrieved documents (green) from the remaining results (gray), which reflects the BM25 scoring function's sensitivity to term frequency and inverse document frequency.

BM25 retrieves documents sharing query terms like "transformer" and "architecture." The scores reflect term overlap, frequency, and document length normalization as defined by the BM25 formula we covered in Part II. Notice that the highest-scoring document explicitly mentions both "transformer" and "architecture," while lower-ranked documents may match only one of these key terms or match through related vocabulary.

Query Encoding vs. Document Encoding

One subtlety that trips up many developers is that queries and documents are not always encoded identically, even when using the same encoder model. In asymmetric search, queries tend to be short, conversational questions while documents are longer, more formal passages. Some dense retrieval models are specifically trained with different prompting strategies for queries versus documents to bridge this distributional gap.

For example, the E5 family of embedding models prepends "query:" to queries and "passage:" to documents before encoding. The model has been trained to understand these prefixes as signals about the type of text being encoded, letting it to represent the query in a way that aligns well with relevant passages rather than with similar-looking queries. Models like BGE use instruction-tuning approaches where you can provide task-specific instructions like "Represent this sentence for searching relevant passages:". These seemingly minor formatting differences can substantially affect retrieval quality on asymmetric tasks, which is exactly the setup typical RAG systems use. When you use a pre-trained retrieval model, always read its documentation carefully to understand how it expects queries and documents to be formatted.

The Generator Component

The generator takes the retrieved documents and the original query, then produces a natural language response. In modern RAG systems, this is almost always a large language model, either an API-based model like GPT-4 or an open-weight model like LLaMA. The generator's role extends beyond simple information extraction: it must understand the query's intent, identify relevant portions of the retrieved documents, resolve any conflicts between sources, and synthesize a coherent, well-structured response. This is fundamentally a reading comprehension and synthesis task, and the generator's effectiveness depends heavily on both its underlying capabilities and how you present the context.

Think of the generator as a lawyer who has been handed a stack of case files and asked to write a brief. The lawyer doesn't just quote the files verbatim; they read critically, synthesize the relevant points, identify the strongest evidence, and construct a coherent argument. When sources conflict, they weigh the evidence. When sources are silent on a key point, a good lawyer acknowledges the gap rather than fabricating. The generator faces the same challenge: read the retrieved documents, identify what's relevant to the query, synthesize a response, and know when to say "I don't have enough information."

The quality of the generator's output depends on three interacting factors: the quality of the retrieved context (garbage in, garbage out), the quality of the prompt (which shapes what the model attends to and how it responds), and the intrinsic capability of the underlying model (which determines what complex synthesis and reasoning are possible). We can control all three, and improving any one of them generally improves the final answer.

Context Integration

The most straightforward approach to context integration is prompt concatenation: retrieved documents are inserted into the prompt before the query, giving the model access to relevant information through its standard attention mechanism. This approach requires no architectural modifications to the language model; we simply provide additional context that the model can reference during generation.

A typical RAG prompt structure looks like:

Context: [Document 1] [Document 2] [Document 3] Question: [User query] Answer based on the context above:

The generator reads this concatenated input and generates a response, attending to both the query and the retrieved documents. This uses the in-context learning capabilities we discussed in Part XXVIII: the model learns to extract and synthesize information from the provided context without any weight updates. The model has been trained on countless examples of reading passages and answering questions, so it naturally applies these learned behaviors to the RAG setting. The point is that prompt concatenation turns the generation task into something the model already knows how to do well. You are not asking the model to learn a new skill; you are giving the raw material and asking it to apply an existing skill.

The context window is both the enabler and the constraint here. With a 128k or 200k token context window, you can retrieve many documents and present them all simultaneously. The model can attend to all of them in a single forward pass, drawing connections across documents that iterative approaches would miss. But larger contexts come with costs: latency increases because the model must process more tokens, and attention costs scale quadratically with sequence length (unless you use efficient attention variants). There is also evidence that models do not use all positions in a long context equally well, with positions in the middle of a very long context receiving systematically less attention than positions at the beginning and end.

Attention Over Retrieved Context

From the model's perspective, retrieved documents are simply additional tokens in the input sequence. There is no special mechanism distinguishing context from query; both are processed uniformly by the same transformer layers. The self-attention mechanism, which we covered extensively in Part XIII, allows every generated token to attend to every token in the context. This means the model can identify which parts of which documents are relevant to the query by computing high attention weights for informative passages and low weights for irrelevant ones. It can synthesize information across multiple documents, combining facts from different sources into a unified answer. When sources conflict or provide complementary information, the model can implicitly judge which source is more authoritative or relevant based on patterns learned during training.

The attention patterns often reveal which documents influenced the response. Tokens in the generated answer attend strongly to the specific passages they're drawing from. This creates an implicit citation mechanism. Researchers have exploited this property to build attribution systems that show which retrieved passages contributed to each part of the generated response. While not a formal guarantee of faithfulness, these attention patterns provide valuable interpretability. Tools like LLM-Grader and various attribution frameworks attempt to use these signals to produce document-level or span-level citations, helping users verify that claims in the generated answer are grounded in the retrieved context rather than hallucinated.

The generator's ability to use retrieved context depends on how well it was trained to do so. Base language models trained only on next-token prediction may not reliably distinguish context-based answers from parametric answers. Models fine-tuned specifically on reading comprehension tasks or instruction-following tend to ground their responses in the provided context more reliably. This is one reason why using a model fine-tuned for question answering or instruction following often improves RAG quality over using a raw language model, even if the base models have similar overall quality.

Prompt Engineering for RAG

The exact prompt template significantly affects generation quality. Small changes in wording can substantially alter the model's behavior, determining whether it faithfully uses the context, hallucinates confidently, or appropriately expresses uncertainty. Prompt engineering for RAG is a non-trivial skill, and the right template often depends on the specific model you're using.

Several considerations shape effective RAG prompts:

  • Instruction clarity: Explicitly telling the model to base its answer on the provided context reduces hallucination. Phrases like "Answer based only on the documents above" or "If the information isn't in the context, say you don't know" provide clear behavioral guidance. Without such instructions, many models will blend retrieved context with their parametric knowledge, which can introduce outdated or incorrect information.
  • Document ordering: Models sometimes exhibit position bias, attending more to documents at the beginning or end of the context. Important documents might be placed in these privileged positions, or the order might be randomized to reduce systematic bias. Some research suggests ordering by decreasing relevance score (most relevant first) helps, while other research suggests the opposite. The safest approach is to experiment on your specific domain and model.
  • Attribution requests: Asking the model to cite which documents it used can improve traceability. Requests like "Reference the document numbers in your answer" encourage the model to explicitly connect claims to sources, which makes it easier to verify generated content.
  • Uncertainty expression: Instructing the model to say "I don't know" when context is insufficient prevents confident hallucinations. Without this instruction, models often generate plausible-sounding answers even when they lack the information to do so correctly. An explicit instruction to express uncertainty shifts the model's behavior from confabulation toward appropriate epistemic humility.
  • Format specification: Telling the model how to structure its response (bullet points, numbered steps, prose paragraphs) improves usability and consistency. Consistent output formats also make downstream parsing more reliable if your application processes the generated text programmatically.
In[8]:
Code
def format_rag_prompt(query: str, documents: list[str]) -> str:
    """Format query and documents into a RAG prompt."""
    context_parts = []
    for i, doc in enumerate(documents, 1):
        context_parts.append(f"[Document {i}]: {doc}")

    context = "\n\n".join(context_parts)

    prompt = f"""Use the following documents to answer the question. If the documents don't contain enough information to answer, say "I cannot answer this based on the provided documents."

{context}

Question: {query}

Answer:"""

    return prompt


# Create RAG prompt with retrieved documents
retrieved_docs = [documents[idx] for idx in top_indices]
rag_prompt = format_rag_prompt(query, retrieved_docs)
Out[9]:
Console
Use the following documents to answer the question. If the documents don't contain enough information to answer, say "I cannot answer this based on the provided documents."

[Document 1]: The transformer architecture was introduced in the paper Attention Is All You Need in 2017.

[Document 2]: The attention mechanism allows models to focus on relevant parts of the input sequence.

[Document 3]: Large language models are typically trained on trillions of tokens from web text.

Question: How does the transformer architecture work?

Answer:

This prompt template makes the task explicit: use the documents, cite uncertainty when appropriate. The numbered document format enables the model to reference specific sources in its response, improving traceability and helping you verify the generated information. Notice how the instruction at the beginning explicitly constrains the model to the provided context, which is important for faithfulness.

System Prompt vs. User Prompt Placement

When using chat-based APIs (GPT-4, Claude, etc.), you have additional choices about where to place context. Retrieved documents can go in the system prompt (which persists across turns in a multi-turn conversation) or in the user message (which is specific to each turn). System prompt placement is useful when you have stable context documents that should inform all responses, like a company's product documentation. User message placement is better for dynamic context that changes per query, which is the typical RAG case. Some practitioners place general instructions in the system prompt ("You are a helpful assistant that answers questions based on provided context") and the actual retrieved documents in the user message alongside the query, striking a balance between reusability and freshness.

The placement decision also affects caching behavior. Some inference providers cache system prompts, so placing frequently-used context there can reduce latency and cost. Prefix caching allows the provider to reuse the KV-cache computation from a shared prefix, avoiding redundant processing. If your application has a fixed document set that many users query, system-prompt placement with prefix caching can significantly reduce latency.

Retrieval Timing

A important architectural decision is when retrieval occurs relative to generation. This timing choice affects latency, complexity, and the types of queries the system can handle effectively. Different timing strategies suit different use cases, and the choice determines the basic capabilities of the system rather than merely optimizing performance.

Think of retrieval timing as the difference between a student who reads all relevant materials before writing an essay (single-shot), one who writes sections and periodically stops to look up needed facts (iterative), and one who has a photographic memory aid that supplies relevant passages for each word they write (token-level). Each approach has different strengths: the first is simpler and faster, the second handles complex multi-part questions better, and the third maintains maximum relevance at maximum cost.

Single-Shot Retrieval

The simplest approach retrieves documents once, before generation begins. The query goes to the retriever, documents come back, and generation proceeds with that fixed context. The flow is linear: your query reaches the retriever, which searches the index and returns relevant documents, then those documents and the original query combine into a prompt that the generator processes to produce the final response.

Single-shot retrieval works well when:

  • The query clearly specifies what information is needed
  • Retrieved documents are likely to contain complete answers
  • Low latency is necessary (one retrieval round-trip)
  • The query does not require multi-hop reasoning across different topics

Most production RAG systems use single-shot retrieval because it handles straightforward queries quickly with a simple architecture. The design is easy to reason about and debug, which also makes optimization more direct. When something goes wrong, you can examine the retrieved documents to determine whether the problem was in retrieval (wrong documents retrieved) or in generation (right documents, wrong answer). This diagnostic clarity is a significant practical advantage.

The limitation of single-shot retrieval becomes apparent with complex queries. Consider "How did the policy changes in 2019 affect the 2020 research outcomes described in the Smith et al. study?" This query requires knowing what the 2019 policy changes were, what research outcomes the study described, and how the former affected the latter. A single retrieval might return some of the relevant documents, but might miss the policy context entirely, or return the study abstract without the relevant methodology section. Single-shot retrieval cannot adapt to what it finds.

Iterative Retrieval

Complex queries sometimes require multiple retrieval steps. The model might need to first retrieve background information, then use that to formulate more specific sub-queries. In iterative retrieval, the process begins with an initial retrieval based on your query. Partial generation may then reveal a gap in the available information, prompting the system to formulate a new, more specific query and retrieve again. The final generation step synthesizes all retrieved content into a complete response.

Consider the query: "Compare the economic impact of the 2008 and 2020 recessions." A single retrieval might not return sufficiently comparable information about both events. The query contains two distinct information needs, and the knowledge base might organize information about each recession separately. Iterative retrieval allows the system to retrieve documents about the 2008 recession first, then retrieve documents about the 2020 recession, and finally synthesize a comparison using both sets of documents.

The trade-off is latency and complexity. Each retrieval adds round-trip time and requires logic to determine when additional retrieval is needed. The system must decide how to decompose complex queries and when it has gathered sufficient information to generate a final answer. In production systems, iterative retrieval often involves a "planning" step where the model first identifies the sub-questions that need answering, then retrieves for each one, and finally synthesizes the collected information. This is closely related to agent-based systems where a language model orchestrates multiple tool calls, with the retriever being one of the available tools.

Iterative retrieval can also implement multi-hop reasoning, where answering one question provides context needed to answer the next. For example, to answer "Who founded the company that makes the most popular Python web framework?", a system might first retrieve "What is the most popular Python web framework?" (answer: Django), then retrieve "Who founded Django?", letting a two-hop answer that neither single query alone could provide. Multi-hop reasoning is particularly important for knowledge-intensive tasks where the answer requires combining information from multiple distinct sources.

Token-Level Retrieval

At the opposite extreme from single-shot, some architectures retrieve new documents for each token generated. The RETRO (Retrieval-Enhanced Transformer) architecture from DeepMind interleaves retrieval throughout generation, attending to freshly retrieved passages at regular intervals. This approach ensures that the retrieved context remains relevant even as generation shifts to new topics or sub-questions.

RETRO divides the input into chunks of a fixed size (typically 64 tokens) and retrieves nearest-neighbor passages from a massive database for each chunk. These retrieved passages are processed by a separate encoder and integrated into the main model through cross-attention layers. The effect is that every portion of the generation process has access to relevant passages from the retrieval database, dynamically updated as the generation proceeds. A passage relevant to the first part of a long answer might not be relevant to the third part, and RETRO can adapt by retrieving different passages for each segment.

This approach can maintain relevance as generation shifts topics, but the computational overhead is substantial. Each retrieval operation adds latency, and performing thousands of retrievals (one per chunk of generation) quickly becomes impractical for interactive applications. Token-level retrieval is primarily a research technique rather than a production pattern, though its insights have influenced more practical architectures. The key insight that retrieval should remain dynamically relevant throughout generation, rather than being fixed at query time, has influenced designs like kNN-LM and more recent memory-augmented approaches.

Out[10]:
Visualization
Gantt-style horizontal bar chart with three rows labeled Single-Shot, Iterative, and Token-Level RETRO, showing orange retrieval blocks and green generation blocks arranged differently for each strategy along a time axis.
Comparison of retrieval timing strategies against generation progress. Single-shot RAG performs one upfront lookup, iterative RAG interleaves multiple retrievals with partial generation, and token-level architectures like RETRO perform continuous retrieval throughout the generation process, trading increasing latency for improved context relevance.

Query-Time vs. Index-Time Retrieval

An orthogonal timing dimension is when document processing occurs. This choice affects the trade-off between preparation time (when documents are added) and response time (when queries arrive).

Index-time processing pre-computes everything possible: chunking, embedding, and indexing happen once when documents are added to the knowledge base. Query-time only computes query encoding and index lookup. This approach minimizes latency for users but requires re-processing documents whenever the chunking strategy or embedding model changes. If you want to upgrade from a 384-dimensional embedding model to a better 1536-dimensional model, you need to re-embed every document in your knowledge base, which can take hours or days for large corpora.

Query-time processing delays some computation until a query arrives. For example, the system might store raw documents and compute embeddings using a query-dependent prompt that incorporates information about the specific query or user context. This enables more sophisticated relevance scoring at the cost of latency. Some hybrid approaches pre-compute base embeddings at index time but compute additional query-specific features at query time, combining the efficiency of pre-computation with the flexibility of dynamic scoring.

Most systems favor aggressive index-time processing to minimize query latency, accepting the overhead of re-indexing when configurations change. This matches the typical read/write pattern: documents are added infrequently, but queries arrive continuously. Optimizing for query latency by accepting higher index-building time is usually the right trade-off.

Architecture Variations

Since the original RAG paper, researchers and practitioners have developed numerous architectural variations. Understanding these helps you choose or design the right architecture for your use case. Simpler designs may sacrifice performance or capabilities, while more capable systems cost more to build and operate. The variations are not mutually exclusive; production systems often combine elements from multiple approaches.

Retrieve-then-Read (Sequential RAG)

The standard architecture we've been describing is sometimes called "retrieve-then-read": retrieve first, then read (generate). Documents pass through a clean interface: the retriever's output becomes the generator's input. This separation creates a clear contract between components: the retriever promises to return relevant documents, and the generator promises to synthesize them into an answer.

In[11]:
Code
class SequentialRAG:
    """Basic retrieve-then-read RAG architecture."""

    def __init__(self, retriever, generator, k=3):
        self.retriever = retriever
        self.generator = generator
        self.k = k

    def answer(self, query: str) -> str:
        # Step 1: Retrieve
        documents = self.retriever.search(query, k=self.k)

        # Step 2: Generate
        prompt = self.format_prompt(query, documents)
        response = self.generator.generate(prompt)

        return response

    def format_prompt(self, query: str, documents: list[str]) -> str:
        context = "\n\n".join(documents)
        return f"Context:\n{context}\n\nQuestion: {query}\n\nAnswer:"

This separation enables independent optimization of each component. You can upgrade the retriever without touching generation logic, or swap in a different LLM without modifying retrieval. This modularity also simplifies debugging: if answers are wrong, you can examine retrieved documents to determine whether the problem is in retrieval (wrong documents) or generation (right documents, wrong answer). The sequential architecture is the right starting point for most projects. Build it first, measure where quality falls short, and then consider more complex architectures only if the sequential version cannot meet your requirements.

Fusion-in-Decoder

The Fusion-in-Decoder (FiD) architecture processes each retrieved document independently through the encoder, then fuses their representations in the decoder. This approach addresses a scalability challenge with the standard concatenation approach.

Instead of concatenating all documents into a single input, which creates one long sequence that grows linearly with the number of documents, FiD keeps encoding costs constant per document by encoding each document-query pair separately. Each encoding produces a fixed-size representation, and these representations are concatenated for the decoder. The decoder then attends over all document representations simultaneously when generating the answer.

This approach scales better with the number of retrieved documents. Standard concatenation creates a sequence of length n×kn \times k where nn is average document length and kk is the number of documents, and attention cost grows quadratically with sequence length. FiD keeps each encoder pass at length nn, with only the decoder needing to handle the combined representations. The decoder's cross-attention mechanism attends over the concatenated encoder outputs, effectively fusing information from all documents. This allows FiD to process many more documents than would be practical with simple concatenation, often improving quality on tasks that require synthesizing information from many sources.

The original FiD paper demonstrated state-of-the-art performance on open-domain question answering benchmarks by retrieving and fusing up to 100 documents per query. At that scale, concatenation would create sequences of tens of thousands of tokens, making generation prohibitively expensive, while FiD kept costs manageable by encoding documents independently. The trade-off is that FiD requires an encoder-decoder architecture (like T5), which makes it less applicable to pure decoder models like LLaMA or GPT, which are increasingly common in modern RAG systems.

Retrieval-Augmented Language Models (REALM and RAG)

The original RAG paper from Facebook AI and the REALM paper from Google both proposed end-to-end trainable retrieval. Rather than using a frozen retriever, these architectures backpropagate gradients through the retrieval step. This means the retriever can learn which documents are similar to the query and which documents are useful for generating correct answers.

The key insight is treating retrieval as a latent variable. In this formulation, we don't commit to using a single retrieved document. Instead, we consider multiple retrieved documents and weight their contributions by how relevant the retriever believes them to be. Given query qq and a set of retrieved documents {d1,…,dn}\{d_1, \dots, d_n\}, the probability of generating answer aa is computed by marginalizing over the documents:

P(a∣q)=∑i=1nP(a∣q,di)⋅P(di∣q)P(a|q) = \sum_{i=1}^{n} P(a|q, d_i) \cdot P(d_i|q)

where:

  • aa: the generated answer we want to produce
  • qq: the input query from the user
  • did_i: the ii-th retrieved document from the knowledge base
  • nn: the total number of retrieved documents under consideration
  • P(a∣q,di)P(a|q, d_i): the probability of the generator creating answer aa given query qq and document did_i, measuring how well this specific document helps generate the correct answer
  • P(di∣q)P(d_i|q): the probability that the retriever assigns to document did_i being relevant to query qq, derived from the similarity score between the query embedding and the document embedding

This formula computes a weighted mixture of generation probabilities, where each document's contribution to the final probability is proportional to how relevant the retriever believes it to be. Think of it as holding a committee vote: each document votes on the answer, and its vote carries weight proportional to its relevance score. If document d1d_1 is highly relevant (P(d1∣q)P(d_1|q) is large) and strongly supports the correct answer (P(a∣q,d1)P(a|q, d_1) is large), it dominates the sum. Documents that are irrelevant or that point toward wrong answers receive low weights and contribute little.

The core advantage of this formulation is differentiability. Because both terms in the product are differentiable functions of model parameters, gradients flow from the generation loss back through the generation probability, through the retriever weights, to the document embeddings. If the generator consistently produces wrong answers when using certain documents, the retriever learns to down-weight those documents for similar queries. The retriever and generator co-adapt: the retriever learns to retrieve what helps the generator, and the generator learns to use whatever the retriever provides.

Training end-to-end is expensive because updating document embeddings requires re-indexing, which can take hours for large knowledge bases. This creates a practical challenge: if you update the document encoder at every gradient step, you need to re-embed and re-index all documents continuously, which is impractical. The original RAG paper uses asynchronous index updates, refreshing document embeddings periodically (e.g., every few thousand gradient steps) rather than after every update. However, the resulting retrievers are specifically optimized for the generation task, often outperforming generic retrievers by substantial margins on knowledge-intensive benchmarks. The trade-off is infrastructure complexity and training cost versus retrieval quality.

RETRO: Chunked Cross-Attention

The RETRO (Retrieval-Enhanced Transformer) architecture from DeepMind takes a different approach, building retrieval into the transformer architecture itself. Rather than treating retrieval as a preprocessing step, RETRO makes retrieval an integral part of the model's forward pass.

RETRO divides the input into chunks and retrieves neighbors for each chunk. These neighbors are documents from the knowledge base that are similar to the current chunk being processed. Retrieved neighbors are processed by a separate encoder, and their representations are integrated through chunked cross-attention layers interleaved with standard self-attention. This means the model alternates between attending to its own representations (self-attention over the input and generated text) and attending to retrieved content (cross-attention over the encoder outputs for retrieved passages). Every chunk of generated text has access to the most relevant passages from the retrieval database, dynamically updated as generation proceeds.

This tight integration allows RETRO to scale retrieval to trillions of tokens while keeping the main model relatively small. The model can effectively "look up" information during generation rather than storing all knowledge in its parameters. A 7B parameter RETRO model can match the performance of a 25B model without retrieval by using a massive retrieval database. This represents a major change in how we think about model capacity: some knowledge lives in parameters (where it's always available but may be outdated), while other knowledge lives in an external database (where it can be updated without retraining) accessed through retrieval. The implication is that the right strategy for scaling might be to scale the retrieval database rather than the model parameters, at least for knowledge-intensive tasks.

Self-RAG: Retrieval as a Learned Decision

Most RAG systems retrieve for every query, but retrieval isn't always necessary or beneficial. Simple queries like "What is 2+2?" don't benefit from retrieval, and retrieving irrelevant documents can hurt performance by introducing noise and confusing the generator. Self-RAG trains models to decide when to retrieve, what to retrieve, and how to use retrieved content, adapting retrieval behavior to the specific query at hand.

The Self-RAG model generates special control tokens during its forward pass, inserting them at appropriate points to indicate its decisions about retrieval and document relevance:

  • [Retrieve]: Signals whether retrieval is needed for this query or generation segment
  • [IsREL]: Assesses whether a retrieved passage is relevant to the current generation context
  • [IsSUP]: Evaluates whether a passage supports the claim being generated
  • [IsUSE]: Determines whether to use a passage in the final generated response

These reflection tokens enable the model to critique its own retrieval usage in real time, filtering irrelevant documents and identifying when it should rely on parametric knowledge instead of retrieved context. The model has been trained through a combination of supervised learning (to generate appropriate reflection tokens) and reinforcement learning (to optimize the quality of final answers), so its retrieval decisions are grounded in observed answer quality.

Self-RAG is more adaptive than systems that blindly retrieve for every query. It achieves competitive quality with selective retrieval, reducing both latency and cost for queries that don't need external information. The training approach also produces models that express more calibrated uncertainty: they retrieve more when they're less sure about the answer and rely on parametric knowledge when they're confident, rather than always deferring to the retriever. This behavior is closer to how a knowledgeable human expert would work: drawing on memory for familiar facts while looking things up for less familiar topics.

Hybrid Retrieval

Rather than choosing between sparse and dense retrieval, many production systems use both and combine their results. The intuition is simple: sparse retrieval is excellent at finding documents with exact term matches (useful when queries use precise technical terminology), while dense retrieval is excellent at finding semantically related documents (useful when queries use everyday language about technical topics). Combining both captures the strengths of each.

The most common combination strategy is Reciprocal Rank Fusion (RRF). Given two ranked lists of documents, one from sparse retrieval and one from dense retrieval, RRF assigns each document a score based on its rank in each list:

RRF(d)=∑r∈R1k+rankr(d)\text{RRF}(d) = \sum_{r \in R} \frac{1}{k + \text{rank}_r(d)}

where:

  • dd: the document being scored
  • RR: the set of retrieval methods (e.g., sparse and dense)
  • rankr(d)\text{rank}_r(d): the rank of document dd in the results from retrieval method rr (lower rank means more relevant)
  • kk: a constant (typically 60) that dampens the influence of very high-ranked documents and prevents any single result from dominating

The key insight behind RRF is that rank is more reliable than score. BM25 scores and cosine similarities live on different scales and cannot be directly compared, but if a document ranks first in both the sparse and dense results, that convergent signal strongly suggests relevance. The constant k=60k = 60 was found empirically to work well across many retrieval tasks by reducing the impact of randomly high rankings. Documents that appear in the top results of both systems receive the highest RRF scores, with the score declining as rank falls.

RRF is simple to implement, requires no model training, and consistently outperforms either retriever alone on most benchmarks. More sophisticated fusion approaches learn weights for each retrieval signal, but RRF is an excellent default that works well across diverse domains.

Worked Example: Tracing a Query Through the Full Pipeline

Let's trace a specific query through every stage of a sequential RAG system to make the abstract components concrete. We'll use the query: "What is the key innovation that allows transformers to train faster than RNNs?"

Stage 1: Query Encoding (Sparse). The query is tokenized and lowercased: ["what", "is", "the", "key", "innovation", "that", "allows", "transformers", "to", "train", "faster", "than", "rnns"]. Stop words ("what", "is", "the", "that", "to", "than") are typically removed for BM25: ["key", "innovation", "allows", "transformers", "train", "faster", "rnns"]. BM25 term weights are computed for each remaining term based on their IDF values across the corpus.

Stage 2: Index Lookup. The retriever looks up each query term in the inverted index. "transformers" might appear in 1,247 documents, "train" in 3,821 documents, "faster" in 892 documents, and "innovation" in 654 documents. The inverted index returns candidate documents for each term, and BM25 scores are computed for each candidate based on the formula we studied in Part II. For our knowledge base, the highest-scoring document contains the sentence "The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training."

Stage 3: Top-k Selection. The retriever ranks all candidate documents by BM25 score and returns the top k=3k = 3. The returned documents discuss transformers, self-attention, and the comparison with RNNs, all directly relevant to the query.

Stage 4: Context Formatting. The three retrieved documents are formatted into numbered context blocks with the query appended, following the template shown earlier. The full prompt is approximately 350 tokens.

Stage 5: Generation. The language model receives the 350-token prompt and generates a response attending to all tokens in the context. The model identifies the passage "replacing recurrence with self-attention" and "process all positions in parallel during training" as the key information. The generated response synthesizes this into a coherent answer: "The key innovation letting transformers to train faster than RNNs is the replacement of sequential recurrence with parallel self-attention. RNNs must process tokens one at a time, preventing parallelization across sequence positions. Transformers compute attention over all positions simultaneously, letting full GPU utilization during training."

Stage 6: Verification. In a production system, you would compare the generated claims against the retrieved documents to verify grounding. The claim about "parallel self-attention" is directly supported by Document 6 in the knowledge base. The claim about "sequential recurrence" preventing parallelization is a reasonable inference from the same passage.

This trace reveals why each component matters. If the retriever had returned documents about cooking instead of transformers, the generator's answer would have been wrong no matter how capable it was. If the generator had ignored the context and hallucinated, it might have said something plausible but incorrect. The final answer quality depends on both components working correctly in sequence.

Building a Complete RAG Pipeline

Let's assemble the components into a working RAG system. We'll use a sparse retriever for simplicity; the next chapter covers dense retrieval in depth. This implementation demonstrates the core patterns you'll use in production systems, even though real systems would connect to actual language model APIs.

In[12]:
Code
from dataclasses import dataclass

import numpy as np
from rank_bm25 import BM25Okapi


@dataclass
class RetrievedDocument:
    """A document with its retrieval score."""

    content: str
    score: float
    doc_id: int


class BM25Retriever:
    """BM25-based sparse retriever."""

    def __init__(self, documents: list[str]):
        self.documents = documents
        self.tokenized_docs = [doc.lower().split() for doc in documents]
        self.bm25 = BM25Okapi(self.tokenized_docs)

    def retrieve(self, query: str, k: int = 3) -> list[RetrievedDocument]:
        """Retrieve top-k documents for a query."""
        tokenized_query = query.lower().split()
        scores = self.bm25.get_scores(tokenized_query)

        top_indices = np.argsort(scores)[::-1][:k]

        return [
            RetrievedDocument(
                content=self.documents[idx], score=scores[idx], doc_id=idx
            )
            for idx in top_indices
        ]
In[13]:
Code
class SimpleGenerator:
    """Placeholder generator (in practice, use an LLM API or local model)."""

    def __init__(self):
        pass

    def generate(self, prompt: str, max_tokens: int = 256) -> str:
        """Generate a response given a prompt.

        This is a placeholder. In practice, you would:
        - Call an API (OpenAI, Anthropic, etc.)
        - Run a local model (llama.cpp, vLLM, etc.)
        """
        return "[Generated response based on retrieved context]"
In[14]:
Code
class RAGPipeline:
    """Complete RAG pipeline combining retriever and generator."""

    def __init__(
        self,
        retriever: BM25Retriever,
        generator: SimpleGenerator,
        k: int = 3,
        include_scores: bool = False,
    ):
        self.retriever = retriever
        self.generator = generator
        self.k = k
        self.include_scores = include_scores

    def format_context(self, docs: list[RetrievedDocument]) -> str:
        """Format retrieved documents into context string."""
        parts = []
        for i, doc in enumerate(docs, 1):
            if self.include_scores:
                parts.append(f"[{i}] (score: {doc.score:.2f}) {doc.content}")
            else:
                parts.append(f"[{i}] {doc.content}")
        return "\n\n".join(parts)

    def build_prompt(self, query: str, context: str) -> str:
        """Construct the full prompt for the generator."""
        return f"""Answer the question based only on the following context. If the context doesn't contain the answer, say "I don't have enough information to answer this question."

Context:
{context}

Question: {query}

Answer:"""

    def answer(self, query: str) -> dict:
        """Run the full RAG pipeline."""
        retrieved_docs = self.retriever.retrieve(query, k=self.k)
        context = self.format_context(retrieved_docs)
        prompt = self.build_prompt(query, context)
        response = self.generator.generate(prompt)

        return {
            "query": query,
            "retrieved_documents": retrieved_docs,
            "prompt": prompt,
            "response": response,
        }
In[15]:
Code
# Build a knowledge base about ML concepts
knowledge_base = [
    "Transformers use self-attention mechanisms that allow each token to attend to all other tokens in the sequence. This enables capturing long-range dependencies without the sequential processing limitations of RNNs.",
    "The attention mechanism computes compatibility scores between queries and keys, then uses these scores to create weighted sums of values. This can be expressed as Attention(Q,K,V) = softmax(QK^T / sqrt(d_k))V.",
    "BERT (Bidirectional Encoder Representations from Transformers) uses masked language modeling for pre-training, randomly masking 15% of input tokens and training the model to predict them.",
    "GPT models are autoregressive, generating one token at a time while attending only to previously generated tokens. This causal attention prevents information leakage from future tokens.",
    "Retrieval-augmented generation grounds language models in external knowledge by retrieving relevant documents before generation. This reduces hallucination and enables access to up-to-date information.",
    "The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training.",
    "Multi-head attention runs several attention operations in parallel, letting the model to capture different types of relationships at different positions.",
    "Position encodings are necessary in transformers because self-attention is permutation-invariant. Sinusoidal encodings and learned position embeddings are common approaches.",
]

retriever = BM25Retriever(knowledge_base)
generator = SimpleGenerator()
rag = RAGPipeline(retriever, generator, k=3, include_scores=True)

result = rag.answer("What is the key innovation of transformers?")
Out[16]:
Console
Query: What is the key innovation of transformers?

============================================================
Retrieved Documents:
============================================================

[Doc 5] Score: 5.025
  The key innovation of transformers is replacing recurrence with self-attention. While RNNs process s...

[Doc 7] Score: 1.061
  Position encodings are necessary in transformers because self-attention is permutation-invariant. Si...

[Doc 0] Score: 0.820
  Transformers use self-attention mechanisms that allow each token to attend to all other tokens in th...

============================================================
Constructed Prompt:
============================================================
Answer the question based only on the following context. If the context doesn't contain the answer, say "I don't have enough information to answer this question."

Context:
[1] (score: 5.02) The key innovation of transformers is replacing recurrence with self-attention. While RNNs process sequences step-by-step, transformers process all positions in parallel during training.

[2] (score: 1.06) Position encodings are necessary in transformers because self-attention is permutation-invariant. Sinusoidal encodings and learned position embeddings are common approaches.

[3] (score: 0.82) Transformers use self-attention mechanisms that allow each token to attend to all other tokens in the sequence. This enables capturing long-range dependencies without the sequential processing limitations of RNNs.

Question: What is the key innovation of transformers?

Answer:

The pipeline retrieves the three most relevant documents (those mentioning "transformers," "attention," and related concepts) and then constructs a prompt instructing the model to answer based on this context. Notice how the BM25 scores reflect the degree of term overlap between the query and each document, with the highest-scoring document containing the exact phrase "key innovation of transformers."

Handling Edge Cases

Production RAG systems need to handle several edge cases gracefully. Not every query will have good matches in the knowledge base, and the system should behave appropriately in these situations rather than hallucinating or returning nonsensical results.

In[17]:
Code
def check_retrieval_quality(
    docs: list[RetrievedDocument], score_threshold: float = 1.0
) -> dict:
    """Assess whether retrieved documents are likely relevant."""
    if not docs:
        return {"status": "no_documents", "action": "fallback_to_parametric"}

    max_score = max(doc.score for doc in docs)
    avg_score = sum(doc.score for doc in docs) / len(docs)

    if max_score < score_threshold:
        return {
            "status": "low_confidence",
            "max_score": max_score,
            "action": "warn_user_or_expand_search",
        }

    return {
        "status": "good",
        "max_score": max_score,
        "avg_score": avg_score,
        "action": "proceed_with_generation",
    }


# Test with a query that has good matches
good_docs = retriever.retrieve("How does attention work in transformers?", k=3)
quality_good = check_retrieval_quality(good_docs)

# Test with a query outside the knowledge base
poor_docs = retriever.retrieve("What is the capital of France?", k=3)
quality_poor = check_retrieval_quality(poor_docs)
Out[18]:
Console
Query with good matches:
  Max score: 1.078
  Status: good
  Action: proceed_with_generation

Query outside knowledge base:
  Max score: 1.722
  Status: good
  Action: proceed_with_generation

When retrieval scores are low, the system might fall back to the model's parametric knowledge, return "I don't know" rather than hallucinating, or trigger a different retrieval strategy (expand query, search different index). The appropriate action depends on your use case. A customer support bot might need to always provide some answer, while a medical information system must be conservative and admit uncertainty. The threshold used in the quality check function should be tuned empirically on representative queries from your target domain, since BM25 scores are not standardized across corpora.

Key Parameters

Understanding the key parameters in a RAG pipeline allows you to tune the system for your specific requirements. These parameters interact with each other, and changing one often requires re-tuning others.

The most important parameters are:

  • k: The number of documents to retrieve. Balances context coverage against noise. Typical values range from 3 to 10, with lower values for focused queries and higher values for broad research questions. Larger kk increases latency both from the retrieval step (more results to rank) and the generation step (longer context to process). Experiment with values from 1 to 20 and measure answer quality to find the sweet spot for your domain.
  • score_threshold: The minimum similarity score required to consider a document relevant. This threshold depends heavily on the retrieval method and corpus, and should be tuned on representative queries. A threshold calibrated on one corpus may not transfer to another.
  • max_tokens: The maximum number of tokens to generate in the response. Longer limits allow more complete answers but increase latency and cost. Consider the typical answer length for your domain when setting this parameter.
  • chunk_size: The size of document chunks in tokens. Smaller chunks (128-256 tokens) are better for precise factual retrieval; larger chunks (512-1024 tokens) preserve more context and work better for questions requiring understanding of extended passages.
  • chunk_overlap: The number of tokens that adjacent chunks share. Overlap of 10-20 percent reduces the chance of splitting relevant content across chunk boundaries at the cost of some index redundancy.

Data Flow Visualization

The following diagram illustrates how data flows through a RAG system:

Out[19]:
Visualization
Flowchart showing query entering retriever, documents being retrieved, context combined with query, and generator creating response.
Data flow within a standard RAG pipeline. A user query triggers document retrieval from an indexed knowledge base, giving the necessary context for the language model generator to synthesize a grounded response. Numbers indicate the sequential steps: query to retriever (1), documents from knowledge base (2), query plus context to generator (3), and final response (4).

The numbers indicate the sequence: (1) the query goes to the retriever, (2) relevant documents return from the knowledge base, (3) the query and retrieved context combine into a prompt, and (4) the generator produces the final response.

Comparing Architecture Variants

Different RAG architectures balance response quality against latency, and their implementation costs also vary. Understanding these trade-offs helps you choose the right approach for your application's requirements.

Out[20]:
Visualization
Grouped bar chart comparing four RAG variants on speed, quality potential, and simplicity metrics using a 1 to 5 scale.
Comparison of RAG architecture variants across speed, quality potential, and simplicity. Sequential RAG provides the best balance of speed and simplicity for production use. More sophisticated architectures like Fusion-in-Decoder and end-to-end REALM/RAG achieve higher quality potential by processing more documents or jointly training the retriever and generator, but at the cost of greater complexity and reduced inference speed.

Sequential RAG dominates in production because it's fast and simple while delivering good quality. Fusion-in-Decoder trades some speed for better multi-document reasoning. Iterative retrieval handles complex queries but adds latency. End-to-end training maximizes quality but requires significant infrastructure investment. The right choice for your application depends on your latency budget, quality requirements, and engineering resources. Start with sequential RAG, measure its performance on your specific queries, and consider more complex architectures only if sequential RAG demonstrably falls short.

Limitations and Considerations

RAG architecture decisions involve inherent trade-offs that you should understand before building systems. These limitations don't make RAG ineffective; rather, they define the boundaries within which RAG excels and help you set appropriate expectations.

Context window constraints remain a basic limitation. Even with modern models supporting 100K or more token contexts, there's a practical limit to how many documents you can retrieve. As we discussed in Part XVII on efficient attention and Part XVIII on context extension, attention costs grow with sequence length, and models can struggle to effectively use information buried deep in long contexts. Research has shown that models tend to focus on information at the beginning and end of the context, with middle sections receiving less attention. This "lost in the middle" phenomenon means retrieval quality matters enormously: retrieving the right 3 documents beats retrieving 30 marginally relevant ones. You can mitigate this by placing the most relevant document first or last in the context, by reranking retrieved documents to prioritize the highest-quality ones, or by using architectures like FiD that avoid concatenation entirely.

The retriever-generator mismatch problem arises because these components are typically trained separately. A retriever optimized for traditional information retrieval metrics may not retrieve what is most useful for generation. A passage that scores highest on BM25 might lack the specific details the generator needs, while a lower-ranked passage contains exactly the right information. The retriever knows about word overlap and semantic similarity, but it doesn't know what the generator needs to answer the specific question it's facing. End-to-end training addresses this mismatch directly, but requires substantial infrastructure. Reranking approaches use a cross-encoder (which jointly encodes query and document) to re-score retrieved documents after initial retrieval, capturing relevance signals that the bi-encoder retriever misses. We'll cover cross-encoders and reranking in later chapters.

Latency accumulates across the pipeline in compound ways. Retrieval adds round-trip time to the knowledge base (potentially remote), and longer contexts increase generation time because the attention mechanism must process more tokens. For real-time applications, this motivates careful optimization: caching frequent queries, using smaller models, or pre-computing responses for common question patterns. Each additional retrieved document adds both retrieval time (more results to rank and return) and generation time (more tokens for the model to attend over), creating compound latency effects. The relationship between kk and total latency is often superlinear in practice, because vector index search time grows with kk and generation cost grows with total context length.

Attribution and faithfulness remain active research areas that no current approach fully solves. When a RAG system generates an answer, how do you verify it came from the retrieved documents rather than from hallucinated parametric knowledge? The model might confidently state a fact that appears nowhere in the retrieved context, drawing instead from knowledge encoded in its weights during pretraining. This is particularly concerning for knowledge that changes over time: a model trained before a major news event might ignore retrieved context about that event and generate answers based on outdated weights. Current approaches include asking models to cite sources, comparing responses with and without context, or using specialized fact-verification models. None are fully reliable, making RAG a tool that reduces rather than eliminates hallucination risk. You should understand that retrieved documents provide evidence for the answer, but the system may still make errors in how it interprets or synthesizes that evidence.

Knowledge base maintenance introduces operational challenges that compound over time. Documents in your knowledge base become outdated as the world changes. Adding new documents requires re-indexing. Removing documents to comply with deletion requests requires removing them from the document store and removing their embeddings from the vector index, which some index types handle poorly. The knowledge base is a living system that requires ongoing maintenance, not a static artifact. In production, you need processes for adding new content, removing stale or incorrect content, and periodically re-evaluating whether your chunking and encoding strategies still serve your users well.

Finally, evaluation is harder than it appears. Measuring RAG quality requires both component-level evaluation (is the retriever finding the right documents?) and end-to-end evaluation (is the final answer correct?). Component metrics like recall at kk and mean reciprocal rank tell you about retrieval quality, but a high recall doesn't guarantee good final answers if the generator fails to synthesize correctly. End-to-end metrics like exact match, F1, or human evaluation tell you about overall quality but don't pinpoint where the system is failing. Building a complete evaluation harness that covers both components, with test cases representing the diversity of queries your users will send, is one of the most important investments you can make in a production RAG system.

Summary

This chapter examined the architecture of retrieval-augmented generation systems, revealing how the seemingly simple "retrieve then generate" pipeline conceals important design decisions at every stage.

The retriever component turns documents into searchable representations (sparse or dense), stores them in specialized indexes, and returns the most relevant passages at query time. The choice of retrieval method significantly affects what information reaches the generator, with sparse methods excelling at exact term matching and dense methods excelling at semantic similarity. Hybrid retrieval combining both approaches often outperforms either alone.

The generator component synthesizes answers from retrieved context using in-context learning. Prompt design matters: clear instructions about grounding responses in context, appropriate document formatting, and explicit uncertainty handling all improve output quality. The generator sees retrieved documents as additional tokens in the prompt and attends to them through the same self-attention mechanism used everywhere in the transformer.

Retrieval timing ranges from single-shot (fast, simple, suitable for most queries) to iterative (handles complex multi-step reasoning) to token-level as in RETRO (maximum relevance at maximum cost). Most production systems use single-shot retrieval with optional iteration for hard queries.

Architecture variations like Fusion-in-Decoder, end-to-end RAG/REALM, RETRO, and Self-RAG demonstrate that the basic sequential pattern is just one point in a large design space. FiD processes documents independently to scale retrieval count. End-to-end training aligns the retriever directly with generation quality. Self-RAG makes retrieval an adaptive, query-dependent decision rather than an unconditional operation.

The limitations are real: context window constraints, retriever-generator mismatch, latency accumulation, and attribution challenges each bound what RAG can reliably achieve. Understanding these boundaries helps you build systems with appropriate safeguards and realistic expectations.

The next chapter dives deep into dense retrieval: how neural networks learn to embed queries and documents into the same vector space, letting semantic matching that goes far beyond keyword overlap. This is the technology that makes modern RAG systems substantially more capable than their sparse retrieval predecessors, and understanding it is needed for building production-quality systems.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about RAG architecture.

RAG Architecture Fundamentals

Question 1 of 70 of 7 completed
Which statement best describes the difference between sparse and dense document representations?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026ragarchitecture, author = {Michael Brenndoerfer}, title = {RAG Architecture: Components, Timing & Design Patterns}, year = {2026}, url = {https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2026). RAG Architecture: Components, Timing & Design Patterns. Retrieved from https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns
MLAAcademic
Michael Brenndoerfer. "RAG Architecture: Components, Timing & Design Patterns." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns>.
CHICAGOAcademic
Michael Brenndoerfer. "RAG Architecture: Components, Timing & Design Patterns." Accessed September 27, 2026. https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns.
HARVARDAcademic
Michael Brenndoerfer (2026) 'RAG Architecture: Components, Timing & Design Patterns'. Available at: https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2026). RAG Architecture: Components, Timing & Design Patterns. https://mbrenndoerfer.com/writing/rag-architecture-retriever-generator-design-patterns

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.