Part of Language AI Handbook
Covers RAG prompt engineering with strategic context placement, citation formats, and truncation strategies to improve LLM accuracy and reduce hallucinations.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
RAG Prompt Engineering
Retrieval-Augmented Generation grounds large language models in external knowledge, but the retrieved context is only useful if the model uses it. As we discussed in the earlier chapters on embedding-based retrieval, we can train models to retrieve relevant documents with high precision, and as we explored in the chapter on reranking, we can refine those results for maximum relevance. Yet none of this pipeline work matters if the prompt construction fails to present that context effectively. A perfectly retrieved document placed in the wrong position, wrapped in ambiguous formatting, or accompanied by vague citation instructions might as well not exist.
Think of RAG prompt engineering as the final mile problem of information retrieval. Your retrieval system has done the hard work of finding the needle in the haystack; prompt engineering is how you hand that needle to the model without it getting lost in the pile again. The model is a sophisticated reader, but one with well-documented quirks: it pays more attention to the beginning and end of what you give it, it can be confused by inconsistent formatting, and it will sometimes ignore explicit instructions when its parametric memory pulls strongly in a different direction. Understanding these quirks, and designing prompts that account for them, is what separates reliable RAG systems from ones that hallucinate despite having the right documents on hand.
RAG prompt engineering combines traditional prompt design with context window management. The model must perform a task while synthesizing information from potentially dozens of retrieved passages, maintaining fidelity to sources, handling irrelevant or contradictory information, and staying within token limits. Prompt structure, context placement, citation formatting, and instructions determine whether the response is grounded and accurate or hallucinated. Every decision you make in the prompt template, from how you label documents to how you phrase the instruction to cite sources, shapes the model's output in measurable ways.
This chapter covers the key engineering decisions in RAG prompt construction: where to place retrieved context in the prompt, how to order multiple documents, what citation formats work best, how to truncate long documents while preserving relevance, and how to write instructions that reliably guide model behavior. We treat these as engineering problems with measurable tradeoffs, not stylistic choices, because your production system's accuracy depends on getting them right. We will build toward a complete prompt-building implementation that you can adapt for your own RAG pipelines, and we will examine the failure modes that haunt even well-designed systems.
The problem of presenting retrieved information to a generative model predates the transformer era. Early information retrieval systems in the 1990s and 2000s faced analogous challenges when presenting search snippets to human readers: how much context to show, whether to bold matching terms, and how to order results. The insight that position matters for attention, whether human or computational, is rooted in classical cognitive psychology research on primacy and recency effects in human memory. Liu et al. (2023) formalized this for language models in their paper "Lost in the Middle: How Language Models Use Long Contexts," which quantitatively demonstrated the U-shaped performance curve we discuss throughout this chapter. That paper became one of the most practically influential findings in the RAG literature because it transformed what had been an intuitive suspicion among practitioners into a measurable, reproducible effect that engineers could explicitly design around.
The Context Placement Problem
When you retrieve documents for RAG, you face an immediate architectural question: where do these documents go in the prompt? Unlike the fixed instructions of a system prompt or the specific query from a user, retrieved context is variable in length, potentially noisy, and important to the final output. The answer is not as simple as "put them before the question" because large language models do not process all positions in their context window with equal attention. Position measurably affects whether the model uses the information you provide.
Think of the context window as a theater where the seats at the front and back are premium. Information placed in those positions gets the most attention from the audience (the model), while information buried in the middle rows often goes unnoticed despite being present. The challenge for RAG systems is that you typically retrieve between 5 and 20 documents, some of which are only marginally relevant, and you need to pack them into a finite window in a way that maximizes the chance the truly important passages get read carefully.
The placement decision interacts with everything else in your prompt. A strong system instruction placed before the context competes for primacy with the documents themselves. A long list of retrieved passages placed immediately before the user query exploits recency bias but risks truncation if the concatenated context is too long. Understanding the mechanics of position bias, and then making deliberate choices about how to fight or exploit it, is the first major skill of RAG prompt engineering.
Recency Bias and the Lost in the Middle Phenomenon
Language models exhibit position bias when processing long contexts. Studies of large language models reveal a U-shaped performance curve: models excel at using information at the very beginning and very end of their context window, but struggle with information buried in the middle. This "lost in the middle" problem means that your most relevant document, if placed poorly, might be effectively ignored.

The bias stems from how attention mechanisms process long sequences. As we explored in the chapter on attention complexity, attention costs grow quadratically with sequence length, and in practice, models employ various approximations and soft attention decay that favor proximal tokens. Even with techniques like rotary position embeddings, the practical reality is that not all positions in a long context receive equal processing fidelity. The model's internal representations are shaped by what is in the context and where each piece of information appears.
To understand this phenomenon mathematically, consider the attention mechanism as a function of position. In a transformer architecture, the attention weight between a query at position and a key at position typically incorporates positional information through either absolute embeddings or relative position biases. As the distance increases, the effective attention capacity diminishes due to several interacting factors.
The first factor is softmax saturation. The attention scores undergo softmax normalization across all positions. In long sequences, the probability mass dilutes across many tokens, reducing the gradient flow to distant positions. Each token in a short context receives a meaningful share of the attention budget; each token in a long context competes with hundreds of others. The second factor is positional embedding degradation. Whether using sinusoidal encodings or learned embeddings, the distinctiveness of positional representations tends to blur for middle positions in very long contexts, creating an information bottleneck where the model struggles to distinguish "token at position 512" from "token at position 513" in a meaningful way. The third factor is context window pressure. Modern large language models often employ attention approximations, such as sparse attention patterns or sliding window attention, which explicitly limit the receptive field of each token to its local neighborhood.
The U-shaped performance curve emerges from the interaction between these architectural constraints and the sequential processing biases inherent in autoregressive generation. Information at the beginning of the context establishes the foundational semantic framework for the model's processing, while information at the end benefits from recency effects in the attention mechanism and proximity to the query tokens. Middle positions suffer from both diminished attention weights and cumulative context drift, where the model's internal state has been transformed by processing preceding tokens, potentially obscuring the relevance of middle content.
The key insight is that the U-shape is not uniform across all models or all context lengths. Smaller models with shorter training context windows show steeper drops in the middle. Newer models trained with techniques specifically designed to improve long-context utilization show shallower drops, but the effect persists even in the most advanced models. You should design your RAG prompt as if the middle will be partially ignored, because even if your specific model handles it better than average, the cost of that assumption being wrong is a hallucinated or grounded-only-on-edge-documents response.
For RAG systems, this creates a strategic dilemma. Beginning placement guarantees the model sees the context early, but may be overshadowed by the final instruction or query when the model generates its response. End placement, placing context just before the query, takes advantage of recency bias but risks truncation if the context is too long. Interleaved placement, weaving context throughout the conversation history, complicates attribution and is rarely worth the implementation complexity. Most production systems use a form of strategic boundary placement that we will explore in the next section.
Strategic Ordering of Retrieved Documents
When you have multiple retrieved passages, their order matters beyond simple position bias. The goal is to arrange documents so that the most important ones occupy the high-attention positions at the beginning and end of the context block, while accepting that documents in the middle will receive less careful processing.
Consider these strategies and their tradeoffs:
-
Relevance-based ordering (Sandwich pattern): Places the highest-scoring document first and the second-highest document last, with lower-confidence documents in the middle. This exploits both primacy and recency effects, sacrificing the marginal documents to ensure the best ones are processed fully. This is the default recommended strategy for most RAG applications.
-
Chronological ordering: Arranges documents by publication date or version, useful when recency matters more than retrieval score. A query about current regulations or recent events may be better served by a slightly less semantically similar but more recent document than by a highly similar but outdated one.
-
Cluster-based grouping: Groups related documents together, reducing the cognitive load on the model to connect dispersed information. This works particularly well when you have retrieved documents that collectively address different facets of a complex question, and you want the model to process each facet as a coherent unit before moving to the next.
-
Authority-based ordering: For domains with hierarchical information authority (legal, medical, regulatory), placing authoritative sources first and supplementary sources later mirrors how human experts read these materials. A Supreme Court ruling before a law review article is not just polite convention; it signals to the model which source to treat as controlling.



The sandwich pattern deserves a closer look because it is the most widely adopted strategy in production systems. The intuition is simple: if you have documents ranked by relevance , you place document 1 at position 1 (primacy), document 2 at position (recency), and fill the middle with documents 3 through in any order. The two most important documents are guaranteed to land in the high-attention zones, and the less important documents absorb the attention deficit of the middle positions. This is a heuristic rather than a formally optimal strategy, but it is well-motivated by the empirical evidence and easy to implement.
The sandwich pattern assumes your relevance ranking is reliable. If your retrieval system returns documents with scores clustered tightly together, say all between 0.75 and 0.80, then the distinction between first and second place is more noise than signal, and you might prefer a different strategy such as random ordering within a relevance tier, or clustering by subtopic so that the model processes related passages together. The best ordering strategy depends on your retrieval quality as much as on model architecture.
Citation Formats and Grounded Generation
A critical requirement for production RAG systems is verifiability. You need to know which retrieved passage supported which claim in the generated response. Without citations, you cannot audit the model's outputs, you cannot build trust with users, and you cannot diagnose failures. This requirement shapes how you instruct the model to cite and how you format the retrieved documents in the first place. The citation format you choose in the context portion of the prompt must match the citation format you request in the output portion, and both must be simple enough that the model can reliably execute them under generation pressure.
Think of the citation system as a contract between you (the prompt author) and the model. You agree to label your sources consistently, and in return you ask the model to reference those labels when it uses the information. The more clearly and consistently you fulfill your side of the contract, the more reliably the model will fulfill its side. Ambiguous labeling, inconsistent formatting across documents, or vague citation instructions break this contract and lead to the hallucinated or omitted citations that plague naive RAG implementations.
The choice of citation format also affects the readability of the output for your end users. A system that produces academic-style footnote references serves a different audience than one that produces inline links or numbered superscripts. It is worth deciding what the output should look like for users before working backward to design the prompt that produces it, rather than starting with whatever is easiest to implement and adjusting later.
Inline Citation Patterns
The most common approach embeds citation markers directly in the context, assigning each document a short identifier:
[Document 1]: The transformer architecture was introduced in 2017.
[Document 2]: Attention mechanisms allow models to focus on relevant parts of the input.
You then instruct the model to cite sources using these markers:
When answering, cite the document numbers that support your claims.
Use the format [Document X] immediately after the relevant information.
This creates a clear mapping between claims and sources, but introduces several challenges. The model must maintain the mapping between bracketed numbers and content across thousands of tokens, and it must learn to use the specific citation syntax you provide. When you have twenty documents, maintaining that mapping under the distributional pressure of language generation becomes error-prone. The model may cite [Document 7] when it means [Document 17], or correctly identify that the claim comes from a specific passage but misremember which number was assigned to it.
A less position-sensitive variant uses content-derived identifiers rather than sequential numbers. If you label your documents with meaningful IDs like [hathitrust_2014] or [campbell_1994], the model can draw on its parametric knowledge of those identifiers to make the association, rather than relying purely on positional counting in the prompt. This works especially well in specialized domains where the model has seen many of these documents during pre-training.
XML and Structured Formats
Systems can also use structured markup that clearly delineates document boundaries:
<sources>
<source id="doc_001" title="Attention Is All You Need" date="2017">
We propose a new simple network architecture, the Transformer,
based solely on attention mechanisms.
</source>
<source id="doc_002" title="Neural Machine Translation" date="2014">
Recurrent neural networks have been used for sequence modeling.
</source>
</sources>
XML-style formatting provides explicit structure that helps the model distinguish document content from its metadata and boundaries. The closing tags create clear delimiters that are less ambiguous than simple newlines or brackets. Modern language models have processed vast amounts of XML during pre-training and handle this structure naturally, often without needing any special instructions about how to interpret the tags.
The metadata attributes within the opening tag serve a dual purpose. They provide the model with provenance information (title, date, source type) that it can use when reasoning about authority and recency, and they give you, the developer, a way to trace exactly which version of a document contributed to a given response. Including a relevance score attribute is particularly useful during debugging, as it lets you verify that the model is preferring high-relevance sources as instructed.
A practical consideration: XML is verbose. Each <source> tag pair consumes tokens, and in a tight context budget, those tokens come at the cost of document content. For very large document sets or very small context windows, you may prefer a lighter delimiter scheme such as triple dashes or numbered headers. The right choice depends on your token budget and how much structural clarity you need.
Citation Instruction Design
The instructions for citation behavior must be precise. Vague instructions lead to inconsistent behavior, where the model sometimes cites, sometimes does not, and occasionally invents citations to documents that do not exist in the prompt. Consider this progression from vague to specific:
Use the provided documents to answer the question.
Answer the question using only the provided documents. For each claim, cite the source document.
Answer the question using only the provided sources. Cite sources for every factual claim using square brackets with the document ID (e.g., [doc_001]). If multiple sources support a claim, cite all of them. If the sources do not contain the answer, respond "I don't have sufficient information to answer this question."
Specificity matters because models are sensitive to instruction following patterns established during fine-tuning. Ambiguous instructions lead to inconsistent citation behavior, either omitting necessary attributions or hallucinating citations to non-existent sources. The specific version eliminates ambiguity about syntax ([doc_001] not doc_001 or (1)), scope (every factual claim, not just major ones), multi-source handling (cite all supporting documents), and fallback behavior (what to say when context is insufficient).
There is a tension between instruction specificity and prompt length. Every additional sentence in your citation instructions consumes tokens that could otherwise hold retrieved content. For applications where every token matters, consider encoding citation instructions as few-shot examples rather than prose descriptions. A single example showing a properly cited response often communicates the pattern more efficiently than three paragraphs explaining it.
Chain-of-Verification for Citations
One technique that significantly improves citation accuracy is instructing the model to verify its citations before finalizing its response. The instruction might look like:
For each citation you include, briefly confirm that the cited document
actually contains the information you attributed to it. If you cannot
confirm, remove the citation and note the uncertainty.
This self-verification step adds output tokens, but it activates a form of retrieval-augmented generation within the generation itself. The model loops back over its claimed citations, checking them against its representation of the source documents, and corrects errors before they reach the user. The technique is most effective with larger models that have sufficient capacity to maintain accurate representations of all source documents simultaneously.
Context Truncation Strategies
Even with efficient retrieval and reranking, you will frequently encounter situations where the combined retrieved documents exceed your model's context window. Context windows have expanded dramatically in recent years, but so has the volume of retrievable content. A system retrieving from a million-document corpus may surface twenty highly relevant passages, each thousands of tokens long, for a total that vastly exceeds even a 128K-token context window. When truncation is necessary, how you truncate matters immensely: a poorly truncated document that cuts off its key sentence is worse than no document at all, because it occupies space while providing no useful information.
Think of truncation as budget allocation under uncertainty. You have a fixed token budget for all retrieved content, and you must distribute it across documents in a way that maximizes the expected relevance of what you keep. Different truncation strategies reflect different assumptions about where useful information lives within a document. The inverted pyramid assumption (information density decreases from beginning to end) favors keeping the head. The conclusions-first assumption (answers live in summaries and conclusions) favors keeping the tail. The query-centric assumption (relevant information can appear anywhere, but clusters around query-similar sentences) favors sliding window extraction.
The mathematical challenge of truncation involves optimizing an information preservation function subject to a constraint on token count. Given a document consisting of tokens , and a maximum token budget , we seek to extract a subsequence that maximizes the relevance function with respect to query , subject to the constraint .
Formally, this is a constrained optimization problem:
where:
- : the extracted subsequence of document tokens, a subset of
- : the semantic relevance of the extracted passage to the query , estimated by embedding similarity or another scoring function
- : the token budget, a hard constraint determined by the context window size minus the space reserved for instructions and the user query
Different truncation strategies represent different approximations of this optimization problem based on assumptions about document structure and information distribution. None of them solve the problem exactly because computing for all possible subsequences is combinatorially intractable. Instead, each strategy uses a structural heuristic to identify a tractable approximate solution.
Head-Only Truncation
The simplest approach keeps the beginning of each document and truncates the end:
# Mock tokenizer for demonstration (replace with transformers tokenizer in production)
class MockTokenizer:
def encode(self, text):
return text.split()
def decode(self, tokens):
return " ".join(tokens)
tokenizer = MockTokenizer()
def truncate_head_only(documents, max_tokens_per_doc):
truncated = []
for doc in documents:
tokens = tokenizer.encode(doc)
truncated.append(tokenizer.decode(tokens[:max_tokens_per_doc]))
return truncatedThis strategy operates on the assumption that information follows an inverted pyramid structure, common in journalistic and academic writing, where the most salient content appears in the introduction. Mathematically, this assumes that the relevance function is monotonically decreasing with position index :
This preserves introductions and thesis statements but loses conclusions and summaries. It works well for academic papers and news articles where the inverted pyramid structure places key information early. It fails badly for technical documentation structured as problem-solution pairs where the solution appears at the end, or for question-answer formatted content where the answer only appears after a lengthy preamble.
Head-only truncation is also computationally simple: it requires no embeddings, no sentence segmentation, and no query awareness. This makes it the default choice for many production systems despite its obvious limitations, because it is the only strategy that adds zero latency. When you are serving thousands of requests per second, even a few milliseconds of additional latency from smarter truncation can matter.
Tail-Only Truncation
Conversely, keeping only the end of documents preserves conclusions and final summaries:
# Mock tokenizer for demonstration (replace with transformers tokenizer in production)
class MockTokenizer:
def encode(self, text):
return text.split()
def decode(self, tokens):
return " ".join(tokens)
tokenizer = MockTokenizer()
def truncate_tail_only(documents, max_tokens_per_doc):
truncated = []
for doc in documents:
tokens = tokenizer.encode(doc)
truncated.append(tokenizer.decode(tokens[-max_tokens_per_doc:]))
return truncatedThis approach inverts the monotonicity assumption, positing that . This suits documents where the final paragraphs contain the synthesis or answer, such as FAQ entries, legal briefs where the conclusion of law appears at the end, or research papers where the results section follows the methodology. The key risk with tail truncation is that it often discards the context necessary to interpret the kept content. A conclusion that says "therefore, the treatment was effective" without the preceding description of what treatment was tested provides little actionable information.
Sliding Window Extraction
For long documents where relevant information might appear anywhere, a sliding window approach extracts segments around query-relevant sentences:
!uv pip install scikit-learn -q
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
# Mock components for demonstration (replace with actual NLTK/spacy/transformers in production)
class MockSentenceSplitter:
def split(self, text):
return [s.strip() for s in text.split('.') if s.strip()]
class MockEmbedder:
def encode(self, sentences):
# Return deterministic dummy embeddings for demonstration
np.random.seed(0)
return np.random.randn(len(sentences), 384)
sentence_splitter = MockSentenceSplitter()
embedder = MockEmbedder()
def extract_relevant_window(document, query_embedding, window_size=512):
sentences = sentence_splitter.split(document)
sentence_embeddings = embedder.encode(sentences)
# Find most similar sentence to query
similarities = cosine_similarity(query_embedding, sentence_embeddings)
center_idx = np.argmax(similarities)
# Extract window around center
start = max(0, center_idx - window_size // 2)
end = min(len(sentences), center_idx + window_size // 2)
return " ".join(sentences[start:end])This approach treats the truncation problem as a center-finding optimization. Given a query embedding vector and sentence embeddings , we calculate cosine similarities:
where:
- : the query embedding vector in
- : the embedding of the -th sentence in the document
- : the Euclidean norm, used to normalize vectors before computing similarity
The optimal center position is determined by:
The extraction window then captures the local context around the maximally relevant sentence, under the assumption that semantic relevance clusters locally in document space. This assumption holds reasonably well for factual documents where a single section discusses any given topic, but breaks down for complex analytical documents where a single sentence might draw on themes established pages earlier.
The sliding window approach requires an additional embedding step at inference time, adding latency proportional to the number of sentences in the document. For documents of a few thousand tokens, this is typically less than 50ms with a fast embedding model. For very long documents, you may want to segment at the paragraph level rather than the sentence level to keep this computation tractable.




Hierarchical Summarization
When individual documents are too long for simple truncation, hierarchical approaches create distillations that preserve more of the semantic content. The idea is to use a fast, cheap model to compress each document before including it in the main prompt, trading some detail for coverage.
The process works in three stages. First, you summarize each long document independently using either an extractive method (selecting the most relevant sentences) or an abstractive method (generating a shorter paraphrase). Second, you use these summaries as the context instead of the full text, allowing you to include information from many more documents within the same token budget. Third, for the top-scoring documents, you may optionally include the full text alongside the summary, giving the model access to fine-grained detail where it matters most.
This approach trades granularity for coverage: you can represent twenty documents instead of five, but each representation is lossy. The most significant risk is that the summarization step introduces its own errors. A summary model that misses a key qualifier, such as "the drug was effective only in patients over 65," can lead the main model to generate a response that omits critical nuance. Summarization before retrieval, at indexing time, is generally more reliable than summarization at inference time because it can be validated and corrected before it affects users.
Instruction Design for RAG
The system prompt or task instruction in a RAG pipeline carries additional responsibility beyond standard task definition. It must establish the relationship between the retrieved context and the generation task, define what counts as an authoritative source, specify citation behavior, and handle edge cases gracefully. A well-designed instruction turns a general-purpose language model into a reliable, attributable research assistant. A poorly designed instruction produces confident-sounding responses that ignore the retrieved context or fabricate plausible-looking citations.
Think of the instruction as a constitution for your RAG system. It specifies the rules that govern how the model treats the retrieved documents: do they take precedence over the model's parametric knowledge? What should the model do when documents contradict each other? What happens when no relevant documents were retrieved? Getting these rules right is what separates a RAG system that can be trusted in production from one that works well on a demo but fails on edge cases.
The instruction also mediates a fundamental tension in RAG design. You want the model to use the retrieved context faithfully, but you do not want it to be so literal that it fails to synthesize, draw inferences, or connect information across documents. A model that only parrots back text from the retrieved passages is not adding value; a model that synthesizes those passages into a coherent, well-reasoned answer is what makes RAG worth the engineering effort.
The Context-Task Contract
Effective RAG instructions establish a clear contract with the model. This contract has four components, each addressing a different failure mode.
The first component is authority: explicitly stating that the retrieved context takes precedence over parametric knowledge prevents the model from substituting its training data when the context says something unexpected or unfamiliar. The second component is scope: defining what the model should do if the context is insufficient prevents hallucination when the retrieval fails to find relevant documents. The third component is constraints: specifying the required format and tone, along with citation rules, ensures the output is usable for downstream processes. The fourth component is fallbacks: providing behavior for contradictory or irrelevant context prevents the model from picking one source arbitrarily or blending conflicting claims without flagging the conflict.
Consider this instruction template:
You are a helpful assistant with access to a knowledge base.
INSTRUCTIONS:
1. Base your answer PRIMARILY on the provided sources below
2. If the sources are insufficient, clearly state what information is missing
3. Cite sources for all factual claims using [source_id] format
4. If sources contradict each other, present both views and explain the disagreement
5. Do not use knowledge outside the provided sources unless specifically asked
SOURCES:
{retrieved_context}
USER QUERY: {query}
Provide a clear, accurate answer following the instructions above.
The numbered list format is intentional. Research on instruction following shows that numbered lists receive more reliable compliance than prose paragraphs, likely because the model has seen many examples of numbered instructions in its training data and has learned to treat each number as a discrete constraint to satisfy. Prose instructions blend together in generation, and the model may satisfy some while forgetting others.
Handling Edge Cases in Instructions
Production RAG systems encounter edge cases that naive prompt designs handle poorly. Addressing these explicitly in the instruction template is far easier than debugging them after deployment.
The empty retrieval edge case occurs when no documents are retrieved above a relevance threshold, typically because the query is out-of-domain or the index does not contain relevant material. Without an explicit instruction, the model will often generate a response from parametric knowledge, making it look like the system found something when it found nothing. The instruction should tell the model to explicitly decline or request clarification: "If no relevant sources were retrieved, respond: 'I could not find information about this topic in the available sources.'"
The contradictory sources edge case occurs when retrieved documents disagree with each other. This is common in rapidly evolving fields, in domains with contested interpretations (law, medicine, policy), and when documents from different time periods coexist in the index. Without an explicit instruction, the model may arbitrarily favor one source, blend the contradictory claims into a false synthesis, or silently ignore one side of the disagreement. The instruction should direct the model to state the contradiction: "If sources disagree, explicitly note the disagreement and cite all conflicting sources."
The outdated information edge case arises when the retrieved documents contain information that was once correct but is no longer current. This is particularly dangerous in fast-moving domains where a two-year-old document about a regulation or a technology product may be significantly misleading. Including timestamp metadata in your source formatting and instructing the model to prefer more recent sources when temporal relevance matters provides a first line of defense, though it does not fully solve the problem.
Worked Example: Legal Document Analysis
Let's walk through a concrete example to see how all these principles come together in a realistic setting. Imagine you are building a RAG system for legal research where users ask about case law precedents.
The query: "What constitutes fair use for educational purposes in the Second Circuit?"
The retrieved documents:
- Campbell v. Acuff-Rose Music (Supreme Court, 1994), 800 tokens discussing transformative use
- NXIVM Corp. v. Ross Institute (2nd Circuit, 2004), 600 tokens on commercial vs. educational use
- Authors Guild v. HathiTrust (2nd Circuit, 2014), 1200 tokens on digital library scanning
- Cambridge University Press v. Patton (11th Circuit, 2016), 900 tokens on e-reserves (a different circuit but relevant)
Total retrieved content: 3500 tokens, which fits comfortably in most context windows. But the documents have heterogeneous authority: two are binding Second Circuit precedent (documents 2 and 3), one is persuasive Supreme Court authority (document 1), and one is persuasive authority from a different circuit (document 4). A naive prompt treats them all equally; a well-designed prompt respects this hierarchy.
Strategy A: Naive Concatenation
The following documents may be relevant:
{doc1}
{doc2}
{doc3}
{doc4}
Question: What constitutes fair use for educational purposes in the Second Circuit?
This strategy has multiple problems. There are no citation markers, so the model cannot attribute claims to specific cases. There is no indication of authority hierarchy, so the model may treat a persuasive 11th Circuit case as equally controlling to Second Circuit precedent. The most relevant documents (2nd Circuit cases) are buried in the middle, right in the lost-in-the-middle zone. And there is no instruction about what to do if the cases conflict.
Strategy B: Structured with Prioritization
You are analyzing legal precedents for fair use in the Second Circuit.
AUTHORITATIVE SOURCES (2nd Circuit precedents - prioritize these):
[Case 1] Authors Guild v. HathiTrust (2014): {first 500 tokens of doc3}
[Case 2] NXIVM Corp. v. Ross Institute (2004): {full doc2}
PERSUASIVE AUTHORITY (Other circuits):
[Case 3] Campbell v. Acuff-Rose (1994): {first 300 tokens of doc1}
[Case 4] Cambridge Univ. Press v. Patton (2016): {first 300 tokens of doc4}
INSTRUCTIONS:
- Answer based primarily on Cases 1 and 2 as controlling precedent
- Cite cases using [Case X] format
- Note circuit boundaries when citing persuasive authority
- If cases conflict, follow the most recent Second Circuit precedent
QUERY: What constitutes fair use for educational purposes in the Second Circuit?
This structure addresses each problem in Strategy A. The most relevant Second Circuit cases are placed at the beginning and end of the sources section (HathiTrust first, NXIVM last), exploiting the primacy and recency zones. Documents are labeled with meaningful identifiers that encode authority level. The instructions explicitly specify the authority hierarchy and citation format. And the fallback behavior for conflicting cases (prefer most recent Second Circuit) removes ambiguity.
The structured version is longer by roughly 100 tokens of overhead (labels, section headers, instructions), but the investment pays for itself in response quality. The model knows exactly which cases are controlling, how to cite them, and what to do when they point in different directions. The result is a response that reads like it was written by a lawyer, not a general-purpose language model.
Numerical breakdown: The total token budget is 4100 (3500 retrieved content + 600 overhead). The sandwich ordering places the 1200-token HathiTrust case (highest relevance, most recent 2nd Circuit precedent) first, which means it lands in the high-attention primacy zone. The 600-token NXIVM case (second-highest relevance, controlling 2nd Circuit precedent) goes last, landing in the recency zone. The 300-token truncated versions of documents 1 and 4 go in the middle, where their lower authority makes the attention deficit less costly. This is the sandwich pattern applied with domain knowledge: the most authoritative and most relevant documents get the premium seats.
Code Implementation: Building a RAG Prompt Template
Let's implement a flexible RAG prompt construction system that incorporates the principles we have discussed throughout this chapter.
from dataclasses import dataclass
from typing import Any, Dict
@dataclass
class RetrievedDocument:
id: str
content: str
score: float
metadata: Dict[str, Any]
def format_xml(self, max_length: int = 1000) -> str:
truncated = self.content[:max_length]
if len(self.content) > max_length:
truncated += "..."
meta_attrs = " ".join([f'{k}="{v}"' for k, v in self.metadata.items()])
return f'<source id="{self.id}" relevance="{self.score:.3f}" {meta_attrs}>\n{truncated}\n</source>'We define a RetrievedDocument class to hold our retrieved content with metadata. The format_xml method creates structured XML representations with truncation. The relevance score is included in the XML attributes so that the model has explicit access to the retrieval confidence for each document, which it can use when deciding how much weight to give each source.
from typing import Dict, List
class RAGPromptBuilder:
def __init__(self, max_context_tokens: int = 4000):
self.max_context_tokens = max_context_tokens
self.system_template = """You are a precise research assistant. Answer using ONLY the provided sources.
INSTRUCTIONS:
1. Cite every factual claim with [source_id] immediately after the claim
2. If multiple sources support a claim, cite all: [id1][id2]
3. If sources are insufficient, state: "The provided sources do not contain sufficient information."
4. When sources disagree, present all viewpoints with citations
5. Prioritize more recent sources when dates differ
AVAILABLE SOURCES:
{sources}
Answer the following question based strictly on the sources above."""
def order_documents(
self, docs: List[RetrievedDocument], strategy: str = "relevance"
) -> List[RetrievedDocument]:
if strategy == "relevance":
# Sort by score descending
sorted_docs = sorted(docs, key=lambda x: x.score, reverse=True)
# Place highest relevance at beginning and end (sandwich pattern)
if len(sorted_docs) > 2:
result = [sorted_docs[0]] # Most relevant first
middle = sorted_docs[2:] # Remaining docs except second highest
result.extend(middle)
result.append(sorted_docs[1]) # Second most relevant last
return result
return sorted_docs
elif strategy == "chronological":
return sorted(
docs, key=lambda x: x.metadata.get("date", ""), reverse=True
)
else:
return docs
def truncate_documents(
self,
docs: List[RetrievedDocument],
truncation_strategy: str = "uniform",
) -> List[str]:
if not docs:
return []
# Estimate tokens (rough heuristic: 4 chars per token)
total_chars = sum(len(d.content) for d in docs)
total_tokens = total_chars // 4
if total_tokens <= self.max_context_tokens:
return [d.format_xml() for d in docs]
if truncation_strategy == "uniform":
# Equal budget for each doc
budget = (self.max_context_tokens * 4) // len(docs)
return [d.format_xml(budget) for d in docs]
elif truncation_strategy == "weighted":
# Allocate by relevance score
total_score = sum(d.score for d in docs)
formatted = []
for doc in docs:
proportion = doc.score / total_score
budget = int(self.max_context_tokens * 4 * proportion)
formatted.append(doc.format_xml(budget))
return formatted
elif truncation_strategy == "head_only":
budget = (self.max_context_tokens * 4) // len(docs)
return [d.format_xml(budget) for d in docs]
else:
return [d.format_xml() for d in docs]
def build_prompt(
self,
query: str,
documents: List[RetrievedDocument],
ordering: str = "relevance",
truncation: str = "weighted",
) -> Dict[str, str]:
# Order documents strategically
ordered_docs = self.order_documents(documents, ordering)
# Truncate to fit context
formatted_sources = self.truncate_documents(ordered_docs, truncation)
sources_text = "\n\n".join(formatted_sources)
# Construct final prompt
system_prompt = self.system_template.format(sources=sources_text)
return {
"system": system_prompt,
"user": f"Question: {query}\n\nProvide a detailed, cited answer.",
}The RAGPromptBuilder class implements document ordering strategies with algorithmic complexity appropriate for production use. The order_documents method with the "relevance" strategy employs a sandwich pattern that operates in time due to the sorting operation, followed by rearrangement. This pattern mathematically optimizes against the U-shaped attention decay curve by positioning the documents with highest relevance scores and at positions and respectively, while placing lower-relevance documents in middle positions where attention fidelity is reduced.
The truncate_documents method implements multiple strategies for handling context overflow. The "weighted" approach solves a resource allocation problem where the token budget is distributed proportionally to relevance scores. For a document with score , the allocated budget is:
where:
- : the token budget allocated to document
- : the total token budget available for all retrieved content
- : the relevance score of document , typically in
- : the sum of all relevance scores, used as a normalizing constant
This ensures that documents with higher relevance receive proportionally more tokens, maximizing the expected information content of the truncated context. The "uniform" strategy instead implements an egalitarian allocation where for all documents, which may be preferable when relevance scores are noisy or when broad coverage is more important than depth for any single document.
The build_prompt method orchestrates the full pipeline with linear complexity relative to document count: ordering documents strategically, truncating to fit context limits, and assembling the final prompt structure with proper XML formatting. The token estimation heuristic, using 4 characters per token, provides a computationally efficient approximation of true token count without requiring expensive tokenizer calls during the budgeting phase.
# Create sample documents
docs = [
RetrievedDocument(
id="doc_001",
content="The Second Circuit in Authors Guild v. HathiTrust held that creating digital copies for search and accessibility constitutes fair use. The court emphasized the transformative nature and public benefit.",
score=0.95,
metadata={"circuit": "2nd", "date": "2014", "type": "case_law"},
),
RetrievedDocument(
id="doc_002",
content="NXIVM Corp. v. Ross Institute established that even commercial use can be fair if sufficiently transformative. The Second Circuit analyzed the four-factor test extensively.",
score=0.88,
metadata={"circuit": "2nd", "date": "2004", "type": "case_law"},
),
RetrievedDocument(
id="doc_003",
content="The Supreme Court in Campbell v. Acuff-Rose Music established that commercial parody can qualify as fair use if transformative. This precedent influences all circuits.",
score=0.72,
metadata={"circuit": "SCOTUS", "date": "1994", "type": "case_law"},
),
RetrievedDocument(
id="doc_004",
content="Recent district court decisions have expanded fair use protections for educational streaming, though these remain controversial and may be overturned.",
score=0.45,
metadata={"circuit": "SDNY", "date": "2023", "type": "case_law"},
),
]
# Build prompt with our system
builder = RAGPromptBuilder(max_context_tokens=1000)
prompt = builder.build_prompt(
query="What constitutes fair use for educational purposes in the Second Circuit?",
documents=docs,
ordering="relevance",
truncation="weighted",
)
# Prepare display strings
system_preview = (
prompt["system"][:1500] + "..."
if len(prompt["system"]) > 1500
else prompt["system"]
)
user_content = prompt["user"]=== SYSTEM PROMPT === You are a precise research assistant. Answer using ONLY the provided sources. INSTRUCTIONS: 1. Cite every factual claim with [source_id] immediately after the claim 2. If multiple sources support a claim, cite all: [id1][id2] 3. If sources are insufficient, state: "The provided sources do not contain sufficient information." 4. When sources disagree, present all viewpoints with citations 5. Prioritize more recent sources when dates differ AVAILABLE SOURCES: <source id="doc_001" relevance="0.950" circuit="2nd" date="2014" type="case_law"> The Second Circuit in Authors Guild v. HathiTrust held that creating digital copies for search and accessibility constitutes fair use. The court emphasized the transformative nature and public benefit. </source> <source id="doc_003" relevance="0.720" circuit="SCOTUS" date="1994" type="case_law"> The Supreme Court in Campbell v. Acuff-Rose Music established that commercial parody can qualify as fair use if transformative. This precedent influences all circuits. </source> <source id="doc_004" relevance="0.450" circuit="SDNY" date="2023" type="case_law"> Recent district court decisions have expanded fair use protections for educational streaming, though these remain controversial and may be overturned. </source> <source id="doc_002" relevance="0.880" circuit="2nd" date="2004" type="case_law"> NXIVM Corp. v. Ross Institute established that even commercial use can be fair if sufficiently transformative. The Second Circuit analyzed the four-fa... === USER PROMPT === Question: What constitutes fair use for educational purposes in the Second Circuit? Provide a detailed, cited answer.
The output shows our structured XML formatting with relevance scores, strategic ordering (highest relevance documents at boundaries), and weighted truncation that allocates more space to higher-scoring documents. Notice how doc_001 (score 0.95) appears first in the sources section and doc_002 (score 0.88) appears last, implementing the sandwich pattern. The weighted truncation gives doc_001 roughly twice as many characters as doc_004. This reflects the 2:1 ratio of their scores.
Key Parameters
The key parameters for the RAG prompt builder form a configuration space that determines the information-theoretic properties of the final context window. These parameters interact in ways that affect both the computational complexity and the semantic fidelity of the retrieval augmentation:
| Parameter | Description | Options | Impact |
|---|---|---|---|
| max_context_tokens | Maximum tokens for retrieved context | 1000-8000+ | Controls total information capacity |
ordering | Document sequence strategy | "relevance", "chronological", "clustered" | Affects attention distribution |
truncation | Token allocation method | "uniform", "weighted", "head_only", "sliding_window" | Determines information preservation |
The parameters interact in non-obvious ways. A large max_context_tokens budget combined with "uniform" truncation may hurt performance if it means every document gets truncated to the same short length, losing the detail in the most relevant documents. A smaller budget with "weighted" truncation may produce better results because the highest-relevance documents receive enough space to present their content coherently. Similarly, "chronological" ordering is only useful when your retrieval scores are not well-calibrated for temporal relevance; if your retriever already down-weights old documents appropriately, chronological ordering just overrides that calibration.
Limitations and Practical Challenges
RAG prompt engineering has fundamental constraints that practitioners must manage carefully. Understanding these limitations is essential for setting realistic expectations, designing appropriate fallback behaviors, and knowing when to invest in additional infrastructure versus accepting the inherent tradeoffs of the approach.
Context Contamination and Attention Dilution
As context length grows, even with careful placement, the fidelity of processing degrades because attention bandwidth is limited. Each additional retrieved document competes for attention weights with every other document and with the query itself. When you include twenty retrieved passages, each one receives, on average, five percent of the attention energy that a single document would receive in isolation. This dilution means that subtle distinctions in source material get smoothed over or lost entirely.

The problem compounds with the nature of retrieved content. Unlike carefully curated training examples, retrieved documents often contain redundant information, boilerplate text, and irrelevant sections. A court case retrieved for its discussion of transformative use might also contain ten pages of procedural history and attorney billing disputes. Without aggressive deduplication and extractive preprocessing, you waste precious attention capacity on repetitive or irrelevant content. This is why the preparation of retrieved content, stripping boilerplate, deduplicating near-identical passages, and normalizing formatting, is often as important as the retrieval quality itself.
The practical implication is to use fewer, better documents rather than many mediocre ones. A RAG system that retrieves 5 high-quality passages will typically outperform one that retrieves 20 mixed-quality passages, because the attention dilution in the 20-document case degrades even the good passages. Set retrieval count limits that reflect your model's effective attention capacity, not just its nominal context window size.
Instruction Override and Parametric Conflicts
Language models possess parametric knowledge acquired during pre-training, and this knowledge can override retrieved context, especially when the retrieval contradicts the model's training data. If your retrieved document states that the capital of Australia is Sydney (incorrect), but the model's training data strongly associates Australia with Canberra, the model may ignore the retrieved text and rely on its parametric knowledge, or worse, blend the two into a confused response.
This parametric override problem is particularly acute in specialized or technical domains where the model has strong prior beliefs from training. A medical RAG system retrieving a document about a newly approved drug dosage may find the model ignoring the correct retrieved dosage because its training data contained a different value. The model's confidence in its parametric knowledge is not calibrated to the date at which its knowledge becomes stale; it simply has a strong association and generates accordingly.
Mitigating this requires explicit instruction design that subordinates parametric knowledge to retrieved context, but this remains imperfect. The model cannot easily "un-know" what it learned during pre-training, and strong factual associations create gravitational wells that pull generated text away from retrieved sources. Techniques like grounding chains of thought ("First, identify the relevant claim in the sources. Then, quote it directly. Then, answer the question based on that quote") can help by forcing the model to explicitly anchor its reasoning in the retrieved text before generating, but they add latency and token cost. Fine-tuning the model on RAG-specific tasks offers the strongest mitigation by teaching it to prefer retrieved context over parametric knowledge in retrieval-augmented settings.
Citation Hallucination
Even with explicit citation instructions, models frequently hallucinate citations, attributing claims to sources that do not support them or citing non-existent document IDs. This stems from the mismatch between the discrete, symbolic nature of citation markers and the continuous, statistical nature of language model generation. The model learns that citations appear after claims, but it does not maintain a perfect lookup table of which source contains which fact during generation. When generation pressure leads to a plausible claim, the model may attach a plausible-sounding citation that turns out to be wrong.
Citation hallucination is particularly insidious because it is hard to detect without post-hoc verification. A response with citations that look syntactically correct may pass a superficial review but contain multiple attribution errors. Users tend to trust cited responses more than uncited ones, which means citation hallucination can increase the harm of model errors by lending them an unwarranted veneer of authority.
Production systems often implement post-hoc verification: using the generated citations to retrieve the claimed sources and checking that they contain the attributed information. This is typically implemented as a secondary pass that runs after generation, extracting all citation markers, looking up the corresponding documents, and verifying that the cited text appears in or is entailed by the source. This adds latency and complexity but significantly improves reliability. For high-stakes applications (legal, medical, financial), this verification step should be considered mandatory rather than optional.
Prompt Engineering at Scale
A subtler limitation is that RAG prompt templates are designed and evaluated on a distribution of queries, but individual queries can fall far outside that distribution. A template optimized for fact-lookup queries may perform poorly when users ask for synthesis, comparison, or multi-hop reasoning that requires connecting information across documents. A template designed for short passages may fail when documents are long and heterogeneous.
This means prompt engineering for RAG is an iterative process: evaluate the system, diagnose failures, and revise the prompt. You need systematic evaluation infrastructure, such as test suites with ground-truth citations, to detect when your prompt template is failing on a class of queries. Building that infrastructure early, before deployment, is far cheaper than discovering failures through user complaints. We will explore evaluation frameworks for RAG systems in the next chapter.
Summary
RAG prompt engineering bridges retrieval and generation, transforming retrieved documents into effective context through careful structural design. The quality of your prompt construction determines whether the retrieved information reaches the model in a form it can use reliably and cite accurately.
Strategic placement combats the lost-in-the-middle problem by positioning the most important documents at the beginning and end of the context window, following the U-shaped attention curve observed in large language models. This guards against a systematic failure mode that affects transformer-based models regardless of their nominal context window size. Designing around the lower attention paid to middle positions is safer than assuming a specific model handles long contexts perfectly.
Structured formatting using XML or clear delimiters helps models distinguish document content from its metadata and boundaries, reducing confusion between consecutive passages. The formatting choice also encodes provenance information: a source ID and date, plus a relevance score that the model can use when reasoning about which sources to trust and how to attribute claims. Consistent formatting within a prompt, using the same structure for every document, makes the model's job easier and reduces citation errors.
Explicit citation instructions establish clear contracts for attribution, though they require specificity and must be paired with verification mechanisms to prevent hallucinated citations. The more precisely you specify the citation format, the conditions under which to cite, and what to do when sources are insufficient, the more reliably the model will follow your instructions. Vague instructions produce vague compliance; specific instructions produce specific, auditable behavior.
Intelligent truncation allocates limited context capacity based on relevance, using strategies like weighted allocation or sliding window extraction to preserve the most query-relevant content. No truncation strategy is universally optimal; the right choice depends on your document corpus structure, your retrieval quality, and your latency budget. Building a system that supports multiple strategies and can be configured per use case is more valuable than optimizing a single strategy.
Edge case handling in system prompts prepares models for empty retrievals, contradictory sources, and authority hierarchies. These edge cases are not rare; in a production system serving diverse queries against a large corpus, they will occur daily. Designing for them upfront, rather than patching them after deployment, is the difference between a reliable system and one that fails in ways that are hard to predict or reproduce.
As we move toward more sophisticated agentic systems in later parts of this book, these RAG prompt engineering techniques form the foundation for systems that retrieve information and synthesize it with tool outputs and reasoning chains. The next chapter explores how we evaluate these complex RAG pipelines to ensure they meet accuracy and reliability standards in production environments.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about RAG prompt engineering.
RAG Prompt Engineering Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!