RAG Evaluation: Metrics for Retrieval and Generation Quality

Michael BrenndoerferJanuary 13, 202565 min read

Part of Language AI Handbook

Covers RAG evaluation with metrics for retrieval quality (Precision@K, NDCG, MRR) and generation faithfulness using the RAGAS framework for AI systems.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

RAG Evaluation

Building a Retrieval-Augmented Generation system involves connecting two complex components: a retriever that fetches relevant context from a knowledge base, and a generator that synthesizes answers from that context. As we explored in earlier chapters on RAG architecture, this two-stage design introduces a unique evaluation challenge. Unlike evaluating a standalone language model or a pure search system, RAG requires us to assess the quality of retrieved documents, the quality of generated text, and how effectively the two components work together. A perfect generator cannot compensate for retrieval that misses necessary information, and flawless retrieval is wasted if the generator hallucinates facts not present in the context.

Think of it like a research team working under a strict deadline. One team member (the retriever) runs to the library and gathers a stack of papers. Another team member (the generator) reads those papers and writes the report. Evaluating only the final report tells you whether the writer is good but nothing about whether the librarian found the right sources. Evaluating only the stack of papers tells you nothing about whether the writer used them correctly. Quality assessment requires examining both tasks and how they interact.

This chapter develops a complete framework for RAG evaluation. We begin with retrieval metrics that measure how effectively our vector search or hybrid retrieval systems surface relevant information. These metrics originate in the information retrieval literature, where researchers have spent decades developing principled ways to measure search engine quality, and we adapt them to the specific constraints of RAG, particularly the limited context window budget. We then examine generation metrics suited for factual synthesis rather than creative writing, explaining why the standard n-gram overlap metrics used in translation and summarization fail for this task and what better alternatives exist.

Finally, we explore end-to-end evaluation approaches that assess whether the system as a whole gives faithful, relevant answers grounded in retrieved evidence. The "RAG Triad" of context relevance, faithfulness, and answer relevance gives a diagnostic framework that lets us pinpoint exactly where a system is failing, whether in retrieval quality, generation integrity, or user-facing utility. Along the way, we implement the RAGAS framework, which uses language models themselves to automate these evaluations at scale, letting the rapid iteration cycles that practical system development requires.

A necessary point deserves emphasis before we dive in: evaluation is an ongoing development tool, not a final validation step. Every design decision in a RAG system, from chunk size to embedding model to retrieval strategy to prompt format, creates tradeoffs that need measurement to optimize. Without rigorous evaluation metrics, you cannot tell which changes help. With them, you can systematically identify bottlenecks and make evidence-based improvements. The investment in evaluation infrastructure pays dividends throughout the system's lifetime.

Historical Context

The metrics we use for RAG evaluation did not emerge from the LLM era but from decades of information retrieval research dating back to the 1960s. Precision and Recall were formalized by Cyril Cleverdon in the Cranfield experiments (1957-1966), which systematically measured indexing effectiveness for scientific literature retrieval. Mean Average Precision became the standard evaluation protocol for TREC (Text REtrieval Conference) competitions starting in 1992. NDCG was proposed by Jarvelin and Kekalainen in 2000 to address the binary nature of earlier metrics. These IR metrics were later adopted by the NLP community for tasks like passage ranking, and they now form the backbone of RAG retrieval evaluation. The generation evaluation side has a different lineage: BLEU emerged from machine translation (Papineni et al., 2002), ROUGE from text summarization (Lin, 2004), and BERTScore from the advent of contextual embeddings (Zhang et al., 2020). RAGAS specifically for RAG evaluation appeared in 2023 (Es et al., 2023). This shows how recently the community recognized that existing metrics were inadequate for this new paradigm.

Evaluating the Retrieval Component

The retrieval stage determines the evidence available to the generator. Dense retrieval, as we covered in earlier chapters, encodes queries and documents into a shared vector space and uses similarity metrics like cosine similarity to find nearest neighbors. But geometric similarity in embedding space does not guarantee semantic relevance for the specific query at hand. A document about a related topic might be a close neighbor without containing the specific information needed. Evaluation metrics for retrieval must account for both whether relevant items appear in the result set and where they appear, since generators typically have limited context windows.

Understanding retrieval performance is foundational because it establishes an upper bound on system capability. If the retrieval component fails to fetch documents containing the answer, even the most advanced generator cannot produce a correct response. This asymmetry matters enormously in practice: you can partially compensate for an imperfect generator with better prompting or fine-tuning, but you cannot compensate for a retriever that fundamentally misses relevant information. Conversely, retrieving relevant documents is necessary but not sufficient; the ranking and presentation of those documents materially influence generation quality. We must therefore evaluate retrieval along multiple dimensions: coverage of relevant information, ranking quality, and the positioning of necessary evidence within the limited context window provided to the generator.

The evaluation of retrieval also requires making decisions about what counts as "relevant." This is more fine-grained than it first appears. A document might contain the answer to a query, which makes it highly relevant. Another might give useful background context that helps the generator reason correctly, which makes it moderately relevant. A third might mention the topic without contributing anything useful. Binary relevance judgments, which classify each document as simply relevant or not, are easier to collect but obscure these important distinctions. Graded relevance judgments offer richer information at the cost of more complex annotation guidelines and higher annotator disagreement. We will cover metrics for both approaches.

Precision and Recall at K

The most intuitive metrics for retrieval quality are Precision@K and Recall@K. These measure the proportion of retrieved items that are relevant, and the proportion of all relevant items that were retrieved, considering only the top KK results. Before diving into formulas, notice what these two metrics capture at an intuitive level. Precision@K asks: "Of the documents I handed to the generator, what fraction were useful?" Recall@K asks: "Of all the useful documents in the knowledge base, what fraction did I manage to find?"

To state this formally, let RKR_K be the set of retrieved documents in the top KK positions, and RrelR_{rel} be the set of all relevant documents for a query in the entire corpus. Then:

Precision@K=∣RK∩Rrel∣∣RK∣\text{Precision@K} = \frac{|R_K \cap R_{rel}|}{|R_K|} Recall@K=∣RK∩Rrel∣∣Rrel∣\text{Recall@K} = \frac{|R_K \cap R_{rel}|}{|R_{rel}|}

where:

  • RKR_K: the set of documents retrieved in the top KK positions
  • RrelR_{rel}: the set of all relevant documents for the query across the full corpus
  • ∣S∣|S|: the cardinality (number of elements) of set SS
  • RK∩RrelR_K \cap R_{rel}: the intersection, meaning documents that were both retrieved and relevant

In RAG systems, KK is typically small, often between 3 and 10, constrained by the generator's context window length and the cost of processing long contexts. High precision ensures the generator receives mostly useful information, while high recall ensures necessary information is not missed. Neither alone is sufficient: a system that always retrieves exactly the one most relevant document has perfect precision but poor recall for complex queries; a system that retrieves every document in the knowledge base has perfect recall but precision near zero.

The key insight is that the context window constraint makes precision especially necessary in RAG, more so than in traditional web search. When you search the web, you can scroll through results and ignore irrelevant ones. When a language model processes a context window, every irrelevant chunk occupies space that could hold necessary evidence, and research suggests that the presence of distracting information can actively degrade generation quality by confusing the model.

These metrics offer complementary perspectives on retrieval effectiveness. Precision@K focuses on the signal-to-noise ratio within the context window. This keeps the generator is not distracted by irrelevant information that might lead to hallucinations or off-topic responses. Recall@K, conversely, measures completeness, showing whether the retrieval system successfully identified all relevant sources in the knowledge base. In practice, these metrics trade off against one another: increasing KK typically improves recall at the expense of precision, as including more documents dilutes the concentration of relevant information.

The Precision-Recall Tradeoff in RAG

Unlike web search where you can scroll through pages, RAG generators process a fixed context window. This makes Precision@K particularly important: irrelevant chunks consume useful tokens that could otherwise accommodate relevant information. However, Recall@K matters for queries requiring information dispersed across multiple documents. The optimal KK depends on the nature of the queries and the density of relevant information in the corpus. For factoid questions with single-document answers, small KK with high precision is ideal. For analytical questions requiring synthesis across sources, larger KK with better recall becomes necessary.

The constraint of limited context windows fundamentally changes how we interpret these metrics compared to traditional information retrieval. In web search, you might tolerate scanning through ten results to find one useful link. In RAG, every irrelevant document actively harms performance by occupying space that could hold necessary evidence. This makes high precision needed, particularly for the earliest positions in the ranking.

Out[3]:
Visualization
Line chart showing precision decreasing and recall increasing as K increases from 1 to 20, with a vertical dashed line at K=5.
Precision@K and Recall@K tradeoff as K increases from 1 to 20, simulated for a scenario with 10 relevant documents out of 100. Precision decreases as more documents are retrieved because later results tend to be less relevant, while recall increases and eventually plateaus when all relevant documents are found. The dashed vertical line at K=5 marks a typical RAG context window cutoff, showing that even at this moderate value, recall is still only partial.

Ranking Metrics: MRR and MAP

Precision and Recall at K treat the retrieved set as unordered, but ranking matters in RAG. Generators tend to focus more on information appearing earlier in the context, a phenomenon known as position bias or the "lost in the middle" effect. Think of the retrieved context as a stack of papers placed in front of the generator: it reads most carefully from the top and may skim or overlook material buried deep in the pile. Ranking metrics account for where relevant items appear, rewarding systems that place the most necessary information near the top.

The position of information within the retrieved context materially impacts generation quality. Research from Liu et al. (2023) demonstrated that language models often pay less attention to content in the middle of long contexts, focusing instead on the beginning and end of the provided documents. This U-shaped attention pattern means that a highly relevant document buried at position eight contributes less to answer quality than the same document at position two. Ranking metrics capture this positional sensitivity, rewarding systems that place the most relevant evidence where the generator will utilize it most effectively.

Understanding the "lost in the middle" effect has direct practical implications for RAG system design. Some practitioners deliberately reorder retrieved documents to place the highest-scored chunks at the very beginning and end of the context window, with less necessary supporting documents in the middle. This "placement strategy" exploits the known attention patterns of language models and can improve answer quality without any changes to the retrieval or generation components themselves. Evaluating systems with ranking-aware metrics ensures you are building this kind of positional awareness into your optimization process.

Out[4]:
Visualization
U-shaped attention curve across context positions 1 through 10, with a dip in performance in the middle positions annotated as 'Middle dip'.
The 'lost in the middle' effect in language model context processing, showing a U-shaped performance curve where information at the beginning and end of the context window is utilized more effectively than information in the middle. The dip centered around position 5-6 is the zone of reduced attention, where relevant documents are more likely to be overlooked by the generator even when present in the retrieved context.

Mean Reciprocal Rank (MRR) measures the position of the first relevant document. For a single query, the reciprocal rank is 1rank\frac{1}{\text{rank}} where rank is the position of the first relevant item. The reciprocal transformation is elegant: if the first relevant document is at position 1, the score is 1.0 (perfect). If it is at position 2, the score is 0.5. At position 4, 0.25. This creates a rapidly diminishing reward for finding relevant material later, which aligns with the generator's position bias described above.

MRR averages the reciprocal rank across all queries in the evaluation set:

MRR=1∣Q∣∑q=1∣Q∣1rankq\text{MRR} = \frac{1}{|Q|} \sum_{q=1}^{|Q|} \frac{1}{\text{rank}_q}

where:

  • ∣Q∣|Q|: the total number of queries in the evaluation set
  • rankq\text{rank}_q: the position (rank) of the first relevant document for query qq, starting from 1
  • 1rankq\frac{1}{\text{rank}_q}: the reciprocal rank for query qq, giving score 1.0 if rank is 1, 0.5 if rank is 2, and so on

If no relevant document appears in the retrieved results for query qq, the reciprocal rank is defined as 0. MRR is appropriate when a single relevant document suffices to answer the query, such as factoid questions with a definitive source. For example, "Who invented the telephone?" requires only finding one document containing the answer; once you have it, additional relevant documents add little value.

MRR gives an intuitive measure of how quickly you can find relevant information. A value of 1.0 indicates that the first retrieved document is always relevant, while a value of 0.5 suggests that the first relevant document typically appears at position 2. This metric is particularly useful for RAG applications involving simple factual lookups where a single authoritative source gives the complete answer. However, MRR has a basic limitation: it only considers the first relevant document. If a query requires synthesizing information from multiple sources, MRR is blind to the quality of later relevant documents.

Mean Average Precision (MAP) extends ranking evaluation to queries with multiple relevant documents. For each query, we calculate Average Precision (AP), which is the mean of precision values calculated at each rank position where a relevant document appears. Think of it this way: every time we encounter a relevant document while scanning down the ranked list, we record our current precision at that position. Then we average all those recorded precision values.

Formally, for a single query:

AP=∑k=1NPrecision@k×rel(k)∣Rrel∣\text{AP} = \frac{\sum_{k=1}^N \text{Precision@k} \times \text{rel}(k)}{|R_{rel}|}

where:

  • NN: the total number of documents retrieved for the query
  • Precision@k\text{Precision@k}: the precision at rank kk, calculated as ∣Rk∩Rrel∣∣Rk∣\frac{|R_k \cap R_{rel}|}{|R_k|}
  • rel(k)\text{rel}(k): a binary relevance indicator that equals 1 if the document at rank kk is relevant, and 0 otherwise
  • ∣Rrel∣|R_{rel}|: the total number of relevant documents for the query (the normalizing constant that makes AP comparable across queries with different numbers of relevant documents)

The rel(k)\text{rel}(k) indicator ensures we only accumulate precision at positions where a relevant document appears. This means that a run of irrelevant documents does not change the running sum, but when a relevant document finally appears at rank kk, we record Precision@k\text{Precision@k} at that moment. The intuition is that each relevant document "votes" for the quality of the ranking at the moment it was found.

MAP then averages AP across all queries in the evaluation set. This metric rewards systems that place multiple relevant documents near the top of the ranking, and it is sensitive to the order in which relevant documents appear. A system that retrieves three relevant documents at positions 1, 2, and 3 will score higher than one that retrieves the same documents at positions 1, 5, and 9, even though both might have identical Precision@10 scores. This sensitivity to ranking quality makes MAP particularly suitable for evaluating RAG systems designed to handle questions requiring synthesis from multiple sources.

MAP offers a more fine-grained view than MRR for complex queries requiring synthesis of multiple sources. Consider a query asking about the causes and effects of a historical event. A system that retrieves three relevant documents at positions 1, 2, and 3 will score higher than one that retrieves the same documents at positions 1, 5, and 9, even though both might have identical Precision@10 scores. This sensitivity to ranking quality makes MAP particularly suitable for evaluating questions with several relevant aspects.

Out[5]:
Visualization
Grouped bar chart comparing MRR, MAP, and NDCG@5 scores across four retrieval scenarios: Perfect, Good Early, Mixed, Poor Early.
Comparison of ranking metrics (MRR, MAP, NDCG@5) across four retrieval scenarios, ranging from perfect early ranking to poor early ranking where all relevant documents are buried. All three metrics decline as relevant documents are placed at lower ranks, but they differ in sensitivity: MRR drops most sharply when the first relevant document is delayed, while MAP and NDCG penalize poor placement of all relevant documents.

Graded Relevance: NDCG

Not all relevant documents are equally relevant. A chunk that directly answers the query is more useful than one that gives tangential background. A document containing a precise statistic is more useful than one that vaguely mentions the same topic. Binary relevance judgments force evaluators to collapse these real distinctions into a single bit, losing information that matters for system quality.

Binary relevance judgments simplify evaluation but obscure important distinctions in information quality. A document containing the exact answer to a query is fundamentally more useful than one that merely mentions related concepts. Think of relevance as a spectrum rather than a switch: at one end, a document that directly and completely answers the query; at the other end, a document that is completely off-topic; and in between, varying degrees of partial relevance, topical relevance, and contextual usefulness. Graded relevance acknowledges these differences, letting evaluators to distinguish between perfect matches, highly relevant documents, marginally relevant background material, and irrelevant content. This richer annotation lets metrics that better reflect the actual utility of retrieved documents.

Normalized Discounted Cumulative Gain (NDCG) handles graded relevance scores (e.g., 0 to 3) rather than binary judgments. The construction of NDCG proceeds in three steps, each building on the previous.

Step 1: Cumulative Gain (CG) at position KK simply sums the graded relevance scores of retrieved documents, ignoring their positions:

CG@K=∑i=1Kreli\text{CG@K} = \sum_{i=1}^K \text{rel}_i

where:

  • reli\text{rel}_i: the graded relevance score (e.g., 0 to 3) of the document at position ii
  • KK: the number of top documents considered in the ranking

CG has an obvious flaw: it assigns equal value to relevant documents regardless of where they appear. A system that returns the most relevant document at position 5 scores identically to one that returns it at position 1. Since position matters for RAG, we need to discount contributions from lower positions.

Step 2: Discounted Cumulative Gain (DCG) penalizes relevant items that appear lower in the ranking using a logarithmic discount factor:

DCG@K=∑i=1K2reli−1log⁡2(i+1)\text{DCG@K} = \sum_{i=1}^K \frac{2^{\text{rel}_i} - 1}{\log_2(i + 1)}

where:

  • reli\text{rel}_i: the graded relevance score of the document at rank ii
  • ii: the position of the document in the ranked list (starting from 1)
  • 2reli−12^{\text{rel}_i} - 1: the exponential gain function, giving exponentially more weight to higher relevance grades
  • log⁡2(i+1)\log_2(i + 1): the logarithmic discount factor that reduces the contribution of documents at lower ranks

The logarithm base ensures that each position contributes less than the previous, but the difference between positions 1 and 2 is larger than between positions 9 and 10. This shows the reality that early positions receive disproportionate attention. We use 2rel−12^{\text{rel}} - 1 to give exponentially more weight to highly relevant documents. With a 0-3 grading scale, the gains are 0, 1, 3, and 7 for grades 0, 1, 2, and 3 respectively. This means a grade-3 document contributes seven times as much gain as a grade-1 document, not just three times as much, which shows the intuition that a perfect answer is substantially more useful than a marginally useful one.

The DCG formula elegantly combines two important principles: the value of relevance and the cost of position. The exponential term in the numerator ensures that finding a highly relevant document (score 3) is substantially more useful than finding a moderately relevant one (score 2), specifically about 7/3≈2.37/3 \approx 2.3 times as useful based on gain values. The logarithmic denominator implements a smooth discounting of lower positions. This shows the diminishing attention that both human users and language models pay to later items in a list.

Step 3: Normalized DCG (NDCG) divides by the Ideal DCG (IDCG), which is the DCG of the perfect ranking where all relevant documents appear at the top in order of decreasing relevance:

NDCG@K=DCG@KIDCG@K\text{NDCG@K} = \frac{\text{DCG@K}}{\text{IDCG@K}}

where:

  • DCG@K\text{DCG@K}: the Discounted Cumulative Gain at position KK for the actual ranking
  • IDCG@K\text{IDCG@K}: the Ideal DCG at position KK, representing the maximum possible DCG if documents were ranked in perfect order of relevance

NDCG ranges from 0 to 1, where 1 is a perfect ranking. The reason is that the normalization by IDCG lets fair comparison across queries of varying difficulty. A query with many highly relevant documents in the corpus will have a high IDCG, while a query with only one relevant document will have a low IDCG. Without normalization, the former would dominate aggregate metrics regardless of retrieval quality. NDCG solves this by measuring the achieved fraction of the theoretically possible gain. This gives a standardized score that shows both the quality and ordering of retrieved documents.

In RAG evaluation, graded relevance is important because context windows are limited: placing a highly relevant chunk at position 8 when it should be at position 1 means it might get truncated or receive less attention from the generator. NDCG captures this cost precisely, which makes it the preferred metric when you have access to graded relevance judgments.

Out[6]:
Visualization
Bar chart with overlaid line plot showing the DCG discount factor 1/log2(i+1) for positions 1 through 10, with annotations at positions 1 and 10.
The logarithmic discount factor in DCG calculation, showing how the gain contribution of each position diminishes as rank increases. Position 1 has a discount factor of 1.00, meaning it contributes full gain, while position 10 has a factor of only 0.29, meaning the same relevance score contributes less than a third as much. This rapid decay shows the reality that users and language models pay decreasing attention to lower-ranked results.

Choosing Between Retrieval Metrics

With Precision@K, Recall@K, MRR, MAP, and NDCG available, the natural question is which to use when. The answer depends on your query types and evaluation goals. For simple factoid questions where a single relevant document suffices, MRR captures the most important information: how quickly does the system return a useful answer? For complex analytical questions requiring multiple sources, MAP or NDCG better reflect whether the system found and ranked all relevant evidence.

If your knowledge base contains documents with varying degrees of relevance (not just relevant or not), NDCG is superior because it uses the full relevance signal. If you only have binary relevance judgments, Precision@K, Recall@K, MRR, and MAP are your options. In practice, many production RAG evaluations report Precision@K and Recall@K alongside MRR or NDCG, giving a complete picture of both coverage and ranking quality.

One practical consideration is the cost of creating relevance judgments. Binary judgments (relevant or not) are faster and cheaper to annotate than graded judgments (0-3). For initial evaluation pipelines, binary judgments with Precision@K and MRR often give enough signal to guide development. Graded judgments and NDCG become more useful once you need to distinguish between systems that already reach decent binary precision and recall but differ in the quality of their ranked order.

Evaluating the Generation Component

Once documents are retrieved, the generator must synthesize them into a coherent answer. Traditional NLP generation metrics like BLEU and ROUGE, which we might recall from machine translation and summarization tasks, have significant limitations when applied to RAG. Understanding why these metrics fail, and what better alternatives look like, helps clarify the basic requirements for evaluating factual synthesis.

The generation component faces a distinct challenge compared to open-ended text generation or translation. Rather than creating original content or creating a predetermined target text, the generator must condense information extracted from the provided context and present it accurately. This is closer to reading comprehension than to creative writing or translation. The evaluation paradigm must therefore prioritize semantic correctness over lexical similarity, and factual grounding over stylistic matching.

Think of evaluating RAG generation as checking whether a student's answer on an open-book exam correctly uses the provided materials. You care about whether the student understood and correctly applied the source material, not whether their answer matches a specific model answer word-for-word. A student who writes "France's capital city is Paris" and one who writes "Paris is the capital of France" have given equally correct answers despite using different words. Standard text generation metrics struggle with this kind of evaluation because they are fundamentally built around lexical similarity to a reference text.

Limitations of N-gram Metrics

BLEU (Bilingual Evaluation Understudy) and ROUGE (Recall-Oriented Understudy for Gisting Evaluation) rely on n-gram overlap between generated and reference texts. BLEU focuses on precision (how many generated n-grams appear in the reference), while ROUGE focuses on recall (how many reference n-grams appear in the generation).

These metrics were developed for scenarios with well-defined target outputs, such as translating a specific sentence or summarizing a specific article. They operate on the assumption that good generated text will closely resemble a reference text at the lexical level, matching specific phrases and word sequences. For these use cases, this is a reasonable assumption: a translation that shares many n-grams with the professional reference translation is likely high quality, and with enough reference translations (BLEU typically uses 4 references), the metric becomes fairly reliable.

These metrics fail for RAG for three specific and important reasons:

  1. Reference dependency: RAG systems often answer open-domain questions where no single reference answer exists. The same factual query can be answered correctly in many ways, using different phrasing, different levels of detail, different amounts of context, and different ways of organizing the information.

  2. Semantic equivalence blindness: Paraphrases like "The treaty was signed in 1919" and "In 1919, the parties signed the agreement" share no bigrams but have identical meanings. N-gram metrics would penalize the second phrasing even though it is equally correct.

  3. Hallucination blindness: N-gram metrics compare against references, not against the retrieved context, so they cannot detect when a model hallucinates information not present in the source documents. A generated answer could reach high BLEU by matching the reference while containing additional fabricated claims that the reference happened not to include.

The reference dependency problem is particularly acute in knowledge-intensive applications. Consider a query about a company's quarterly earnings. A valid answer might state revenue in millions or billions, might round to different decimal places, might present figures in different orders, or might emphasize different metrics, all while being factually correct. No single reference text can capture this valid variation, making n-gram overlap a poor proxy for quality.

The hallucination blindness problem may be the most serious limitation for RAG specifically. The entire motivation for RAG is to ground model outputs in specific, verifiable source documents. If our evaluation metric cannot detect violations of this grounding property, we are measuring the wrong thing entirely. A system that frequently hallucinates but happens to produce fluent text resembling the reference answer could score well on BLEU while being entirely untrustworthy in deployment.

Semantic Similarity Metrics

BERTScore addresses the paraphrase problem by using contextual embeddings. Instead of comparing n-grams, it computes pairwise cosine similarity between token embeddings from the generated and reference texts, using a pre-trained model like BERT or RoBERTa.

BERTScore is a change from lexical matching to semantic matching. By embedding tokens in a high-dimensional vector space where semantically similar words cluster together, it recognizes that "automobile" and "car" are in effect equivalent in most contexts, even though they share no characters. This aligns better with human judgment of text quality, which focuses on meaning rather than exact wording.

For each token tgt_g in the generated text, we find its most similar token trt_r in the reference by computing cosine similarity between their contextual embeddings:

Similarity(tg,tr)=etg⋅etr∥etg∥∥etr∥\text{Similarity}(t_g, t_r) = \frac{\mathbf{e}_{t_g} \cdot \mathbf{e}_{t_r}}{\|\mathbf{e}_{t_g}\| \|\mathbf{e}_{t_r}\|}

where:

  • tgt_g: a token from the generated text
  • trt_r: a token from the reference text
  • etg\mathbf{e}_{t_g}: the contextual embedding vector for token tgt_g (from a pre-trained model like BERT), which encodes the token's meaning in context
  • etr\mathbf{e}_{t_r}: the contextual embedding vector for token trt_r
  • ⋅\cdot: the dot product between two vectors
  • ∥e∥\|\mathbf{e}\|: the Euclidean norm (magnitude) of embedding vector e\mathbf{e}

The cosine similarity ranges from -1 to 1, with values near 1 showing that the two tokens appear in similar contexts in the training data, suggesting they carry similar meanings. BERTScore then builds precision, recall, and F1 scores based on these token-level similarities, following the same structure as ROUGE but using embedding similarity instead of exact overlap.

BERTScore's precision score assigns each generated token to its most similar reference token and averages those maximum similarities. The recall score does the reverse, assigning each reference token to its most similar generated token. The F1 score is the harmonic mean of these two. This gives an overall semantic similarity measure.

While better than n-gram metrics for capturing semantic equivalence, BERTScore still requires reference answers and does not verify factual grounding in retrieved contexts. The cosine similarity calculation measures the angle between embedding vectors, with values near 1 showing that tokens appear in similar contexts in the training data and thus likely share meanings. This allows BERTScore to recognize synonyms, hypernyms, and paraphrases that n-gram metrics would miss. However, the limitation remains that BERTScore only compares the generated text to a reference, not to the source documents that should ground the answer. A generated text could reach high BERTScore by closely matching a reference while containing hallucinated details or ignoring the retrieved context entirely.

Factual Consistency Metrics

For RAG, we need metrics that compare generated text against the retrieved context, not just against reference answers. This is a fundamentally different evaluation task: instead of asking "does this answer resemble the gold standard answer?", we ask "does this answer accurately stand for what the source documents say?"

QuestEval and similar question-answering-based approaches use question generation and answering to verify consistency between generated text and source documents. The process works as follows: first, generate specific factual questions from the candidate answer (for example, "When was the treaty signed?" from "The treaty was signed in 1919"). Second, attempt to answer those questions using only the retrieved context. Third, check whether the answers obtained from the context match the original claims in the generated text.

If the context cannot support the claims made in the answer, the system flags a potential hallucination or faithfulness violation. This approach avoids the need for reference answers by treating evaluation as a consistency checking task. By generating specific questions about factual claims, then attempting to answer those questions using only the retrieved documents, the system verifies whether the evidence supports the conclusions drawn by the generator.

The question-answering framework has strong theoretical appeal: it is asking, "could a reasonable reader extract this information from the retrieved sources?" If the answer is no, the generated claim is either a hallucination or an unwarranted inference. If yes, the claim is grounded. This framing aligns closely with how domain experts would evaluate answer quality in a fact-checking context.

A related approach uses Natural Language Inference (NLI) models directly. NLI models are trained to classify whether a "hypothesis" sentence is entailed by (logically follows from), contradicted by, or neutral with respect to a "premise" sentence. For faithfulness evaluation, the premise is a retrieved context passage and the hypothesis is a claim extracted from the generated answer. If all claims in the answer are entailed by at least one context passage, the answer is faithful; if any claims are contradicted, the answer contains errors.

End-to-End Evaluation

The ultimate goal is evaluating whether the RAG system as a whole answers your queries accurately and faithfully. This requires three distinct dimensions of assessment, often called the "RAG Triad": context relevance, faithfulness (or attribution), and answer relevance. No single dimension captures the full picture. A system could excel at all three independently designed components while still failing users if those components do not integrate effectively.

End-to-end evaluation recognizes that RAG systems are more than the sum of their parts. A retrieval system might reach perfect precision and recall on benchmark queries, and a generator might produce fluent text on all test cases, yet the combination might fail to answer user needs because the retrieval and generation are tuned to different distributions or because the prompt format does not effectively communicate the retrieved context to the generator. Conversely, a system with imperfect retrieval might still succeed if the generator can synthesize partial information effectively. We must therefore evaluate whether the integrated pipeline produces useful answers that are accurate and supported by evidence.

The RAG Triad also is a diagnostic tool. When a system fails to produce a satisfactory answer, the triad helps identify which component is responsible. Low context relevance points to retrieval failures: the system is not finding the right information. Low faithfulness points to generation failures: the system is not using the retrieved information correctly, instead fabricating or misrepresenting it. Low answer relevance points to alignment failures: the system may be giving accurate and grounded information, but it is not the information the user requested. Each failure mode suggests a different remediation strategy. This makes the triad a practical framework for targeted improvement.

Context Relevance

Context Relevance measures whether the retrieved chunks contain information necessary to answer the query. Think of it as asking: "Did the librarian bring back books that address my question?" Unlike retrieval precision, which requires human judgments of relevance collected before seeing the generated answer, context relevance for RAG can be assessed by checking whether the generator uses the retrieved information.

Formally, we can measure context relevance through attribution: for each sentence in the generated answer, we determine if it can be attributed to specific sentences in the retrieved context. High context relevance means retrieved documents contributed to the answer; low context relevance suggests the documents were off-topic or the generator ignored them in favor of parametric knowledge from its training data.

This attribution-based approach to context relevance has an important implication: it evaluates whether the retrieved documents contain useful information and whether that information appears in the generated answer. A retrieval system might return a highly relevant document, but if the generator's context window format, prompt design, or attention patterns cause it to ignore that document, context relevance will still be low. Context relevance therefore covers both retrieval quality and generator-retrieval integration, while retrieval precision measures only the former.

Attribution analysis gives a mechanistic understanding of how the generator interacts with retrieved documents. By tracing specific claims in the output back to specific passages in the input, we can determine whether the retrieval component successfully identified useful sources. This is particularly important for debugging: if the generator produces correct answers but shows low attribution to retrieved documents, it may be relying on parametric knowledge rather than the provided context. This creates a false sense of reliability that will fail when the knowledge base contains more recent or specific information than the model's training data.

Faithfulness and Attribution

Faithfulness (also called groundedness or attribution) measures whether the generated answer is factually consistent with the retrieved context. A faithful answer contains only information supported by the evidence, while an unfaithful answer hallucinates facts or contradicts the context.

Evaluating faithfulness requires claim-level verification. Modern approaches use entailment models or LLM prompting to check if claims in the answer are entailed by, contradicted by, or neutral to the retrieved passages. For a claim cc and context passage pp, an NLI model predicts one of three categories:

  • Entailment: pp supports cc (the claim follows logically from the passage)
  • Contradiction: pp contradicts cc (the passage asserts something incompatible with the claim)
  • Neutral: pp gives no information about cc (the passage is unrelated to the claim)

A faithful answer should have all claims entailed by the union of retrieved passages. Any claims classified as contradicted stand for direct factual errors. Claims classified as neutral are ambiguous: they might be general knowledge not present in the corpus, or they might be hallucinations. Many evaluation frameworks flag neutral claims as potential faithfulness violations and require human review for ambiguous cases.

Faithfulness is the cornerstone of trustworthy RAG systems. Unlike general language models that might generate plausible-sounding but unverified information, RAG systems promise to ground their outputs in retrievable evidence. Violations of faithfulness break this contract, creating hallucinations that appear authoritative because they are presented alongside real documents. Natural Language Inference models are particularly well-suited for this task because they are trained specifically to recognize textual entailment relationships, distinguishing between cases where one text supports, contradicts, or is unrelated to another.

The practical difficulty of faithfulness evaluation lies in claim extraction. Before we can verify claims, we must identify the claims made by the generated answer. Some claims are explicit and easily extracted: "The company was founded in 1995." Others are implicit: an answer that says "the technique was later improved" implicitly claims that an earlier version existed. Thorough faithfulness evaluation requires identifying both explicit and implicit claims, which is non-trivial for complex multi-sentence answers. LLM-based claim extractors handle this better than rule-based approaches, but they introduce their own potential for errors and inconsistencies.

Answer Relevance

Answer Relevance (or answer usefulness) measures whether the generated response addresses your query. A system could retrieve perfect context and cite it faithfully, yet fail to answer the question asked. This failure mode is more common than it might seem: language models sometimes respond to the topic of a question rather than the specific information requested. This gives general background when a specific answer is needed.

This often happens with overly cautious models that give relevant background information without the specific answer requested. Consider a query asking for the CEO of a company. A system might retrieve relevant documents and faithfully report that "The company was founded in 1990 and has grown to employ 50,000 people worldwide," all while never stating who the current CEO is. The answer is faithful and context-relevant but scores zero on answer relevance.

Answer relevance can be evaluated through several approaches:

  1. Checking if the answer contains the specific information type requested (a date, a name, an explanation, a comparison).
  2. Measuring semantic similarity between the query and answer, with the intuition that a relevant answer should address the same concepts as the question.
  3. Using LLM judges to assess whether the answer responds to the question, drawing on the model's understanding of pragmatic discourse structure.

An elegant automated approach from RAGAS works by reverse-engineering the query. Given the generated answer, prompt an LLM to generate questions that the answer would appropriately respond to. Then measure the semantic similarity between these reverse-engineered questions and the original query. If the answer is truly relevant, the generated questions should closely resemble the original query. If the answer talks around the question without addressing it, the generated questions will diverge from the original.

Answer relevance captures the pragmatic dimension of system performance. A query asking for the current stock price of a company might receive a response that explains the company's business model, recent news, and market position, all accurately cited from retrieved documents. Such a response might score perfectly on faithfulness and context relevance while completely failing on answer relevance because it never states the stock price. This metric ensures that the system remains focused on user intent rather than merely generating related content.

The RAGAS Framework

Manually evaluating the RAG Triad for thousands of queries is impractical. The RAGAS (Retrieval-Augmented Generation Assessment) framework automates RAG evaluation using LLMs as judges. This gives metrics that approximate human judgment without requiring reference answers for most metrics. RAGAS was introduced by Es et al. in 2023 specifically to address the evaluation bottleneck in RAG development, recognizing that existing evaluation tools from NLP were designed for different tasks.

The automation of evaluation is important for practical RAG development. As systems iterate through different retrieval models, chunking strategies, embedding models, and prompt templates, developers need rapid feedback on which changes improve performance. Manual evaluation creates a bottleneck that slows iteration; automated metrics enable systematic optimization and regression testing. RAGAS uses the reasoning capabilities of large language models to simulate human judgment at scale, turning the evaluation problem into a collection of natural language tasks that LLMs can perform with reasonable accuracy.

The key insight behind RAGAS is that if you have a model capable enough to evaluate RAG system outputs, you can use it as an automated judge. This requires the judge model to be more capable than the system being evaluated, and it requires careful prompt design to minimize judge biases. When these conditions hold, automated evaluation can scale to hundreds of thousands of query-answer pairs, letting statistical analyses that would be impossible with human annotation.

Automated Metric Computation

RAGAS defines four key metrics, each targeting a specific dimension of system quality and computed via carefully designed LLM prompts.

Context Precision measures the signal-to-noise ratio in retrieved chunks. For each chunk in the top-KK results, an LLM is asked whether it contains information relevant to answering the query. Precision is then the fraction of chunks judged relevant, computed as a weighted average that accounts for chunk position. The prompts ask the judge model to make binary relevance judgments for each chunk independently, then aggregate these judgments into a single score.

Context Recall estimates coverage by asking an LLM to identify which sentences from a ground-truth answer are supported by the retrieved context. It is the fraction of ground-truth sentences that can be attributed to at least one retrieved chunk. This is the one RAGAS metric that requires reference answers, which makes it more expensive to compute but also more reliable as a measure of retrieval completeness. The distinction between context precision and context recall mirrors the precision-recall tradeoff we discussed earlier, but here both are measured from the perspective of informational coverage rather than document-level judgments.

Faithfulness uses an LLM to extract factual claims from the generated answer, then for each claim, determine if it can be inferred from the retrieved context. The score is the ratio of supported claims to total claims:

Faithfulness=Number of supported claimsTotal number of claims\text{Faithfulness} = \frac{\text{Number of supported claims}}{\text{Total number of claims}}

where:

  • Number of supported claims\text{Number of supported claims}: the count of factual claims in the generated answer that are entailed by the retrieved context
  • Total number of claims\text{Total number of claims}: the total count of factual claims extracted from the generated answer by the LLM judge

A faithfulness score of 1.0 means every claim in the answer is grounded in the retrieved context. A score of 0.5 means half the claims are hallucinated or unverifiable from the context. This metric does not require reference answers because it only checks consistency between the generated answer and the retrieved context.

Answer Relevancy evaluates how directly the answer addresses the question. The LLM generates several potential questions that the given answer would appropriately respond to, then measures the mean cosine similarity between these generated questions and the original query using sentence embeddings. Low similarity indicates the answer is off-topic or incomplete. This metric also does not require reference answers, which makes it suitable for open-domain evaluation.

The metric definitions reflect a careful decomposition of system quality. Context Precision focuses on retrieval efficiency. This keeps limited context windows are not wasted. Context Recall addresses completeness, verifying that the retrieval system found all necessary information. Faithfulness checks the integrity of the generation process, while Answer Relevancy confirms utility to the end user. Together, these metrics give a complete dashboard for system health without requiring expensive human annotation for most queries.

Implementation Strategy

RAGAS operates through a carefully orchestrated pipeline of LLM calls. For faithfulness evaluation, the process involves three distinct stages that transform a raw query-answer-context triple into a numerical score.

In the first stage, claim extraction, the system prompts the LLM with the generated answer and asks it to decompose the answer into atomic factual statements. An atomic claim is a minimal, self-contained factual assertion: "The Eiffel Tower is located in Paris" rather than "The Eiffel Tower, which is located in Paris, was built in 1889 and stands 330 meters tall." Breaking compound statements into atomic claims ensures that each claim can be independently verified against specific context sentences.

In the second stage, claim verification, the system presents each extracted claim alongside the full retrieved context and asks the LLM to classify whether the context supports, contradicts, or does not address the claim. This requires careful prompt design to prevent the judge model from hallucinating its own judgments or being influenced by the fluency of the generated answer rather than its grounding.

In the third stage, score aggregation, the system counts the fraction of verified (supported) claims and reports this as the faithfulness score. When implementing this pipeline, you should log intermediate results (extracted claims, individual verdicts) to let debugging when scores seem incorrect or inconsistent.

This approach assumes the LLM judge is more capable than the generator being evaluated, and that it can accurately assess entailment without hallucinating its own judgments. The implementation relies on the zero-shot reasoning capabilities of modern language models. By carefully crafting prompts that instruct the model to perform specific evaluation subtasks, RAGAS turns the evaluation problem into a series of classification and generation tasks that capable LLMs perform reliably.

The dependency on LLM judges introduces considerations about judge quality: the evaluating model should ideally be more capable than the system being evaluated, and prompts must be designed to minimize the judge's own biases and hallucinations. Practically, many teams use a stronger model (e.g., GPT-4) to evaluate outputs from a smaller model (e.g., GPT-3.5 or Llama), though this creates cost considerations that must be balanced against the evaluation budget.

Out[7]:
Visualization
Polar radar chart comparing RAGAS scores for Query A, B, and C across four axes: Context Precision, Context Recall, Faithfulness, and Answer Relevancy.
RAGAS evaluation scores for three sample queries displayed as a radar chart across four dimensions: Context Precision, Context Recall, Faithfulness, and Answer Relevancy. Query A shows a well-balanced, high-performing profile. Query B shows a performance gap specifically in context precision, suggesting that irrelevant chunks are being retrieved despite good generation quality. Query C shows materially lower faithfulness and answer relevancy, showing that the generator is creating off-topic or unsupported content.

Worked Example: Step-by-Step Metric Calculation

Let us walk through a complete concrete example that calculates every metric we have discussed, showing each numerical step. This kind of hands-on calculation is the best way to develop intuition for how the metrics behave and where they agree or disagree in their verdicts.

Setup: A user asks "What carbon pricing mechanism did Canada implement in 2023?"

Our retriever returns five chunks, ranked in order:

  • Chunk 1 (rank 1): "The European Union's carbon border adjustment mechanism affects imports from countries without carbon pricing..." (Irrelevant to Canada 2023)
  • Chunk 2 (rank 2): "Canada implemented a federal carbon pricing system in 2019 that applies a charge on fossil fuels..." (Relevant but discusses 2019, not 2023)
  • Chunk 3 (rank 3): "In 2023, Canada introduced a carbon contract for difference program to guarantee carbon prices for industrial emitters..." (Highly relevant: directly answers the question)
  • Chunk 4 (rank 4): "Carbon pricing mechanisms vary across jurisdictions, including carbon taxes, cap-and-trade, and contracts for difference..." (Marginally relevant: explains the mechanism type but not specific to Canada)
  • Chunk 5 (rank 5): "Canadian provinces have implemented their own clean fuel standards separate from federal carbon pricing..." (Marginally relevant: related to Canada's climate policy generally)

We assign graded relevance: Chunk 1 = 0, Chunk 2 = 2, Chunk 3 = 3, Chunk 4 = 1, Chunk 5 = 1. Binary relevance: Chunks 2, 3, 4, 5 are relevant (grades > 0), Chunk 1 is irrelevant.

Step 1: Precision@K and Recall@K

At K=3K=3: retrieved set is {Chunk1, Chunk2, Chunk3}, relevant set in top 3 is {Chunk2, Chunk3} (2 items).

Precision@3=∣{C2,C3}∣3=23≈0.667\text{Precision@3} = \frac{|\{C2, C3\}|}{3} = \frac{2}{3} \approx 0.667

Total relevant items in corpus (grades > 0): 4. Relevant items in top 3: 2.

Recall@3=24=0.5\text{Recall@3} = \frac{2}{4} = 0.5

At K=5K=5: all 5 chunks retrieved, 4 are relevant.

Precision@5=45=0.8Recall@5=44=1.0\text{Precision@5} = \frac{4}{5} = 0.8 \qquad \text{Recall@5} = \frac{4}{4} = 1.0

Notice the tradeoff: increasing KK from 3 to 5 improved recall from 0.5 to 1.0 and also improved precision from 0.667 to 0.8 in this example (because chunks 4 and 5 happen to be relevant). In general, precision would be expected to decrease or stay flat as K increases.

Step 2: MRR

The first relevant chunk is at rank 2 (Chunk 2 is relevant with grade 2).

MRR=1rank1=12=0.5\text{MRR} = \frac{1}{\text{rank}_1} = \frac{1}{2} = 0.5

An MRR of 0.5 shows that the system placed an irrelevant EU document at the top position, making the user "wait" until position 2 to find something useful.

Step 3: NDCG@5

Using the graded relevance scores [0, 2, 3, 1, 1] for ranks 1 through 5:

DCG@5=20−1log⁡2(2)+22−1log⁡2(3)+23−1log⁡2(4)+21−1log⁡2(5)+21−1log⁡2(6)=01+31.585+72+12.322+12.585=0+1.893+3.5+0.431+0.387=6.211\begin{aligned} \text{DCG@5} &= \frac{2^0 - 1}{\log_2(2)} + \frac{2^2 - 1}{\log_2(3)} + \frac{2^3 - 1}{\log_2(4)} + \frac{2^1 - 1}{\log_2(5)} + \frac{2^1 - 1}{\log_2(6)} \\ &= \frac{0}{1} + \frac{3}{1.585} + \frac{7}{2} + \frac{1}{2.322} + \frac{1}{2.585} \\ &= 0 + 1.893 + 3.5 + 0.431 + 0.387 \\ &= 6.211 \end{aligned}

The ideal ranking places scores in descending order: [3, 2, 1, 1, 0].

IDCG@5=23−1log⁡2(2)+22−1log⁡2(3)+21−1log⁡2(4)+21−1log⁡2(5)+20−1log⁡2(6)=71+31.585+12+12.322+02.585=7+1.893+0.5+0.431+0=9.824\begin{aligned} \text{IDCG@5} &= \frac{2^3 - 1}{\log_2(2)} + \frac{2^2 - 1}{\log_2(3)} + \frac{2^1 - 1}{\log_2(4)} + \frac{2^1 - 1}{\log_2(5)} + \frac{2^0 - 1}{\log_2(6)} \\ &= \frac{7}{1} + \frac{3}{1.585} + \frac{1}{2} + \frac{1}{2.322} + \frac{0}{2.585} \\ &= 7 + 1.893 + 0.5 + 0.431 + 0 \\ &= 9.824 \end{aligned} NDCG@5=6.2119.824≈0.632\text{NDCG@5} = \frac{6.211}{9.824} \approx 0.632

The NDCG of 0.632 tells us the system achieved 63.2% of the theoretically achievable DCG. The main losses came from placing the highly relevant Chunk 3 at rank 3 instead of rank 1, and placing the irrelevant EU document at rank 1.

Step 4: End-to-End Metrics

The generator, using all five chunks, produces: "Canada introduced a carbon contract for difference program in 2023 to give price certainty for industrial carbon emitters. This builds on Canada's earlier carbon pricing system established in 2019."

We evaluate the claims in this answer:

  • Claim 1 ("Canada introduced a carbon contract for difference program in 2023"): supported by Chunk 3 (direct match).
  • Claim 2 ("to give price certainty for industrial carbon emitters"): supported by Chunk 3 ("guarantee carbon prices for industrial emitters").
  • Claim 3 ("This builds on Canada's earlier carbon pricing system established in 2019"): supported by Chunk 2 ("Canada implemented a federal carbon pricing system in 2019").

All three claims are supported. Faithfulness = 3/3 = 1.0.

Context Precision at K=5K=5: 4 of 5 chunks were relevant (all except Chunk 1). Context Precision = 4/5 = 0.8.

Answer Relevancy: The answer directly addresses the mechanism (carbon contracts for difference) and the year (2023). Answer Relevancy would be high, approximately 0.9 on a 0-1 scale.

Synthesis: This example reveals a fine-grained picture. Retrieval quality is suboptimal (MRR = 0.5, NDCG = 0.632) because the most relevant document was placed at rank 3 and an irrelevant document at rank 1. Yet the generator achieved perfect faithfulness and high answer relevancy by correctly identifying and using the right information from the context. This shows how reliable generation can partially compensate for imperfect retrieval, though ideally both components would perform optimally. In a production system, the low MRR and suboptimal NDCG would motivate improvements to the retrieval component, particularly better discrimination between the EU carbon mechanism (irrelevant) and Canada-specific content.

Code Implementation

Let us implement these evaluation metrics from first principles, then show the RAGAS framework approach. We will create a minimal RAG evaluation pipeline and compute all the metrics discussed.

First, we establish our imports and set up the environment:

Now we define the retrieval metrics. We implement Precision@K, Recall@K, MRR, and MAP as standalone functions that operate on lists of document identifiers:

In[9]:
Code
from typing import List, Set


def precision_at_k(retrieved: List[str], relevant: Set[str], k: int) -> float:
    """Calculate Precision@K for binary relevance judgments."""
    if k == 0:
        return 0.0
    retrieved_k = set(retrieved[:k])
    relevant_retrieved = retrieved_k.intersection(relevant)
    return len(relevant_retrieved) / len(retrieved_k)


def recall_at_k(retrieved: List[str], relevant: Set[str], k: int) -> float:
    """Calculate Recall@K."""
    if not relevant:
        return 0.0
    retrieved_k = set(retrieved[:k])
    relevant_retrieved = retrieved_k.intersection(relevant)
    return len(relevant_retrieved) / len(relevant)


def mean_reciprocal_rank(
    retrieved_list: List[List[str]], relevant_list: List[Set[str]]
) -> float:
    """Calculate MRR across multiple queries."""
    rr_sum = 0.0

    for retrieved, relevant in zip(retrieved_list, relevant_list):
        rank = None
        for i, doc in enumerate(retrieved, 1):
            if doc in relevant:
                rank = i
                break
        if rank:
            rr_sum += 1.0 / rank

    return rr_sum / len(retrieved_list) if retrieved_list else 0.0


def average_precision(retrieved: List[str], relevant: Set[str]) -> float:
    """Calculate Average Precision for a single query."""
    if not relevant:
        return 0.0

    hits = 0
    sum_precisions = 0.0

    for i, doc in enumerate(retrieved, 1):
        if doc in relevant:
            hits += 1
            sum_precisions += hits / i

    return sum_precisions / len(relevant) if relevant else 0.0


def mean_average_precision(
    retrieved_list: List[List[str]], relevant_list: List[Set[str]]
) -> float:
    """Calculate MAP across multiple queries."""
    ap_scores = [
        average_precision(ret, rel)
        for ret, rel in zip(retrieved_list, relevant_list)
    ]
    return np.mean(ap_scores) if ap_scores else 0.0

The MAP implementation illustrates the key design decision: we accumulate precision at each relevant document position and normalize by the total number of relevant documents (not the number retrieved). This means systems that miss relevant documents are penalized even if their retrieved set has perfect precision.

Next, we implement NDCG for graded relevance judgments. This requires relevance scores rather than binary labels:

In[10]:
Code
from typing import Dict, List


def dcg_at_k(relevance_scores: List[float], k: int) -> float:
    """Calculate Discounted Cumulative Gain."""
    dcg = 0.0
    for i, rel in enumerate(relevance_scores[:k], 1):
        dcg += (2**rel - 1) / np.log2(i + 1)
    return dcg


def ndcg_at_k(
    retrieved: List[str], relevance_map: Dict[str, float], k: int
) -> float:
    """Calculate Normalized DCG@K."""
    # Get relevance scores for retrieved items (0 if not in map)
    retrieved_scores = [relevance_map.get(doc, 0.0) for doc in retrieved[:k]]

    # Calculate actual DCG
    actual_dcg = dcg_at_k(retrieved_scores, k)

    # Calculate ideal DCG (perfect ranking of all relevant documents)
    ideal_scores = sorted(relevance_map.values(), reverse=True)[:k]
    ideal_dcg = dcg_at_k(ideal_scores, k)

    return actual_dcg / ideal_dcg if ideal_dcg > 0 else 0.0

The ndcg_at_k function handles an important edge case: if no relevant documents exist (IDCG is zero), we return 0 to avoid division by zero. The ideal ranking uses all relevant documents across the full corpus, not just those that were retrieved. This keeps systems missing relevant documents are penalized.

Now let us create a simulated RAG evaluation dataset and compute these metrics:

In[11]:
Code
# Simulated evaluation data for the Canada carbon pricing example
# Retrieved document IDs (in order of retrieval rank)
retrieved_docs = [
    "doc_eu_carbon",  # Rank 1: Irrelevant (EU, not Canada)
    "doc_canada_2019",  # Rank 2: Relevant (Canada 2019 system)
    "doc_canada_2023",  # Rank 3: Highly relevant (Canada 2023 contract)
    "doc_carbon_mechanisms",  # Rank 4: Marginally relevant
    "doc_canada_provinces",  # Rank 5: Marginally relevant
]

# Binary relevance judgments
relevant_binary = {
    "doc_canada_2019",
    "doc_canada_2023",
    "doc_carbon_mechanisms",
    "doc_canada_provinces",
}

# Graded relevance (0-3 scale)
relevance_graded = {
    "doc_canada_2019": 2.0,  # Relevant but not 2023-specific
    "doc_canada_2023": 3.0,  # Directly answers the question
    "doc_carbon_mechanisms": 1.0,  # Background context only
    "doc_canada_provinces": 1.0,  # Related but not answering
}
Out[12]:
Console
K=3: Precision=0.667, Recall=0.500, NDCG=0.574
K=5: Precision=0.800, Recall=1.000, NDCG=0.632
MRR: 0.500

We can see that the MRR of 0.5 shows the irrelevant EU document at rank 1, while NDCG captures both this ranking problem and the sub-optimal placement of the most relevant document at rank 3. These metrics together identify the retrieval weaknesses more precisely than any single number could.

Now we implement a simplified faithfulness checker using an entailment heuristic. In practice, this would use NLI models or LLM prompting, but the word-overlap heuristic shows the core logic:

In[13]:
Code
def extract_claims(answer: str) -> List[str]:
    """Simple sentence segmentation to extract claims."""
    import re

    sentences = re.split(r"(?<=[.!?])\s+", answer)
    return [s.strip() for s in sentences if len(s.strip()) > 10]


def check_entailment_simple(claim: str, context: str) -> bool:
    """Simplified entailment check using word overlap.

    In production, replace with an NLI model or LLM judge.
    Word overlap is a rough proxy that works for clear entailment
    and clear contradiction but struggles with paraphrases.
    """
    claim_words = set(claim.lower().split())
    context_words = set(context.lower().split())

    if not claim_words:
        return False

    overlap = len(claim_words.intersection(context_words))
    return overlap / len(claim_words) > 0.7  # 70% word overlap heuristic


def faithfulness_score(answer: str, contexts: List[str]) -> float:
    """Calculate faithfulness as ratio of supported claims."""
    claims = extract_claims(answer)
    if not claims:
        return 0.0

    supported = 0
    for claim in claims:
        # Check if claim is supported by ANY context chunk
        if any(check_entailment_simple(claim, ctx) for ctx in contexts):
            supported += 1

    return supported / len(claims)
Out[14]:
Console
Faithful answer score: 1.00
Hallucinated answer score: 0.00

The faithfulness checker correctly distinguishes between these cases: the faithful answer shares key terms ("Paris", "capital", "France", "987") with the context, while the hallucinated answer claims 1789 (a date not present in the context). In practice, the 70% word overlap threshold would be replaced by a proper entailment model that handles paraphrasing and implicit entailment.

Now let us show a complete RAGAS-style evaluation pipeline using an LLM judge simulation based on TF-IDF similarity:

In[15]:
Code
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity


class SimpleRAGASEvaluator:
    """Simplified RAGAS evaluator using heuristic judges.

    This class simulates the RAGAS framework using lightweight
    heuristics. For production use, replace each method's
    similarity computation with actual LLM judge calls.
    """

    def context_precision(self, query: str, retrieved: List[str]) -> float:
        """Fraction of retrieved chunks relevant to query."""
        # Simulate LLM judgment with keyword overlap
        query_words = set(query.lower().split())
        relevant_count = 0

        for doc in retrieved:
            doc_words = set(doc.lower().split())
            overlap = len(query_words.intersection(doc_words))
            if overlap > 0:
                relevant_count += 1

        return relevant_count / len(retrieved) if retrieved else 0.0

    def answer_relevance(self, query: str, answer: str) -> float:
        """Semantic similarity between query and answer via TF-IDF cosine."""
        vectorizer = TfidfVectorizer()
        try:
            vectors = vectorizer.fit_transform([query, answer])
            similarity = cosine_similarity(vectors[0:1], vectors[1:2])[0][0]
            return float(similarity)
        except Exception:
            return 0.0

    def faithfulness(self, answer: str, contexts: List[str]) -> float:
        """Check if answer is grounded in contexts using word overlap."""
        return faithfulness_score(answer, contexts)
Out[16]:
Console
RAGAS-style Evaluation:
  Context Precision: 1.00
  Answer Relevance:  0.72
  Faithfulness:      1.00

Interpretation:
  All 3 chunks mention 'France', so context precision is 1.0.
  The answer directly addresses the query, so answer relevance is high.
  The answer is grounded in the first retrieved chunk.
Out[17]:
Visualization
4x4 heatmap with red-to-green color scale showing combined end-to-end RAG scores at different retrieval and generation quality levels, with annotated cell values.
End-to-end RAG performance as a function of retrieval and generation quality, displayed as a heatmap. Each cell shows the combined system score when retrieval quality (rows) and generation quality (columns) are at specified levels. The multiplicative interaction structure means that even excellent generation cannot fully compensate for poor retrieval (left column remains low), and excellent retrieval cannot save poor generation (bottom row remains low). Both components must perform well together for high end-to-end performance.

The heatmap visualization makes the multiplicative nature of RAG performance viscerally clear. Even perfect generation (right column, generation quality 1.0) cannot rescue poor retrieval: with retrieval quality 0.2, end-to-end performance is capped at 0.18. Similarly, perfect retrieval cannot save poor generation. Both components must work well simultaneously for the system to deliver high-quality answers.

Evaluation in Practice: Building an Evaluation Pipeline

Understanding individual metrics is necessary but not sufficient for running effective evaluations in a production environment. Building an end-to-end evaluation pipeline requires decisions about query sampling, annotation workflows, metric selection, and result interpretation.

The most important practical decision is what query distribution to evaluate against. Evaluation queries should reflect real user queries, including their difficulty distribution, topic distribution, and linguistic variety. If your system will be queried by domain experts, your evaluation set should include expert-level queries. If users will ask both simple factoid questions and complex analytical questions, both should be represented. A benchmark that only tests easy queries will mask failures on hard ones.

Creating a good evaluation dataset for RAG requires collecting or constructing query-context-answer triples. For retrieval evaluation, you need queries and relevance judgments for documents in the corpus. These can come from several sources: human annotators who manually assess relevance, existing question-answering datasets that can be aligned with your corpus, or synthetic queries generated by LLMs from corpus documents. Each source has tradeoffs. Human annotation is most reliable but expensive and slow. Existing datasets may not match your corpus or use case. Synthetic queries can be generated at scale but may not reflect real user query patterns.

For generation evaluation, you have additional choices. Reference-based metrics (BLEU, ROUGE, BERTScore) require collecting gold standard answers, which is expensive but lets comparison against a known correct answer. Reference-free metrics (RAGAS faithfulness, answer relevancy) can be computed on any query without pre-collected answers, letting evaluation on live user queries. A practical pipeline typically combines both: reference-based evaluation on a curated benchmark for rigorous tracking, plus reference-free evaluation on production traffic for continuous monitoring.

Monitoring over time is as important as point-in-time evaluation. RAG system performance degrades for several reasons: knowledge base drift (the documents become outdated), query distribution shift (users start asking different kinds of questions), component model updates (embedding models or generators get replaced), and corpus growth (the knowledge base expands, increasing retrieval difficulty). A good evaluation pipeline runs automatically on a regular schedule, tracks metric trends over time, and alerts when metrics drop below acceptable thresholds.

Limitations and Challenges

RAG evaluation presents unique challenges that standard NLP benchmarks do not address. The primary difficulty lies in the absence of ground truth for open-domain applications. Unlike closed-domain question answering where answers come from a limited knowledge base with pre-defined answer spans, RAG systems often answer questions where the correct answer depends on the specific documents in the corpus, which makes it impossible to pre-annotate reference answers for all possible queries.

This absence of ground truth fundamentally changes the evaluation paradigm. Traditional machine learning relies on labeled datasets with clear correct answers, letting straightforward calculation of accuracy. In open-domain RAG, the knowledge base might contain millions of documents, and the space of possible queries is in effect infinite. We cannot create a test set covering all possible questions, and the correct answer to any specific question might change as the knowledge base is updated with newer information. This necessitates reference-free evaluation methods or dynamic evaluation pipelines that can assess quality without static gold standard answers.

A second major challenge is error compounding. Retrieval errors and generation errors interact in complex ways that are difficult to disentangle. A retrieval system might fetch the correct document but rank it at position 8, beyond the effective context window. Alternatively, perfect retrieval cannot save a generator prone to hallucination or one that ignores the provided context in favor of outdated parametric knowledge. The multiplicative interaction between retrieval and generation quality, as illustrated in our heatmap, means that small improvements to a weak component produce larger gains than equivalent improvements to an already strong component. But identifying which component is the weak link requires careful controlled experiments.

Understanding these interactions is important for system optimization. If retrieval reaches high NDCG but the generator produces unfaithful answers, the problem lies in generation, suggesting a need for better prompting, fine-tuning, or a stronger base model. If retrieval scores are low but the generator performs well, the model may be relying heavily on parametric knowledge, creating a false sense of security that will fail when queries target information outside the training distribution. Isolating these effects requires controlled ablation studies where components are evaluated independently and in combination.

The third major challenge is evaluator reliability. When using LLMs as judges in frameworks like RAGAS, the quality of the evaluation depends critically on the quality of the judge model and its prompts. LLM judges can be inconsistent across runs due to sampling temperature, biased toward confident-sounding but incorrect answers, sensitive to prompt phrasing, and calibrated differently on different query types. Research has shown that LLM judges correlate reasonably well with human judgment on average but have significant error rates on individual examples, particularly for ambiguous cases or queries requiring domain expertise. This means RAGAS scores should be treated as approximate indicators rather than ground truth, requiring periodic validation against human judgments.

A fourth challenge specific to RAG is the tension between faithfulness and completeness. A system that only generates answers when it finds strong supporting evidence will be highly faithful but may refuse to answer many valid queries. A system that generates answers aggressively will answer more queries but with lower faithfulness. Evaluation metrics typically measure faithfulness and answer relevancy independently, but the right balance depends on the application. In medical or legal contexts, faithfulness should be nearly absolute even at the cost of lower answer coverage. In general-purpose assistants, users may prefer a helpful but potentially less precise answer to no answer at all.

The computational cost of evaluation is also non-trivial. LLM-based evaluation can be as expensive as the generation step itself, especially for frameworks that make multiple LLM calls per query (claim extraction, claim verification, question generation). For large-scale evaluation across thousands of queries, this cost can become significant. Some teams use cheaper models for automated evaluation and reserve expensive models for spot-checking. Others develop lightweight proxy metrics that correlate with RAGAS scores on their specific domain, letting fast evaluation without full RAGAS computation for every change.

These challenges highlight the importance of human-in-the-loop validation. Automated metrics give necessary scalability for development and regression testing, but periodic human evaluation remains needed for validating the automated metrics themselves. If an LLM judge gradually drifts in its assessment criteria, or if prompt templates become ineffective for new types of queries, only human oversight can detect these systematic failures. The field continues to evolve toward hybrid approaches that combine automated screening at scale with human judgment on challenging or borderline cases, using active learning techniques to select the most informative cases for human review.

Out[18]:
Visualization
Scatter plot showing evaluation metrics positioned by computational cost on the x-axis and human alignment on the y-axis, with annotations for each metric.
Comparison of evaluation metric computational cost versus alignment with human judgment, illustrating the tradeoff between evaluation speed and quality. Simple metrics like Precision@K are fast and cheap but have lower human alignment. RAGAS metrics with LLM judges have high human alignment but require multiple LLM API calls per query. The ideal evaluation pipeline uses cheap metrics for high-frequency monitoring and expensive metrics for periodic deep evaluation.

Summary

RAG evaluation must assess retrieval quality, generation faithfulness, and end-to-end answer relevance because failures at any stage compound across the pipeline. We have explored how traditional information retrieval metrics, including Precision@K, Recall@K, MRR, MAP, and NDCG, measure the quality of document retrieval with particular attention to the position-sensitive nature of generator context windows. For generation, we moved beyond n-gram overlap metrics to semantic similarity approaches and factual consistency checking against retrieved contexts, understanding specifically why BLEU and ROUGE fail for this task.

The end-to-end evaluation framework stresses three pillars: context relevance (do retrieved documents contain the answer?), faithfulness (is the generated answer supported by the documents?), and answer relevance (does the response address the query?). Each pillar targets a distinct failure mode, making the framework useful for measuring system quality and diagnosing where problems originate. The RAGAS framework automates these assessments using LLM judges, enabling scalable evaluation without reference answers for most metrics, though at the cost of potential judge bias and computational expense.

Effective RAG evaluation ultimately requires human validation of automated metrics, careful analysis of error modes, and an evaluation strategy that shows real query distributions rather than artificial benchmarks. Understanding whether failures stem from retrieval quality, generation faithfulness, or their interaction guides targeted improvements, whether that means adopting better embedding models, implementing reranking, refining prompt engineering strategies, or investing in a stronger base generation model. As RAG systems grow more complex, incorporating reranking, query expansion, and multi-step retrieval, evaluation frameworks must evolve accordingly, measuring final answer quality and the quality of each intermediate reasoning step.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about RAG Evaluation.

RAG Evaluation Quiz

Question 1 of 70 of 7 completed
Which formula correctly defines Precision@K for retrieval evaluation?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025ragevaluation, author = {Michael Brenndoerfer}, title = {RAG Evaluation: Metrics for Retrieval and Generation Quality}, year = {2025}, url = {https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). RAG Evaluation: Metrics for Retrieval and Generation Quality. Retrieved from https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation
MLAAcademic
Michael Brenndoerfer. "RAG Evaluation: Metrics for Retrieval and Generation Quality." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation>.
CHICAGOAcademic
Michael Brenndoerfer. "RAG Evaluation: Metrics for Retrieval and Generation Quality." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation.
HARVARDAcademic
Michael Brenndoerfer (2025) 'RAG Evaluation: Metrics for Retrieval and Generation Quality'. Available at: https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). RAG Evaluation: Metrics for Retrieval and Generation Quality. https://mbrenndoerfer.com/writing/rag-evaluation-metrics-retrieval-generation

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.