Part of Language AI Handbook
Explains how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Attribution and Citation
When a language model tells you that "the Eiffel Tower is 330 meters tall" or "researchers at Stanford found that coffee reduces Alzheimer's risk," how do you know if it is true? More importantly, how do you know where that information came from? Attribution and citation address exactly this problem: they are the mechanisms by which AI-generated text is connected back to its supporting sources.
In the previous chapters of this part, we explored hallucination types, detection methods, and causes. We saw how models generate fluent-sounding text that can be factually wrong, and we studied the roots of that failure. Mitigation strategies like retrieval-augmented generation (RAG) help ground model outputs in actual documents. But grounding alone is not enough. Even when a model retrieves relevant documents, it may misattribute quotes, conflate claims from different sources, or silently omit supporting evidence for assertions that needed it. Attribution is the accountability layer: it answers not just "is this true?" but "what evidence supports this, and where does that evidence come from?"
The distinction is subtle but important. A model that only retrieves documents has improved the probability that its outputs are correct. A model that attributes its claims has created verifiability, which is a stronger guarantee. Verifiability means that a human reader can follow the chain from claim to evidence, evaluate the quality of that evidence, and decide whether the claim is well-supported. Without attribution, the reader can only trust the model's output or discard it entirely. With attribution, they have a third option: check.
This chapter covers the full scope of attribution in language models. We begin with what attribution means and why it matters, then move through the mechanisms by which models produce citations, the metrics used to evaluate whether those citations are correct, and the practical challenges that make attribution hard to get right. We close with an implementation of an NLI-based attribution evaluation system that you can extend for your own use cases.
Attribution is the act of connecting a model-generated statement to a specific source document or passage that supports that statement. A fully attributed system does not just produce correct text; it provides evidence that the text is correct.
Why Attribution Matters
The value of attribution goes well beyond academic citation conventions. Consider three scenarios that illustrate why the accountability layer matters in practice.
A medical professional uses an AI assistant to summarize recent drug trial data. The assistant synthesizes information from dozens of recent papers and produces a clean, readable summary. If that summary contains an error, the consequences could be severe. The professional needs to know which specific study or paper supports each claim, because they need to go back to the source when something looks unexpected. A summary without attribution is not useful in this setting; it is a risk.
A journalist uses an LLM to research a news story about a complex geopolitical topic. The model synthesizes information from three conflicting sources without indicating which source says what. The story's accuracy cannot be verified by an editor, and conflicting claims are silently blended into a seemingly coherent narrative. The journalist has no way to check whether the synthesis is accurate or which sources have been weighted most heavily.
A legal researcher uses an AI tool to draft a brief. The brief cites case law. If the citations are hallucinated or point to the wrong rulings, the brief becomes useless and professionally dangerous. Courts expect accurate citation, and an attorney who submits a brief with fabricated citations faces serious professional consequences.
In all three cases, attribution serves a trust function. It allows a human to verify, dispute, or build on what the model has said. Without attribution, model output must either be trusted blindly or re-verified from scratch, eliminating much of the efficiency benefit that motivated using the model in the first place.
Attribution also plays a role in accountability. When a model claims something controversial, attribution reveals whether that claim has legitimate backing or whether it is synthesized from thin air. This is increasingly important as AI-generated content appears in high-stakes applications including education, law, medicine, and journalism. The absence of attribution is not neutral: it implicitly asserts that the model's outputs are reliable without providing any mechanism for the user to check that assertion.
There is also a more subtle benefit: attribution shapes how users interact with AI outputs. When a user sees a citation alongside a claim, they naturally approach the claim as something to be checked, rather than as an established fact. This shifts the epistemic responsibility back toward the human in the loop, which is precisely where it should be in high-stakes settings. A model that cannot attribute its claims implicitly asks the user to take them on faith. A model that cites its sources invites scrutiny, and that invitation, even when the citations are only partially helpful, is educationally and epistemically valuable.
This shift in user behavior is not accidental. Research on human-AI interaction has consistently found that presenting AI outputs with explicit uncertainty or supporting evidence leads to more appropriate reliance: users are more likely to defer to the model when it is correct and more likely to override it when it is wrong. Attribution is one mechanism that enables this calibrated reliance. A cited claim naturally prompts the question "does this source actually say that?", while an uncited claim prompts only "do I believe this model?" The first question is answerable; the second is not.
Types of Attribution
Not all attribution looks the same. There are several distinct forms, each with different granularity, form, and use case. Understanding these distinctions is important because the right form of attribution depends on the application, the available retrieval infrastructure, and the level of transparency you need to provide to your users.
Inline Citation
Inline citation is the most direct form of attribution. A claim appears in the model's output, and immediately alongside or after it, a reference to the supporting source is provided. This mirrors the convention in academic writing, where a sentence like "Transformers outperform RNNs on most NLP benchmarks (Vaswani et al., 2017)" carries its evidence right with it.
In the context of retrieval-augmented generation, inline citation typically takes the form of numbered references linked to retrieved document passages. A system might produce:
"The Mars Science Laboratory Curiosity rover landed on August 6, 2012 [1]. It has since traveled more than 28 kilometers across Gale Crater [2]."
Here, [1] and [2] point to specific passages from source documents. The model's task is to produce a correct sentence and ensure that [1] contains a passage that supports the landing date, and that [2] contains a passage about distance traveled.
Inline citation is demanding because it requires the model to track provenance at the claim level. A single paragraph might contain five distinct factual claims, each requiring its own source. This is substantially harder than simply retrieving a document and summarizing it. The model must simultaneously attend to the content it is generating, the retrieved context it is drawing from, and the mapping between the two. When this tracking breaks down, the model may produce citations that point to the right document but the wrong passage, or cite a topically related passage that does not contain the specific information claimed.
One important nuance is that inline citation requires a deliberate training signal. A model that has not been exposed to examples of attributed generation during fine-tuning will rarely produce inline citations spontaneously, even if the retrieved passages are clearly numbered in the context. Models like Perplexity AI and various RAG-augmented systems that produce reliable inline citations have been fine-tuned on large collections of cited responses, often with human-annotated citation chains.
Post-Hoc Attribution
Some systems separate generation from attribution. The model first produces a response, then a separate component attempts to find supporting evidence for each claim in the response. This is sometimes called post-hoc attribution because the sourcing happens after generation rather than during it.
Post-hoc attribution is easier to implement because it decouples the generation and retrieval problems. A standard language model can produce its normal output, and then a retrieval system searches a corpus for passages that match each claim. The passages found become the citations, and the system presents them alongside the original response.
However, this approach introduces a structural risk: if the model generates a claim for which no support exists, the post-hoc search either returns incorrect sources (hallucinated attribution) or fails to find any source (leaving the claim unsupported). In either case, the generation step has already produced potentially wrong content, and post-hoc attribution is trying to patch a problem that began at generation time.
This failure mode is qualitatively different from generation-integrated attribution. When attribution happens during generation, the model is constrained to produce claims that its current context supports. When attribution is retrofitted, there is no such constraint, and the mismatch between what was generated and what evidence exists becomes apparent only after the fact. Some post-hoc systems handle this by flagging unsupported claims rather than forcing a citation, which at least makes the gap visible to the user. This flagging behavior is useful: it turns unsupported claims from invisible failures into explicit warnings, allowing the user to investigate further.
Post-hoc attribution is also sensitive to the quality of the claim decomposition step. Before searching for supporting evidence, the system needs to identify the claims in the response. This is itself a non-trivial natural language understanding problem. A sentence like "The Transformer architecture, introduced in 2017, revolutionized NLP by replacing recurrent networks with self-attention" contains multiple claims, each of which needs independent verification. Automated claim decomposition can miss some claims, merge distinct claims, or identify boundaries incorrectly, introducing noise into the downstream attribution search.
Document-Level Attribution
At the coarser end of the spectrum, document-level attribution connects an entire response to one or more source documents, without specifying which parts of the response come from which document. This is common in summarization systems, where a model might produce a multi-sentence summary and attribute it to a provided article.
Document-level attribution is less informative than inline citation but is easier to evaluate and implement. It answers "did the model use this document?" rather than "which specific claim comes from which specific passage?" For many applications, this level of granularity is sufficient. If you are summarizing a single known document, document-level attribution simply confirms that the summary is derived from the intended source. If you are comparing multiple documents, it tells you which documents contributed to the response.
The limitation becomes apparent in multi-document settings. If a response draws on ten retrieved documents and only three of them contributed relevant information, document-level attribution that lists all ten gives the reader a misleading picture of what the response is based on. More importantly, if the response contains an error, document-level attribution cannot tell you which source the error came from, making debugging and correction much harder.
Span-Level Attribution
Span-level attribution is the most fine-grained form. Each span of generated text, whether a phrase, sentence, or clause, is individually linked to the specific passage in the source that supports it. This provides the most informative attribution signal, but it is also the most technically challenging to produce and evaluate.
Research systems like ALCE (Attributed LM Comprehension Evaluation) have formalized span-level attribution evaluation, making it possible to assess whether a model provides citations and whether each cited passage supports the corresponding claim. ALCE treats attribution as a property of individual sentences, asking: for each sentence in the generated response, is there at least one cited passage that entails that sentence?
In practice, span-level attribution requires the model to reason simultaneously about what to say and which piece of its provided context licenses that statement. This is a form of grounded generation that pushes against language models' natural tendency toward fluent synthesis. Models trained with supervised fine-tuning on attributed corpora, where human annotators have linked each sentence to its source, can learn this behavior. But it requires careful data collection and substantial annotation effort to produce the training signal.
The annotation effort is particularly challenging because span-level attribution ground truth is ambiguous in many cases. A sentence may be supported by multiple passages simultaneously. A claim may be partially supported by one passage and partially by another. Human annotators asked to identify the single best supporting passage for each sentence will disagree at rates that are high enough to make inter-annotator agreement a serious concern. These ambiguities flow through to evaluation: a model that identifies a valid supporting passage that differs from the annotator's preferred passage should not be penalized, but simplistic exact-match evaluation would penalize it.
Attribution in RAG Systems
The most common practical setting for attribution is retrieval-augmented generation. As we saw in earlier chapters, RAG systems retrieve relevant documents from a corpus before generating a response, giving the model access to external knowledge that can ground its outputs. Attribution in RAG is the additional requirement that the model indicate which retrieved passages support which parts of its output.
A RAG-with-attribution pipeline has several interconnected stages, each of which affects the quality of the final attribution.
The retrieval stage fetches candidate documents or passages from a knowledge base using a query. This is typically done with a dense retriever like a bi-encoder, a sparse retriever like BM25, or a hybrid of both. The quality of attribution is bounded by the quality of retrieval: if the most relevant passages are not retrieved, the model cannot attribute its claims to them, even if those claims happen to be correct.
The prompt construction stage inserts the retrieved passages into the model's context, usually in a structured format. A common pattern labels each passage with a numeric identifier:
[1] <passage from document A>
[2] <passage from document B>
[3] <passage from document C>
Based on the above passages, answer the following question and cite your sources:
...
The way passages are formatted, ordered, and labeled matters more than it might seem. Models are sensitive to position: passages presented earlier in the context receive more attention and are more likely to be cited. Numbering schemes and delimiters affect how reliably the model refers back to specific passages. Prompt engineering for attribution is an active area of practical development, with significant variation in citation quality across different formatting conventions.
The generation stage produces a response. A well-instructed model should produce text that cites the relevant passage numbers inline, like "According to [2], the population of Brazil reached 215 million in 2022." The instruction-following quality of the model, the length and complexity of the passages, and the difficulty of the query all affect how well the model maintains attribution discipline across the full length of its response.
The attribution verification stage checks whether the citations in the generated text are supported by the cited passages. This is the evaluation step, and it is logically separate from generation. Even a model that produces citations for every claim may produce incorrect ones, and the verification stage catches these errors. For production systems, attribution verification can be done automatically using NLI-based methods and used to trigger re-generation, response filtering, or user warnings when citation quality falls below a threshold.
This pipeline architecture reveals an important design choice. Attribution can be built as a property of the generation model (the model learns to cite during fine-tuning) or as a property of the pipeline (a separate component adds and verifies citations post-generation). The former tends to produce better integrated attribution; the latter is more modular and easier to update independently.
Attribution Accuracy
Producing citations is necessary but not sufficient. A model can produce citations that are wrong in several distinct ways, and each failure mode has different consequences for the user.
Unsupported citations: The model cites a passage that does not support the claim. For example, the model says "The heart has four chambers [1]" but passage [1] discusses blood pressure, not cardiac anatomy. The citation is present but irrelevant. This is perhaps the most dangerous failure mode because it looks like correct attribution while providing no verification value.
Overclaiming: The model makes a claim that goes beyond what the source says. A source might say "in some experimental conditions, treatment X reduced inflammation by 20% in mice," and the model summarizes this as "treatment X reduces inflammation," citing the same source. The direction of the claim is correct, but it omits essential caveats: the experimental conditions, the magnitude, and the species. The citation technically points to a related source, but the claim is stronger than what the evidence supports.
Missing citations: The model makes factual claims without any citation. This is not strictly wrong if the claims are common knowledge, but for specific or technical assertions, missing citations leave the reader unable to verify. A model that claims "the GDP of Germany in 2023 was $4.1 trillion" without citing a source has made a specific empirical claim that requires evidence, even if the claim happens to be accurate.
Citation to wrong source: In a multi-document RAG setting, the model supports a claim by citing document 3, but the supporting information is in document 1. The claim may be correct, and the information may even be consistent with document 3, but the specific citation is wrong. This matters when the user wants to follow up: they go to document 3 and cannot find the passage they were supposed to verify.
Copied but unattributed: The model reproduces a specific phrase or finding verbatim from a source without citing it. This is a form of attribution failure that may also constitute a copyright concern. Models trained on large text corpora sometimes reproduce training data near-verbatim, and this reproduction may or may not be accompanied by proper attribution.
Each of these failure modes requires different detection strategies. Unsupported citations and overclaiming can be detected by NLI-based entailment checking. Missing citations require identifying claims and comparing them against a citation registry. Citations to wrong sources require checking that the cited passage is the one from which the relevant information was drawn, not just that it is topically related. Copied but unattributed text can be detected using plagiarism detection approaches.
Evaluating Attribution
Evaluating attribution requires determining whether a generated statement is supported by its cited source. This is itself a difficult natural language inference (NLI) problem: given a claim and a passage, does the passage support, contradict, or neither support nor contradict the claim?
The challenge is deeper than it might appear. Entailment is not a simple semantic similarity check. A passage that is topically related to a claim may not entail it. A passage that entails a claim from a different angle than the model's phrasing may be difficult for an NLI model to classify correctly. Negation, conditionals, and quantifiers all complicate entailment judgments in ways that challenge current NLI models.
Several formal evaluation frameworks and metrics have been developed for this task.
AutoAIS and AIS
The Attributable to Identified Sources (AIS) framework defines a binary judgment for each statement in a model's output: is this statement fully attributable to the provided source material? A statement is AIS-compliant if a human, reading only the cited sources, could verify the statement as true.
This definition is deliberately strict. It requires that the verification be possible from the cited sources alone, without prior knowledge. A statement like "Albert Einstein was born in 1879" is common knowledge, but if the cited source does not contain this information, the statement fails the AIS check in the strict sense. This strictness is intentional: it forces systems to be explicit about what knowledge they are drawing on, rather than relying on implicit assumptions that the reader may or may not share.
AutoAIS uses a trained NLI model to automate this binary judgment. The AIS framework is notable because it is grounded in a precise definition of attribution: it is not enough that the statement is true, it must be verifiable from the cited sources. A true statement with no citation, or a citation that does not contain the supporting evidence, fails the AIS check.
The NLI model used for AutoAIS is typically a cross-encoder, meaning it takes both the passage and the claim as joint input and outputs a single entailment judgment. Cross-encoders are more accurate than bi-encoders for this task because they can model the interaction between premise and hypothesis directly, but they are also more computationally expensive since they cannot pre-compute passage representations independently.
ALCE
The ALCE benchmark (Attributed LM Comprehension Evaluation) is an evaluation framework that tests whether language models can produce attributed responses across multiple tasks including open-domain question answering, slot filling, and document summarization.
ALCE measures two primary properties:
- Correctness: Does the response contain the right information, assessed against a gold-standard answer?
- Citation quality: Are the citations accurate, meaning does each cited passage support the accompanying claim?
ALCE uses a combination of automatic NLI-based evaluation and human annotation to assess citation quality. The framework's key insight is that correctness and attribution are independent properties: a model can be correct but uncited, incorrect but well-cited (citing a wrong source that happens to say something related), or correct and well-cited. The ideal system achieves both, but real systems often trade one for the other.
ALCE revealed a counter-intuitive pattern when first published: larger language models did not necessarily produce better attribution than smaller ones. A model might improve in factual accuracy as it scales, but if it was not specifically trained on attributed generation tasks, its citation quality could remain poor regardless of model size. This shows that attribution is a behavior that must be learned, not an automatic consequence of knowing more facts.
Citation Precision and Recall
A more metric-oriented evaluation framework treats attribution evaluation like an information retrieval problem. For each claim in the generated text, we compute two complementary quantities.
Citation precision measures, among all the passage citations the model provided, what fraction support the claim. If is the set of cited passages and is the subset that supports the claim, then:
where:
- : the set of all passages cited by the model for a given claim
- : the subset of those cited passages that support the claim (as judged by NLI or human annotation)
- and : the cardinalities of those sets
A precision of 1.0 means every citation the model provided supports its associated claim. A precision of 0.5 means half of the provided citations are irrelevant or incorrect. High precision is achieved when the model is selective: it only cites a passage when it is confident the passage supports the claim.
Citation recall measures, among all passages in the retrieved set that support the claim, what fraction are cited by the model. If is the set of all retrieved passages that support the claim, then:
where:
- : all retrieved passages (regardless of whether they were cited) that support the claim
- : the supporting passages that the model cited
A recall of 1.0 means the model cited every retrieved passage that was supportive. A recall of 0.5 means the model cited only half of the available supporting evidence, leaving relevant corroboration on the table.
These metrics can be computed at the sentence or claim level. They parallel standard IR precision/recall and share the same tradeoffs: systems can optimize for one at the cost of the other. A model that cites every retrieved passage will have high recall but low precision; a model that only cites when highly confident will have high precision but may miss supporting evidence.
The F1 score over citation precision and recall provides a balanced summary metric that equally weights both:
This formula is undefined when both precision and recall are zero, in which case by convention.
The choice of whether to optimize precision, recall, or F1 depends on the application. In high-stakes settings where a false citation is actively dangerous, precision matters more. In research assistance settings where comprehensiveness is important, recall matters more. Most practical attribution evaluations report all three, allowing system designers to choose the operating point that fits their use case.
Consistency-Based Evaluation
A different approach to attribution evaluation focuses on internal consistency rather than source matching. The idea is that a well-attributed claim should be reproducible: if you ask the model the same question multiple times with different retrieved contexts, a claim that is supported by evidence should appear consistently, while a hallucinated claim should be inconsistent across samples.
This approach, related to the consistency checking methods we discussed in the hallucination detection chapter, does not require a human to evaluate source documents. Instead, it relies on the model's own variance as a signal of reliability. Claims that appear in every sampled response are likely to be well-grounded; claims that appear only occasionally may be hallucinated.
The practical implementation uses sampling-based generation. Given a query and multiple different retrieved contexts (sampled from the same corpus), you generate multiple responses and look for claims that appear consistently across all of them. Consistent claims are attributed to the model's knowledge; inconsistent claims are flagged as potentially unreliable.
The limitation of consistency-based evaluation is that a consistently wrong model will score well, and a consistently correct model that appropriately varies its response based on context may score worse than it should. More fundamentally, this approach measures consistency, not truth. A model that has memorized an incorrect fact from training data and always repeats it will score perfectly on consistency metrics while being systematically wrong. Consistency-based evaluation is therefore most useful as a complement to source-based evaluation, not a replacement.
Source Verification
Attribution tells you where a claim came from. Source verification asks a harder question: is the source itself reliable? This distinction matters because a language model can produce a well-attributed statement based on a low-quality or factually incorrect source.
Imagine a RAG system that retrieves from a large web corpus without quality filtering. The model produces a claim like "according to [1], the optimal daily protein intake is 400 grams per kilogram of body weight." The citation is correct: passage [1] does say this. The attribution evaluation will give this a perfect score. But passage [1] is from a fitness blog with no scientific backing, and the claim is wrong by more than an order of magnitude. The attribution layer has done its job; the source verification layer has not.
Source verification goes beyond what language models themselves can easily accomplish. It typically requires integration with external systems:
- Domain authority checks: Does the cited source come from a recognized, authoritative domain? A medical claim cited from a peer-reviewed journal is different from the same claim cited from a blog. Systems can encode these authority hierarchies, assigning higher trust scores to citations from curated sources.
- Publication date checks: Is the cited source current? In rapidly evolving fields, a citation from five years ago may have been superseded by more recent research. A source from 2018 about transformer architectures will lack significant developments introduced in subsequent years.
- Cross-reference verification: Does the claim appear in multiple independent sources? Triangulation across sources reduces the risk that a single incorrect source is misleading the model. If five independent papers all report the same finding, the probability that the finding is correct is much higher than if only one source says it.
- Structured knowledge base lookup: Can the claim be verified against a structured knowledge base like Wikidata or a medical ontology? Structured knowledge is generally more reliable than free-text documents because it is curated and validated. Wikidata claims are linked to external sources and undergo community review; they carry a different epistemic status than a passage from an arbitrary web document.
In practice, RAG systems often use curated document corpora with known quality properties. A production system might restrict retrieval to documents from a trusted set of publishers, journals, or databases, giving a degree of implicit source verification at the corpus level. This architectural choice, essentially treating corpus quality as a proxy for source quality, is simpler than runtime source verification but offers weaker guarantees: even trusted sources contain errors, and outdated information from trusted sources is still outdated.
The interaction between attribution and source verification creates a two-layer trust model. The first layer asks "does the model's output match the sources it cites?" The second layer asks "are the sources the model cites trustworthy?" Only systems that pass both layers can be said to provide reliable attributed outputs. Most current attribution evaluation frameworks operate at the first layer and leave source verification as an architectural concern.
The Attribution-Fluency Tradeoff
One subtle tension in attributed generation is the tradeoff between attribution accuracy and response fluency. Accurate attribution requires the model to closely track what specific passages say and only claim what those passages support. But language models are trained to produce fluent, coherent responses, which often requires synthesizing, paraphrasing, and interpolating across multiple sources.
This synthesis process is exactly what makes language models useful for response generation, but it is also what makes attribution hard. When a model produces a sentence that synthesizes information from three different passages, no single citation fully supports that sentence. The synthesized claim may be more informative than any individual source, but it is also harder to attribute accurately. The model has done cognitive work to combine the sources, and that work cannot easily be traced back to any one passage.
Consider a concrete example. Passage 1 says "Transformer models use self-attention mechanisms." Passage 2 says "Self-attention scales quadratically with sequence length." The model produces the sentence: "Transformer models use self-attention, which scales quadratically with sequence length." This sentence is correct and is a natural synthesis of passages 1 and 2. But it requires two citations, and the attribution logic must recognize that the first clause comes from passage 1 and the second from passage 2. A model that cites only one passage, or cites both for the entire sentence rather than assigning each part to its source, has technically produced incorrect attribution even though the synthesized claim is fully accurate.
Several approaches attempt to manage this tradeoff:
Extractive generation: Rather than having the model freely generate attributed text, constrain it to produce responses that are maximally close to exact extracts from source documents. This makes attribution trivial because the text is essentially the citation, but it produces less fluent, more fragmented responses. Users often find extractive responses harder to read than fluently synthesized ones, even when the information content is the same.
Claim decomposition: Decompose complex claims into atomic sub-claims, each of which can be independently verified and attributed. The model might first generate a free-form response, then a second step decomposes each sentence into its constituent atomic claims, and a third step matches each claim to a supporting passage. This improves attribution precision at the cost of response coherence: the final attributed response may feel choppy if it closely follows the claim-by-claim structure.
Citation confidence thresholding: Have the model provide citations only when it has high confidence that a passage supports a claim, leaving some claims uncited rather than providing uncertain citations. This trades citation recall for citation precision: the user will see fewer citations, but those they do see are more likely to be correct. For applications where false citations are particularly harmful, this conservative strategy may be preferable.
Constrained decoding: At inference time, restrict the model's generation to only produce claims that can be grounded in the current context. This is related to the constrained decoding methods discussed in the GPT architecture chapter. By constraining the generation space, you prevent the model from making claims that cannot be attributed, at the cost of potentially less fluent or less informative responses.
Each of these strategies represents a different point on the attribution-fluency frontier. The right choice depends on the application, the users' ability to tolerate reduced fluency in exchange for increased reliability, and the costs associated with false citations versus missing citations.
Code Implementation
In this section, we build a simple attribution evaluation system. We simulate a RAG response with citations and use NLI-based scoring to evaluate whether each cited passage supports the corresponding claim.
We need transformers and torch for the NLI model. Install them if needed:
# subprocess.run(["uv", "pip", "install", "transformers", "torch"], check=False, capture_output=True)
# Run once locally, then comment outSetting Up the NLI Scorer
The core of attribution evaluation is natural language inference: given a premise (the cited passage) and a hypothesis (the generated claim), determine whether the premise entails the hypothesis. We use a pretrained NLI model for this. The cross-encoder architecture is important here: it encodes both the passage and the claim jointly, allowing the model to attend to their interaction directly rather than comparing independently encoded representations.
from transformers import pipeline
# Use a lightweight cross-encoder NLI model
nli_model = pipeline(
"text-classification",
model="cross-encoder/nli-deberta-v3-small",
device=-1, # CPU
)NLI model loaded successfully. Model: cross-encoder/nli-deberta-v3-small
Defining the Attribution Evaluator
We define a function that takes a claim and a list of cited passages, and returns an entailment score for each passage. A high score indicates strong support; a low score indicates the passage does not support the claim. The function returns all three NLI scores (entailment, neutral, contradiction) so callers can apply whatever threshold logic fits their application.
def evaluate_citation(claim: str, passage: str, model) -> dict:
"""
Evaluate whether a passage supports a claim using NLI.
Returns entailment, neutral, and contradiction probabilities.
"""
result = model(f"{passage} [SEP] {claim}", top_k=None)
# result is a list of dicts: [{"label": ..., "score": ...}, ...]
# When batched, it may be wrapped in an extra list; unwrap if needed
if result and isinstance(result[0], list):
result = result[0]
scores = {item["label"].lower(): item["score"] for item in result}
return {
"entailment": scores.get("entailment", 0.0),
"neutral": scores.get("neutral", 0.0),
"contradiction": scores.get("contradiction", 0.0),
"is_supported": scores.get("entailment", 0.0) > 0.5,
}
def evaluate_attributed_response(
claims_with_citations: list, passages: dict, model
) -> list:
"""
Evaluate all claim-citation pairs in a response.
claims_with_citations: list of {"claim": str, "citation_ids": list[int]}
passages: dict mapping passage ID to passage text
"""
results = []
for item in claims_with_citations:
claim = item["claim"]
cited_passages = [
passages[cid] for cid in item["citation_ids"] if cid in passages
]
claim_results = []
for idx, passage in zip(item["citation_ids"], cited_passages):
score = evaluate_citation(claim, passage, model)
claim_results.append(
{
"passage_id": idx,
"passage_preview": passage[:80] + "..."
if len(passage) > 80
else passage,
**score,
}
)
# A claim is considered supported if at least one cited passage entails it
any_supported = any(r["is_supported"] for r in claim_results)
results.append(
{
"claim": claim,
"citation_results": claim_results,
"claim_supported": any_supported,
}
)
return resultsExample Attribution Evaluation
We create a simulated RAG scenario with three passages and a model response containing three claims. The first claim is well-supported. The second contains a date error that the source contradicts. The third cites a completely unrelated passage, simulating hallucinated attribution.
# Simulated retrieved passages (as a RAG system would provide)
passages = {
1: "The transformer architecture was introduced in the paper 'Attention is All You Need' by Vaswani et al. in 2017. It replaced recurrent neural networks with self-attention mechanisms.",
2: "BERT (Bidirectional Encoder Representations from Transformers) was published by Devlin et al. in 2019. It uses masked language modeling for pretraining.",
3: "The Amazon rainforest covers approximately 5.5 million square kilometers and is home to 10% of all species on Earth.",
}
# Simulated attributed response claims (what the model generated with citations)
claims_with_citations = [
{
"claim": "The transformer model was proposed by Vaswani et al. in 2017.",
"citation_ids": [1], # Correctly cited
},
{
"claim": "BERT was first introduced in 2017.", # Wrong date!
"citation_ids": [
2
], # Passage 2 actually says 2019, but claim says 2017
},
{
"claim": "Transformers use self-attention instead of recurrence.",
"citation_ids": [
3
], # Completely wrong citation; passage 3 is about the Amazon
},
]
# Evaluate attribution
evaluation_results = evaluate_attributed_response(
claims_with_citations, passages, nli_model
)Attribution Evaluation Results ============================================================ Claim: "The transformer model was proposed by Vaswani et al. in 2017." Supported: YES [Passage 1] Entailment: 0.995 | Contradiction: 0.002 Preview: The transformer architecture was introduced in the paper 'Attention is All You N... Claim: "BERT was first introduced in 2017." Supported: NO [Passage 2] Entailment: 0.000 | Contradiction: 0.960 Preview: BERT (Bidirectional Encoder Representations from Transformers) was published by ... Claim: "Transformers use self-attention instead of recurrence." Supported: NO [Passage 3] Entailment: 0.000 | Contradiction: 1.000 Preview: The Amazon rainforest covers approximately 5.5 million square kilometers and is ...
The results show the behavior we want from an attribution evaluator. The first claim (transformer proposed in 2017) is well-supported by passage 1 and receives a high entailment score. The second claim (BERT in 2017) gets a lower entailment score because passage 2 says 2019, not 2017. Notice that the passage does not merely fail to confirm the date: it actively contradicts it, so the contradiction score should be higher. The third claim (about transformers and self-attention) receives near-zero entailment from passage 3, which is about the Amazon rainforest. The NLI model correctly recognizes that this passage provides no support for a claim about neural architectures.
Computing Attribution Metrics
We compute citation precision and recall across the evaluation results. These aggregate metrics summarize the quality of the attribution system in terms the precision-recall framework we defined earlier.
def compute_citation_metrics(evaluation_results: list) -> dict:
"""
Compute citation-level precision, recall, and F1.
Precision: fraction of cited passages that support the claim.
Recall: fraction of supported claims that have at least one supporting citation.
"""
total_citations = 0
supported_citations = 0
total_claims = len(evaluation_results)
supported_claims = 0
for res in evaluation_results:
for cr in res["citation_results"]:
total_citations += 1
if cr["is_supported"]:
supported_citations += 1
if res["claim_supported"]:
supported_claims += 1
citation_precision = (
supported_citations / total_citations if total_citations > 0 else 0.0
)
claim_recall = supported_claims / total_claims if total_claims > 0 else 0.0
f1 = (
2
* citation_precision
* claim_recall
/ (citation_precision + claim_recall)
if (citation_precision + claim_recall) > 0
else 0.0
)
return {
"citation_precision": citation_precision,
"claim_recall": claim_recall,
"f1": f1,
"total_citations": total_citations,
"supported_citations": supported_citations,
"total_claims": total_claims,
"supported_claims": supported_claims,
}
metrics = compute_citation_metrics(evaluation_results)Citation Metrics ---------------------------------------- Citation Precision: 0.333 (1/3 citations are actually supportive) Claim Recall: 0.333 (1/3 claims are supported by at least one citation) F1 Score: 0.333
These metrics translate directly to system quality. A system with citation precision 0.33 means two-thirds of the citations it provides do not support the claim they are associated with. A system with claim recall 0.33 means two-thirds of the claims it makes are not supported by any of the provided citations. Both failures erode trust in the attributed response in different ways: low precision makes citations unreliable, while low recall means the response has claims that are essentially assertions made without evidence.
Key Parameters
The key parameters for attribution evaluation are:
- Entailment threshold: The minimum entailment probability required to consider a passage supportive. Higher thresholds produce fewer false positives but increase citation misses. The 0.5 threshold used in our example is a reasonable starting point but should be tuned on validation data for each application.
- NLI model: The quality of attribution evaluation is bounded by the NLI model's accuracy. Larger, better models (e.g., DeBERTa-v3-large) will evaluate attribution more accurately than smaller ones. For high-stakes applications, using a larger NLI model for evaluation is worthwhile even if the generation model itself is fast.
- Evaluation granularity: Evaluating at the sentence level is more precise than the paragraph level, but requires accurate sentence segmentation of the generated text. Sentence segmenters introduce their own errors, particularly on lists, numbered items, and sentences containing quotations or parenthetical remarks.
- Passage truncation: Long passages may need to be truncated before being passed to the NLI model, since most NLI models have a maximum input length. Truncating a passage can remove the very sentence that contains the supporting evidence, causing false negatives. Strategies for handling long passages include sliding-window evaluation (evaluate the claim against multiple overlapping windows of the passage and take the maximum score) or passage chunking.
Visualizations
Let us visualize citation evaluation behavior across a range of entailment thresholds and examine how different claim types behave under NLI-based evaluation.



Comparing Attribution Approaches
The four attribution forms differ in granularity and in the tradeoffs they impose on system designers. Document-level attribution is cheapest to implement and evaluate but gives users little actionable information about which specific passage supports which specific claim. Span-level attribution provides maximum transparency but demands careful training data and imposes a synthesis constraint on generation that can reduce fluency.
When choosing an attribution strategy for a production system, you need to weigh four practical dimensions: implementation effort, the user experience you want to provide, the cost of false citations in your domain, and the capabilities of the underlying language model.
For internal tooling where users are sophisticated and primarily care about whether to trust a response, document-level attribution often suffices. For public-facing products in regulated domains like medicine or law, inline or span-level attribution is almost always worth the additional implementation effort, because users need to be able to verify individual claims. For research assistance tools, the balance depends on whether the primary use case is exploration (where fluency and comprehensiveness matter most) or fact verification (where precision matters most).

Notice that transparency and complexity track together: the approaches that give users the most useful information also require the most engineering investment. This is not a coincidence. Providing per-claim attribution requires the system to maintain claim-level provenance throughout the generation process, which is structurally more complex than producing a response and then attaching document-level metadata to it.
Training Models for Attribution
Producing reliable attribution requires more than inference-time changes; it requires models that have been trained to attribute their claims. This training dimension is often underappreciated in discussions of attribution evaluation, which focus on how to measure attribution quality without asking how to produce it in the first place.
The primary approach is supervised fine-tuning on attributed corpora: large collections of (query, retrieved passages, attributed response) triples where the response contains inline citations that have been verified as correct. These training examples teach the model the behavioral pattern of citing passages as it generates text, rather than generating freely and adding citations as an afterthought.
Collecting high-quality attributed training data is expensive. Human annotators must read passages and generated responses, identify which passage supports which claim, and verify that the attributions are accurate. This annotation process requires careful instruction because the judgments are difficult: what counts as sufficient support for a claim? When is a paraphrase close enough to the original passage to count as attributed? When does the model's synthesis go beyond what any single passage supports?
One approach to reducing annotation costs is to use existing attributed text as supervision. Academic papers with in-text citations are a natural source of attributed text: each sentence with a citation is an example of a claim linked to a source. Preprocessing these papers, extracting the cited passages, and treating each in-text citation as a training example creates a large supervised dataset at relatively low cost. The challenge is that academic citation conventions differ from the in-context RAG attribution we typically want: academic citations point to external papers rather than to retrieved passages in the model's context, and academic authors sometimes cite papers for background rather than for specific factual support.
Reinforcement learning approaches can also be used to train for attribution. The reward signal is attribution accuracy: given a response with citations and the corresponding retrieved passages, an NLI-based judge assigns a reward based on how well each citation is supported. The model is trained to maximize this reward using policy gradient methods. This approach sidesteps the need for human annotation of training examples, since the NLI judge provides the training signal automatically. The limitation is that the NLI judge's errors become the model's training signal, so systematic NLI errors will be reinforced rather than corrected.
A third approach is constitutional AI methods applied to attribution: the model generates responses, critiques its own citation quality using explicit attribution criteria ("does this citation actually support this claim?"), and revises its response to improve attribution. This iterative self-critique can improve citation quality without requiring external supervision, though it depends on the model's ability to accurately assess its own attribution accuracy, which is itself a form of the evaluation problem we are trying to solve.
The Role of Prompt Engineering
Even without fine-tuning, prompt engineering can significantly affect how well a language model attributes its claims. Models that have been instruction-tuned to follow detailed prompts can be directed to produce attributed responses by specifying the citation behavior explicitly in the prompt.
Several prompt patterns have been found effective in practice. Providing numbered passages with clear delimiters, then explicitly instructing the model to cite passage numbers, typically produces better attribution than simply providing context without instructions. Including a formatting example in the prompt (few-shot attribution examples) further improves citation consistency. Asking the model to explain its reasoning before giving its final answer can also improve attribution quality, because the reasoning step forces the model to identify which passages it is drawing on before it writes the cited response.
However, prompt engineering for attribution has limits. A model that has not been trained on attributed generation tasks will frequently cite passages selectively, misattribute claims, or produce citations that are syntactically correct but semantically wrong. The model may have learned the surface form of citation (putting [1] after a claim) without learning the underlying behavior (ensuring the claim is entailed by the cited passage). Distinguishing surface-form citation from valid attribution behavior is one reason evaluation frameworks like ALCE and AIS are important: they look past the presence of citations to assess whether those citations are correct.
Limitations and Practical Challenges
Attribution and citation evaluation face a set of persistent technical and conceptual challenges that current systems have not fully solved.
NLI models are imperfect judges. The automatic evaluation of whether a passage supports a claim depends on NLI models, and these models have their own error rates. They may fail on claims that require background knowledge not present in the passage, on claims that involve negation (a passage that says "X does NOT cause Y" can be misclassified as supporting "X causes Y"), or on claims that are correct by common knowledge but go beyond what any single passage states. NLI models trained on standard entailment datasets may also be poorly calibrated for the specific types of claims that appear in specialized domains like medicine, law, or finance.
Multi-hop attribution is hard. Some claims require combining information from multiple passages to verify. "Compound A was synthesized by the team that also discovered Compound B" requires knowing who synthesized A and who discovered B from potentially different passages. Current NLI-based evaluation treats each passage-claim pair independently and cannot handle this kind of multi-hop reasoning. Systems that aggregate evidence across passages are more capable but much harder to implement and evaluate.
Granularity mismatch. Attribution evaluation works best when claims are simple, atomic statements. Real model outputs often contain complex, multi-part sentences that mix claims from different sources. Automatically segmenting a generated paragraph into evaluable claims is a non-trivial task that introduces its own errors. Sentence boundaries do not always correspond to claim boundaries: a single sentence can contain multiple independent factual claims, and a claim can sometimes span multiple sentences.
Coverage vs. accuracy. Models can improve their measured attribution metrics by being conservative: cite fewer things, and the things you do cite are more likely to be accurate. But a highly conservative system that cites nothing achieves perfect precision while providing zero value. Evaluation frameworks must penalize both over-citation and under-citation, which requires defining what "adequate citation coverage" looks like for a given query, a question that does not have a universal answer.
Source quality is out of scope for attribution. Attribution evaluation checks whether a claim matches its cited source; it does not check whether the source is correct. A perfectly attributed response that cites a paper with fabricated data will score well on attribution metrics while misleading the reader. This is why source curation and corpus quality are architectural concerns, not just evaluation concerns.
Temporal decay. Even a correctly attributed claim may become outdated. A source from 2019 may have been superseded by 2024 research. Attribution evaluation that ignores the temporal dimension of sources will undervalue the importance of citation freshness. In fast-moving fields, a well-attributed but stale response can be just as misleading as a hallucinated one. Systems that track citation dates and flag outdated sources add a valuable additional layer to attribution quality.
The annotation agreement problem. Human evaluation of attribution quality suffers from lower inter-annotator agreement than might be expected. Judgments about whether a passage "supports" a claim involve interpretive choices: how strong does the evidence need to be? Does a passage that is consistent with a claim without directly stating it count as support? Different annotators apply different standards, which makes it difficult to create gold-standard evaluation datasets with high confidence that the labels are correct. This ambiguity flows through to the NLI models trained on these labels, which inherit the annotation disagreements as model uncertainty.
Despite these limitations, attribution evaluation is substantially more informative than accuracy evaluation alone. A system that is accurate but uncited is just as unverifiable as a system that hallucinates, from the user's perspective. Attribution provides the transparency layer that separates a trustworthy AI assistant from an oracle that demands blind faith. The imperfections of current attribution methods do not undermine the importance of the goal; they motivate continued work on better evaluation frameworks, better training methods, and better source curation strategies.
Summary
Attribution and citation form the accountability layer of language AI systems. The core insight is that correctness and verifiability are different properties: a model can be correct without being verifiable, and an attributed model provides verifiability as an additional guarantee on top of correctness.
Key takeaways from this chapter:
- Attribution connects generated claims to supporting sources. Forms range from document-level (which sources were used) to span-level (exactly which passage supports each phrase), with different granularities offering different tradeoffs between implementation complexity and user transparency.
- In RAG systems, attribution is the additional requirement that models indicate which retrieved passages support which parts of their output. The pipeline has distinct stages for retrieval, prompt construction, generation, and attribution verification.
- Attribution accuracy failures take several forms: unsupported citations, overclaiming, missing citations, citations to the wrong source, and unattributed copying. Each failure mode requires different detection strategies and has different consequences for users.
- Evaluation frameworks like AIS and ALCE assess both claim correctness and citation quality. AIS defines attribution as verifiability from cited sources alone, while ALCE evaluates attribution across multiple NLP tasks using a combination of automatic and human judgment.
- Citation precision and recall measure, respectively, the fraction of citations that are supportive and the fraction of supported claims that are cited. The F1 score summarizes the tradeoff between the two.
- NLI-based attribution evaluation automates the judgment of whether a passage entails a claim, but NLI models have their own error rates and cannot handle multi-hop inference. The entailment threshold is a key parameter that governs the precision-recall tradeoff.
- Source verification is conceptually distinct from attribution evaluation: attribution checks whether a claim matches its source, not whether the source is itself correct. Both layers are needed for full reliability.
- Training models for attribution requires supervised data with verified citation chains. Prompt engineering can improve attribution behavior at inference time, but does not fully substitute for training-time exposure to attributed generation examples.
- Key tensions include the fluency-attribution tradeoff (synthesis improves response quality but complicates attribution), the precision-recall tradeoff (conservative citation improves precision but reduces coverage), and the granularity tradeoff (finer attribution is more informative but harder to produce and evaluate).
The next chapter covers uncertainty quantification: how models can communicate their own confidence in their outputs, which complements attribution by adding calibrated probability estimates alongside source evidence.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about attribution and citation in language models.
Attribution and Citation Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!