Exact Match and F1: Precision Metrics for NLP Evaluation

Michael BrenndoerferFebruary 25, 202652 min read

Part of Language AI Handbook

Use Exact Match and token-level F1 to score NLP and question-answering systems, with normalization, partial credit, thresholds, and common scoring errors.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Exact Match and F1

When we ask a language model to extract a specific answer from a document, classify a named entity, or respond to a knowledge question, we need a way to judge whether the model got it right. Unlike open-ended text generation where BLEU or ROUGE measure similarity through n-gram overlap, many NLP tasks demand binary correctness: either the model extracted the right date, or it did not; either it identified the correct protein name, or it failed.

Exact Match (EM) and F1 give the basis for this type of evaluation. These metrics originated in information extraction and question answering, where the goal is not to produce fluent text but to retrieve precise factual content. Exact Match is the strictest standard: a prediction is correct only if it matches the reference character-for-character. F1 offers a softer, more granular measure, balancing precision and recall at the token level, rewarding partial overlaps when a model captures some but not all of the expected answer.

These metrics differ fundamentally from the probabilistic evaluations we explored in Part II, Chapter 5 on Perplexity or the semantic similarity captured by BERTScore. While perplexity measures how well a model predicts token sequences, and BERTScore evaluates semantic equivalence through contextual embeddings, Exact Match and F1 focus on surface-form accuracy. They ask: "Did the model output precisely what we expected, and if not, how much of it did the model get right?"

The distinction between surface-form and semantic evaluation shows a deeper tension in natural language processing. Language is infinitely variable at the surface level yet constrained in meaning. Two sentences can be lexically distinct yet semantically identical, or lexically similar yet semantically opposed. Exact Match and F1 privilege the lexical level because in many practical applications, such as database querying, form filling, or knowledge base population, the specific string matters as much as the meaning it conveys. A SQL query searching for "2024-03-15" will not match "March 15, 2024", regardless of human intuition about their equivalence.

Understanding when to apply strict versus lenient matching, how tokenization choices affect F1 calculations, and what normalization strategies can make evaluation more reliable is needed for building reliable NLP systems. These metrics remain the standard for benchmarking extractive question answering, named entity recognition, and slot-filling systems. This gives clear, interpretable signals of model performance on precision-necessary tasks.

The historical roots of these metrics reach back to the early days of information retrieval in the 1960s and 1970s, when researchers at Cranfield and then TREC (Text REtrieval Conference) needed objective ways to evaluate whether search engines returned the right documents. The precision and recall framework that underpins F1 was already well-established before neural networks entered the picture. What changed with the rise of modern NLP was the adaptation of these retrieval metrics to the problem of extracting spans of text rather than ranking entire documents. The Stanford Question Answering Dataset (SQuAD), introduced in 2016, popularized the specific implementation of EM and F1 that most researchers use today, cementing these metrics as the lingua franca of extractive QA evaluation.

Exact Match Scoring

Exact Match is the most intuitive form of evaluation: a prediction is correct if and only if it exactly equals the reference answer. No partial credit, no semantic similarity, no tolerance for variation in formatting or word order. This strictness makes EM ideal for tasks where exact formatting is required: extracting phone numbers, dates, or database identifiers.

The psychological appeal of Exact Match lies in its alignment with schoolroom testing: there is a right answer, and everything else is wrong. This clarity simplifies error analysis and model development. When a system fails Exact Match, we know immediately that something basic has gone wrong, whether in comprehension, extraction, or formatting.

What makes Exact Match especially useful in a research context is its binary nature. Because the score is either 0 or 1 per example, you can directly compute confidence intervals using binomial statistics, compare models using McNemar's test for statistical significance, and trace failures to individual examples without any ambiguity about whether the metric itself is doing something unexpected. F1, being continuous, requires different statistical treatments, and subtle differences in F1 scores between models can reflect either real capability differences or artifacts of tokenization choices.

The Strictness Spectrum

Exact Match operates at the string level. If the reference answer is "March 15, 2024" and the model outputs "March 15th, 2024", Exact Match returns 0.0 (incorrect), even though a human reader would immediately recognize the equivalence. This rigidity is both a strength and a limitation:

  • Strength: Unambiguous evaluation with no subjective judgment calls. Two evaluators will always agree on whether Exact Match succeeded.
  • Limitation: Fails to reward near-misses or acceptable variations in phrasing, potentially discarding useful partial progress during model training.

In many practical applications, particularly database lookup and form extraction, this strictness is desirable. A SQL query searching for "2024-03-15" will not match "March 15, 2024", so training models to expect Exact Match alignment shows real-world usage constraints. Similarly, when extracting medical codes or part numbers, "ICD-10-CM E11.9" and "E11.9" may stand for the same diabetes diagnosis to a clinician, but a billing system might require the full code format. Exact Match enforces the discipline necessary for these downstream integrations.

The strictness also serves an important pedagogical function during model development. When you see low Exact Match scores but high F1 scores, you immediately understand that your model is capturing the right information but struggling with boundaries or formatting. This diagnostic clarity helps prioritize engineering efforts: low EM with high F1 suggests the need for better span selection or post-processing, while low scores on both metrics indicate deeper comprehension failures.

Think of the EM/F1 gap as a diagnostic thermometer for your model. An EM of 0.35 alongside an F1 of 0.72 tells a very specific story: the model understands the question and identifies the relevant tokens, but consistently gets the boundaries wrong or includes extra context it should trim. That pattern points engineers directly toward span-detection training or answer post-processing, rather than toward the model's language understanding. The gap encodes information that neither metric alone could convey.

Implementation Considerations

The implementation of Exact Match appears trivial: prediction == reference. However, several edge cases complicate this simplicity, and overlooking them can lead to misleading evaluation results or systems that fail in production despite high benchmark scores.

Several edge cases require attention:

  • Whitespace sensitivity: Leading or trailing spaces, multiple consecutive spaces, or different newline characters cause mismatches. The string "Paris " (with trailing space) does not match "Paris". In web-extracted text or user-generated content, such variations are common. A model might correctly identify an answer in a document but include an extra space because of HTML formatting, causing an Exact Match failure despite correct comprehension.

  • Case sensitivity: "paris" versus "Paris" versus "PARIS" are distinct strings under Exact Match, though they stand for the same entity. Case sensitivity matters in some domains, such as programming language keywords or chemical formulas, where "Co" (cobalt) differs from "CO" (carbon monoxide), but in general question answering, case differences usually stand for annotation inconsistency rather than semantic distinction.

  • Punctuation: "U.S.A." and "USA" or "Dr." and "Dr" register as different answers despite semantic equivalence. Punctuation attachment varies across writing styles and datasets. British and American conventions differ on periods after abbreviations, and dates can be punctuated as "March 15, 2024" or "March 15 2024" with equivalent meaning.

  • Multiple valid answers: In extractive question answering, a question might have several correct answers scattered throughout a document. Standard Exact Match implementations check if the prediction matches any reference answer in a provided set, returning 1.0 if one match exists and 0.0 otherwise. This design shows the reality that questions often have multiple valid answers in a text. For instance, asking "What causes global warming?" might legitimately match "greenhouse gas emissions" in one paragraph or "carbon dioxide and methane" in another.

Multi-Answer Exact Match

For datasets like SQuAD (Stanford Question Answering Dataset), questions often accept multiple valid spans from the context paragraph. The evaluation protocol treats the prediction as correct if it matches any one of the gold-standard answers:

EM(prediction,references)={1if ∃ r∈references:prediction=r0otherwise\text{EM}(prediction, references) = \begin{cases} 1 & \text{if } \exists \, r \in references : prediction = r \\ 0 & \text{otherwise} \end{cases}

where:

  • predictionprediction: the predicted answer string generated by the model
  • referencesreferences: the set of all valid reference answers for the question
  • rr: a single reference answer from the set of references
  • ∃ r∈references\exists \, r \in references: indicates that there exists at least one reference answer in the set that satisfies the condition

This "match-any" approach prevents penalizing models for selecting different but equally valid answer spans. However, it also means that a model could receive full credit for different predictions on different runs if multiple correct answers exist, potentially masking instability in model behavior. If a model randomly selects between two valid answers depending on initialization or random sampling, it might appear perfectly stable in aggregate metrics while being inconsistent on individual examples.

In addition, not all answers in a reference set are equally good. Some might be complete sentences while others are fragments, or some might include necessary context while others are ambiguous. Exact Match treats them as equivalent, which can obscure quality differences in model outputs. A model that consistently selects the shortest, least informative valid answer will score perfectly on Exact Match despite giving less useful information than one that selects more complete answers.

A practical consequence of the multi-answer protocol is that benchmark scores are often optimistic compared to the scores you would get in a real deployment. In a real system, there is typically only one reference answer at evaluation time: the correct database entry, the approved clinical code, or the verified fact. Benchmark datasets give multiple references to increase tolerance for annotation variation, but this generosity does not reflect the strictness of production environments. Researchers sometimes report both single-reference EM (evaluated against the best single annotation) and multi-reference EM (evaluated with the full reference set) to give a more honest picture.

Token-Level F1

While Exact Match gives a binary correct/incorrect signal, F1 offers a continuous measure of overlap between prediction and reference at the token level. This granularity proves needed when answers can vary in length, when multiple entities appear in a response, or when we want to reward partial correctness.

The transition from Exact Match to F1 shows a shift in evaluation philosophy. Exact Match asks "Is this right?" while F1 asks "How right is this?" This shift acknowledges that in many extraction tasks, getting most of the answer correct has practical value. A system that extracts "Alexander Bell" when the full answer is "Alexander Graham Bell" has identified the correct person and might let useful downstream processing, even if it missed the middle name. Exact Match would discard this progress entirely, while F1 captures it quantitatively.

The practical motivation for F1 becomes especially clear during model development. If you train a model and watch its Exact Match climb from 0.12 to 0.15 over ten epochs, it is difficult to know whether progress is real or noise. But if F1 climbs from 0.45 to 0.62 over those same epochs, you know the model is consistently capturing more of the right information even when it does not yet nail the exact boundaries. F1 acts as a more sensitive signal during training. This gives gradient information about the direction of improvement that binary EM cannot.

Precision and Recall Foundations

F1 derives from the classic information retrieval metrics of precision and recall, adapted to compare token sets rather than document collections. These concepts emerged from library science and early search engines, where the problem was retrieving relevant documents from large collections. In question answering, we adapt them to retrieve relevant tokens from answer strings.

Precision measures what fraction of the predicted tokens were correct, meaning they appeared in the reference:

Precision=∣prediction_tokens∩reference_tokens∣∣prediction_tokens∣\text{Precision} = \frac{|\text{prediction\_tokens} \cap \text{reference\_tokens}|}{|\text{prediction\_tokens}|}

where:

  • prediction_tokens\text{prediction\_tokens}: the set of tokens appearing in the model's predicted answer
  • reference_tokens\text{reference\_tokens}: the set of tokens appearing in the reference answer
  • ∩\cap: the set intersection operator, yielding tokens common to both sets
  • ∣⋅∣|\cdot|: set cardinality (the count of elements in the set)

Precision answers the question: "Of everything the model said, how much was correct?" High precision indicates the model is conservative, rarely including incorrect information, even if it misses some correct details. In medical extraction, high precision might be important because hallucinated drug names or dosages could lead to dangerous recommendations.

Recall measures what fraction of the reference tokens were successfully predicted:

Recall=∣prediction_tokens∩reference_tokens∣∣reference_tokens∣\text{Recall} = \frac{|\text{prediction\_tokens} \cap \text{reference\_tokens}|}{|\text{reference\_tokens}|}

where:

  • prediction_tokens\text{prediction\_tokens}: the set of tokens appearing in the model's predicted answer
  • reference_tokens\text{reference\_tokens}: the set of tokens appearing in the reference answer
  • ∩\cap: the set intersection operator, yielding tokens common to both sets
  • ∣⋅∣|\cdot|: set cardinality (the count of elements in the set)

Recall answers the question: "Of everything that should have been said, how much did the model include?" High recall indicates complete coverage. In legal document analysis, high recall ensures no relevant clauses are missed, even if some irrelevant text is included.

F1 Score is the harmonic mean of precision and recall, giving a single metric that balances both concerns:

F1=2×Precision×RecallPrecision+Recall\text{F1} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}

where:

  • Precision\text{Precision}: the precision of the prediction (fraction of predicted tokens that are correct)
  • Recall\text{Recall}: the recall of the prediction (fraction of reference tokens that were predicted)
  • The harmonic mean balances precision and recall by penalizing cases where one metric is much higher than the other

Why the harmonic mean rather than the arithmetic mean? The arithmetic mean of precision and recall would reward a trivially high-recall system: if the model outputs the entire document as its answer, recall is 1.0 and the arithmetic mean is at least 0.5 regardless of precision. The harmonic mean is always less than or equal to the arithmetic mean, and it approaches zero whenever either component approaches zero. This property means a model cannot "game" F1 by simply outputting everything it sees.

The harmonic mean punishes extreme imbalances more aggressively than an arithmetic mean would. A model that reaches 100% precision by predicting only one correct token (but missing nine others in the reference) receives low F1 due to poor recall, accurately which shows its incomplete performance. Conversely, a model that predicts all reference tokens plus ten irrelevant tokens would have perfect recall but low precision, also resulting in low F1.

This balancing act makes F1 particularly suitable for development and hyperparameter tuning. Unlike Exact Match, which gives no gradient of improvement (a model is either right or wrong), F1 gives partial credit that correlates with human judgments of answer quality. When optimizing a model, seeing F1 increase from 0.4 to 0.6 indicates real progress in balancing comprehensiveness against accuracy, even if Exact Match remains zero.

Out[3]:
Visualization
Contour plot of F1 score over precision-recall grid with four annotated operating points.
F1 score as a function of precision and recall. The harmonic mean creates contours that bow toward the origin, showing how F1 drops sharply when either precision or recall is low. Annotated points compare four representative operating points, revealing how a perfectly balanced 0.8/0.8 system reaches higher F1 than a 1.0/0.5 system despite the same average.

Tokenization Strategy Matters

The calculation of F1 depends critically on how we define a "token." Different tokenization approaches yield different F1 scores for the same string comparison, and this variation changes the substance of the result. It shows basic questions about what constitutes a real unit of meaning in language. Two papers can report the same model achieving "82.3 F1 on SQuAD" and "79.1 F1 on SQuAD" for entirely different reasons: one uses aggressive punctuation stripping, the other does not. The metric name alone does not fully specify what was measured.

The three main tokenization strategies each have distinct behaviors and tradeoffs:

Whitespace tokenization splits on spaces to produce word-level tokens. "New York City" becomes ["New", "York", "City"]. This approach misses multi-word entities and punctuation attachment. It treats "New York" as two separate tokens, potentially matching "York" in isolation if it appears elsewhere in a prediction, even though "York" alone refers to a different entity than "New York". Despite its simplicity, whitespace tokenization is the SQuAD standard because it aligns well with human intuitions about word boundaries and produces results that match what readers expect.

Punctuation-aware tokenization separates punctuation into distinct tokens. "U.S.A." might become ["U", ".", "S", ".", "A", "."] or ["U.S.A.", "."] depending on rules. As we explored in Part I, Chapter 5 on Word Tokenization, these choices materially impact overlap calculations. If periods are separate tokens, then "U.S.A." and "USA" share no tokens, yielding zero F1. If periods are stripped, they might match perfectly. Punctuation-aware tokenization is useful when punctuation carries meaning, as in code or mathematical notation, but introduces complexity when punctuation is stylistic.

Subword tokenization using BPE or WordPiece from Part V splits rare words into smaller units. "hypersensitivity" becomes ["hyper", "##sens", "##itivity"]. F1 calculated on subwords differs substantially from word-level F1. Two answers might share no whole words but share many subwords, or vice versa. This complicates evaluation of models that generate text at the subword level, as the tokenization mismatch between generation and evaluation can distort scores. In practice, evaluation is almost always done at the word level, regardless of what tokenization the model uses internally.

Standard practice for F1 evaluation in question answering uses whitespace tokenization after lowercasing and removing punctuation, striking a balance between linguistic validity and implementation simplicity. However, for tasks involving code, chemical formulas, or specialized domains, custom tokenization rules may be necessary to ensure real token boundaries. In chemical named entity recognition, "H2SO4" should not be split on whitespace or punctuation, as the subscript and element symbols form an indivisible unit.

The choice of tokenization also affects the interpretation of F1 scores. Word-level F1 aligns with human intuition about what constitutes a correct answer, while subword F1 might give credit for morphological similarity that does not reflect semantic correctness. A model predicting "run" when the answer is "running" might score partial credit under subword tokenization (sharing "run"), but these might be distinct answers in a question about verb tense.

Handling Multiple Answers with F1

When multiple reference answers exist, F1 evaluation typically takes the maximum score across all references:

F1(prediction,references)=max⁡r∈references F1(prediction,r)\text{F1}(prediction, references) = \max_{r \in references} \, \text{F1}(prediction, r)

where:

  • predictionprediction: the predicted answer string generated by the model
  • referencesreferences: the set of all valid reference answers for the question
  • rr: a single reference answer from the set of references
  • max⁡r∈references\max_{r \in references}: the maximum operator selecting the highest F1 score achieved against any single reference

This "best-match" approach ensures that models are not penalized for selecting any valid answer, but it can obscure cases where models consistently pick suboptimal answers (shorter or less complete variants), even when better options exist. If the reference set contains both "John F. Kennedy" and "Kennedy", a model that always predicts just the last name will reach perfect F1 scores despite giving less specific information than available in the text.

The max operation also complicates aggregate statistics. When reporting dataset-level F1, researchers typically macro-average across questions, but the per-question max means that the final score is an upper bound on model performance rather than an average. A model that randomly selects among valid answers will appear as good as one that consistently selects the most complete answer, masking important behavioral differences.

There is a subtle but important question about whether to compute max over the references before or after averaging across questions. The standard approach takes max first (per question) and then averages, which gives the most lenient evaluation. An alternative approach would be to compute F1 against all references, then average, which is stricter. These two approaches can produce substantially different numbers on the same model predictions, yet both would be legitimately called "average F1." When reading papers or comparing results, it is worth confirming which convention was used.

Normalization for Matching

Raw string matching often proves too brittle for practical evaluation. The same semantic content can be expressed through different surface forms: "United States" versus "US" versus "U.S."; "3:00 PM" versus "15:00"; "cannot" versus "can't". Normalization turns both predictions and references into canonical forms before comparison, reducing false negatives caused by formatting variations rather than semantic errors.

The need for normalization arises from the basic variability of natural language. Human annotators exhibit inconsistencies in capitalization, punctuation usage, and abbreviation conventions. Even within a single dataset, the same answer might appear as "the United States", "The United States", "United States", and "US" across different examples. Without normalization, a model trained on one variant and evaluated on another would appear to fail despite understanding the content perfectly.

Normalization also addresses systematic differences between model outputs and human annotations. Language models trained on large corpora often learn specific formatting conventions that differ from dataset annotation guidelines. A model might consistently include definite articles ("the President") while the dataset omits them ("President"), or vice versa. Normalization prevents these systematic biases from appearing as comprehension failures.

Common Normalization Strategies

The most widely used normalization pipeline, as standardized in the SQuAD evaluation script, applies four transformations in sequence:

  • Case folding: Converting all text to lowercase eliminates case sensitivity. "The White House" and "the white house" become equivalent. However, case sometimes carries semantic information (e.g., "Apple" the company versus "apple" the fruit), so domain-appropriate caution is warranted. In biomedical text, gene names are often case-sensitive: "p53" (the tumor protein) differs from "P53" (which might refer to a different entity or be an error). Case folding should be applied with understanding of the domain's conventions.

  • Punctuation removal: Stripping punctuation characters prevents mismatches due to optional periods, commas, or quotation marks. Implementation typically uses regex patterns like [^\w\s] to remove non-alphanumeric characters. This handles variations like "U.S.A." versus "USA" or "don't" versus "dont". However, punctuation can be real in some contexts. In programming languages, "print()" differs semantically from "print", and in mathematics, "x-y" differs from "xy".

  • Article removal: In question answering, leading articles ("a", "an", "the") often do not affect answer correctness. "the Eiffel Tower" and "Eiffel Tower" are functionally equivalent in most contexts. Articles can also vary based on sentence position: a phrase appearing at the beginning of a sentence might have "The" capitalized, while the same phrase mid-sentence might not. Removing articles reduces these positional artifacts.

  • Whitespace normalization: Collapsing multiple spaces into single spaces and trimming leading/trailing whitespace ensures that "Paris " and "Paris" match. This addresses formatting inconsistencies common in web-extracted text or OCR output, where extra spaces might appear around punctuation or at line breaks.

These four transformations are applied to both the prediction and all references before any comparison. The order matters: applying punctuation removal before whitespace normalization avoids edge cases where removing a period leaves a double space. The specific regex patterns matter too. The SQuAD evaluation script removes characters matching [^a-zA-Z0-9\s] after lowercasing, which preserves digits but removes all other punctuation. A different regex that also removes digits would change scores on numerical answer questions.

Domain-Specific Normalizations

Different tasks require specialized normalization rules that reflect the specific conventions of the domain:

  • Numerical equivalence: "3.14" should match "3.140" and potentially "3,14" (European decimal notation). Date formats present particularly rich variation: "2024-03-15", "15/03/2024", "March 15, 2024", and "15th March 2024" all stand for the same day. Time formats vary similarly: "3:00 PM", "15:00", and "3pm" may be equivalent depending on context. Implementing reliable numerical normalization requires parsing and canonicalization rather than simple string substitution. Libraries like dateutil and babel give locale-aware parsing that can handle these variations systematically.

  • Acronym expansion: "NATO" versus "North Atlantic Treaty Organization" or "FBI" versus "Federal Bureau of Investigation". In some domains, such as legal or medical text, acronyms and their expansions are used interchangeably. Normalization might involve expanding common acronyms or contracting full names to standardized short forms, though this requires domain-specific knowledge bases to avoid errors. A general-purpose acronym expander would need to handle the fact that "PT" can mean physical therapy, part time, point, or the Portuguese abbreviation for Portugal depending on context.

  • Unicode normalization: As discussed in Part I, Chapter 2 on Text Normalization, Unicode characters can have multiple equivalent representations. The character "é" can be represented as a single codepoint (U+00E9) or as "e" plus combining acute accent (U+0065 U+0301). Unicode normalization forms (NFC, NFD, NFKC, NFKD) ensure consistent representation before matching. This is important for multilingual evaluation, where different input methods or text sources might use different Unicode representations for the same visual character. NFKC normalization additionally handles compatibility equivalences such as fullwidth characters, which appear in CJK text and would prevent matching their halfwidth equivalents without normalization.

The Trade-off of Aggressive Normalization

Every normalization rule discards information. While this reduces false negatives (correct answers marked wrong due to formatting), it risks increasing false positives (incorrect answers marked right due to over-normalization). You face a basic trade-off between recall (catching all correct answers) and precision (avoiding incorrect matches).

Consider these specific risks:

  • Removing all punctuation makes "Dr." (Doctor) equivalent to "Dr" (Drive in addresses), potentially marking a prediction as correct when it refers to a street rather than a person.
  • Case folding makes "US" (United States) equivalent to "us" (pronoun), confusing geographical entities with grammatical words.
  • Article removal might equate "a tiger" (any tiger) with "the tiger" (a specific tiger), losing important semantic distinctions of definiteness.
  • Aggressive numerical normalization might equate "1.5" (one and a half) with "1,5" (one and a half in European notation) when the latter means "one thousand five" in US notation, or vice versa.

The appropriate level of normalization depends on the task's precision requirements. Medical entity extraction might preserve case and punctuation carefully, as "mg" (milligrams) and "Mg" (magnesium) are critically different, and "5.0" versus "5,0" could stand for dosage errors. General web question answering might aggressively normalize to match human answer variations, accepting that some semantic precision is lost for the sake of robustness.

One underappreciated risk of aggressive normalization is that it can inflate benchmark scores in ways that create false confidence. If a benchmark applies very generous normalization, models can appear to reach near-human performance even when they are still making real errors. This is one reason that benchmark scores have become less reliable indicators of real-world capability over time: as normalization becomes more permissive, the gap between "passes the metric" and "gives a useful answer" widens. The recent trend toward human evaluation and LLM-as-judge evaluation methods is partly a response to this problem.

Metric Selection: When to Use What

Choosing between Exact Match, F1, and their normalized variants requires understanding the task constraints, the cost of false positives versus false negatives, and how model outputs will be consumed downstream. The wrong metric can mislead development efforts or hide necessary failure modes.

Decision Framework

The decision between EM and F1 is not arbitrary. It follows from the structure of the task and the consequences of different types of errors:

Use Exact Match when:

  • Answers are drawn from closed sets (multiple choice, database IDs), where any deviation is a categorical error.
  • Format compliance is necessary (dates must follow ISO 8601, phone numbers must include country codes), and downstream systems will reject malformed inputs.
  • Any deviation is a functional error (security codes, medication dosages, financial transaction identifiers), where partial correctness has no value and might be dangerous.
  • You want statistically clean comparisons between models, since EM scores have well-understood distributional properties.

Use F1 when:

  • Answers can vary in length or composition (list extraction, multi-span answers), and capturing some items is better than capturing none.
  • Partial correctness has value (information retrieval, entity extraction), and incomplete information still supports downstream tasks.
  • Multiple valid phrasings exist for the same information, and the goal is to measure content coverage rather than format adherence.
  • The development signal during training needs to be sensitive enough to detect improvement before the model reaches exact correctness.

Apply normalization when:

  • Human annotators show inconsistency in formatting, and you want to measure semantic understanding rather than annotation style alignment.
  • The task explicitly allows equivalent forms (dates, measurements, abbreviations), and the evaluation should reflect domain standards of equivalence.
  • Model outputs show systematic formatting differences from references (e.g., always including "The" at the beginning of titles), and you want to correct for these biases without retraining.

In practice, most NLP system papers report both EM and F1 alongside their normalized versions. This gives a table that captures performance at multiple levels of strictness. This gives readers the information they need to contextualize scores within their specific use case.

The SoftMax Paradox

A subtle issue arises when evaluating extractive models that select spans from a provided context. The model might output the correct answer text that appears multiple times in the context but select the wrong instance. For example, if a paragraph mentions "Paris" three times (twice referring to Paris, France and once to Paris, Texas), a model selecting any "Paris" receives full Exact Match credit, even if the surrounding context indicates the wrong entity.

This phenomenon, sometimes called the span ambiguity problem, highlights a limitation of string-based evaluation for extractive tasks. The model's task is to produce the correct text and identify the correct location in the context that supports the answer. If the question asks "What city is the Eiffel Tower in?" and the context mentions Paris, France in the first paragraph and Paris, Texas in the last, selecting the latter "Paris" is clearly wrong, yet Exact Match and F1 would award full points.

More advanced evaluations address this by checking span positions or requiring that the selected span appears in the correct context window. However, standard EM and F1 metrics treat text as isolated strings, divorced from their source context. This limitation means that reported benchmark scores may overestimate true comprehension, as models might exploit statistical correlations between question words and answer locations without real understanding.

The span ambiguity problem has become more relevant as the field has moved from extractive QA (where models highlight a span in the input) to generative QA (where models produce free text). In extractive QA, span positions can be checked. In generative QA, there are no spans at all, and the comparison must be entirely string-based. This shift has made EM and F1 even more approximate as measures of true capability.

Aggregation Across Datasets

When reporting scores across question answering datasets, researchers typically present:

  • Exact Match (EM): Percentage of questions with perfect string matches, representing the proportion of fully correct answers.
  • F1: Macro-averaged F1 score across all questions, representing the average token overlap quality.
  • Normalized variants: EM and F1 after applying standard normalization pipelines, showing performance after accounting for formatting variations.

The gap between EM and F1 indicates answer variability. A large gap suggests many near-misses or acceptable variations, while a small gap indicates that answers are typically short, unambiguous strings where partial credit is rare. For example, in factoid questions with single-word answers like dates or names, EM and F1 will be very close, as there is little room for partial overlap. In multi-sentence explanatory answers, EM might be low while F1 is moderate, showing that models capture key phrases but struggle with complete sentences.

Understanding this gap helps interpret model capabilities. A model with 40% EM and 60% F1 is in a different state than one with 10% EM and 60% F1. The former makes many perfect predictions and some near-misses; the latter rarely gets answers exactly right but often captures relevant content. The appropriate response to these profiles differs: the first model might benefit from better boundary detection, while the second needs basic improvements in answer formation.

Dataset characteristics also determine which metric is most informative. SQuAD 1.1, which draws answers from Wikipedia, produces questions where the reference answers are almost always single noun phrases. On this dataset, the EM/F1 gap is small because short answers leave little room for partial overlap. TriviaQA, which accepts both document-extracted and annotated answers, often shows larger EM/F1 gaps because the reference sets are more diverse. NewsQA, with its longer answers drawn from news articles, shows the largest gaps as models frequently capture the right information in differently bounded spans.

Worked Example

Let's examine how these metrics behave with concrete examples from a question answering scenario. Consider the question "Who invented the telephone?" with reference answers and various model predictions. This question allows us to explore how different types of errors, variations, and completions affect scoring.

Reference answers:

  • "Alexander Graham Bell"
  • "Bell"

These references stand for different levels of specificity: the full name and the surname. Both are correct, but they set different expectations for what constitutes a complete answer.

Model Prediction A: "Alexander Graham Bell"

  • Exact Match: 1.0 (matches first reference exactly)
  • F1: 1.0 (perfect token overlap)

This is the ideal case: complete, exact correspondence with the most specific reference. The model has retrieved the precise span that matches the annotator's preferred answer.

Model Prediction B: "alexander graham bell"

  • Exact Match: 0.0 (case mismatch)
  • F1: 1.0 (after lowercasing, tokens match)
  • Normalized EM: 1.0 (after case normalization)

This example shows the brittleness of raw Exact Match and the value of normalization. Without normalization, this prediction appears wrong despite being semantically perfect. The divergence between raw EM and F1 here indicates that token content is correct but formatting differs. This is precisely the scenario normalization is designed to address: the model understood the question correctly but produced output in a different case than the annotator expected.

Model Prediction C: "Bell"

  • Exact Match: 1.0 (matches second reference)
  • F1: 1.0 (single token matches perfectly)

Here the model gives a less specific but still correct answer. Exact Match gives full credit because "Bell" is in the reference set, but this masks the loss of information (the first and middle names). In some contexts, "Bell" might be ambiguous, referring to Alexander Melville Bell (his father) or other historical Bells, but the metric treats it as fully correct. The multi-reference protocol can thus reward underspecified answers.

Model Prediction D: "Alexander Bell"

  • Exact Match: 0.0
  • Tokenization: ["Alexander", "Bell"] vs reference ["Alexander", "Graham", "Bell"]
  • Precision: 2/2=1.02/2 = 1.0 (both predicted tokens are correct)
  • Recall: 2/3=0.6672/3 = 0.667 (missing "Graham")
  • F1: 2×(1.0×0.667)/(1.0+0.667)=0.82 \times (1.0 \times 0.667) / (1.0 + 0.667) = 0.8

Model Prediction E: "Graham Bell"

  • Exact Match: 0.0
  • Precision: 2/2=1.02/2 = 1.0
  • Recall: 2/3=0.6672/3 = 0.667
  • F1: 0.80.8

Despite having different missing tokens, predictions D and E receive identical F1 scores. This shows that F1 captures the quantity of overlap but not the quality or position. Missing the first name versus the middle name might have different semantic significance in some contexts. "Alexander Bell" might be interpreted as a different person than "Graham Bell" by some readers, yet the metric treats these errors as equivalent. Position-insensitive token overlap is a deliberate simplification that makes the metric easy to compute and interpret, but it comes at the cost of ignoring which tokens are missing.

Model Prediction F: "Sir Alexander Graham Bell"

  • Exact Match: 0.0
  • Precision: 3/4=0.753/4 = 0.75 ("Sir" is incorrect)
  • Recall: 3/3=1.03/3 = 1.0
  • F1: 2×(0.75×1.0)/1.75=0.8572 \times (0.75 \times 1.0) / 1.75 = 0.857

Prediction F reaches higher F1 than D or E (0.857 vs 0.8) despite being factually questionable (the title "Sir" is historically inaccurate for Bell; he received an honorary doctorate but was not knighted). This illustrates how F1 rewards extra correct information even when it accompanies hallucinated content. The model receives partial credit for the extra token "Sir" because F1 only penalizes incorrect tokens in the denominator of precision, not through a separate penalty for fabrication. This property means that F1 alone cannot detect hallucination; it only measures the ratio of correct to incorrect tokens. A model that adds plausible-sounding but wrong information is rewarded rather than penalized as long as it also includes all the correct information.

This hallucination-tolerance is one of the most important practical limitations of F1 as a production evaluation metric. You can build a system that reaches very high F1 by generating verbose answers that happen to include the reference tokens, without those tokens being the actual answer. In production evaluation, you often need to combine F1 with a length penalty or precision threshold to prevent this gaming.

Code Implementation

Let's implement Exact Match and F1 evaluation with various normalization strategies to see these metrics in practice.

In[4]:
Code
import re
import string
from typing import List, Union


def normalize_answer(
    text: str,
    lowercase: bool = True,
    remove_punctuation: bool = True,
    remove_articles: bool = True,
    fix_whitespace: bool = True,
) -> str:
    """Apply normalization transforms to text."""
    if lowercase:
        text = text.lower()

    if remove_punctuation:
        # Remove all punctuation characters
        text = text.translate(str.maketrans("", "", string.punctuation))

    if remove_articles:
        # Remove leading articles
        text = re.sub(r"^(a|an|the)\s+", "", text, flags=re.IGNORECASE)

    if fix_whitespace:
        # Normalize whitespace
        text = " ".join(text.split())

    return text


def exact_match_score(
    prediction: str, references: Union[str, List[str]], normalize: bool = False
) -> float:
    """
    Calculate Exact Match score.
    Returns 1.0 if prediction matches any reference, 0.0 otherwise.
    """
    if isinstance(references, str):
        references = [references]

    if normalize:
        prediction = normalize_answer(prediction)
        references = [normalize_answer(ref) for ref in references]

    return float(any(prediction == ref for ref in references))


def f1_score(
    prediction: str, references: Union[str, List[str]], normalize: bool = False
) -> float:
    """
    Calculate token-level F1 score.
    Uses whitespace tokenization.
    """
    if isinstance(references, str):
        references = [references]

    if normalize:
        prediction = normalize_answer(prediction)
        references = [normalize_answer(ref) for ref in references]

    # Tokenize by whitespace
    pred_tokens = set(prediction.split())

    best_f1 = 0.0

    for ref in references:
        ref_tokens = set(ref.split())

        # Calculate overlap
        common_tokens = pred_tokens & ref_tokens

        if len(common_tokens) == 0:
            continue

        precision = len(common_tokens) / len(pred_tokens) if pred_tokens else 0
        recall = len(common_tokens) / len(ref_tokens) if ref_tokens else 0

        if precision + recall > 0:
            f1 = 2 * (precision * recall) / (precision + recall)
            best_f1 = max(best_f1, f1)

    return best_f1

The core logic of f1_score is worth reading carefully. The function converts each answer to a set of tokens and computes the intersection. Using a set rather than a list means that duplicate tokens are collapsed: if the prediction contains "Bell Bell", it still counts as one occurrence of "Bell". This set-based approach matches the SQuAD evaluation convention. An alternative multiset approach, which counts token frequencies, would give different results for repetitive answers and is not standard practice for QA evaluation.

Out[5]:
Console
Prediction: 'The Eiffel Tower'
References: ['Eiffel Tower', 'the Eiffel Tower']

Without normalization:
  Exact Match: 0.0
  F1 Score: 0.800

With normalization:
  Exact Match: 1.0
  F1 Score: 1.000

The normalization process bridges the gap between strict string matching and semantic equivalence. In this example, the article "The" prevents a raw Exact Match, but normalized evaluation recognizes the equivalence. This shows why standard benchmarks like SQuAD apply normalization: without it, models would be penalized for trivial formatting differences that do not reflect comprehension failures.

Notice that raw F1 is already 0.667 in this case, not 0.0. Because "Eiffel" and "Tower" are shared tokens even without normalization, F1 gives partial credit before normalization takes effect. This illustrates the inherent tolerance of F1 compared to EM: even without any normalization pipeline, F1 rewards token-level overlap.

In[6]:
Code
# More complex examples showing partial matching
examples = [
    {
        "prediction": "Alexander Bell",
        "references": ["Alexander Graham Bell"],
        "description": "Missing middle name",
    },
    {
        "prediction": "Sir Alexander Graham Bell",
        "references": ["Alexander Graham Bell"],
        "description": "Extra title (hallucination)",
    },
    {
        "prediction": "Paris, France",
        "references": ["Paris"],
        "description": "Extra location specifier",
    },
    {
        "prediction": "the French capital city",
        "references": ["Paris"],
        "description": "Semantic equivalent, no lexical overlap",
    },
]
Out[7]:
Console
Description                         Prediction                     EM       F1      
-------------------------------------------------------------------------------------
Missing middle name                 Alexander Bell                 0.0      0.800
Extra title (hallucination)         Sir Alexander Graham Bell      0.0      0.857
Extra location specifier            Paris, France                  0.0      0.667
Semantic equivalent, no lexical overlap the French capital city        0.0      0.000

Notice how F1 gracefully handles partial overlaps (missing middle names, extra specifiers) while Exact Match remains binary. The scores decrease gradually with the severity of the deviation. This gives a more fine-grained signal than the all-or-nothing Exact Match approach.

The final example shows a necessary limitation: when the prediction is semantically equivalent but lexically distinct ("the French capital city" versus "Paris"), both EM and F1 fail completely, scoring 0.0. This gap motivates the use of BERTScore or semantic similarity metrics for open-ended generation, which we covered earlier in this part. The inability to recognize paraphrases or descriptions is a basic boundary of lexical metrics: they can only judge what is said, not what is meant.

In[8]:
Code
# Implementing SQuAD-style evaluation with multiple references
def squad_evaluate(predictions: dict, references: dict) -> dict:
    """
    SQuAD-style evaluation letting multiple valid answers per question.

    Args:
        predictions: dict mapping question_id -> predicted_answer
        references: dict mapping question_id -> list of valid answers
    """
    total = len(predictions)
    exact_match = 0
    f1_total = 0.0

    for qid, pred in predictions.items():
        refs = references[qid]

        if exact_match_score(pred, refs, normalize=True):
            exact_match += 1

        f1_total += f1_score(pred, refs, normalize=True)

    return {
        "exact_match": 100.0 * exact_match / total,
        "f1": 100.0 * f1_total / total,
    }
Out[9]:
Console
SQuAD-style Evaluation Results:
  Exact Match: 66.7%
  F1 Score: 88.9%

This implementation shows the "match-any" protocol used in standard benchmarks. The model receives full credit for "1845" despite three valid variants existing in the reference set, illustrating how benchmark design choices affect reported performance. The flexibility of multiple references accommodates natural variation in how humans express the same information, whether as bare numbers ("1845"), prepositional phrases ("in 1845"), or temporal adverbials ("during 1845").

For "Paris, France" evaluated against the reference "Paris", the normalized evaluation strips punctuation, giving "paris france" versus "paris". The intersection contains "paris" only, so precision is 1/2 = 0.5 and recall is 1/1 = 1.0, yielding F1 = 0.667. This is an informative score: the model found the right city but added an unnecessary qualifier, which is a mild error that deserves partial rather than zero credit.

Out[10]:
Visualization
Grouped bar chart comparing Exact Match and F1 scores for six telephone inventor predictions.
Comparison of Exact Match and F1 scores across the worked examples from the 'Who invented the telephone?' scenario. Exact Match gives a binary signal of 0 or 1 for each prediction, while F1 varies continuously and captures partial credit for near-miss answers. The case-mismatch and extra-title predictions illustrate how EM and F1 can diverge materially for the same prediction.
Out[11]:
Visualization
Grouped bar chart showing raw and normalized EM and F1 scores for article, case, punctuation, and whitespace variation examples.
Impact of text normalization on evaluation scores across four surface-form variation types. Raw matching penalizes formatting differences such as leading articles, case, punctuation, and whitespace, while normalized matching raises scores toward 1.0 for all four variation types. The article case is most large: raw EM is 0 but normalized EM is 1, while raw F1 is already 0.667 due to partial token overlap.

Key Parameters

The key parameters for the evaluation functions are:

  • lowercase: Convert text to lowercase to eliminate case sensitivity. Essential for matching "Paris" with "paris", but should be used cautiously in domains where case carries meaning, such as chemical formulas or programming languages.
  • remove_punctuation: Strip punctuation characters that may vary between predictions and references (e.g., "U.S.A." vs "USA"). Useful for general text, but potentially harmful for code or mathematical expressions where punctuation is semantic.
  • remove_articles: Remove leading articles ("a", "an", "the") that do not affect answer correctness in most question answering contexts. This handles variations like "the President" versus "President".
  • fix_whitespace: Collapse multiple spaces and trim leading/trailing whitespace to handle inconsistent spacing from OCR errors, HTML extraction, or copy-paste artifacts.
  • normalize: Master toggle letting all normalization steps. When False, performs raw string matching; when True, applies the full normalization pipeline. Development often uses normalized scores to measure comprehension, while production systems might require strict matching to ensure downstream compatibility.

Relationship to Named Entity Recognition Evaluation

Exact Match and token-level F1 appear in question answering and in the evaluation of Named Entity Recognition (NER) systems, where they take on slightly different interpretations. In NER, the task is to identify spans of text that correspond to entities of specified types (person, organization, location, date, and so on). Evaluation requires token overlap and entity boundaries that match the reference annotation.

Two evaluation protocols are common in NER:

Exact span match requires that the predicted entity span exactly matches the reference span in both position and type. If the reference annotates "General Electric" as an ORGANIZATION and the model predicts "Electric" as an ORGANIZATION, the prediction receives zero credit under exact span evaluation, even though it identified the right entity type and the right sub-span. This strict protocol is used in the CoNLL-2003 shared task evaluation and is in effect Exact Match adapted to the span level.

Token-level F1 computes precision and recall at the level of individual entity tokens, giving partial credit for spans that overlap but do not perfectly align. Under this protocol, predicting "Electric" for "General Electric" would yield recall = 0.5 (one of two tokens found) and precision = 1.0 (the one token predicted is correct), giving F1 = 0.667. Token-level F1 is more forgiving of boundary errors and is often used during development to diagnose whether a model finds the right entity regions even before it gets boundaries precisely right.

The distinction matters because NER errors often fall into predictable patterns. Models frequently find the right entity but predict a slightly shorter or longer span, a phenomenon called "boundary errors." These are qualitatively different from "type errors" (wrong entity type for the right span) or "false positives" (predicting an entity where none exists). By reporting both exact span F1 and token-level F1, you can distinguish these error types and direct engineering effort accordingly.

Partial Match in NER

The NLP community has not fully standardized whether "partial match" in NER means token-level overlap (any shared tokens count) or span-level overlap (the predicted span must start or end at the same position as the reference). The seqeval library, widely used for NER evaluation, implements entity-level F1 where the entire span must match. The SQuAD evaluation script implements token-level F1. These are different metrics even though both are called "F1 in NER."

Adversarial Examples and Metric Gaming

Any evaluation metric that models are trained to optimize will eventually be gamed. Exact Match and F1 are no exception, and understanding how they can be gamed is important for understanding what they measure.

For Exact Match, the most straightforward gaming strategy is to memorize the distribution of reference answer formats in the training set. If training answers always use the short form "US" rather than "United States", a model can learn to always produce "US" and reach high EM on the training distribution, even if it would produce "United States" on a different distribution. This is a form of distributional overfitting that EM cannot detect on its own.

For F1, the gaming strategy is more subtle. Because F1 rewards every correct token regardless of incorrect tokens, a model can increase F1 by generating longer answers that include all the reference tokens plus additional text. In the limit, outputting the entire document as the answer reaches recall = 1.0 for any question, degrading precision but often still achieving moderate F1. Benchmark evaluations that use short reference answers are particularly susceptible to this, as a model that generates three sentences while the reference is one noun phrase will reach zero EM but possibly 30-50% F1 by chance.

Adversarial datasets have been specifically designed to expose these weaknesses. The AddSent adversarial QA dataset adds confusing sentences to context paragraphs that contain answer-like strings. Models that rely on surface-form matching rather than real comprehension sharply degrade on such datasets. Exact Match drops more than F1 on adversarial examples, because models can still capture some correct tokens even when fooled about the right answer span.

These vulnerabilities explain why modern NLP evaluation has moved increasingly toward human evaluation, LLM-as-judge protocols, and multi-metric evaluations that combine EM and F1 with semantic similarity scores. No single metric captures all dimensions of answer quality, and gaming one metric typically degrades performance on complementary metrics. Reporting an ensemble of metrics makes gaming harder and gives a more reliable picture of model capability.

Limitations and Impact

Exact Match and F1 have shaped the development of extractive NLP systems, but their limitations reveal important tensions in how we define "correctness" for language understanding. These metrics reflect a specific philosophical stance: that language understanding can be evaluated through string comparison, and that the goal of NLP systems is to reproduce human-like text rather than to reach semantic goals.

The Lexical Bias Problem

Both metrics are fundamentally lexical: they compare strings, not meanings. This creates a mismatch between metric optimization and human judgment. A model that outputs "The City of Light" when the reference says "Paris" receives zero credit despite being factually correct. Conversely, a model that outputs "Paris, Texas" when asked about French landmarks receives perfect Exact Match if "Paris" appears in the references, even though the answer is wrong in context.

This limitation motivated the development of BERTScore, which uses contextual embeddings to judge semantic similarity rather than surface form overlap. Modern evaluation pipelines often report both F1 (for strict accountability) and BERTScore (for semantic adequacy), particularly for open-domain question answering where paraphrased answers are common.

The lexical bias also affects model development. When you optimize solely for F1, models learn to produce text that matches the lexical patterns of training data rather than focusing on semantic correctness. This can lead to "surface form overfitting," where models memorize specific phrasings from training data rather than learning to generate equivalent paraphrases. The resulting models perform well on held-out data from the same distribution but fail to generalize when question or answer phrasing shifts.

The Tokenization Trap

The choice of tokenization materially affects F1 scores in ways that can obscure true model capabilities. Consider the answer "co-operation" versus "cooperation":

  • Whitespace tokenization treats these as single distinct tokens (no overlap), yielding zero F1 despite the clear semantic equivalence.
  • Character-level tokenization finds substantial overlap, potentially yielding high F1.
  • BPE tokenization from Part V might split these differently depending on training data, leading to unpredictable scoring.

This sensitivity means that F1 scores are not comparable across studies using different preprocessing pipelines. A model achieving 85.2 F1 on SQuAD with aggressive punctuation removal might score 82.1 with strict tokenization, even with identical underlying predictions. When comparing results across papers, it is needed to verify the metric name and the specific tokenization and normalization implementation.

The reproducibility problem created by tokenization variation is more serious than it might appear. A 3-point difference in F1 can determine which model a researcher selects for a downstream task, how a paper ranks in a leaderboard, or whether a model is deemed to have surpassed human performance. If that 3-point difference shows tokenization choices rather than real capability differences, the entire comparison is misleading. The community has made progress toward standardized evaluation scripts, but implementation differences persist, particularly in non-English languages where standard English tokenization rules do not apply.

Order Independence and Structured Data

F1 treats predictions as unordered sets of tokens, which suits entity extraction but fails for structured outputs where word order determines meaning. If asked "Who did what to whom?" with reference "Alice gave Bob a book", the prediction "Bob gave Alice a book" reaches high F1 through token overlap (Alice, Bob, gave, book), despite reversing the semantic roles completely. The prediction describes an entirely different event, yet F1 suggests strong performance.

For such structured prediction tasks, specialized metrics like TACRED relation extraction scoring or tree-based similarity measures become necessary. These metrics consider the relational structure between entities, not just their presence in the text. In coreference resolution, temporal ordering tasks, and relation extraction, the word order carries needed semantic information that token-set comparison discards entirely.

Normalization as Hidden Complexity

The apparent simplicity of Exact Match belies the complexity hidden in normalization pipelines. Different benchmarks apply different normalization rules, making direct comparison difficult. SQuAD 1.1 uses lowercasing and whitespace normalization; SQuAD 2.0 adds punctuation removal; some implementations strip articles while others preserve them. When comparing models across papers, it is needed to verify that metric implementations match, as "F1" is not a single standardized quantity but a family of measures parameterized by tokenization and normalization choices.

This hidden complexity can lead to reproducibility issues. A reported F1 score of 80.0 might stand for very different underlying capabilities depending on whether the implementation used aggressive normalization (making the task easier) or strict matching (which makes it harder). The research community has moved toward standardized evaluation scripts, but variation remains common, particularly in legacy work or domain-specific applications.

Despite these limitations, Exact Match and F1 remain indispensable for development and debugging. Their interpretability allows researchers to inspect specific failure modes: a model consistently achieving 0.0 EM but 0.8 F1 is capturing the right entities but failing on exact boundaries or formatting, suggesting issues with span selection mechanisms rather than comprehension. This diagnostic clarity explains why these metrics persist as standard benchmarks even as more advanced evaluation methods emerge.

The Reference Quality Problem

One often-overlooked limitation is that both EM and F1 treat the reference answers as ground truth, yet reference quality varies considerably across datasets and annotators. Studies on SQuAD have found inter-annotator agreement on answer spans of around 90%, meaning that roughly 10% of questions have ambiguous or contested gold answers. When you evaluate a model against these contested references, you are not measuring whether the model found the "correct" answer, but whether it found the answer one particular annotator preferred.

This reference quality problem compounds with the multi-answer protocol. Datasets that give multiple references per question acknowledge annotator disagreement, but the number of references provided is often small (SQuAD gives 3-5 answers per question from the development set). If a model finds a sixth equally valid answer, it receives no credit. The reported EM and F1 scores thus reflect model capability and the completeness of the reference set, which varies by question and by dataset construction methodology.

Summary

Exact Match and F1 give the basis for evaluating precision-necessary language tasks. It offers complementary perspectives on model performance. Exact Match enforces strict compliance, accepting no deviation from reference answers. This makes it ideal for structured extraction and closed-set classification where format and content must both be perfect. F1 introduces flexibility through token-level precision and recall, rewarding partial correctness and accommodating natural variation in answer phrasing.

The key conceptual distinction is directional: Exact Match is diagnostic at the binary level, revealing whether a model has fully mastered a task, while F1 is diagnostic at the continuous level, revealing how close a model is and in which direction (precision or recall) it needs improvement. In practice, both metrics together tell a fuller story than either alone, and the gap between them is itself informative.

The implementation of these metrics requires careful attention to tokenization strategies and normalization pipelines. Whitespace tokenization, case folding, and punctuation removal are standard practices, but the specific choices materially affect reported scores and cross-study comparability. Domain-appropriate normalization can reduce false negatives caused by formatting variations, though aggressive normalization risks masking real distinctions or creating false equivalences between semantically distinct strings.

When selecting evaluation metrics, consider the downstream cost of errors and the nature of the task. Use Exact Match when format compliance is necessary or when answers come from closed vocabularies. Use F1 when partial information has value or when multiple valid phrasings exist. For semantic equivalence beyond surface forms, complement these metrics with embedding-based approaches like BERTScore. For tasks with order-sensitive structure or complex relational semantics, consider task-specific metrics that capture what EM and F1 cannot.

These metrics trace their lineage to classical information retrieval, yet they remain relevant for modern LLM evaluation. Even as models generate increasingly complex outputs, the need to extract precise facts, identify entities, and answer specific questions ensures that Exact Match and F1 will continue serving as needed tools in the language AI evaluation toolkit. They give the necessary rigor for precision-necessary applications while their limitations highlight the ongoing challenge of bridging lexical form and semantic meaning in natural language understanding. In the next chapter, we will explore calibration metrics that assess whether models are right and how well their confidence scores reflect their actual accuracy.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Exact Match and F1 evaluation metrics.

Exact Match and F1 Evaluation

Question 1 of 70 of 7 completed
What is the Exact Match score for prediction 'paris' versus reference 'Paris' without any normalization?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026exactmatch, author = {Michael Brenndoerfer}, title = {Exact Match and F1: Precision Metrics for NLP Evaluation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Exact Match and F1: Precision Metrics for NLP Evaluation. Retrieved from https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics
MLAAcademic
Michael Brenndoerfer. "Exact Match and F1: Precision Metrics for NLP Evaluation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics>.
CHICAGOAcademic
Michael Brenndoerfer. "Exact Match and F1: Precision Metrics for NLP Evaluation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Exact Match and F1: Precision Metrics for NLP Evaluation'. Available at: https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Exact Match and F1: Precision Metrics for NLP Evaluation. https://mbrenndoerfer.com/writing/exact-match-f1-nlp-evaluation-metrics

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.