Part of Language AI Handbook
BERTScore evaluates text generation using contextual embeddings to measure semantic similarity.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
BERTScore: Semantic Text Evaluation Using BERT Embeddings
Traditional n-gram metrics like BLEU and ROUGE dominated automatic evaluation for decades, but they share a basic limitation: they measure surface form overlap. When a candidate translation contains "feline" instead of "cat," or "automobile" instead of "car," these metrics penalize the variation despite perfect semantic equivalence. As we discussed in Part XXIV: BERT and Variants, contextualized embeddings capture meaning in a way that bag-of-words representations cannot. BERTScore applies this insight by using pre-trained contextual embeddings to evaluate text generation based on semantic similarity rather than lexical overlap.
This move from counting matching words to measuring meaning is a significant advance in how we evaluate natural language generation. Before the advent of deep contextualized representations, automatic evaluation was in effect a string matching exercise. You needed an exact word or n-gram hit to receive any credit. The key insight behind BERTScore is that pre-trained language models have already learned rich semantic relationships from large text corpora, and we can tap into that knowledge to evaluate whether two pieces of text convey the same meaning, regardless of the specific words chosen. This approach aligns automatic evaluation more closely with how humans assess quality: we care about meaning preservation, not just word overlap.
BERTScore operates on a simple yet powerful premise: if two sentences convey the same meaning, their constituent tokens should have similar contextual representations in a pre-trained language model. By computing pairwise cosine similarities between tokens in the reference and candidate texts, then aligning them greedily, BERTScore produces precision, recall, and F1 scores that correlate remarkably well with human judgments of quality. The metric was introduced by Zhang et al. in 2019 and has since become a standard component of evaluation suites across machine translation, summarization, dialogue generation, and other natural language generation tasks.
Understanding why BERTScore works requires understanding what contextual embeddings represent. Unlike static word embeddings such as Word2Vec or GloVe, which assign a single fixed vector to each word, contextual embeddings assign a different vector to each word token depending on the surrounding sentence context. The word "bank" receives one representation in "she sat by the river bank" and a completely different one in "she deposited money at the bank." This context-dependence means that contextual representations capture a word's general meaning and its sentence-specific meaning, making them ideal for semantic similarity comparisons.
From Surface Form to Semantic Similarity
To understand why BERTScore is a change in evaluation, consider this example:
Reference: The scientist conducted rigorous experiments to validate the hypothesis.
Candidate: The researcher performed strict tests to confirm the theory.
A BLEU score would see almost no matching n-grams ("scientist" ≠ "researcher", "conducted" ≠ "performed", etc.), assigning a near-zero score despite the sentences being semantically identical. BERTScore, however, recognizes that "scientist" and "researcher" share similar contexts in the training data, as do "experiments" and "tests," "validate" and "confirm," and "hypothesis" and "theory."
This example illustrates a important limitation of surface-form metrics: they cannot recognize paraphrasing, synonym substitution, or morphological variation. In real-world translation and summarization tasks, competent human translators rarely produce word-for-word matches with reference texts. Instead, they convey the same meaning using different linguistic choices. Traditional metrics penalize this valid variation, creating a disconnect between automatic scores and human judgment. BERTScore bridges this gap by operating in the semantic space rather than the lexical space.
The metric works by extracting contextual embeddings from a pre-trained model (typically BERT or RoBERTa), computing a similarity matrix between all pairs of tokens from the reference and candidate, and then finding the optimal alignment between tokens. This process turns the discrete problem of comparing word sequences into a continuous problem of comparing vectors in high-dimensional space, where semantic similarity corresponds to geometric proximity.
Consider what happens when a human expert evaluates a machine translation. The expert does not check whether the candidate shares the same exact words as the reference. Instead, they read both sentences and evaluate the overall result: does the candidate preserve the meaning, tone, and information of the reference? BERTScore approximates this process computationally by using a language model's learned representations as a proxy for human semantic understanding. The model has been trained on large corpora to understand when two expressions are interchangeable, and BERTScore exploits this knowledge without any fine-tuning or task-specific supervision.
Embedding Extraction
BERTScore uses the hidden states of a pre-trained transformer to obtain contextualized representations. For a sentence with tokens , we pass it through BERT to obtain hidden states:
where:
- : the input sentence represented as a sequence of tokens
- : the number of tokens in the sentence
- : the matrix of hidden states from the final BERT layer, with each row corresponding to a token
- : the hidden dimension size (768 for BERT-base)
The process of extracting these embeddings involves feeding the tokenized input through the transformer's encoder stack. Each layer of the transformer refines the representations, incorporating context from surrounding tokens through the self-attention mechanism. The lower layers tend to capture surface features such as part-of-speech tags and syntactic dependencies, while the upper layers capture more abstract semantic and contextual relationships. By selecting specific layers, we can control whether we want representations that are more syntactic or more semantic in nature.
Research shows that intermediate layers (typically layers 9-10 for BERT-base) capture the most transferable semantic information, balancing surface syntax in lower layers and task-specific features in upper layers. BERTScore typically uses the second-to-last layer by default.
The choice of layer involves an important trade-off. Lower layers preserve more lexical and syntactic information, which can be useful for detecting grammatical errors or exact word matches. Higher layers contain more task-specific abstractions that might miss fine-grained semantic distinctions. Empirical studies have shown that layers in the middle-to-upper range (roughly layers 8-10 for BERT-base) strike the best balance for semantic similarity tasks, capturing enough context to recognize synonyms and paraphrases while maintaining stable, generalizable representations across different domains.
The embedding for token is denoted as:
where:
- : the -th token in the input sequence
- : the -dimensional contextual embedding vector for token
An important subtlety arises from the subword tokenization used by BERT. As we covered in Part V: Subword Tokenization, BERT uses WordPiece tokenization, which splits uncommon words into sub-units. The word "unbelievable" might be tokenized as ["un", "##believe", "##able"], and "running" might remain a single token while "unhappiness" gets split into ["un", "##happiness"]. This means BERTScore does not operate on word-level tokens but on subword tokens, which has practical implications. When computing scores, each subword receives its own contextual embedding and participates independently in the similarity computation. Rare words that get split into subwords effectively contribute more tokens to the score computation, which may slightly overweight them relative to common single-token words.
Another subtlety involves the special tokens that BERT prepends and appends to every sequence: [CLS] at the beginning and [SEP] at the end. These tokens participate in the self-attention computation and influence the contextual representations of surrounding tokens, but they do not correspond to any semantic content in the text. BERTScore implementations filter out these special tokens before computing similarities. This keeps the score shows only the actual content tokens in each sentence.
Similarity Computation
Given a reference sentence with tokens and a candidate sentence with tokens, we construct a similarity matrix where each entry is the cosine similarity between token embeddings:
where:
- : the cosine similarity between the -th reference token and -th candidate token
- : the contextual embedding of the -th token from the reference sentence
- : the contextual embedding of the -th token from the candidate sentence
- : the dot product between two vectors
- : the L2 norm (magnitude) of a vector
The cosine similarity ranges from -1 to 1, but since BERT embeddings are non-negative in practice (due to the GELU activation functions and layer normalization), the similarities typically fall between 0 and 1.
Cosine similarity is particularly well-suited for this task because it measures the angle between vectors rather than their absolute magnitude. In the context of language model embeddings, the magnitude of a vector often correlates with the confidence or specificity of the representation, while the direction encodes the semantic meaning. By normalizing out the magnitude, cosine similarity focuses purely on semantic alignment. This means that rare, specific words and common function words can be compared on equal footing based on their directional similarity in the embedding space.
Token Alignment
Once we have the similarity matrix, we need to determine how well the candidate tokens cover the reference tokens (recall) and how much of the candidate is supported by the reference (precision). BERTScore uses a greedy matching approach:
For Recall: Each reference token is matched to its most similar candidate token :
where:
- : the index of the candidate token most similar to reference token
- : the operator that returns the index maximizing the similarity
- : the cosine similarity between reference token and candidate token
For Precision: Each candidate token is matched to its most similar reference token :
where:
- : the index of the reference token most similar to candidate token
- : the operator that returns the index maximizing the similarity
- : the cosine similarity between reference token and candidate token
This greedy approach differs from optimal matching algorithms like the Earth Mover's Distance used in MoverScore, but it is computationally efficient and works well in practice.
The choice of greedy matching shows a practical compromise between computational complexity and alignment quality. In an optimal matching scenario, we might want to find the global assignment of tokens that maximizes total similarity. This keeps each token in one sentence is matched to a distinct token in the other. However, such optimal matching problems, often framed as linear assignment problems, require algorithms like the Hungarian method that scale cubically with the sequence length. For long sentences with dozens of tokens, this computational overhead becomes significant, especially when evaluating thousands of sentence pairs. Greedy matching reduces this to linear complexity while still capturing the needed semantic correspondence between texts. In practice, because semantic similarity matrices tend to have clear diagonal or near-diagonal structures (where corresponding words align), the greedy approach usually finds alignments very close to the optimal solution.
Score Aggregation
The final scores aggregate these maximum similarities:
Precision measures how much of the candidate is supported by the reference:
where:
- : the precision score (average of maximum similarities for each candidate token)
- : the number of tokens in the candidate sentence
- : the maximum similarity between candidate token and any reference token
Recall measures how much of the reference is covered by the candidate:
where:
- : the recall score (average of maximum similarities for each reference token)
- : the number of tokens in the reference sentence
- : the maximum similarity between reference token and any candidate token
F1 balances both:
where:
- : the harmonic mean of precision and recall
- : the precision score
- : the recall score
The aggregation strategy treats each token as an independent contributor to the overall semantic content of the sentence. When we average the maximum similarities, we are in effect asking: on average, how well does each token in the candidate find a semantic counterpart in the reference (precision), and how well does each reference token find a match in the candidate (recall)? This token-level granularity allows BERTScore to give fine-grained feedback about which specific concepts are preserved or missing in the candidate text.
The use of precision, recall, and F1 mirrors the structure of information retrieval metrics and shows similar intuitions. High precision means that the candidate does not contain content unsupported by the reference, which makes it a useful check for hallucination: if the candidate invents facts not present in the reference, those tokens will have low maximum similarities and will pull down the precision score. High recall means that the candidate covers the key concepts present in the reference, which makes it a useful check for completeness: if the candidate omits important information from the reference, those reference tokens will have no good match and will reduce the recall score. F1 balances these two concerns into a single number, rewarding candidates that are both faithful and complete.
Importance Weighting
By default, BERTScore treats all tokens equally. However, we can incorporate Inverse Document Frequency (IDF) weights to give less importance to common words like "the" or "and" (recall our discussion of TF-IDF in Part II).
The motivation for IDF weighting stems from the observation that content words (nouns, verbs, adjectives) carry more semantic information than function words (determiners, prepositions, conjunctions). In the sentence "The cat sat on the mat," the words "cat," "sat," and "mat" convey the core meaning, while "the" and "on" give grammatical structure. Without weighting, a candidate sentence that correctly captures all the function words but misses the content words could reach misleadingly high similarity scores. IDF weighting addresses this by assigning higher weights to rare, informative words and lower weights to common, structural words.
With IDF weighting, the scores become:
where:
- : the IDF-weighted precision score
- : the inverse document frequency weight for candidate token
- : the maximum similarity between candidate token and any reference token
Similarly, the IDF-weighted recall is computed as:
where:
- : the IDF-weighted recall score
- : the inverse document frequency weight for reference token
- : the maximum similarity between reference token and any candidate token
The IDF weights are typically computed over the reference corpus, with rare words receiving higher weights. This approach ensures that correctly matching specialized terminology or rare concepts contributes more to the final score than matching common function words. In practice, IDF weighting most noticeably improves BERTScore's correlation with human judgments when evaluating texts with rich domain-specific vocabulary, such as scientific abstracts or legal documents. For general text with a mix of function and content words, the improvement is often modest but consistent.
Computing IDF requires a reference corpus, which introduces a practical consideration: the choice of corpus affects the weights. Using a large general corpus like Wikipedia produces stable, broadly applicable weights but may not reflect the term frequency patterns of your specific evaluation domain. Using only the provided references gives more domain-relevant weights but can be unstable with small sample sizes. In the original BERTScore paper, Zhang et al. recommend using the provided references as the IDF corpus and report marginal improvements in human correlation, particularly for longer texts with varied vocabulary.
Worked Example
Let's walk through a concrete example to see how BERTScore handles semantic equivalence that BLEU misses. We'll step through the full computation by hand so the mechanics are completely transparent before moving to code.
Consider:
- Reference: "The cat sat"
- Candidate: "The feline rested"
We'll use simplified 3-dimensional embeddings for illustration (in practice, these are 768-dimensional). The embeddings are chosen to reflect a key property of real contextual representations: synonymous words have nearly identical vectors, while unrelated words occupy very different regions of the space:
| Token | Embedding |
|---|---|
| cat | [0.5, 0.2, 0.1] |
| feline | [0.48, 0.22, 0.12] |
| sat | [0.1, 0.8, 0.3] |
| rested | [0.12, 0.75, 0.35] |
| the | [0.9, 0.1, 0.1] |
First, we compute cosine similarities between all pairs (excluding "the" for brevity):
where:
- : the embedding vector for "cat" ()
- : the embedding vector for "feline" ()
- : the dot product operator
- : the Euclidean norm (magnitude) of the vector
The similarity matrix (excluding "the") is:
| feline | rested | |
|---|---|---|
| cat | 0.99 | 0.35 |
| sat | 0.28 | 0.97 |
Recall calculation:
- "cat" is matched to "feline" (0.99)
- "sat" is matched to "rested" (0.97)
Precision calculation:
- "feline" is matched to "cat" (0.99)
- "rested" is matched to "sat" (0.97)
Even though no content words match exactly, BERTScore recognizes the semantic equivalence and assigns a high score. BLEU would assign 0 to unigram precision for the content words (only "The" matches) and fail to capture the quality of the candidate.
Now consider a different scenario: a candidate that is too long and introduces extra information. Suppose the candidate is "The feline rested quietly in the sun." The reference tokens "cat" and "sat" still match well to "feline" and "rested," preserving high recall. But now the candidate has four extra tokens ("quietly," "in," "the," "sun") that are not present in the reference. Each of these tokens will find its best match in the reference, but those matches will be poor. This pulls down the precision score. This shows that the candidate introduced content not supported by the reference. The F1 score will therefore be lower than for the lean, exact synonym substitution.
This example illustrates the interpretive value of looking at precision and recall separately. High recall with low precision indicates a candidate that covers the reference but adds irrelevant or hallucinated content. Low recall with high precision indicates a candidate that is faithful but incomplete. F1 gives a single balanced summary for ranking purposes, but precision and recall offer diagnostic insight when debugging system outputs.
This example shows the power of distributional semantics: words that appear in similar contexts in the training data (like "cat" and "feline," or "sat" and "rested") occupy nearby regions in the embedding space. The high cosine similarities reflect this distributional proximity, letting BERTScore to recognize that these sentences describe the same event using different vocabulary.

Code Implementation
Let's implement BERTScore from scratch using Hugging Face transformers to understand the mechanics, then compare with the optimized library implementation.
First, we'll load the necessary libraries and a pre-trained BERT model:
from transformers import BertModel, BertTokenizer
tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertModel.from_pretrained("bert-base-uncased")
model.eval()
reference = (
"The scientist conducted rigorous experiments to validate the hypothesis"
)
candidate = "The researcher performed strict tests to confirm the theory"Reference: The scientist conducted rigorous experiments to validate the hypothesis Candidate: The researcher performed strict tests to confirm the theory
We'll tokenize both sentences and extract embeddings from layer 9 (one of the middle layers that captures good semantic information):
import torch
def get_embeddings(text, model, tokenizer, layer=9):
"""Extract contextual embeddings from specified BERT layer."""
inputs = tokenizer(text, return_tensors="pt", padding=True, truncation=True)
with torch.no_grad():
outputs = model(**inputs, output_hidden_states=True)
# hidden_states is a tuple of 13 tensors (embedding layer + 12 BERT layers)
# We use the specified layer (index 9 corresponds to layer 10 in 1-indexed)
hidden_states = outputs.hidden_states[layer]
# Remove [CLS] and [SEP] tokens for scoring
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
embeddings = hidden_states[0].numpy()
# Filter out special tokens
valid_indices = [
i for i, t in enumerate(tokens) if t not in ["[CLS]", "[SEP]", "[PAD]"]
]
filtered_tokens = [tokens[i] for i in valid_indices]
filtered_embeddings = embeddings[valid_indices]
return filtered_tokens, filtered_embeddings
# Get embeddings for both sentences
ref_tokens, ref_embeddings = get_embeddings(reference, model, tokenizer)
cand_tokens, cand_embeddings = get_embeddings(candidate, model, tokenizer)Let's examine the tokens BERT produced:
Reference tokens (10): ['the', 'scientist', 'conducted', 'rigorous', 'experiments', 'to', 'valid', '##ate', 'the', 'hypothesis'] Candidate tokens (9): ['the', 'researcher', 'performed', 'strict', 'tests', 'to', 'confirm', 'the', 'theory']
Notice that "scientist" and "researcher" are single tokens, while some words might be split into subwords (WordPiece tokenization, as we covered in Part V, Chapter 3: WordPiece).
Now we compute the cosine similarity matrix between all pairs of tokens:
from sklearn.metrics.pairwise import cosine_similarity
# Compute similarity matrix
similarity_matrix = cosine_similarity(ref_embeddings, cand_embeddings)Let's visualize this similarity matrix to see which tokens align:


The heatmap reveals strong similarities between semantically related words. Now let's implement the greedy matching algorithm to compute precision and recall:
# Greedy matching for Recall: each reference token matched to best candidate
recall_matches = np.max(similarity_matrix, axis=1)
recall = np.mean(recall_matches)
# Greedy matching for Precision: each candidate token matched to best reference
precision_matches = np.max(similarity_matrix, axis=0)
precision = np.mean(precision_matches)
# F1 calculation
f1 = 2 * (precision * recall) / (precision + recall)Precision: 0.857 Recall: 0.841 F1: 0.849 Top semantic alignments (Reference → Candidate): the → the (similarity: 0.939) scientist → researcher (similarity: 0.815) conducted → performed (similarity: 0.912) rigorous → strict (similarity: 0.818) experiments → tests (similarity: 0.807) to → to (similarity: 0.936) valid → confirm (similarity: 0.688) ##ate → confirm (similarity: 0.701) the → the (similarity: 0.910) hypothesis → theory (similarity: 0.879)
Even though only "the" appears in both sentences, BERTScore recognizes the semantic equivalence between word pairs like "scientist"-"researcher" and "validate"-"confirm", yielding high scores.
For production use, the bert-score library gives optimized implementations with batching and GPU support:
# Install bert-score if needed
# uv pip install bert-score
# BERTScore library handles tokenization, embedding extraction, and scoring
try:
from bert_score import score
P, R, F1 = score([candidate], [reference], lang="en", verbose=False)
except ModuleNotFoundError:
P = np.array([precision])
R = np.array([recall])
F1 = np.array([f1])Library implementation: Precision: 0.953 Recall: 0.963 F1: 0.958
The library results should closely match our from-scratch implementation, with minor differences due to layer selection and the library's exact tokenization handling. The library implementation also applies baseline rescaling by default in some configurations, which can shift the absolute values.
Key Parameters
Understanding the key parameters for BERTScore is needed for applying the metric effectively across different domains and use cases. The choices you make regarding model selection, layer extraction, and weighting schemes can materially impact the resulting scores and their interpretation.
Layer Selection
Which transformer layer to extract embeddings from (typically 9-10 for BERT-base) has a noticeable effect on evaluation quality. Intermediate layers capture the most transferable semantic information, balancing surface syntax with task-specific features from higher layers.
The layer selection parameter controls the depth within the transformer network from which we extract token representations. BERT-base consists of 12 transformer layers, each processing the output of the previous one. Early layers (1-4) focus primarily on local syntactic patterns and part-of-speech information. Middle layers (5-8) begin to integrate broader context and semantic relationships. Upper layers (9-12) capture high-level semantic and pragmatic features, though the final layer often becomes specialized for the pre-training objectives (masked language modeling and next sentence prediction). Research suggests that layers 9 or 10 give the best balance for semantic similarity tasks, capturing rich contextual meaning while avoiding the task-specific biases of the final layer. When using RoBERTa or other variants, optimal layers may differ slightly due to architectural variations, so experimentation with your specific domain is recommended.

Model Selection
The pre-trained model you choose fundamentally determines the semantic space in which similarity is measured. Different models yield different score distributions and semantic sensitivities.
The choice of pre-trained model BERT-base-uncased gives a good general-purpose baseline with reasonable computational requirements. RoBERTa models, trained with optimized hyperparameters and larger batches, often produce more reliable semantic representations. Larger variants (bert-large, roberta-large) offer richer representations but require more memory and computation. Domain-specific models such as SciBERT for scientific text, BioBERT for biomedical literature, or CodeBERT for programming languages can materially improve evaluation accuracy when working with specialized vocabulary. The casing strategy (cased vs. uncased) also matters: cased models preserve capitalization information, which can be important for distinguishing proper nouns and acronyms, while uncased models treat "Apple" and "apple" identically.
Language Configuration
The language code for optimized library implementations (e.g., 'en' for English) automatically selects appropriate pre-trained models for the target language.
The language parameter ensures that the appropriate tokenizer and model architecture are used for the target language. While multilingual models like BERT-multilingual or XLM-RoBERTa can handle many languages, monolingual models often give better representations for specific languages. When evaluating non-English text, ensure that the model you select has been trained on sufficient data in that language to produce real embeddings. For low-resource languages, multilingual models may be the only viable option, though their semantic representations may be less precise than those for high-resource languages.
IDF Weighting Strategy
Applying IDF weighting requires decisions about the reference corpus used to compute word frequencies. Using a large, diverse corpus ensures stable IDF estimates but may not reflect the specific domain of your evaluation task. Using only the provided references can lead to unstable estimates with small sample sizes but ensures domain relevance. Some implementations offer options to use pre-computed IDF weights from standard corpora or to compute them dynamically. The choice affects whether common function words are effectively discounted and whether rare, domain-specific terms receive appropriate emphasis.
Baseline Rescaling
Raw BERTScore similarities tend to be high even for unrelated sentences because BERT embeddings are dense and cosine similarity rarely approaches zero between arbitrary text pairs. Baseline rescaling parameters allow you to compute scores against random sentence pairs and normalize the results, which makes it easier to distinguish between mediocre and high-quality generations. This is particularly important when comparing scores across different model choices or when establishing absolute quality thresholds.
BERTScore Variants
Several extensions to the original BERTScore address specific limitations or application domains. These variants explore different alignment strategies, training paradigms, and normalization techniques to improve upon the basic approach.
MoverScore
While BERTScore uses greedy matching, MoverScore employs optimal transport (specifically the Earth Mover's Distance) to find the minimal cost of "moving" probability mass from candidate tokens to reference tokens. This considers the global structure of the similarity matrix rather than independent greedy matches. MoverScore often correlates better with human judgments but is computationally more expensive, scaling cubically with sequence length in the worst case compared to BERTScore's linear complexity.
The key insight behind MoverScore is that greedy matching may create suboptimal alignments when multiple tokens in one sentence could reasonably match a single token in the other. For example, in comparing "the big dog" to "the large canine," greedy matching might align both "big" and "dog" to "canine" if their similarities are locally maximal, leaving "large" unmatched even though a global view would prefer matching "big" to "large" and "dog" to "canine." Earth Mover's Distance solves this by finding the optimal flow that minimizes total transportation cost, effectively letting fractional matches and considering the entire similarity distribution. This global optimization often produces alignments that better reflect human intuition about semantic correspondence, particularly for sentences with complex syntactic structures or multiple clauses. However, the computational cost limits its applicability in high-throughput scenarios or when evaluating very long documents.
BLEURT
BLEURT (Bilingual Evaluation Understudy with Representations from Transformers) takes a different approach: instead of computing token-level similarities, it fine-tunes BERT on human ratings to directly predict quality scores. This requires supervised training data but can learn task-specific evaluation criteria. Unlike BERTScore, which is reference-based by design, BLEURT learns to predict human judgments end-to-end.
BLEURT is a move from unsupervised to supervised evaluation metrics. The model is fine-tuned on collections of human judgments, learning to map the embedding representations directly to quality scores. This allows BLEURT to capture fine-grained aspects of quality that might not emerge from simple similarity calculations, such as fluency, style appropriateness, or factual accuracy in specific domains. However, this supervised approach requires large labeled training data and may not generalize well to domains or tasks different from those in the training set. BLEURT also operates at the sequence level rather than the token level. This gives a single overall score rather than the detailed precision and recall breakdown offered by BERTScore. This makes BERTScore more suitable for diagnostic analysis, while BLEURT may give better absolute quality predictions when the training domain matches the evaluation domain.
BERTScore with Rescaling
Raw BERTScore similarities tend to be high (often above 0.5) even for unrelated sentences because BERT embeddings are dense and cosine similarity rarely drops to zero between arbitrary text pairs. To make scores more interpretable and comparable across different model choices, practitioners often apply baseline rescaling: computing scores against random sentences and subtracting this baseline, or using normalized scores based on empirical distributions.
Rescaling addresses the interpretability challenge posed by the absolute values of BERTScore. Because pre-trained language models produce embeddings that occupy a relatively constrained region of the high-dimensional space, even random or unrelated sentences may exhibit cosine similarities of 0.3 to 0.5. This compression of the similarity range makes it difficult to set thresholds for quality or to compare scores across different pre-trained models, which may have different baseline similarity levels. Rescaling techniques typically involve computing the mean and standard deviation of similarity scores on a corpus of unrelated sentence pairs, then z-score normalizing the actual evaluation scores against this distribution. This transformation produces scores centered around zero for random text, with positive values showing better-than-random semantic alignment and higher values showing stronger semantic equivalence.

Domain-Specific Variants
Using domain-specific pre-trained models (SciBERT for scientific text, BioBERT for biomedical text, CodeBERT for programming languages) can improve BERTScore's sensitivity to domain terminology. For example, in evaluating code generation, CodeBERT embeddings will better capture that "append" and "push" are semantically related operations in list manipulation contexts.
Domain adaptation is important because general-purpose language models may not adequately stand for specialized vocabulary or domain-specific semantic relationships. In scientific writing, terms like "oxidation" and "reduction" have specific technical meanings distinct from their everyday usage. General BERT models might not capture the precise semantic relationships between such terms, whereas SciBERT, trained on scientific literature, encodes these relationships more accurately. Similarly, in medical contexts, BioBERT better understands that "myocardial infarction" and "heart attack" refer to the same condition, while in legal text, domain-specific models recognize that "plaintiff" and "claimant" are synonymous. For code generation evaluation, CodeBERT and similar programming-language models understand that different API calls or algorithmic approaches can reach the same computational goal, recognizing semantic equivalence between syntactically different but functionally equivalent code snippets.
Comparison with BLEU
As we explored in Part LIII, Chapter 3: BLEU Score, BLEU relies on precision of n-gram matches with a brevity penalty. The basic differences between these metrics illuminate the move from statistical to neural evaluation paradigms. Understanding when to prefer each metric requires appreciating what each one measures.
Surface Form vs. Semantic Content
BLEU requires exact string matches. "Ran quickly" and "sprinted" produce zero BLEU overlap despite being synonymous. BERTScore captures these paraphrases through embedding similarity.
This distinction is the core philosophical difference between the two approaches. BLEU operates on the assumption that good translations will share specific word sequences with reference translations, an assumption that holds reasonably well for closely related languages or when multiple references are available, but breaks down for distant language pairs or creative paraphrasing. BERTScore operates on the assumption that meaning is preserved when semantic content is maintained, regardless of surface realization. This makes BERTScore more reliable to the variability inherent in human translation and summarization, where competent professionals routinely produce texts that convey the same meaning using entirely different phrasing.

Granularity
BLEU operates at the n-gram level (typically up to 4-grams), while BERTScore operates at the token embedding level, effectively capturing meaning regardless of phrase boundaries or word order variations that preserve meaning.
The granularity difference affects what each metric can detect. BLEU is sensitive to local word order and collocation: "quickly ran" and "ran quickly" may receive different n-gram scores even though they are semantically equivalent. BERTScore's token-level approach focuses on semantic content rather than syntactic arrangement. This allows BERTScore to recognize that "the theory was confirmed by the researcher" and "the researcher confirmed the theory" convey the same meaning despite different syntactic structures, while BLEU might penalize the word order differences. However, this granularity also means BERTScore may miss certain types of errors that BLEU catches, such as incorrect function word usage or grammatical agreement issues that don't materially shift semantic embeddings but stand for quality problems.
Robustness to Morphological Variation
BLEU is sensitive to inflections ("run" vs. "ran" vs. "running"). BERTScore handles these gracefully because contextual embeddings cluster morphological variants of the same lemma.
Morphological variation poses a significant challenge for n-gram metrics. A candidate translation using "running" where the reference uses "ran" receives no credit from BLEU, even though this is a minor grammatical variation rather than a semantic error. Stemming or lemmatization can mitigate this, but such preprocessing introduces its own errors and complexity. BERTScore naturally handles inflectional variation because pre-trained language models learn that "run," "ran," and "running" appear in similar contextual distributions, placing their embeddings in nearby regions of the vector space. This property extends to derivational morphology as well, letting BERTScore to recognize relationships between "decide," "decision," and "decisive."
Reference Requirements
BLEU requires multiple reference translations to be reliable, as a single reference cannot cover all valid paraphrases. BERTScore is more reliable with single references because it generalizes to synonymous expressions.
The cost of collecting multiple reference translations is substantial, making BLEU expensive to apply in new domains or low-resource settings. BERTScore's ability to generalize from single references materially reduces the annotation burden. A single reference sentence, combined with BERTScore's semantic generalization capabilities, can effectively evaluate a wide range of valid paraphrases that would require dozens of reference translations to cover adequately with n-gram matching. This makes BERTScore particularly useful for niche domains or specialized content where collecting multiple professional translations is prohibitively expensive.
However, BLEU retains advantages in specific scenarios. It is deterministic, model-agnostic, and computationally trivial compared to BERTScore's transformer inference. BLEU also directly measures lexical fidelity, which matters when exact terminology is required (e.g., medical or legal translations where "hypertension" cannot be paraphrased as "high blood pressure" depending on context).
In applications requiring strict terminological consistency, such as pharmaceutical documentation or patent translation, the exact word matching that BLEU enforces may be preferable to BERTScore's semantic flexibility. When translating drug labels or safety warnings, using the precise approved terminology may be legally required, and a metric that penalizes even semantically equivalent paraphrases serves a quality control function. Similarly, BLEU's computational efficiency makes it suitable for high-volume screening tasks, such as filtering candidate translations during data mining or preliminary quality assessment, where the overhead of neural inference would be prohibitive.
Score Interpretation and Calibration
One of the practical challenges in using BERTScore is interpreting raw scores. Unlike BLEU, which has well-established reference points (0.30 is "understandable" translation, 0.60 is near human quality), BERTScore values depend heavily on the model and configuration used. A score of 0.85 using BERT-base-uncased is not directly comparable to a score of 0.85 using RoBERTa-large.
The key insight for interpretation is that BERTScore should usually be evaluated in relative terms: which of two candidate translations scores higher for a given reference? The absolute value matters less than the ranking. When you need absolute thresholds, the recommended approach is to calibrate against human judgments on a sample of your specific domain. Compute BERTScore for a set of candidate-reference pairs that have been rated by human evaluators, then find the score threshold that best separates acceptable from unacceptable outputs in your domain.
Baseline rescaling gives a partial solution to the interpretability problem. The bert-score library supports rescaling with a pre-computed baseline: for each model and language, the library has computed the mean and standard deviation of BERTScore on a set of randomly paired sentence fragments. The rescaled score is:
where:
- : the rescaled score, centered around zero for random text
- : the raw BERTScore value
- : the mean BERTScore on random sentence pairs for this model
- : the standard deviation of BERTScore on random sentence pairs
After rescaling, a score near zero indicates similarity comparable to random text, positive scores indicate real semantic alignment, and scores above 1 indicate strong paraphrase-level similarity. This makes it much easier to set quality thresholds and to compare results across different studies that use the same rescaling procedure.
Even after rescaling, you should be cautious about interpreting small differences in BERTScore as real. The metric has measurement noise, particularly for short sentences where a single token alignment can materially affect the overall score. For reliable evaluation, average BERTScore over many samples and use appropriate statistical tests when comparing system outputs.
Cross-Lingual BERTScore
One natural extension of BERTScore is its application to cross-lingual evaluation, where the reference and candidate texts are in different languages. This is particularly relevant for machine translation: instead of comparing a French candidate against a French reference, you might compare a French candidate directly against an English source, eliminating the need for target-language references entirely.
Cross-lingual BERTScore relies on multilingual pre-trained models such as mBERT (multilingual BERT) or XLM-RoBERTa. These models are trained on text in dozens of languages simultaneously and learn representations that are, to a significant degree, language-agnostic. Words with similar meanings in different languages tend to cluster in nearby regions of the multilingual embedding space, even without explicit parallel supervision. For example, the English word "dog" and the French word "chien" are not the same token, but their contextual embeddings in mBERT are close because they appear in similar sentence contexts across the training data.
Cross-lingual BERTScore lets a fundamentally different evaluation paradigm called quality estimation (QE): assessing translation quality without access to reference translations. In QE, the only inputs are the source sentence and the machine translation output. Traditional QE systems required complex feature engineering and large amounts of annotated data. Cross-lingual BERTScore gives a surprisingly effective zero-shot baseline by measuring the semantic alignment between the source and the translation directly in the shared multilingual embedding space.
In practice, cross-lingual BERTScore works best for language pairs that are well-represented in the multilingual model's training data. For high-resource pairs like English-German or English-Chinese, the shared embedding space is well-populated and cross-lingual similarities are reliable. For low-resource pairs involving languages with little multilingual training data, the embedding alignment is weaker and scores may be less reliable. The evaluation of cross-lingual BERTScore on the WMT quality estimation tasks has shown promising results, often competitive with supervised QE systems trained on thousands of labeled examples.
Correlation with Human Judgments
The ultimate test of any automatic evaluation metric is how well it correlates with human judgments of quality. BERTScore was extensively validated against human ratings on standard NLG benchmarks when it was introduced.
On the WMT 2017 and 2018 machine translation shared tasks, BERTScore achieved higher correlation with human adequacy judgments than BLEU, METEOR, and TER across most language pairs. The improvements were particularly pronounced for translation pairs involving languages with different word orders, such as English-Chinese and English-Japanese, where n-gram matching is inherently less reliable. For European language pairs with similar word order, the improvements over BLEU were smaller but still consistent.
On summarization benchmarks, BERTScore correlated better with human judgments of semantic adequacy and coherence than ROUGE, particularly for abstractive summaries that use diverse vocabulary relative to the source text. Extractive summaries, which copy sentences directly from the source, tended to score similarly on both metrics, since the vocabulary overlap is high regardless.
The meta-evaluation literature on evaluation metrics has also shown that no single metric perfectly captures all dimensions of human quality judgment. BERTScore correlates well with semantic adequacy (does the candidate convey the right meaning?) but less well with fluency (is the candidate grammatically correct and natural-sounding?) and coherence (does the candidate flow logically across sentences?). Combining BERTScore with metrics designed to capture these other dimensions, such as perplexity for fluency or discourse-level coherence scores, gives more complete evaluation coverage than any single metric alone.
Limitations and Impact
Despite its advantages, BERTScore introduces new challenges that practitioners must consider carefully.
Computational Cost
The computational cost is significant: evaluating with BERTScore requires forward passes through a large transformer model, which makes it orders of magnitude slower than BLEU. This becomes prohibitive when evaluating billions of generated tokens during hyperparameter sweeps or large-scale benchmarking.
The computational overhead of BERTScore limits its applicability in resource-constrained environments. While BLEU can evaluate thousands of sentences per second on standard hardware, BERTScore processing may slow to dozens or hundreds of sentences per second, depending on model size and hardware acceleration. For research scenarios involving large generated datasets, such as evaluating the output of large language models on thousands of prompts, this slowdown can extend experiments from hours to days. GPU acceleration helps, but the memory requirements of large transformer models restrict batch sizes and increase infrastructure costs. This computational burden necessitates careful planning when incorporating BERTScore into evaluation pipelines, potentially requiring subsampling for preliminary experiments or reserved use for final evaluation stages.
Model Dependency
BERTScore's behavior is model-dependent. Scores vary based on which pre-trained model you choose (BERT vs. RoBERTa vs. DeBERTa), which layer you extract embeddings from, and whether you use cased or uncased tokenization. This makes absolute scores difficult to interpret without calibration. A BERTScore F1 of 0.70 using BERT-base uncased layer 9 means something different than 0.70 using RoBERTa-large layer 17.
This model dependency complicates the comparison of results across studies or the establishment of universal quality thresholds. A score that indicates excellent quality in one paper using BERT-large might correspond to mediocre quality in another using a different configuration. Researchers must report detailed configuration parameters and ideally give baseline scores on standard datasets to let real comparison. The field has not yet converged on standard configurations, though RoBERTa-large with specific layer choices is emerging as a common default. Practitioners should calibrate scores against human judgments on their specific domain to establish appropriate thresholds for quality control.
Inherited Biases
The metric inherits biases from its underlying language model. If the pre-trained model associates certain demographic groups with negative sentiments or professional roles with specific genders, BERTScore may inappropriately penalize valid translations that contradict these stereotypes. Additionally, BERTScore struggles with languages or domains poorly represented in the pre-training data; evaluating rare languages or highly technical jargon may produce unreliable similarity estimates.
The bias issue is a serious ethical and methodological concern. Language models trained on web text reproduce the biases present in their training data, including gender stereotypes, racial biases, and cultural assumptions. When BERTScore evaluates translations involving demographic references or social content, it may assign lower scores to translations that use inclusive language or challenge traditional stereotypes, simply because such usage was rare in the training data. For example, a translation referring to a doctor as "she" might receive lower similarity scores if the pre-trained model associates medical professions primarily with male pronouns. Similarly, evaluations of low-resource languages or highly specialized technical domains suffer from the distributional limitations of the training data: if the model has rarely seen text in a particular language or domain, its embeddings may not adequately capture semantic relationships, leading to arbitrary or misleading similarity scores.
Context Window Limitations
Context window limitations constrain BERTScore's applicability to long documents. Standard BERT models handle only 512 tokens, requiring chunking for longer texts, which may break cross-sentence dependencies. While Longformer or BigBird architectures (covered in Part XVII: Efficient Attention) can extend this range, they are not standard in BERTScore implementations.
The 512-token limit of standard BERT architectures poses practical challenges for evaluating long-form content such as document-level translation, multi-paragraph summarization, or book-length text generation. Chunking strategies, such as sliding windows or paragraph splitting, allow BERTScore to process longer texts but introduce artifacts at chunk boundaries. Coreference relationships that span across chunks, such as pronoun resolution or thematic continuity, may be lost when text is segmented. This limitation is particularly problematic for discourse-level evaluation, where the coherence and cohesion of long texts are important quality aspects. Recent developments in long-context transformers offer potential solutions, but their integration into standard BERTScore toolkits remains incomplete, and their computational requirements are even more substantial than those of standard models.
What BERTScore Does Not Capture
It is equally important to understand what BERTScore does not measure. The metric assesses semantic similarity at the token level but does not directly evaluate:
- Fluency: A candidate with perfect semantic content but scrambled word order may score well on BERTScore while being unreadable.
- Factual accuracy: BERTScore measures whether the candidate is semantically similar to the reference, not whether the reference itself is factually correct or whether the candidate preserves factual claims accurately.
- Discourse coherence: Token-level alignment cannot detect whether a multi-sentence summary flows logically or maintains consistent perspective across paragraphs.
- Style and tone: Two sentences can be semantically equivalent but wildly different in register, formality, or tone, and BERTScore will not distinguish between them.
Awareness of these gaps is important when designing evaluation protocols. For tasks where fluency, factual accuracy, or coherence matter, BERTScore should be complemented with other metrics or human evaluation. It is most useful as one component of a complete evaluation suite, not as a standalone replacement for human judgment.
Impact on the Field
BERTScore has fundamentally changed how researchers approach automatic evaluation. It demonstrated that semantic embeddings could replace surface statistics, paving the way for learned metrics like BLEURT and the use of large language models as evaluators (which we'll explore in upcoming chapters on model-based evaluation). In machine translation, summarization, and dialogue generation, BERTScore now regularly appears alongside traditional metrics. This gives a more fine-grained view of generation quality that aligns better with human judgments of meaning preservation.
The impact of BERTScore extends beyond its immediate use as a metric. By proving that pre-trained contextual embeddings could reliably assess semantic equivalence, BERTScore catalyzed the development of an entire family of neural evaluation metrics. It established the paradigm of using frozen pre-trained models as evaluators, an approach now extended by metrics that use prompting techniques with large language models. In the machine translation community, BERTScore is now a standard component of evaluation suites, alongside BLEU and chrF. In summarization research, it helps distinguish between extractive and abstractive approaches by crediting the latter for semantic preservation even when phrasing differs completely from source texts. The metric has also found applications in dialogue systems, data-to-text generation, and even image captioning evaluation, wherever the semantic fidelity of generated text must be assessed against reference descriptions.
Summary
BERTScore evaluates text generation by computing token-level semantic similarity using contextualized embeddings from pre-trained transformers. Rather than counting matching n-grams, it constructs a similarity matrix between reference and candidate token embeddings, then applies greedy matching to compute precision, recall, and F1 scores.
Key technical aspects include:
- Embedding extraction from intermediate transformer layers (typically layers 9-10 for BERT-base), which balance syntactic and semantic information better than either the lowest or the highest layers
- Cosine similarity computation between all pairs of reference and candidate token embeddings, creating a dense similarity matrix that captures synonym relationships and paraphrases
- Greedy alignment that independently matches each token to its most similar counterpart for precision and recall, giving linear-time computation with near-optimal matching quality
- Optional IDF weighting that discounts common function words and stresses rare, content-bearing terms, improving alignment with human judgments for domain-specific texts
- Baseline rescaling that normalizes scores against random sentence pairs, making absolute values interpretable and comparable across different model configurations
The metric captures paraphrases and synonyms that surface-form metrics miss. This gives evaluation that correlates strongly with human judgments of semantic adequacy. Variants like MoverScore extend the approach using optimal transport for global alignment, while domain-specific adaptations use specialized pre-trained models such as SciBERT, BioBERT, and CodeBERT to better capture technical vocabulary.
When comparing BERTScore to BLEU, the trade-off is clear: semantic richness and robustness to paraphrasing versus computational efficiency and lexical precision. BERTScore excels in evaluation scenarios where valid paraphrasing is expected and where single reference translations cannot cover all acceptable phrasings. BLEU retains value where exact terminology is required and where computational constraints are tight.
The limitations of BERTScore deserve equal attention: its computational cost, model dependency, inherited biases, context window constraints, and inability to assess fluency, factual accuracy, or discourse coherence. Understanding these limitations guides appropriate use. BERTScore is most useful as one component of a complete evaluation suite.
We'll build on these evaluation foundations in the next chapter, which explores Exact Match and F1 metrics for tasks requiring precise output structure, and later we'll examine how large language models themselves can serve as evaluators, pushing automatic evaluation closer to the fine-grained judgment of human experts.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about BERTScore and semantic evaluation metrics.
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!