Uncertainty Quantification in Language Models

Michael BrenndoerferFebruary 24, 202657 min read

Part of Language AI Handbook

Measure and communicate LLM confidence through calibration, verbalized uncertainty, semantic entropy, and trust-calibrated user interfaces.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Uncertainty Quantification

Language models generate text with apparent confidence. They produce fluent sentences, cite statistics, name sources, and answer questions in a tone that reads as authoritative. But that surface confidence is not the same as actual reliability. A model can be wrong and sound certain, or correct and sound hesitant. The two things, confidence and accuracy, are entirely independent unless something specifically ties them together.

Uncertainty quantification (UQ) is the discipline of measuring and communicating how much a model knows. In traditional machine learning, this means producing probability estimates that accurately reflect the likelihood of being correct. A classifier that says "90% confident" on 100 examples should be right about 90 of them. That property is called calibration, and well-calibrated models are useful precisely because their confidence scores are actionable. When a model's score is high, you can trust the answer. When it is low, you know to send it for review.

Language models complicate this picture in several ways. First, they output distributions over entire token sequences, not single-class probabilities. A classification model has a small, fixed output space and can directly assign a probability to each class. A language model's output space is astronomically large: every possible sequence of tokens is a valid output, and measuring confidence over that space requires either clever approximations or accepting a necessarily incomplete picture. Second, large language models have been trained through reinforcement learning from human feedback (RLHF), which systematically pushes them toward confident-sounding responses because human raters often prefer decisive, authoritative answers over hedged, qualified ones. This is a direct calibration-corrupting force built into the training process itself. Third, they operate on open-ended tasks where "correctness" is hard to define and harder to measure. The result is that modern large language models are frequently miscalibrated: they express more confidence than their actual accuracy warrants, especially on questions near the edge of their knowledge.

Consider a concrete example. You ask a language model who won the 1987 Nobel Prize in Chemistry. The model answers confidently and fluently. But unless something in its training process specifically tied that confident tone to verified correct answers, the fluency is evidence only of good language modeling, not factual accuracy. The model generates the words "Donald Cram, Jean-Marie Lehn, and Charles Pedersen" with high token probability not because it checked a database but because those tokens follow plausibly from the context in a way that matches its training distribution. Fluency and accuracy are not the same thing, and well-calibrated uncertainty is what bridges the gap between them.

As we explored in the previous chapters on hallucination types, detection, causes, and mitigation, the problem of hallucination is fundamentally about a model producing output that doesn't match reality. Uncertainty quantification provides a complementary lens: instead of asking "did this output hallucinate?", it asks "how confident should we be in this output before we even deliver it?" Properly quantified uncertainty can trigger retrieval augmentation, send answers for human review, or communicate limitations directly to users. The previous chapter on attribution and citation handles the question of "where did this information come from?". Uncertainty quantification handles "how certain is the model that the information is correct?"

This chapter covers the four major pillars of uncertainty quantification for language models:

  • Confidence calibration: measuring and correcting the gap between expressed and actual confidence
  • Verbalized uncertainty: training models to express their uncertainty in natural language
  • Sampling-based uncertainty: estimating uncertainty by observing variance across multiple model outputs
  • Uncertainty communication: designing interfaces and language to convey uncertainty to end users effectively
Calibration vs. Uncertainty Quantification

These terms are related but distinct. Calibration refers specifically to the alignment between confidence scores and actual accuracy. Uncertainty quantification is the broader practice of estimating, representing, and communicating model uncertainty, of which calibration is one component.

Confidence Calibration

Calibration is the foundational concept in uncertainty quantification. A model is well-calibrated if its stated confidence probabilities match the frequency of being correct. Understanding calibration requires first understanding how confidence is defined for language models, then understanding how that confidence deviates from ideal behavior, and finally seeing how it can be corrected.

Token-Level and Sequence-Level Probability

Language models output a probability distribution over vocabulary tokens at each decoding step. Given a context cc (the input prompt plus all tokens generated so far), the model assigns a probability to each possible next token tt:

P(t∣c)=exp⁡(zt/T)∑v∈Vexp⁡(zv/T)P(t \mid c) = \frac{\exp(z_t / T)}{\sum_{v \in V} \exp(z_v / T)}

where:

  • ztz_t is the logit (pre-softmax score) for token tt, computed from the model's final linear layer
  • VV is the full vocabulary (commonly 32,000 to 128,000 tokens in modern LLMs)
  • TT is the temperature parameter, defaulting to T=1T = 1 at inference time

The logit ztz_t is a raw score proportional to the model's internal "preference" for token tt. The softmax function converts these scores into a valid probability distribution by exponentiating and normalizing. The temperature TT controls how sharp or flat the resulting distribution is: at T=1T = 1 the distribution reflects the raw model preferences, at T<1T < 1 it concentrates mass on the top token (more deterministic), and at T>1T > 1 it spreads mass more evenly (more random).

This gives us token-level probabilities. For a complete response y=(t1,t2,…,tn)y = (t_1, t_2, \ldots, t_n), the sequence probability follows from the chain rule of probability, factoring the joint distribution over the autoregressive generation process:

P(y∣x)=∏i=1nP(ti∣x,t1,…,ti−1)P(y \mid x) = \prod_{i=1}^{n} P(t_i \mid x, t_1, \ldots, t_{i-1})

where xx is the original input prompt. Each factor is the probability of generating the ii-th token given the prompt and all previously generated tokens.

Sequence probabilities are products of many values below 1, so they shrink exponentially with length. A response of 20 tokens might have a sequence probability of 10−3010^{-30}, making direct comparison across different-length responses misleading. A short, confident answer would appear more probable than an equally correct but longer one simply because it has fewer multiplicative terms.

The log-probability resolves this by converting the product into a sum:

log⁡P(y∣x)=∑i=1nlog⁡P(ti∣x,t1,…,ti−1)\log P(y \mid x) = \sum_{i=1}^{n} \log P(t_i \mid x, t_1, \ldots, t_{i-1})

Since log⁡P(ti∣⋅)\log P(t_i \mid \cdot) is negative (probabilities are between 0 and 1, so their log is ≤0\leq 0), the log-probability is a non-positive number: less negative means higher confidence. Length-normalized log-probability divides by nn to make scores comparable across responses of different lengths:

score(y∣x)=1n∑i=1nlog⁡P(ti∣x,t1,…,ti−1)\text{score}(y \mid x) = \frac{1}{n} \sum_{i=1}^{n} \log P(t_i \mid x, t_1, \ldots, t_{i-1})

This normalized score is the most commonly used single-number summary of how "confident" a model is about a particular response. It represents the average per-token log-probability: a value near 0 means the model assigned high probability to each token it generated (confident), while a very negative value means many tokens had low probability (uncertain).

There is one important subtlety here. The length-normalized log-probability measures fluency, not accuracy. A model can fluently generate a confident-sounding false statement with high per-token probability. The score reflects how consistent the output is with the model's training distribution, not how consistent it is with ground truth. This disconnect is the central challenge that all calibration methods attempt to address.

What Is Calibration?

A classifier is perfectly calibrated if, among all predictions made with confidence pp, exactly a fraction pp are correct. Formally:

P(y^=y∣p^=p)=p∀p∈[0,1]P(\hat{y} = y \mid \hat{p} = p) = p \quad \forall p \in [0, 1]

where y^\hat{y} is the predicted class, yy is the true label, and p^\hat{p} is the model's stated confidence.

You can measure calibration empirically by grouping predictions into bins by confidence value, computing accuracy within each bin, and comparing the two. A reliability diagram (also called a calibration curve) plots predicted confidence on the x-axis against observed accuracy on the y-axis. Perfect calibration is the diagonal line y=xy = x: a model predicting 70% confidence should be right 70% of the time.

Reliability diagrams are powerful diagnostic tools. If the curve lies below the diagonal, the model is overconfident: it claims higher confidence than its actual accuracy. If the curve lies above the diagonal, the model is underconfident: its stated confidence is lower than its actual accuracy. The shape of the deviation reveals the type of miscalibration. A model that is overconfident across all confidence levels has a curve uniformly below the diagonal. A model that is overconfident only at high confidence values (but well-calibrated at low confidence) has a curve that tracks the diagonal until about 0.7, then bends below it. These different patterns call for different correction methods.

Expected Calibration Error (ECE)

ECE is the standard metric for measuring calibration quality. It partitions predictions into MM equal-width confidence bins and computes the weighted average of the gap between confidence and accuracy in each bin:

ECE=∑m=1M∣Bm∣n∣acc(Bm)−conf(Bm)∣\text{ECE} = \sum_{m=1}^{M} \frac{|B_m|}{n} \left| \text{acc}(B_m) - \text{conf}(B_m) \right|

where:

  • MM is the total number of bins (typically 10 or 15)
  • BmB_m is the set of predictions whose confidence falls in bin mm
  • ∣Bm∣|B_m| is the number of predictions in bin mm
  • nn is the total number of predictions
  • acc(Bm)\text{acc}(B_m) is the fraction of predictions in bin mm that are correct
  • conf(Bm)\text{conf}(B_m) is the mean predicted confidence in bin mm

The weight ∣Bm∣/n|B_m|/n ensures that larger bins contribute more to the overall ECE. Lower ECE is better; a perfectly calibrated model achieves ECE = 0.

ECE has a known limitation: if most predictions fall in a few bins, those bins dominate the metric even if the model is well-calibrated in the sparse bins. Adaptive calibration error (ACE) addresses this by using equal-frequency bins rather than equal-width bins, making sure each bin contains roughly the same number of predictions. Maximum calibration error (MCE) reports the worst-case bin gap rather than the average. For most purposes, standard ECE with 10-15 bins provides an adequate first-pass diagnostic.

Overconfidence in Language Models

Most large language models are overconfident: their confidence scores are systematically higher than their actual accuracy. If you ask a model 100 factual questions where it expresses "95% confidence," you might find it correct on only 70. The calibration curve bends below the diagonal throughout the range.

This happens for several distinct reasons that reinforce each other:

Training objective mismatch. The next-token prediction objective optimizes for placing probability mass on the correct token during training. On the training data, the model repeatedly encounters the correct answer and learns to push high probability toward it. But this process doesn't teach the model to reserve lower probabilities for questions it's less certain about. During inference on novel questions, the model still outputs high probabilities because that's what it was optimized to do, even when the answer is uncertain.

RLHF calibration shift. Reinforcement learning from human feedback trains models to produce responses that humans rate highly. Confident, decisive answers systematically receive higher ratings than hedged, qualified ones. Annotators typically find "The Battle of Hastings was fought in 1066" more helpful than "The Battle of Hastings was fought around 1066, though I'm not entirely certain of the exact year." This creates a training signal that directly penalizes verbal uncertainty and rewards overconfidence.

Distribution shift. Models are calibrated on their training distribution but evaluated on a much wider range of queries. When a model encounters a question near the edge of its training knowledge, it doesn't necessarily recognize that it's at the edge, so its confidence doesn't drop accordingly. A model trained primarily on English Wikipedia may be quite confident on topics that happen to be well-covered there, even when its knowledge is shallow or biased by the articles' own errors.

Temperature effects at inference. Many production systems use low inference temperatures (often T<1T < 1) to produce more deterministic, less repetitive outputs. Low temperature sharpens the probability distribution, pushing top-token probabilities higher without necessarily improving accuracy. A model running at T=0.5T = 0.5 will report higher token probabilities than the same model at T=1.0T = 1.0 even when generating the same output.

Lack of calibration objectives during training. Standard language model training has no explicit calibration objective. The model is never asked "how often are you right when you say you're 90% confident?" Calibration can only emerge implicitly from the training data, and since most training text doesn't include explicit uncertainty labels, the model has no direct signal to learn from.

Understanding these causes matters because they suggest different remedies. RLHF miscalibration calls for changing the reward function. Temperature effects call for post-hoc scaling. Distribution shift calls for domain-specific calibration. No single method addresses all causes simultaneously.

Temperature Scaling

Temperature scaling is the simplest and most effective post-hoc calibration technique. After training, you learn a single scalar temperature T∗T^* on a held-out calibration set by minimizing the negative log-likelihood:

T∗=arg⁡min⁡T  −∑ilog⁡exp⁡(zyi/T)∑vexp⁡(zv/T)T^* = \arg\min_{T} \; -\sum_{i} \log \frac{\exp(z_{y_i} / T)}{\sum_v \exp(z_v / T)}

where:

  • TT is the temperature parameter (the single scalar we are optimizing)
  • zyiz_{y_i} is the logit of the true class for example ii
  • The sum in the denominator is over all vocabulary items vv

A temperature T>1T > 1 softens the distribution (reduces confidence), while T<1T < 1 sharpens it. For overconfident models, the optimal T∗>1T^* > 1. The key insight is that the optimization is over just one parameter, making it fast and stable even with small calibration sets.

Temperature scaling is appealing because it modifies only the confidence values, not the ranking of predictions. The most likely class remains the same; only the probability it's assigned changes. This means you can apply temperature scaling to an already-deployed model without changing its output quality, adding a calibration layer on top of the existing system.

The main limitation is that temperature scaling applies the same correction uniformly across all inputs and confidence levels. If a model is overconfident on some question types but underconfident on others, a single temperature can't correct both simultaneously. More flexible methods are needed for heterogeneous miscalibration.

Platt Scaling and Isotonic Regression

More flexible calibration methods learn non-linear mappings from raw confidence scores to calibrated probabilities.

Platt scaling fits a logistic regression model to the raw scores:

p^=σ(A⋅s+B)\hat{p} = \sigma(A \cdot s + B)

where ss is the raw confidence score, σ\sigma is the sigmoid function, and AA and BB are learned parameters. With two parameters instead of one, Platt scaling can handle asymmetric miscalibration: for example, a model that is overconfident at high scores but underconfident at low scores. The logistic function also ensures the output is a valid probability between 0 and 1.

Isotonic regression fits a piecewise constant monotone function from raw scores to probabilities. The monotonicity constraint preserves the model's original ranking (if score A>A > score BB before calibration, then P(A)>P(B)P(A) > P(B) after), while the piecewise constant shape allows arbitrary flexible correction within each piece. Isotonic regression can correct any form of miscalibration as long as it is monotone, but it requires more calibration data to avoid overfitting. With small calibration sets, the piecewise function may overfit to the specific examples seen during calibration.

Both methods require a labeled calibration set, which is harder to construct for language models than for classifiers. You need questions with known ground-truth answers, model confidence scores, and correctness labels. For factual QA, this can be approximated by sampling questions from a knowledge base and automatically checking answers. For more complex tasks, human annotation is required.

Calibration for Open-Ended Generation

Calibration for open-ended generation is qualitatively harder than for classification. A factual QA task has clear correct answers, but for summarization or open-ended dialogue, "correct" is a spectrum. Two practical approaches have emerged:

Binary correctness framing. Define correctness as "does the response match the ground truth?", use an automated evaluator (such as an NLI model or an LLM-as-judge) to label responses, and treat the problem as binary classification. This lets standard calibration metrics apply directly. The limitation is that the quality of the calibration estimate is bounded by the quality of the automated evaluator. If the evaluator makes systematic errors, those errors propagate into the calibration measurement.

Semantic clustering. Run the model multiple times on the same question, cluster responses by meaning, and treat the frequency of the most common cluster as a proxy for confidence. Well-calibrated models should produce high-frequency clusters when they're correct and more diverse clusters when uncertain. This approach does not require ground-truth answers, making it applicable even when automated evaluation is not possible. We cover this in detail in the sampling-based uncertainty section below.

Verbalized Uncertainty

Beyond numerical confidence scores, models can express uncertainty in natural language. Rather than producing a probability value, the model says something like "I'm not certain about this, but..." or "I believe this is correct, though you should verify with a primary source." This is called verbalized uncertainty or linguistic uncertainty expression.

Why Verbalized Uncertainty Matters

Numerical confidence scores are not human-readable by default. When you deploy an LLM in a product, the end user receives text, not logits. Verbalized uncertainty lets models communicate the same information in a form users can act on, directly in the response.

Verbalized uncertainty can convey more detail than a single probability. A model can distinguish between different types of uncertainty within a single response:

  • "I'm confident this happened but unsure of the exact date" (factual uncertainty with partial knowledge)
  • "This is a contested topic among experts" (epistemic disagreement in the world)
  • "This falls outside my knowledge cutoff" (temporal uncertainty about recent events)
  • "I may be confusing this with a similar event" (confusion or conflation uncertainty)

These distinctions matter for how a user should respond. Knowing to check a specific date is different from knowing to consult an expert on a contested topic, which is different again from knowing that the answer simply isn't knowable from an LLM's training data. A single numeric confidence score collapses all of these into a single value and loses the actionable structure.

There is also a practical benefit in user experience. Research on trust in AI systems shows that users who receive explicit uncertainty information make better decisions than users who receive confident answers without uncertainty markers. When a model says "I believe this is correct, but you should verify," users check more often and catch errors more reliably. The uncertainty expression acts as a metacognitive cue that activates verification behavior.

Epistemic vs. Aleatoric Uncertainty

The uncertainty quantification field distinguishes two fundamental types of uncertainty that appear throughout statistics and machine learning:

Epistemic vs. Aleatoric Uncertainty

Epistemic uncertainty (model uncertainty) arises from limited knowledge. It could in principle be reduced by providing more information or training data. "I don't know who invented the transistor" is epistemic: the answer exists and is known, but this particular model doesn't have it reliably encoded.

Aleatoric uncertainty (data uncertainty) arises from inherent randomness in the world that cannot be reduced by more information. "I can't predict exactly when the next earthquake will occur" is aleatoric: even a perfect model with complete knowledge of the world cannot reduce this uncertainty because it reflects inherent randomness in the underlying process.

In language models, the distinction is blurry in practice. The model cannot distinguish "I don't know because I wasn't trained on this" from "nobody knows because it is inherently uncertain." A question about a contested historical interpretation and a question about a fact the model simply wasn't trained on look the same from the model's perspective: both produce uncertainty in the output, but for very different reasons.

The distinction matters practically because the appropriate response differs. Epistemic uncertainty can often be resolved by retrieval or tool use: if the model doesn't know something, fetching a document might provide the answer. Aleatoric uncertainty cannot be resolved this way: no amount of retrieval will tell you exactly when an earthquake will happen. A well-designed system that can distinguish the two would route differently: retrieval augmentation for epistemic uncertainty, clear communication of irreducible uncertainty for aleatoric cases.

Some researchers have proposed using sampling-based methods to proxy this distinction. The intuition is that epistemic uncertainty (things the model learned inconsistently across different parts of training) would produce high sampling variance, while aleatoric uncertainty (inherently uncertain facts) would produce consistently uncertain outputs. This intuition doesn't always hold in practice, but it motivates semantic entropy and related techniques as approximate probes of epistemic uncertainty.

Training for Verbal Calibration

The challenge with verbalized uncertainty is that models must learn to say "I don't know" when they do not know, not just when prompted to hedge. Several training approaches have been developed:

Supervised fine-tuning on uncertainty-labeled data. Collect question-answer pairs where the model's answers are scored for correctness. Fine-tune the model to produce hedged responses when it gets the answer wrong and confident responses when correct. This requires a labeled dataset of model errors, which is expensive to construct. The model must learn what to say and when saying "I'm not sure" is appropriate, which depends on having a reliable internal signal of its own knowledge boundaries.

RLHF with calibration reward. Include a calibration signal in the reward function during RLHF. Responses that accurately hedge on difficult questions get higher reward; overconfident wrong responses get penalized. The challenge is operationalizing "accurate hedging" as a reward. One approach: use an external verifier to check factual claims, then reward hedging when claims are wrong and penalize hedging when claims are correct.

Prompting. Large models can be prompted to verbalize uncertainty without fine-tuning. A system prompt like "If you are not certain about something, say so explicitly and explain why" can improve verbal calibration for models with sufficient instruction-following ability. The effect is stronger in larger models that have better instruction comprehension. This is the most practically accessible approach for most deployment teams, requiring no additional training infrastructure.

RLAIF (reinforcement learning from AI feedback). An AI judge evaluates whether the model's expressed uncertainty matches its actual correctness. This creates a scalable signal for calibration without requiring human labeling of every question. The judge assesses: "This model said it was 'fairly confident' and was correct. Did the expressed confidence match the actual outcome?" Aggregated over many examples, this signal trains toward verbal calibration.

Limitations of Verbalized Uncertainty

Verbalized uncertainty has a fundamental problem: the model can say "I'm not sure" without being uncertain internally. It's a surface-level behavior that can be decoupled from the model's internal confidence state.

Studies have found that LLMs often produce calibrated-sounding hedges without having lower token-level confidence on the content they're hedging. The model has learned that certain phrasings ("it may be", "I believe") are stylistically appropriate in certain contexts. This is sometimes called sycophantic hedging: the model hedges to seem appropriately humble rather than because its confidence is lower. If you measure the token-level probability on responses where the model verbally hedges versus responses where it doesn't hedge, you often find very little difference.

A related problem is topic-based hedging: models learn to hedge on topics that are stereotypically uncertain (emerging science, predictions about the future, contested political topics) and to express confidence on topics that are stereotypically factual (historical dates, mathematical facts), regardless of whether they know the specific answer. A model might say "I'm confident that World War II ended in 1945" (correct) but also "I'm confident that the Battle of Maldon was fought in 981" (incorrect) because historical dates pattern-match to the confident register, not because the model has verified the date.

A practical consequence: you cannot take verbalized uncertainty at face value. If a model says "I'm quite certain that...", that phrase may reflect the model's writing style for that topic type more than its actual reliability on that specific claim. Sampling-based methods and calibration analysis provide stronger evidence because they measure behavioral consistency rather than surface language patterns.

Sampling-Based Uncertainty

Instead of trying to extract a confidence score from a single model run, sampling-based methods estimate uncertainty by running the model multiple times on the same query and measuring how much the outputs vary. High variance across samples suggests the model is uncertain; low variance suggests confidence.

This approach sidesteps several problems with token-level probability as a confidence measure. Token probabilities are affected by generation artifacts (temperature, sampling parameters), training biases, and RLHF shifts that don't necessarily reflect factual uncertainty. Behavioral variance across samples is a more direct measure of whether the model would give consistent answers in deployment: a model that gives the same answer ten times out of ten when sampled is behaving consistently, regardless of what probabilities it assigns internally.

Monte Carlo Sampling

The basic idea is Monte Carlo estimation: draw NN samples from the model's distribution and compute a statistic over them. The variance of that statistic estimates uncertainty.

For a yes/no question, this is straightforward:

  1. Generate NN responses with temperature T>0T > 0 (to introduce sampling randomness)
  2. Count how many responses say "yes" versus "no"
  3. Let pp be the fraction of "yes" responses: p=∣{i:yi="yes"}∣/Np = |\{i : y_i = \text{"yes"}\}| / N
  4. Compute binary entropy as the uncertainty measure:
H(p)=−plog⁡p−(1−p)log⁡(1−p)H(p) = -p \log p - (1-p) \log(1-p)

where:

  • pp is the proportion of responses in the majority class
  • H(p)H(p) reaches its maximum of log⁡2≈0.693\log 2 \approx 0.693 nats (or 1 bit) when p=0.5p = 0.5, indicating maximum uncertainty
  • H(p)=0H(p) = 0 when p∈{0,1}p \in \{0, 1\}, indicating complete agreement and minimum uncertainty

Higher entropy (answers split near 50/50) means higher uncertainty. Near-unanimous answers indicate confidence. The beauty of this approach is that it requires no access to model internals: you only need to observe the model's outputs.

The computational cost is the main drawback. Running NN samples means NN forward passes through the model. For a model with a single generation costing $0.01 and N=10N = 10 samples per query, the uncertainty estimation costs $0.10 per query, ten times more than a single generation. For high-throughput production systems, this cost is often prohibitive.

Semantic Consistency

For open-ended responses, simple string comparison is too strict. "Paris is the capital of France" and "France's capital city is Paris" express the same fact but would be counted as different answers under exact match. Semantic consistency measures whether multiple responses convey the same meaning, independent of surface form.

One approach: use an NLI (natural language inference) model to check whether pairs of sampled responses entail each other. If response AA entails BB and BB entails AA, they are semantically equivalent. The fraction of response pairs that are semantically equivalent is the consistency score. Low consistency means high uncertainty.

Consistency(y1,…,yN)=1(N2)∑i<j1[NLI(yi,yj)=entail  ∧  NLI(yj,yi)=entail]\text{Consistency}(y_1, \ldots, y_N) = \frac{1}{\binom{N}{2}} \sum_{i < j} \mathbb{1}[\text{NLI}(y_i, y_j) = \text{entail} \;\land\; \text{NLI}(y_j, y_i) = \text{entail}]

where NLI(a,b)=entail\text{NLI}(a, b) = \text{entail} means "response aa entails response bb". The bidirectional entailment check (both AA entails BB and BB entails AA) is a proxy for semantic equivalence. NLI models are fast compared to generation, so the consistency computation adds minimal overhead once samples are available.

The limitation is that NLI models are themselves imperfect and can make errors on semantically subtle cases. Responses that are paraphrases of the same claim may fail the entailment check if they use different vocabulary. This is an active area of improvement in the field.

Self-Consistency Prompting

Self-consistency (Wang et al., 2022) was originally proposed to improve reasoning accuracy but also provides a natural uncertainty estimate. The method generates multiple chain-of-thought reasoning paths and uses majority voting to select the final answer.

The process works as follows. The model is prompted to think step by step, and this prompt is sent NN times with non-zero temperature to produce NN different reasoning paths. Each path arrives at a final answer. Majority voting selects the most common final answer as the output. The fraction of samples agreeing with the majority answer is a confidence score. If 9 out of 10 sampled paths arrive at the same answer, confidence is 0.9. If only 6 out of 10 agree, confidence is 0.6.

Self-consistency works well for structured tasks (math problems, multiple-choice QA) where you can extract a clean final answer from each reasoning path. For open-ended generation, you need semantic clustering instead of exact match voting: run all sampled responses through an NLI model to cluster them by meaning, then take the fraction belonging to the largest cluster as the confidence estimate.

The insight behind self-consistency is that different reasoning paths through a problem serve as approximate independent samples from the model's belief distribution. When the model reliably knows the answer, most paths converge on it. When the model is uncertain, paths diverge into different regions of the answer space. The variance across paths encodes the model's behavioral uncertainty.

Semantic Entropy

Semantic entropy (Kuhn et al., 2023) formalizes sampling-based uncertainty using information theory. Rather than computing entropy over raw text samples (which conflates surface variation with semantic variation), it computes entropy over semantic clusters.

The procedure has four steps:

  1. Sample NN responses y1,…,yNy_1, \ldots, y_N from the model with non-zero temperature
  2. Group responses into semantic equivalence classes C1,C2,…,CKC_1, C_2, \ldots, C_K using bidirectional NLI entailment
  3. Estimate the probability of each class: P(Ck)≈∣Ck∣/NP(C_k) \approx |C_k| / N
  4. Compute semantic entropy:
SE=−∑k=1KP(Ck)log⁡P(Ck)\text{SE} = -\sum_{k=1}^{K} P(C_k) \log P(C_k)

where KK is the number of distinct semantic clusters, CkC_k is the kk-th cluster, ∣Ck∣|C_k| is the number of responses assigned to it, and NN is the total number of sampled responses. The log⁡\log here is the natural logarithm, so SE is in units of nats.

Semantic entropy is zero when all responses are semantically identical (certain) and increases as responses become more diverse in meaning. It is invariant to surface-form variation: if all responses say the same thing in different words, they form one cluster and SE is zero. It correlates more strongly with factual accuracy than token-level probability entropy, which is noisier because it treats surface variation as signal.

Why Semantic Entropy Is Better Than Token Entropy

Token-level entropy over responses is noisy because different surface forms of the same meaning contribute independently. "The capital is Paris" and "Paris is the capital" are different character sequences but identical semantic content. Computing entropy over them as distinct strings artificially inflates the uncertainty estimate. Semantic entropy groups by meaning before computing entropy, removing this noise and producing a cleaner, more accurate uncertainty signal. The Kuhn et al. paper showed that semantic entropy substantially outperforms token-level probability measures at predicting factual accuracy on open-domain QA benchmarks.

Dropout-Based Uncertainty

In neural networks, Monte Carlo Dropout approximates Bayesian inference by keeping dropout active at test time. Multiple forward passes through the network, each with a different random dropout mask, produce a distribution over predictions. The variance of this distribution estimates epistemic uncertainty.

The theoretical foundation is that a neural network with dropout at every layer is mathematically equivalent to a deep Gaussian process under certain conditions (Gal and Ghahramani, 2016). Each forward pass with a different dropout mask can be interpreted as sampling from an approximate posterior distribution over the network's weights. The variance across these samples therefore estimates the posterior uncertainty about the model's predictions.

For language models, this has limited practical use because large language models typically do not use dropout during inference, and enabling it changes the model's behavior unpredictably. The sheer scale of modern LLMs makes the Gaussian process approximation less precise, and activating dropout at inference time often degrades output quality substantially. This makes it difficult to separate "uncertainty due to knowledge limits" from "degradation due to random dropped activations."

But the conceptual framework carries over: any source of stochasticity (temperature sampling, different random seeds, different prompt phrasings) can be used to generate a distribution from which uncertainty is estimated. Temperature sampling is the most direct analogue. It introduces randomness proportional to the model's own uncertainty in its token predictions, and sampling multiple times with non-zero temperature produces a distribution over outputs that encodes the model's behavioral uncertainty about what to say.

Consistency-Based Uncertainty via Prompt Perturbation

A related technique uses paraphrased versions of the same question to probe consistency. If the model gives different answers to "What year was Einstein born?" and "When was Einstein born?", the inconsistency suggests uncertainty or knowledge gaps. Formally:

  1. Generate KK semantically equivalent paraphrases of the question
  2. Obtain the model's answer to each paraphrase
  3. Measure semantic consistency across answers
  4. Treat inconsistency as a proxy for uncertainty

This approach is robust to the specific sampling temperature and doesn't require running the model multiple times on the exact same input (which can be dominated by sampling artifacts from a single phrasing).

The intuition is that a model with stable knowledge should retrieve consistent answers regardless of how a question is phrased. When knowledge is solid, the retrieval process is robust to surface-form variation. When it is shaky, small changes in phrasing tip the model toward different answers, because the model is pattern-matching against surface forms rather than retrieving a stable fact. Sensitivity to question phrasing is a symptom of shallow or inconsistent encoding.

Prompt perturbation is especially useful for detecting factual inconsistency: the model might answer "Einstein was born in 1879" when asked directly, but answer "Einstein was born in Ulm in 1878" when a slightly different phrasing activates a different context window. Neither answer carries explicit uncertainty; they contradict each other. Standard log-probability methods might rate both answers as high-confidence because the model generates each fluently. Only consistency checking reveals the problem.

The practical challenge is generating high-quality paraphrases automatically. Large models can be prompted to paraphrase, but this adds latency and cost. Rule-based perturbations (changing word order, substituting synonyms, adding neutral context) are faster but less semantically faithful. Research on this technique has shown that consistency correlates with factual correctness in the 0.6-0.7 range, which is meaningful but not sufficient as a standalone reliability signal.

Worked Example: Measuring and Correcting Miscalibration

To make these concepts concrete, let's trace through a calibration analysis step by step using a simplified but realistic scenario.

Suppose you have a factual QA system built on a large language model. You have 200 held-out questions with known ground-truth answers, and you've extracted the model's confidence score for each response (here, the length-normalized log-probability converted to a probability via an exponential transform). You want to know: how miscalibrated is this model, and how much does temperature scaling help?

Step 1: Group predictions into bins. Divide the 200 predictions into 10 equal-width confidence bins: [0,0.1)[0, 0.1), [0.1,0.2)[0.1, 0.2), ..., [0.9,1.0][0.9, 1.0]. For each bin, compute the mean confidence and the fraction of predictions in that bin that are correct. If you plot these points (mean confidence on x-axis, fraction correct on y-axis), you get the reliability diagram.

Step 2: Read the diagram. Suppose the result shows that predictions in the [0.9,1.0)[0.9, 1.0) bin have a mean confidence of 0.93 but an actual accuracy of only 0.74. Predictions in the [0.7,0.8)[0.7, 0.8) bin have mean confidence 0.74 but accuracy 0.62. The curve lies below the diagonal throughout the high-confidence range, confirming systematic overconfidence.

Step 3: Compute ECE. Plug the bin statistics into the ECE formula. Suppose the [0.9,1.0)[0.9, 1.0) bin contains 80 of the 200 predictions (40% of the total). That bin contributes 0.40×∣0.93−0.74∣=0.40×0.19=0.0760.40 \times |0.93 - 0.74| = 0.40 \times 0.19 = 0.076 to the ECE. Summing over all bins gives a total ECE of, say, 0.11. This is high: a well-calibrated model should have ECE below 0.05.

Step 4: Learn temperature. Take the model's raw logits (before softmax) for each of the 200 predictions and learn the temperature T∗T^* that minimizes negative log-likelihood on this calibration set. The optimizer finds T∗=1.8T^* = 1.8, meaning the model is nearly twice as overconfident as it should be. Applying this temperature to rescale all confidence scores gives a new set of calibrated probabilities.

Step 5: Re-check calibration. Recompute the reliability diagram and ECE with the calibrated probabilities. The curve now tracks much closer to the diagonal, and the ECE drops to 0.03. The model is now well-calibrated: when it says "70% confident," it's right about 70% of the time.

This is the core calibration loop in practice: extract confidence scores, measure miscalibration via ECE and reliability diagram, apply a correction method, and verify improvement. The implementation section below demonstrates this loop in code.

Code Implementation

Let's build a practical uncertainty quantification system. We'll implement token-level confidence extraction, temperature scaling calibration, and sampling-based uncertainty estimation.

Setup and Data

We start with an implementation that works with any model through the Hugging Face Transformers API. For this example, we'll use a small model and simulate a factual QA scenario to illustrate the key ideas.

In[3]:
Code
import warnings

import numpy as np

warnings.filterwarnings("ignore")

# Reproducibility
np.random.seed(42)

Simulating Model Confidence Scores

In a real deployment, you'd extract logits from a language model. Here we simulate the behavior of a well-documented phenomenon: overconfident LLMs that assign high probabilities even when wrong.

In[4]:
Code
def simulate_lm_confidences(n_samples=500, overconfidence_factor=2.0):
    """
    Simulate an overconfident language model.
    True probabilities are drawn from a beta distribution (realistic spread).
    Observed confidences are a monotone transform that stretches toward 1.
    """
    # True underlying probabilities
    true_probs = np.random.beta(3, 2, n_samples)  # skewed toward higher values

    # Overconfident model: compress toward 1 using a power < 1
    # overconfidence_factor > 1 means logit is scaled up
    logit_true = np.log(true_probs / (1 - true_probs + 1e-9))
    logit_observed = logit_true * overconfidence_factor
    observed_confidences = 1 / (1 + np.exp(-logit_observed))

    # Generate binary outcomes: correct (1) or incorrect (0)
    correct = (np.random.uniform(0, 1, n_samples) < true_probs).astype(int)

    return observed_confidences, correct, true_probs


# Generate data
confidences, correct, true_probs = simulate_lm_confidences(n_samples=1000)
Out[5]:
Console
Mean model confidence: 0.644
Actual accuracy:       0.616
Overconfidence gap:    0.028
Samples:               1000

The gap between mean confidence and actual accuracy quantifies the overconfidence. In this simulation, the model claims an average confidence well above its actual accuracy rate, reproducing what has been observed in empirical studies of large language models.

Calibration Curve and ECE

Let's compute and visualize the calibration curve.

In[6]:
Code
def compute_ece(confidences, correct, n_bins=10):
    """Compute Expected Calibration Error (ECE)."""
    bin_boundaries = np.linspace(0, 1, n_bins + 1)
    ece = 0.0
    n = len(confidences)

    for i in range(n_bins):
        lower, upper = bin_boundaries[i], bin_boundaries[i + 1]
        mask = (confidences >= lower) & (confidences < upper)
        if mask.sum() == 0:
            continue
        bin_conf = confidences[mask].mean()
        bin_acc = correct[mask].mean()
        bin_weight = mask.sum() / n
        ece += bin_weight * abs(bin_conf - bin_acc)

    return ece


ece_before = compute_ece(confidences, correct)
Out[7]:
Console
ECE before calibration: 0.1026

Temperature Scaling

Now we implement temperature scaling to correct the overconfidence. Temperature scaling finds the single scalar T∗T^* that minimizes the negative log-likelihood on the calibration set.

In[8]:
Code
from scipy.optimize import minimize_scalar


def apply_temperature(logits, temperature):
    """Apply temperature scaling to logits and return probabilities."""
    scaled = logits / temperature
    return 1 / (1 + np.exp(-scaled))  # sigmoid for binary


def calibrate_temperature(confidences, correct):
    """
    Find optimal temperature by minimizing negative log-likelihood
    on a calibration set (here we use the full dataset for illustration).
    """
    # Convert confidences back to approximate logits
    eps = 1e-7
    logits = np.log(confidences / (1 - confidences + eps) + eps)

    def nll(temperature):
        if temperature <= 0:
            return 1e9
        probs = apply_temperature(logits, temperature)
        probs = np.clip(probs, eps, 1 - eps)
        return -np.mean(
            correct * np.log(probs) + (1 - correct) * np.log(1 - probs)
        )

    result = minimize_scalar(nll, bounds=(0.1, 10.0), method="bounded")
    return result.x


logits = np.log(confidences / (1 - confidences + 1e-7) + 1e-7)
optimal_temp = calibrate_temperature(confidences, correct)
calibrated_confidences = apply_temperature(logits, optimal_temp)
ece_after = compute_ece(calibrated_confidences, correct)
Out[9]:
Console
Optimal temperature:    2.107
ECE before calibration: 0.1026
ECE after calibration:  0.0531
ECE reduction:          48.3%

The optimal temperature is greater than 1.0, confirming the model was overconfident (temperature above 1 softens probabilities). The ECE drops substantially after calibration.

Sampling-Based Uncertainty

Now let's implement sampling-based uncertainty estimation. We simulate multiple samples from a model on the same question and measure semantic consistency.

In[10]:
Code
from scipy.stats import entropy


def simulate_model_samples(question_difficulty, n_samples=10, noise_level=0.3):
    """
    Simulate model answers for a question of given difficulty.

    difficulty: 0 (easy) to 1 (hard)
    Returns: list of 0/1 answers (0=wrong, 1=correct)
    """
    # Move from a certain correct answer toward a maximally uncertain 50/50 split.
    true_prob = 1.0 - 0.5 * question_difficulty
    # Add sampling noise
    sample_probs = np.clip(
        true_prob + np.random.normal(0, noise_level, n_samples), 0.05, 0.95
    )
    answers = (np.random.uniform(0, 1, n_samples) < sample_probs).astype(int)
    return answers


def sampling_uncertainty(answers):
    """
    Compute uncertainty metrics from binary (0/1) answer samples.
    Returns confidence (agreement fraction) and entropy.
    """
    p_correct = np.mean(answers)
    agreement = max(p_correct, 1 - p_correct)
    # Binary entropy
    h = (
        entropy([p_correct, 1 - p_correct], base=2)
        if 0 < p_correct < 1
        else 0.0
    )
    return agreement, h


# Test across different difficulty levels
difficulties = np.linspace(0, 1, 100)
n_samples = 20

agreements = []
entropies = []
for d in difficulties:
    answers = simulate_model_samples(d, n_samples=n_samples)
    agreement, h = sampling_uncertainty(answers)
    agreements.append(agreement)
    entropies.append(h)
Out[11]:
Console
Easy question (diff=0.1): agreement=0.80, entropy=0.722
Medium question (diff=0.5): agreement=0.80, entropy=0.722
Hard question (diff=0.9): agreement=0.60, entropy=0.971

Self-Consistency Score

Self-consistency computes confidence from majority voting across chain-of-thought samples. The confidence score is simply the fraction of samples that agree with the majority answer.

In[12]:
Code
from collections import Counter


def self_consistency_confidence(answers):
    """
    Given a list of sampled answers (any hashable type),
    return the self-consistency confidence (majority fraction).
    """
    counts = Counter(answers)
    majority_count = counts.most_common(1)[0][1]
    return majority_count / len(answers)


def run_self_consistency_experiment(n_questions=200):
    """
    Simulate self-consistency over questions with varying difficulty.
    Returns a list of dicts with difficulty, sc_confidence, and correctness.
    """
    results = []
    difficulties = np.random.uniform(0, 1, n_questions)

    for diff in difficulties:
        answers = simulate_model_samples(diff, n_samples=10)
        sc_conf = self_consistency_confidence(answers)
        majority_answer = Counter(answers).most_common(1)[0][0]
        results.append(
            {
                "difficulty": diff,
                "sc_confidence": sc_conf,
                "correct": int(majority_answer == 1),
            }
        )

    return results


sc_results = run_self_consistency_experiment(n_questions=500)
sc_confidences = np.array([r["sc_confidence"] for r in sc_results])
sc_correct = np.array([r["correct"] for r in sc_results])
sc_ece = compute_ece(sc_confidences, sc_correct)
Out[13]:
Console
Self-consistency results:
  Mean SC confidence: 0.714
  Overall accuracy:   0.854
  ECE:                0.1396

Semantic Entropy Simulation

Let's implement a simplified version of semantic entropy for open-ended responses. In a real system, the clustering step uses NLI entailment; here we simulate clusters directly.

In[14]:
Code
def compute_semantic_entropy(response_clusters):
    """
    Compute semantic entropy from cluster assignment counts.

    response_clusters: list of cluster IDs (one per sampled response)
    Returns: semantic entropy in nats
    """
    counts = Counter(response_clusters)
    total = len(response_clusters)
    probs = np.array([count / total for count in counts.values()])
    return entropy(probs)  # natural log entropy


def simulate_semantic_responses(question_difficulty, n_samples=10):
    """
    Simulate semantic cluster assignments for an open-ended question.

    For easy questions: most responses fall in the same cluster.
    For hard questions: responses are spread across many clusters.
    """
    # Number of distinct clusters increases with difficulty
    n_clusters_expected = 1 + question_difficulty * 5
    cluster_probs = np.ones(int(np.ceil(n_clusters_expected)))
    # Weight toward first cluster (most common answer) inversely with difficulty
    cluster_probs[0] = (1.0 - question_difficulty) * 10 + 1
    cluster_probs /= cluster_probs.sum()

    clusters = np.random.choice(
        len(cluster_probs), size=n_samples, p=cluster_probs
    )
    return clusters


# Compute semantic entropy across difficulty levels
sem_entropies = []
for d in difficulties:
    clusters = simulate_semantic_responses(d, n_samples=20)
    se = compute_semantic_entropy(clusters)
    sem_entropies.append(se)
Out[15]:
Console
Semantic entropy by difficulty:
  Easy (diff=0.1):   0.325 nats
  Medium (diff=0.5): 0.613 nats
  Hard (diff=0.9):   1.330 nats

Higher semantic entropy at higher difficulty confirms that more diverse (uncertain) answers correspond to harder questions. The overall upward relationship between difficulty and semantic entropy validates semantic entropy as an uncertainty signal that tracks model uncertainty rather than surface variation.

Key Parameters

The key parameters for the uncertainty quantification methods implemented above are:

  • n_bins (ECE): Number of confidence bins for calibration error computation. More bins give finer-grained measurement but require more data for reliable estimates. 10-15 bins is standard.
  • overconfidence_factor (temperature scaling): The logit scaling factor that produces overconfident outputs. Values greater than 1.0 stretch probabilities toward certainty, simulating the systematic bias seen in RLHF-trained models.
  • temperature (calibration): The divisor applied to logits before the softmax. Learned values above 1.0 soften an overconfident model; values below 1.0 sharpen an underconfident one.
  • n_samples (sampling-based methods): Number of model runs per query. More samples produce a more accurate uncertainty estimate but multiply inference cost proportionally.
  • n_questions (self-consistency): Number of questions in the evaluation set. Larger evaluation sets produce more reliable ECE estimates.

Visualizations

Let's visualize the key concepts with figures that make the ideas concrete.

Out[16]:
Visualization
Reliability diagram comparing overconfident and temperature-scaled model calibration against the ideal diagonal.
Reliability diagrams comparing the overconfident model (coral) against the temperature-scaled model (blue) and perfect calibration (dashed diagonal). The overconfident model's curve lies below the diagonal, meaning it claims higher confidence than its accuracy justifies. After temperature scaling, the calibrated model tracks the diagonal much more closely, especially in the high-confidence region where most predictions fall.
Out[17]:
Visualization
Line plot of agreement and entropy vs question difficulty, with agreement decreasing and entropy increasing toward the hardest questions.
Agreement score and binary entropy as a function of question difficulty. High-difficulty questions show lower agreement and higher entropy, confirming that sampling-based methods correctly identify uncertain questions. Entropy peaks for the hardest questions, where sampled answers approach an even split.
Noisy line plot of semantic entropy vs question difficulty, showing a clear upward trend.
Semantic entropy generally increases with question difficulty despite sampling noise. Unlike token-level entropy, semantic entropy is invariant to surface-form variation and directly captures the spread of distinct meaning clusters across sampled responses, making it a cleaner uncertainty signal.
Out[18]:
Visualization
Bar chart comparing ECE and calibration gap for uncalibrated, temperature-scaled, and Platt-scaled models.
ECE and calibration gap across three methods: uncalibrated, temperature-scaled, and Platt-scaled. Both calibration methods substantially reduce ECE compared to the raw overconfident baseline, with the calibration gap (mean confidence minus mean accuracy) also dropping toward zero. Temperature scaling uses one parameter; Platt scaling uses two and can correct asymmetric miscalibration.

Uncertainty Communication

Measuring uncertainty is only half the problem. The other half is communicating it to users in a way that influences behavior. This is where uncertainty quantification meets human factors, interface design, and cognitive science. A technically accurate uncertainty estimate that is poorly communicated might as well not exist: users will either ignore it, misinterpret it, or find it paralyzing.

The Communication Challenge

Humans are poor at reasoning about probabilities in the abstract. If you tell a user that a model response has "72% confidence", most users won't know what to do with that number. Is 72% high? Low? Does it mean they should verify or just proceed? Research on probability communication consistently shows that numerical probabilities are difficult to interpret and often misread. People tend to treat probabilities above 50% as "yes" and below 50% as "no," discarding the continuous information the number contains.

Communication also has asymmetric consequences. If you communicate too much uncertainty, users stop trusting the system even when it's correct, leading to over-verification that wastes time and undermines the value of the AI system. If you communicate too little, users over-rely on incorrect answers, with potentially serious downstream consequences. Getting the calibration of communication right matters for both accuracy and user trust.

There is also a framing effect in how uncertainty is communicated. "There is a 30% chance this is wrong" and "There is a 70% chance this is right" are mathematically identical, but users perceive them very differently. The first frames the response as mostly reliable; the second frames it as risky. Which framing is appropriate depends on the stakes: for a low-stakes recommendation, 70% right is good enough to act on. For a medical diagnosis, 30% wrong is alarming enough to require verification.

Verbal Uncertainty Scales

Standardized verbal uncertainty scales map numeric probabilities to natural language phrases. The IPCC (Intergovernmental Panel on Climate Change) uses a well-known example for communicating scientific findings:

IPCC verbal probability scale (AR6).
Likelihood termProbability
Virtually certain99-100%
Extremely likely95-100%
Very likely90-100%
Likely66-100%
About as likely as not33-66%
Unlikely0-33%
Very unlikely0-10%
Extremely unlikely0-5%

The key insight is that agreed-upon, consistent mappings are more communicable than ad-hoc hedges. If your system always uses "I'm fairly confident" to mean 70-80%, users can calibrate to that phrase over time. Inconsistent hedging teaches users nothing and prevents them from building accurate mental models of the system's reliability.

Verbal scales have well-known limitations. Different cultures and individuals bring different prior interpretations to words like "likely" or "probable." Research has found that "probable" is interpreted as meaning anywhere from 55% to 85% across different populations. This spread means that verbal scales introduce noise even when the underlying probability estimate is accurate. The IPCC addresses this by publishing explicit numeric mappings and referring readers to them, a practice that AI systems serving technical users might adopt.

Interface Design for Uncertainty

Beyond verbal language, interfaces can communicate uncertainty through visual design:

Confidence indicators. A visual bar, color coding (green/yellow/red), or star rating system adjacent to a response provides a quick uncertainty signal without interrupting the flow of text. Color is the most immediately readable but requires careful design for colorblind users.

Expandable uncertainty details. "Click to see why I'm unsure" keeps the primary response clean while giving power users access to detailed uncertainty information such as which specific claims are uncertain or what sources are available.

Source display. Showing which sources support a claim implicitly communicates confidence: more high-quality sources means higher confidence. A response backed by three authoritative citations reads as more reliable than the same content without citations, even if the user doesn't check the citations.

Differentiated formatting. Some systems use italic or muted text for uncertain claims within a response, drawing attention to which specific statements are shaky. This is particularly useful for responses that mix high-confidence background information with uncertain specific claims.

Explicit uncertainty sections. For longer responses, a dedicated "What I'm not sure about" paragraph at the end surfaces all uncertainties together. This is less granular than inline markers but more readable for users who want a high-level reliability assessment before reading the full response.

Hedging Phrases and Epistemic Markers

Natural language contains a rich vocabulary of epistemic markers: words and phrases that signal a speaker's confidence level. In English:

High confidence markers include: "is," "definitely," "certainly," "I'm sure that," "it's clear that."

Medium confidence markers include: "appears to be," "seems likely," "I believe," "probably," "generally speaking."

Low confidence markers include: "might be," "could be," "I'm not certain but," "I may be wrong," "possibly."

Uncertainty markers often include: "I don't know," "I'm not sure," "this is outside my knowledge," "you should verify this."

Well-designed LLM systems calibrate their use of these markers to the actual confidence level of the response. Miscalibrated systems use high-confidence language for uncertain responses, or hedge unnecessarily on confident answers, both of which erode trust over time.

The challenge is that epistemic markers in natural language are not universally interpreted. Different readers bring different calibration priors. "Probably" might mean 60% to one reader and 90% to another. This is precisely the problem that motivated the development of standardized verbal probability scales: by agreeing in advance that "likely" maps to 66-100%, a consistent communication channel is established.

Research in psycholinguistics shows that epistemic markers also interact with domain context. The phrase "I believe" in an academic paper signals a stronger degree of confidence than the same phrase in casual conversation. An AI system serving multiple contexts (customer support, scientific writing, medical triage) would ideally calibrate its hedging vocabulary to the norms of each domain. A medical AI saying "it is probable that this is benign" carries different weight than a shopping assistant saying the same phrase about a product recommendation.

Another subtlety is the difference between hedging about facts and hedging about opinions. "I think this policy is misguided" is an opinion hedge, not a factual uncertainty marker. Well-designed systems distinguish between "I am uncertain whether X is true" and "X is a matter on which reasonable people disagree." Conflating the two misrepresents the nature of the uncertainty and can mislead users about whether verification is useful.

Actionability of Uncertainty

Useful uncertainty communication tells users what to do, not just how uncertain the model is. Raw uncertainty numbers do not guide action; uncertainty combined with recourse does. This means designing uncertainty communication around four principles:

Specificity. "I'm not sure about this" is less useful than "I'm not sure about the exact year; the range is probably 1880-1900." Specific uncertainty helps users decide exactly what to check.

Recourse. "You should verify this with a current source" is more actionable than "I might be wrong." Recourse tells users what action to take, not just that action is needed.

Differentiation. "This fact is contested among historians" helps the user more than a blanket confidence percentage because it explains the nature of the uncertainty and suggests who to consult.

Scope. "I'm confident about the general trend but not the specific numbers" lets the user act on the high-confidence part while verifying the uncertain part. Partial confidence is still useful if it's communicated with scope.

The design goal is trust calibration: users should trust the model's outputs to the degree they deserve trust, neither more nor less. Achieving this requires both accurate uncertainty quantification (measuring the right thing) and effective uncertainty communication (conveying it clearly).

Uncertainty in Downstream Tasks

How uncertainty should be handled depends heavily on the downstream task and the stakes involved:

High-stakes decisions. In medical, legal, or financial contexts, any uncertainty above a threshold should trigger human review. The uncertainty score becomes a routing signal: below threshold, deliver the answer; above threshold, escalate to a human reviewer. This architectural pattern converts the uncertainty estimate from a display concern into an operational routing decision.

Retrieval augmentation. Low confidence can trigger a retrieval step. The model's uncertainty about a factual question prompts a search, and the retrieved context enables a more confident and accurate answer. Retrieval-augmented generation (RAG) systems often use this pattern: generate first, measure confidence, retrieve and regenerate if confidence is low.

Interactive clarification. In dialogue systems, uncertainty can trigger a clarifying question: "I'm not sure which John Smith you mean. Can you give me more context?" This is more useful than delivering an uncertain answer because it involves the user in resolving the ambiguity rather than leaving them to discover the error themselves.

Ensemble outputs. When uncertainty is high, present multiple candidate answers with their confidence levels rather than picking one. This is common in information retrieval and recommendation systems where giving users a ranked list is more appropriate than committing to a single recommendation.

The common thread across all these patterns is that uncertainty is treated as an actionable signal, more than a display quantity. Systems that use uncertainty to route, retrieve, or clarify achieve better end-to-end accuracy than systems that generate a single confident answer and hope for the best.

Limitations and Impact

Limitations of Uncertainty Quantification in LLMs

Uncertainty quantification methods face several hard challenges when applied to large language models, and being clear about these limitations is important for responsible deployment.

The confidence-accuracy decoupling problem. In classification models, confidence scores have a clear reference: the probability assigned to the predicted class. For LLMs, the mapping between internal representations of confidence and actual output reliability is loose. The model's "confidence" as measured by token probabilities may not correlate well with factual correctness, because the generation process is optimized for fluency and coherence rather than accuracy. A model can fluently generate false statements with high token probability. Semantic entropy and sampling-based methods partially address this by measuring behavioral uncertainty rather than internal representations, but they are proxies, not direct measures of factual accuracy.

Scaling and cost. Sampling-based methods require running the model NN times per query. At inference time, this multiplies compute costs by a factor of NN. For large models serving millions of requests, this cost may be prohibitive. Temperature scaling is cheap (it modifies only the output probabilities) but requires a labeled calibration set and only addresses single-number confidence, not semantic uncertainty. The trade-off between uncertainty quality and inference cost is a central engineering constraint in production systems.

Coverage of uncertainty types. None of the methods discussed here cleanly separates epistemic uncertainty (the model doesn't know) from aleatoric uncertainty (the world is inherently uncertain). This matters because the appropriate response differs: retrieve more information for epistemic uncertainty, but communicate irreducible uncertainty for aleatoric cases. Conflating the two can lead to retrieval attempts on questions that are inherently unanswerable, or to failing to retrieve information that would resolve a gap in the model's knowledge.

Distribution shift. Calibration is always relative to a distribution. A model calibrated on general knowledge QA may be poorly calibrated on specialized domain questions or on questions about recent events after its knowledge cutoff. Post-hoc calibration methods like temperature scaling can only correct for the distribution shift they are evaluated on. A temperature learned on a Wikipedia QA benchmark may be completely wrong for medical or legal questions. This argues for domain-specific calibration datasets, but those are expensive to construct and may quickly become outdated.

Verbal calibration training. Teaching models to verbalize uncertainty accurately is difficult because it requires labeled data about model errors, and models may learn surface patterns (always hedge on certain question types) rather than accurate uncertainty awareness. Research has shown that verbalized uncertainty often correlates more with question topic than with actual model knowledge. A model may hedge consistently on questions about recent events (because it's been trained to say it doesn't know about things after its cutoff) while confidently asserting incorrect facts about historical topics.

The verbalization-internalization gap. Even when a model is trained to verbalize uncertainty accurately, there is no guarantee that the verbalized uncertainty reflects an internal uncertainty state that the model acts on. A model might say "I'm not sure about this, but..." and then proceed to generate the rest of the response as if it were completely sure. The uncertainty expression becomes a prefix that satisfies the calibration training objective without changing the generation process that follows.

Impact and Applications

Despite these limitations, uncertainty quantification has had significant practical impact in making language AI systems more reliable and trustworthy.

Automated verification pipelines. Systems that flag high-uncertainty responses for verification, rather than delivering them directly to users, have been shown to reduce factual errors in production. The uncertainty signal is used as a routing criterion, not as a confidence label delivered to end users. This architectural pattern separates "measuring uncertainty" from "communicating uncertainty," allowing different strategies for each.

RAG triggering. Using uncertainty as a signal to trigger retrieval augmented generation is now a standard design pattern. When the model's semantic entropy on a question is high, the system retrieves relevant documents and conditions the answer on them. This hybrid approach achieves better accuracy than pure generation or pure retrieval alone, because retrieval is applied selectively where it provides the most value.

Scientific and medical applications. In high-stakes domains, uncertainty communication is not optional, it is required. Medical AI systems must communicate what they don't know. Scientific text generation tools must distinguish claims with different levels of evidence. These use cases have driven methodological advances in verbal uncertainty expression and have pushed the field toward more rigorous evaluation of calibration in specialized domains.

Model evaluation. Calibration metrics such as ECE and reliability diagrams have become standard components of LLM evaluation benchmarks. The TruthfulQA benchmark evaluates accuracy and whether models confidently assert false things, measuring the conjunction of truthfulness and calibration. This evaluation culture has created incentives to improve calibration during training rather than treating it as a post-hoc afterthought.

Training improvements. Research has begun to integrate calibration objectives directly into pretraining and fine-tuning. Rather than post-hoc calibration, these approaches train models to be well-calibrated from the start by including calibration loss terms alongside the standard language modeling objective. Early results suggest that models trained with calibration objectives show better ECE without significant accuracy degradation, though this research is still young.

The field continues to evolve toward tighter integration between uncertainty quantification and model training. We will explore the broader challenges of safety and reliability in the next part of this handbook, which covers safety risks, red teaming, and guardrails that build on the foundation of factuality and uncertainty we've developed here.

Summary

Uncertainty quantification gives language models a way to signal the limits of their knowledge, and gives users a way to know when to trust model outputs. The key ideas from this chapter:

  • Calibration measures whether confidence scores align with actual accuracy. A perfectly calibrated model that says "80% confident" is right 80% of the time. Most LLMs are systematically overconfident due to training objectives, RLHF, and temperature effects during inference.

  • Temperature scaling is the simplest and most effective post-hoc calibration method. A single learned temperature parameter scales logits to reduce overconfidence. Expected Calibration Error (ECE) quantifies how well-calibrated a model is by measuring the weighted average gap between confidence and accuracy across bins.

  • Verbalized uncertainty lets models express confidence in natural language. This is more user-friendly than numeric scores but risks sycophantic hedging, where models learn to hedge for stylistic reasons rather than actual uncertainty. You cannot rely on verbal hedges as direct evidence of model uncertainty without also verifying through behavioral methods.

  • Sampling-based methods estimate uncertainty by running the model multiple times and measuring consistency. Self-consistency uses majority voting; semantic entropy groups responses by meaning and computes entropy over semantic clusters. High semantic entropy indicates substantial uncertainty: the model produces diverse and semantically inconsistent answers when it doesn't know the right one.

  • Prompt perturbation probes consistency by paraphrasing the same question. A model with solid knowledge answers consistently regardless of phrasing. A model with shallow or inconsistent knowledge is sensitive to surface-form variation, revealing uncertainty that token probabilities might miss.

  • Uncertainty communication is as important as uncertainty measurement. Verbal scales, visual indicators, and actionable recourse design help users calibrate their trust appropriately. The goal is to neither over-trust nor under-trust model outputs, and to give users enough information to take the right action when the model is uncertain.

Together, these techniques form a practical toolkit for building language AI systems that are honest about the boundaries of their knowledge, a prerequisite for responsible deployment in any high-stakes context.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about uncertainty quantification in language models.

Uncertainty Quantification Quiz

Question 1 of 80 of 8 completed
What does it mean for a language model to be 'well-calibrated'?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026uncertaintyquantification, author = {Michael Brenndoerfer}, title = {Uncertainty Quantification in Language Models}, year = {2026}, url = {https://mbrenndoerfer.com/writing/uncertainty-quantification}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Uncertainty Quantification in Language Models. Retrieved from https://mbrenndoerfer.com/writing/uncertainty-quantification
MLAAcademic
Michael Brenndoerfer. "Uncertainty Quantification in Language Models." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/uncertainty-quantification>.
CHICAGOAcademic
Michael Brenndoerfer. "Uncertainty Quantification in Language Models." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/uncertainty-quantification.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Uncertainty Quantification in Language Models'. Available at: https://mbrenndoerfer.com/writing/uncertainty-quantification (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Uncertainty Quantification in Language Models. https://mbrenndoerfer.com/writing/uncertainty-quantification

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.