Calibration in Machine Learning: Confidence, Accuracy & ECE

Michael BrenndoerferFebruary 26, 202655 min read

Part of Language AI Handbook

Covers calibration in machine learning and how model confidence aligns with actual accuracy. Examines ECE, temperature scaling.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Calibration

When you ask a language model a question, it returns an answer and a probability. That probability represents the model's confidence: "I am 90% sure the answer is Paris." The practical question is whether you should believe it. If the model says it is 90% confident one hundred times, should it be correct about ninety of those answers? Calibration is the study of this alignment between stated confidence and observed accuracy. Understanding calibration matters because it determines whether we can trust a model's uncertainty estimates when making decisions. In high-stakes applications such as medical diagnosis, legal advice, or autonomous systems, a model's confidence score directly influences how humans act upon its predictions. If the confidence is misleading, you may either disregard correct predictions due to apparent uncertainty or trust erroneous predictions because of false certainty.

A perfectly calibrated model is one where the confidence scores match the empirical probabilities of correctness. If a weather forecaster predicts a 70% chance of rain every day for a month, and it rains on exactly 70% of those days, that forecaster is well-calibrated. Similarly, if a language model assigns 0.8 probability to its top prediction, and that prediction is correct exactly 80% of the time across many instances, the model is calibrated at that confidence level. This concept extends beyond individual predictions to the entire distribution of the model's outputs. Consistency across the full range of confidence values ensures that the probabilities generated by the model carry meaningful semantic content about uncertainty. A model that is calibrated at every confidence level simultaneously possesses what researchers call "global calibration," meaning there is no systematic bias in how it reports uncertainty anywhere on the probability scale.

Calibration differs from accuracy. A model can achieve high accuracy while being poorly calibrated. It might be correct most of the time but systematically overconfident, declaring 99% certainty even when it is wrong half the time. Conversely, a model can be perfectly calibrated but have low accuracy if it consistently expresses low confidence. As we discussed in Cross-Entropy Loss, the loss function encourages the model to assign high probability to correct answers, but it does not strictly enforce that these probabilities reflect true likelihoods. The training objective is discriminative rather than probabilistic: it rewards the model for ranking the correct answer highest, not for estimating how often it will be right. Modern large language models, despite their impressive performance on benchmarks we will explore in later chapters, often exhibit significant miscalibration, particularly becoming overconfident as scale increases.

To understand why calibration matters practically, consider the difference between a model used as an oracle versus one used as a collaborator. When a model acts as an oracle, you either accept or reject its outputs wholesale. But in most real-world deployments, models serve as collaborators where the confidence score determines how much human oversight is applied. A doctor using an AI diagnostic system might review every prediction where confidence falls below 0.85, and accept those above without further scrutiny. If that threshold was calibrated on a well-calibrated model, this policy works. If the model is overconfident, the doctor is inadvertently accepting many incorrect diagnoses because the model's stated 87% confidence corresponds to only 60% accuracy in practice. The downstream harm of this miscalibration is not just statistical: it translates directly into misallocated human attention and worse patient outcomes.

The Calibration Problem

To understand calibration concretely, consider a binary classification scenario. Let Y^\hat{Y} be the predicted label and P^\hat{P} be the predicted probability (confidence) for that label. A model is perfectly calibrated if, for any confidence value p∈[0,1]p \in [0, 1]:

P(Y^=Y∣P^=p)=pP(\hat{Y} = Y \mid \hat{P} = p) = p

where:

  • Y^\hat{Y}: the predicted label
  • YY: the true label
  • P^\hat{P}: the predicted probability (confidence) for the predicted label
  • pp: a specific confidence value in [0,1][0, 1]

This definition captures the intuitive notion that predictions carrying a specific confidence level should materialize with exactly that frequency. Mathematically, it states that the conditional probability of correctness given a confidence score must equal the confidence score itself. This means that among all predictions where the model claims probability pp, the fraction that are correct should equal pp. If the model outputs confidence 0.9 for a thousand different predictions, exactly nine hundred of them should be correct.

The definition looks simple, but achieving it during training is difficult. Notice that the formula says nothing about individual predictions: it is a statement about the statistical behavior of the model over many predictions at each confidence level. A single prediction with confidence 0.9 can be either right or wrong without violating calibration. What matters is the long-run empirical frequency. This is fundamentally a frequentist notion of probability: the model's stated probabilities must match empirical frequencies when aggregated across many predictions with similar confidence values.

Multiclass Calibration

In practice, we work with multiclass classification where language models output a distribution over the vocabulary. For a generated token or classification label, we typically consider the maximum probability (top-1 confidence) or the probability assigned to the specific target class. The extension from binary to multiclass settings requires careful consideration of which probability mass to evaluate.

The three most common formulations of multiclass calibration are:

  • Top-1 (confidence) calibration: The probability assigned to the predicted class (the argmax class) should match how often that prediction is correct. This is the most natural extension of binary calibration and the one most commonly measured in practice.
  • Classwise calibration: Each class's probability distribution should be calibrated independently. If a model assigns probability 0.7 to class "cat" across many predictions, cat should appear 70% of the time among those predictions. This is a stronger requirement than top-1 calibration because it demands calibration for every class simultaneously.
  • Canonical calibration: The full predicted probability vector should satisfy a multi-dimensional calibration condition. This is the strongest requirement and the hardest to achieve or measure.

For most practical purposes, top-1 calibration is what practitioners care about because it governs decision thresholds for acceptance and rejection. When you configure an automated system to "act on predictions above 0.8 confidence," you need top-1 calibration to be accurate. If you care about the full distribution for reasons such as ranking multiple candidates or computing entropy-based uncertainty, classwise calibration becomes more important.

Confidence vs. Probability

In deep learning, the terms are often used interchangeably, but technically:

  • Probability: The raw softmax output for a class
  • Confidence: Often refers to the maximum probability (top-1) or the probability assigned to the predicted class

For calibration, we care about whether these probability values are statistically meaningful. A raw softmax output of 0.9 only matters if the model is correct 90% of the time when it outputs that value.

Why Neural Networks Are Poorly Calibrated

The modern deep learning training regime systematically produces overconfident models. This might seem counterintuitive: training on cross-entropy loss directly involves probabilities, so why would the learned probabilities be wrong?

The answer lies in the asymmetry of the cross-entropy gradient. When the model assigns probability pp to the correct class, the gradient of cross-entropy with respect to the logit for that class is −(1−p)-(1 - p). This gradient is large when pp is small (the model is uncertain) and shrinks toward zero as pp approaches 1.0. The gradient is zero at perfect confidence. This means the training signal always pushes the model toward higher confidence in correct answers, and the signal becomes weaker as confidence increases. The optimization naturally gravitates toward high-confidence predictions because that is where the loss approaches zero. There is no corresponding force that prevents confidence from being calibrated; the loss only cares about discrimination, not about whether the confidence levels are meaningful.

Early neural network literature noted calibration as a property that networks should naturally exhibit, but empirical observations eventually contradicted this optimism. Guo et al. (2017) famously documented that modern deep neural networks, despite being far more accurate than their predecessors, are significantly less calibrated. They found that a ResNet trained on CIFAR-100 achieved 70% accuracy with only modest calibration error in 1990s-era networks but severe overconfidence in modern deep residual networks. The culprit was a combination of increased model capacity, batch normalization, and weight decay applied without calibration-awareness. Each advance that improved accuracy tended to worsen calibration.

For large language models, the calibration picture is more complex because the models produce probabilities over thousands of vocabulary tokens at each position, and the relationship between token-level probabilities and downstream task accuracy is indirect. A model's confidence in its next token does not directly correspond to its confidence in a complete multi-step answer. This gap between token-level calibration and answer-level calibration is an active area of research and has implications for how we should interpret language model confidence estimates in practice.

Measuring Calibration

Since we cannot verify the exact equality P(Y^=Y∣P^=p)=pP(\hat{Y} = Y \mid \hat{P} = p) = p for continuous pp, we measure calibration error by grouping predictions into bins based on their confidence scores. This discretization allows us to empirically estimate the relationship between predicted confidence and observed accuracy. Without binning, we would rarely see multiple predictions with exactly the same confidence value. This makes statistical estimation impossible. The binning strategy turns the continuous calibration problem into a discrete estimation task where we can compare average confidence against empirical accuracy within each bin.

The fundamental trade-off in binning is resolution versus reliability. With many fine-grained bins, you can detect localized calibration errors, but each bin contains few samples, making the accuracy estimate noisy. With few coarse bins, each estimate is more stable but may average over meaningful within-bin variation. There is no universally optimal number of bins; the choice depends on the total number of predictions available and how finely you need to characterize the calibration curve.

Expected Calibration Error (ECE)

The Expected Calibration Error (ECE) is the standard metric for quantifying miscalibration. It partitions predictions into MM equally-spaced bins and computes the weighted average of the absolute difference between accuracy and confidence in each bin. This metric provides a single scalar value representing the overall calibration quality of the model, where lower values indicate better alignment between confidence and accuracy.

For bin BmB_m containing nmn_m predictions, we compute the empirical accuracy:

acc(Bm)=1nm∑i∈Bm1(y^i=yi)\text{acc}(B_m) = \frac{1}{n_m} \sum_{i \in B_m} \mathbb{1}(\hat{y}_i = y_i)

where:

  • BmB_m: the set of predictions falling into bin mm
  • nmn_m: the number of predictions in bin mm
  • y^i\hat{y}_i: the predicted label for the ii-th sample
  • yiy_i: the true label for the ii-th sample
  • 1(⋅)\mathbb{1}(\cdot): the indicator function, equal to 1 when the condition is true and 0 otherwise

We also compute the average confidence within the bin:

conf(Bm)=1nm∑i∈Bmp^i\text{conf}(B_m) = \frac{1}{n_m} \sum_{i \in B_m} \hat{p}_i

where:

  • p^i\hat{p}_i: the predicted confidence (probability) for the ii-th sample

The ECE is then the weighted sum of absolute differences across all bins:

ECE=∑m=1Mnmn∣acc(Bm)−conf(Bm)∣\text{ECE} = \sum_{m=1}^{M} \frac{n_m}{n} \left| \text{acc}(B_m) - \text{conf}(B_m) \right|

where:

  • MM: the total number of bins
  • nn: the total number of samples across all bins
  • nmn\frac{n_m}{n}: the proportion of samples in bin mm, serving as the weight for that bin's contribution

The weighting by nmn\frac{n_m}{n} is critical and reflects a sensible design choice: calibration errors in densely populated bins matter more than errors in sparse ones. If a model is slightly off in a bin containing 2% of predictions, that matters far less than being off in a bin containing 40% of predictions. An ECE of 0 indicates perfect calibration; higher values indicate worse calibration. The absolute value ensures that overconfidence and underconfidence both count as errors rather than canceling out.

The choice of MM matters. Too few bins lose resolution and may mask local calibration errors by averaging over wide confidence ranges. Too many bins create noisy estimates with few samples per bin, leading to high variance in the accuracy estimates. Common choices are 10 or 15 bins, balancing granularity with statistical reliability. Research comparing different binning choices finds that the optimal number depends on dataset size: for datasets with thousands of predictions, 15 bins works well; for smaller datasets, 5 to 10 bins provide more stable estimates.

One subtle issue with ECE is that it measures the expected calibration error for the top-1 confidence. In multiclass settings with many classes, the model could be perfectly calibrated at the top-1 level while being systematically wrong about the probability mass assigned to non-top classes. If you care about the full distribution, such as when using the model's outputs as inputs to another system, you would need to measure classwise calibration separately for each class.

Maximum Calibration Error (MCE)

While ECE captures average calibration across all confidence levels, the Maximum Calibration Error (MCE) identifies the worst-case deviation:

MCE=max⁡m∈{1,…,M}∣acc(Bm)−conf(Bm)∣\text{MCE} = \max_{m \in \{1, \ldots, M\}} \left| \text{acc}(B_m) - \text{conf}(B_m) \right|

where:

  • max⁡\max: the maximum operator, selecting the largest absolute deviation between accuracy and confidence across all bins
  • mm: the bin index, ranging from 11 to MM
  • MM: the total number of bins (as defined previously for ECE)
  • acc(Bm)\text{acc}(B_m): the accuracy of predictions in bin mm (fraction correct)
  • conf(Bm)\text{conf}(B_m): the average confidence of predictions in bin mm

MCE is particularly important in safety-critical applications where even occasional overconfidence can be dangerous. While a model might exhibit low average calibration error, a specific confidence range where the model is severely overconfident could lead to catastrophic failures. For instance, if a medical diagnosis model is perfectly calibrated at low and medium confidence but severely overconfident when predicting above 95% probability, clinicians might be misled into trusting incorrect high-confidence diagnoses. In such a system, having a small average ECE provides false reassurance: what matters is that the high-stakes high-confidence region is calibrated, even if it contains relatively few predictions overall.

The trade-off between ECE and MCE reflects a choice between optimizing for average behavior versus worst-case behavior. In most evaluation contexts, ECE is reported because it summarizes overall calibration quality. When deploying models in high-stakes settings, MCE deserves more attention because it reveals the most dangerous miscalibration that could occur.

Adaptive Binning (Equal-Mass Binning)

Equal-width binning can leave some bins empty or sparsely populated if the model's confidence distribution is skewed. Many neural networks tend toward extreme confidence values, clustering predictions near 0 and 1, leaving middle ranges empty. Adaptive binning, also called equal-mass binning, ensures each bin contains approximately the same number of samples by adjusting bin boundaries based on the empirical distribution of confidence scores. This approach prevents empty bins and ensures that each accuracy estimate is computed over a sufficient sample size.

The procedure for adaptive binning is straightforward: sort all predictions by their confidence scores, then divide them into MM groups of equal size. The boundaries between bins are the confidence values at the k/Mk/M quantiles for k=1,…,M−1k = 1, \ldots, M-1. Every bin thus contains exactly ⌊n/M⌋\lfloor n/M \rfloor or ⌈n/M⌉\lceil n/M \rceil predictions. This ensures that the accuracy estimate in each bin is equally statistically reliable, a property that equal-width binning lacks when the confidence distribution is non-uniform.

The trade-off is interpretability. With equal-width binning, the bin for "predictions between 0.7 and 0.8" has an obvious semantic meaning. With adaptive binning, the boundaries are data-dependent and harder to communicate. For model comparison, adaptive binning is often more reliable; for explaining calibration behavior to non-technical stakeholders, equal-width binning is easier to explain. This often provides more reliable ECE estimates, especially for models that are rarely uncertain or rarely highly confident.

The Reliability Diagram as a Diagnostic Tool

Beyond computing a single ECE number, it is informative to examine the full calibration curve through a reliability diagram. The ECE aggregates calibration errors into a single number, necessarily losing information about where in the confidence spectrum the errors occur. Two models with identical ECE values can have very different calibration profiles: one might be overconfident at medium confidence levels but perfectly calibrated at high confidence, while another shows the opposite pattern. Understanding the shape of the miscalibration helps choose an appropriate correction method.

When examining reliability diagrams, look for these patterns:

  • Uniform overconfidence: The curve sits below the diagonal at all confidence levels. This is the most common pattern in neural networks and responds well to temperature scaling.
  • Non-uniform overconfidence: The curve crosses the diagonal, being overconfident in some ranges and underconfident in others. This pattern requires a non-parametric correction like isotonic regression.
  • Confidence collapse: Most predictions cluster near 0.5, indicating the model is uncertain about nearly everything. This can happen with under-trained models or models facing extreme out-of-distribution inputs.
  • Extreme overconfidence: Predictions cluster near 1.0 regardless of true accuracy. This is common in models trained with very large datasets where the model has memorized training examples and generalizes that certainty to test data inappropriately.

Calibration Plots (Reliability Diagrams)

A calibration plot, or reliability diagram, visualizes the relationship between predicted confidence and actual accuracy. This graphical representation makes it easy to identify systematic patterns of overconfidence or underconfidence across the confidence spectrum. To construct one:

  1. Divide the confidence range [0,1][0, 1] into MM bins
  2. For each bin, plot the mean confidence on the x-axis and the empirical accuracy on the y-axis
  3. Draw a diagonal reference line representing perfect calibration

Points falling below the diagonal indicate overconfidence (the model is less accurate than its confidence suggests). Points above the diagonal indicate underconfidence. The gap between the calibration curve and the diagonal line, weighted by bin size, corresponds to the ECE.

Modern reliability diagrams often include histograms showing the distribution of predictions across confidence values, revealing whether the model tends toward extreme confidences (common in overconfident models) or expresses more moderate uncertainty. A histogram showing most mass near 1.0 suggests an overconfident model that rarely admits uncertainty, while a uniform distribution indicates the model uses the full range of confidence values.

Out[3]:
Visualization
Line plot showing perfect calibration diagonal.
Perfect calibration: accuracy equals confidence at all levels, shown by the diagonal line where predicted probabilities match empirical frequencies exactly.
Line plot showing overconfidence curve below diagonal with shaded gap.
Overconfident model: accuracy consistently falls below the confidence level (typical of neural networks), with the red shaded gap indicating the magnitude of overconfidence.
Line plot showing underconfidence curve above diagonal with shaded gap.
Underconfident model: accuracy exceeds confidence levels, showing the model is more accurate than its predicted probabilities suggest, with the green shaded gap representing underconfidence.

Sources of Miscalibration

Why do neural networks, including large language models, become miscalibrated? Several factors contribute, often interacting in complex ways during training and deployment. Understanding the sources helps you anticipate when calibration problems are likely and which remedies will be most effective.

Overfitting on the training objective. The cross-entropy loss (which we covered in Cross-Entropy Loss) optimizes for discrimination ability, not probability calibration. A model can minimize loss by being correct with high probability without learning the true likelihood of correctness. The training objective only penalizes the model for assigning low probability to the correct answer, but does not penalize assigning excessively high probability when the model is correct. This asymmetry encourages the model to push probabilities toward 1.0 whenever possible, regardless of the actual uncertainty. As a result, models that fit the training data well tend to be overconfident because the optimization pressure always pushes confidence upward on correct examples.

Model capacity and overconfidence. As models grow larger, they tend to become more overconfident. The increased capacity allows them to fit the training data more aggressively, pushing softmax probabilities toward 0 and 1 even when the true uncertainty remains high. Large models can memorize training examples or fit spurious correlations, leading to unjustified certainty in their predictions. This phenomenon appears paradoxical because larger models often generalize better in terms of accuracy, yet their confidence calibration degrades. The resolution is that generalization in terms of getting the right answer and generalization in terms of knowing how often you will get the right answer are different capabilities, and scale improvements thus far have primarily benefited accuracy over calibration.

Label smoothing effects. During training, label smoothing distributes some probability mass to incorrect classes, which can improve calibration by preventing the model from becoming too confident. A hard target of 1.0 for the correct class creates infinite gradient pressure to push confidence toward 1.0 indefinitely. A smoothed target of, say, 0.9 provides a finite ceiling: once the model reaches that confidence level on training data, the gradient approaches zero. This prevents the logits from growing without bound and forces the model to maintain some uncertainty. However, finding the optimal smoothing parameter requires careful tuning, and using too much smoothing can hurt both accuracy and the interpretability of confidence values.

Batch normalization interactions. Batch normalization, which rescales intermediate activations to have unit variance, can interact with calibration in unexpected ways. By controlling the scale of activations throughout the network, batch normalization changes how sensitive the final logit scale is to individual training examples. This can make the temperature of the softmax output layer effectively arbitrary: the logit values represent relative ranking more than absolute magnitude, meaning the raw softmax outputs may be systematically over- or under-confident in ways that are difficult to predict without post-hoc calibration.

Distribution shift. Models are often poorly calibrated on out-of-distribution data or domains different from their training distribution, even when they maintain reasonable accuracy. When the input distribution changes, the relationship between model confidence and true probability often breaks down. A model trained on news articles may be overconfident when processing medical texts, even if it sometimes produces correct answers. This compound challenge of maintaining both accuracy and calibration under distribution shift remains an active research area. Calibration under distribution shift is significantly harder than calibration within the training distribution because post-hoc calibration methods are typically learned on validation data from the same distribution as the training data.

RLHF and fine-tuning effects. Language models fine-tuned with reinforcement learning from human feedback (RLHF) can develop unusual calibration properties because the training signal rewards producing outputs that humans rate highly, not necessarily outputs where stated confidence matches accuracy. When a model has been trained to sound confident and authoritative because humans prefer assertive-sounding responses, the resulting confidence scores may be systematically inflated across all outputs regardless of actual accuracy. This creates a misalignment between the communicative function of confidence (reassuring the user) and the statistical function of calibration (informing accurate decision-making).

Calibration Methods

When a model is poorly calibrated, we can apply post-hoc calibration methods that adjust the output probabilities without retraining the model or changing its architecture. These methods treat the trained model as a fixed feature extractor and learn a transformation that maps uncalibrated probabilities to calibrated ones. This approach is computationally efficient because it requires only a small validation set to learn the calibration parameters, avoiding the expense of retraining the entire model.

The standard workflow for post-hoc calibration is:

  1. Train the model normally on the training set
  2. Collect predictions (logits or probabilities) on a held-out calibration set
  3. Learn the calibration mapping on the calibration set using maximum likelihood or mean squared error
  4. Apply the learned mapping to all future predictions

The calibration set should be large enough to reliably estimate calibration parameters and diverse enough to cover the confidence range you care about. Using the validation set for both hyperparameter selection and calibration is a common practice, but it risks slight overfitting to the validation distribution. For the best practice, maintain a separate calibration set that was never used for any other purpose during model development.

Temperature Scaling

Temperature scaling is the simplest and most widely used calibration method for neural networks. It works by dividing the logits (pre-softmax scores) by a scalar temperature parameter T>0T > 0 before applying the softmax function. This scaling operation softens or sharpens the probability distribution without changing the relative ordering of the logits, meaning the model's predictions (which class it chooses) remain unchanged while only the confidence values are adjusted.

Given logits z\mathbf{z} for each class, the calibrated probability for class ii becomes:

pi=exp⁡(zi/T)∑jexp⁡(zj/T)p_i = \frac{\exp(z_i / T)}{\sum_{j} \exp(z_j / T)}

where:

  • pip_i: the calibrated probability for class ii
  • ziz_i: the logit (pre-softmax score) for class ii
  • TT: the temperature parameter (T>0T > 0)
  • jj: index iterating over all classes

When T=1T = 1, we recover the original model. When T>1T > 1, the probability distribution becomes softer, meaning more uniform, reducing confidence. When T<1T < 1, the distribution becomes sharper, increasing confidence. Geometrically, increasing the temperature pulls the probabilities toward the uniform distribution 1/K1/K for KK classes, while decreasing it pushes them toward one-hot vectors.

The intuition behind temperature scaling connects to the physics of thermodynamics, where the term "temperature" originates. In a physical system, higher temperature means more random thermal fluctuations, leading to more uniform distributions over states. Lower temperature means the system is more likely to settle into low-energy states, creating a peaked distribution. The Boltzmann distribution pi∝exp⁡(−Ei/T)p_i \propto \exp(-E_i / T) is exactly the softmax applied to negative energies scaled by temperature. Neural network logits play the role of negative energies, and temperature scaling applies the same principle to probability distributions over classes.

The optimal temperature TT is learned on a validation set by minimizing the negative log-likelihood (NLL). This optimization is univariate (only one parameter to fit), has a unique minimum under mild conditions, and can be solved efficiently with any scalar optimization method. Since temperature scaling preserves the ranking of logits, it does not change the model's predictions (argmax), only the confidence values. This makes it ideal for calibration without sacrificing accuracy.

Temperature scaling works remarkably well for many deep learning models, particularly those using cross-entropy loss, because it addresses the general overconfidence problem with a single parameter. The universality of overconfidence in neural networks means that a single scalar adjustment often suffices to significantly improve calibration across the entire confidence range. Guo et al. (2017) found that temperature scaling achieved competitive or superior calibration performance compared to more complex methods on many standard benchmarks, an impressive result given its simplicity.

The limitation of temperature scaling is its single degree of freedom. It assumes that miscalibration is uniform: the same temperature corrects the model's confidence at all confidence levels. If the model is overconfident at high confidence but underconfident at medium confidence, temperature scaling cannot fix both problems simultaneously. In practice, this limitation is minor for many models because overconfidence tends to be the dominant direction across the full confidence range.

Out[4]:
Visualization
Grouped bar chart showing how different temperature values change the probability distribution across classes.
Effect of temperature parameter T on softmax probability distribution over five classes with fixed logits. Higher temperatures (T > 1) soften the distribution by reducing peak confidence and spreading mass to other classes, while lower temperatures (T < 1) sharpen it further toward a one-hot prediction. Temperature T = 1 recovers the original softmax output.

Platt Scaling

Platt scaling, originally developed for Support Vector Machines by John Platt in 1999, fits a logistic regression model to the classifier outputs. Where temperature scaling uses a single parameter to rescale all logits uniformly, Platt scaling uses two parameters: a slope and a bias. This two-parameter affine transformation can correct for both the scale of logits (like temperature scaling) and a systematic offset in the logit values. For binary classification:

p(y=1∣z)=σ(az+b)=11+exp⁡(−(az+b))p(y=1 \mid z) = \sigma(az + b) = \frac{1}{1 + \exp(-(az + b))}

where:

  • p(y=1∣z)p(y=1 \mid z): the calibrated probability of the positive class given logit zz
  • zz: the model's logit (or confidence score) for the positive class
  • σ(⋅)\sigma(\cdot): the sigmoid function
  • aa: the slope parameter (learned on the calibration set)
  • bb: the bias parameter (learned on the calibration set)

The bias parameter bb allows Platt scaling to correct for models that are systematically biased toward one class, a situation that pure temperature scaling cannot address because it preserves the relative ordering and marginal distribution of predictions. For multiclass problems, Platt scaling can be applied to each class independently in a one-vs-rest fashion, fitting a separate (ak,bk)(a_k, b_k) pair for each class kk.

Unlike temperature scaling, Platt scaling can change the ranking of predictions because the affine transformation with a non-unit slope and non-zero bias does not preserve all orderings. It is more flexible than temperature scaling but requires more parameters and is therefore more prone to overfitting on small validation sets. The additional degrees of freedom allow Platt scaling to correct more complex miscalibration patterns, but this flexibility requires careful regularization or larger calibration sets to avoid learning noise in the calibration data.

When should you prefer Platt scaling over temperature scaling? If you have evidence that the model's logits have a systematic bias, such as always predicting too confidently in the positive direction for one class, Platt scaling can address this through its bias term. If the miscalibration is purely a matter of the logit scale being off (the model ranks examples correctly but assigns unreliable confidence scores), temperature scaling is simpler and less prone to overfitting.

Isotonic Regression

Isotonic regression is a non-parametric calibration method that learns a piecewise constant, monotonically increasing function mapping uncalibrated probabilities to calibrated ones. This approach makes no assumptions about the functional form of the calibration error beyond the reasonable constraint that higher uncalibrated confidence should map to higher calibrated probability. It finds the function ff that minimizes the mean squared error between calibrated outputs and true outcomes:

min⁡f∑i=1n(f(p^i)−yi)2\min_{f} \sum_{i=1}^{n} (f(\hat{p}_i) - y_i)^2

where:

  • ff: the isotonic (non-decreasing) calibration function being learned
  • p^i\hat{p}_i: the uncalibrated predicted probability for the ii-th sample
  • yiy_i: the true binary outcome (0 or 1) for the ii-th sample
  • nn: the number of validation samples

subject to the constraint that ff is isotonic (non-decreasing): f(p^i)≤f(p^j)f(\hat{p}_i) \leq f(\hat{p}_j) whenever p^i≤p^j\hat{p}_i \leq \hat{p}_j.

The monotonicity constraint is important and reasonable: if the model assigns higher confidence to one prediction than another, the calibrated probability should also be at least as high. This preserves the relative ordering of predictions, which is the fundamental information we want to keep while correcting the absolute probability values.

The isotonic regression problem is solved by the pool adjacent violators algorithm (PAVA), a linear-time algorithm that merges adjacent bins whenever the monotonicity constraint is violated. Starting from the sorted list of predictions, PAVA identifies pairs of adjacent predictions that violate the monotone constraint (where the fitted value decreases) and pools them by averaging. This pooling continues until no violations remain.

Isotonic regression makes no assumptions about the shape of the calibration error beyond monotonicity, making it more flexible than parametric methods. However, it requires more data to avoid overfitting and does not generalize well beyond the range of confidence values observed during calibration. If the validation set contains no examples with confidence below 0.2, the isotonic regression function may behave unpredictably when encountering such low-confidence predictions at test time. It is particularly useful when the calibration curve has a complex, non-linear shape that temperature scaling cannot capture, such as when the model is underconfident at low probabilities but overconfident at high probabilities.

Beta Calibration

Beta calibration extends Platt scaling by using the Beta distribution's cumulative distribution function, which is more naturally suited to probabilities bounded in [0,1][0, 1]. Standard logistic regression assumes logits can range over the entire real line, but probabilities naturally live in a bounded interval. The Beta distribution respects these natural bounds and can model asymmetric calibration curves more effectively.

Beta calibration applies a parametric transformation to the uncalibrated probability p^\hat{p}:

log⁡p1−p=a⋅log⁡p^−b⋅log⁡(1−p^)+c\log \frac{p}{1-p} = a \cdot \log \hat{p} - b \cdot \log(1 - \hat{p}) + c

where:

  • pp: the calibrated probability
  • p^\hat{p}: the uncalibrated probability
  • a,ba, b: shape parameters controlling the steepness of calibration correction for low and high probabilities respectively
  • cc: an intercept term

The parameters aa, bb, and cc are learned by maximizing likelihood on the calibration set. The Beta calibration model reduces to Platt scaling when a=ba = b, making it a strict generalization. The asymmetry between aa and bb is the key advantage: the model can apply stronger correction at high confidence levels without distorting low confidence levels, or vice versa.

This approach respects the natural bounds of probability values and can handle asymmetric calibration curves better than standard Platt scaling. When the calibration error differs between low and high confidence regions, Beta calibration can capture these nuances through the flexibility of its separate shape parameters for each end of the probability spectrum.

Vector Scaling and Matrix Scaling

For models where different classes have different calibration profiles, class-independent temperature scaling is insufficient. Vector scaling generalizes temperature scaling by learning a separate temperature for each class, applying an element-wise multiplication of logits by a learned vector:

p=softmax(W⊙z)\mathbf{p} = \text{softmax}(\mathbf{W} \odot \mathbf{z})

where W\mathbf{W} is a diagonal matrix of per-class temperature parameters and ⊙\odot denotes element-wise multiplication. This allows the model to be made more confident about some classes and less confident about others based on where the calibration errors lie.

Matrix scaling takes this further by learning a full K×KK \times K matrix transformation of the logits. This is the most expressive parametric post-hoc calibration method but also the most prone to overfitting when the calibration set is small relative to the number of classes. For models with thousands of vocabulary tokens, matrix scaling is impractical: a full matrix for a 50,000-vocabulary model would require 50,0002=2.550{,}000^2 = 2.5 billion parameters. Vector scaling (with KK parameters) is more tractable but still requires a calibration set with many examples per class.

Label Smoothing as Calibration

While technically a training modification rather than post-hoc calibration, label smoothing deserves mention here because it directly addresses calibration during training. By replacing hard targets (0 and 1) with soft targets (e.g., 0.1 and 0.9), label smoothing prevents the model from becoming overconfident during training. This regularization technique encourages the model to maintain some uncertainty even when it correctly classifies training examples. A standard cross-entropy loss with label smoothing ϵ\epsilon uses the modified target:

yismooth=(1−ϵ)⋅yi+ϵKy_i^{\text{smooth}} = (1 - \epsilon) \cdot y_i + \frac{\epsilon}{K}

where yiy_i is the original one-hot target, ϵ\epsilon is the smoothing parameter (typically 0.1), and KK is the number of classes. The model is now rewarded for assigning at most (1−ϵ+ϵ/K)(1 - \epsilon + \epsilon/K) probability to the correct class, which prevents logits from growing without bound.

Label smoothing often results in better-calibrated models out of the box, with the trade-off of a potential small cost to accuracy. The optimal level of smoothing depends on model size, dataset size, and the specific task. Research on large language model calibration suggests that some amount of label smoothing is almost always beneficial for calibration while having minimal impact on accuracy when the smoothing parameter is in the range of 0.05 to 0.2.

Worked Example

We will walk through a concrete example to see calibration in action. Imagine a sentiment classifier that outputs probabilities for Positive and Negative classes. We evaluate it on 10 samples:

Example predictions from a sentiment classifier showing confidence scores and correctness.
SampleTrue LabelPredicted LabelConfidenceCorrect?
1PositivePositive0.95Yes
2PositivePositive0.90Yes
3NegativeNegative0.85Yes
4PositivePositive0.80No
5NegativeNegative0.75Yes
6PositivePositive0.70No
7NegativeNegative0.65Yes
8PositivePositive0.60Yes
9NegativeNegative0.55No
10PositivePositive0.50Yes

Examining this table, we see the model makes predictions across the confidence spectrum from 0.50 to 0.95. Out of 10 predictions, 7 are correct, giving an overall accuracy of 70%. However, the average confidence is higher than 70%, suggesting overconfidence. To quantify this miscalibration, we need to group these predictions into bins and compare the accuracy and confidence within each bin.

Using 5 equal-width bins: [0.5-0.6), [0.6-0.7), [0.7-0.8), [0.8-0.9), [0.9-1.0]

  • Bin [0.5-0.6): 2 samples (0.55, 0.50), confidence = 0.525, accuracy = 0.5 (1 correct out of 2)
  • Bin [0.6-0.7): 2 samples (0.65, 0.60), confidence = 0.625, accuracy = 1.0 (2 correct out of 2)
  • Bin [0.7-0.8): 2 samples (0.75, 0.70), confidence = 0.725, accuracy = 0.5 (1 correct out of 2)
  • Bin [0.8-0.9): 2 samples (0.85, 0.80), confidence = 0.825, accuracy = 0.5 (1 correct out of 2)
  • Bin [0.9-1.0]: 2 samples (0.95, 0.90), confidence = 0.925, accuracy = 1.0 (2 correct out of 2)

Notice how the accuracy fluctuates relative to confidence. In the [0.6-0.7) bin, the model is underconfident, achieving perfect accuracy while expressing only 62.5% confidence. However, in the [0.7-0.8) and [0.8-0.9) bins, the model is overconfident, claiming 72.5% and 82.5% confidence respectively while only achieving 50% accuracy.

The ECE calculation proceeds by summing the weighted absolute gaps across all bins:

  • Bin 1: ∣0.5−0.525∣×0.2=0.005|0.5 - 0.525| \times 0.2 = 0.005
  • Bin 2: ∣1.0−0.625∣×0.2=0.075|1.0 - 0.625| \times 0.2 = 0.075
  • Bin 3: ∣0.5−0.725∣×0.2=0.045|0.5 - 0.725| \times 0.2 = 0.045
  • Bin 4: ∣0.5−0.825∣×0.2=0.065|0.5 - 0.825| \times 0.2 = 0.065
  • Bin 5: ∣1.0−0.925∣×0.2=0.015|1.0 - 0.925| \times 0.2 = 0.015

Total ECE = 0.205 or 20.5%

This indicates significant miscalibration, particularly in the middle confidence ranges where the model is overconfident. The ECE of 20.5% suggests that, on average, the model's confidence deviates from its accuracy by over 20 percentage points across the confidence spectrum. An ECE of 20.5% is quite high; in practice, well-calibrated models achieve ECE values below 5%, and many production systems aim for ECE below 2%.

Notice also that the MCE in this example is the maximum bin gap, which is bin 2: ∣1.0−0.625∣=0.375|1.0 - 0.625| = 0.375. This means that in the confidence range [0.6-0.7), the model's accuracy is 37.5 percentage points higher than its confidence, indicating substantial underconfidence in that range. If this were a real model, you would want to understand why the model is so certain about its errors in the middle confidence range while being correct when expressing slightly lower confidence.

With only 10 samples, these estimates are highly noisy. Each bin contains only 2 predictions, making the accuracy estimate either 0.0, 0.5, or 1.0 with no intermediate values. This example illustrates why ECE estimation requires large datasets: with a typical real evaluation containing thousands of predictions, the accuracy estimates within each bin become much more reliable and the resulting ECE value more meaningful.

Code Implementation

We implement calibration metrics and visualization using Python. We will create synthetic classification data with known calibration properties, then apply calibration methods.

In[5]:
Code
import numpy as np
import torch

# Set random seed for reproducibility
np.random.seed(42)
torch.manual_seed(42)

# Generate synthetic data: 1000 samples, 5 classes
n_samples = 1000
n_classes = 5

# True labels
true_labels = np.random.randint(0, n_classes, n_samples)

# Simulate uncalibrated model outputs (overconfident)
# The model tends to be correct but with inflated confidence
logits = np.random.randn(n_samples, n_classes) * 2.0
# Boost the true class logit to simulate overconfidence
for i in range(n_samples):
    logits[i, true_labels[i]] += 3.0

# Convert to probabilities
probs = torch.softmax(torch.tensor(logits), dim=1).numpy()
pred_labels = np.argmax(probs, axis=1)
confidences = np.max(probs, axis=1)

We have created synthetic logits where the correct class receives an artificial boost, simulating the overconfidence common in neural networks. We examine the accuracy and confidence distribution.

Out[6]:
Console
Accuracy: 0.678
Mean Confidence: 0.740
Gap (Confidence - Accuracy): 0.062

The model shows higher average confidence than accuracy, indicating overconfidence. This gap between mean confidence and accuracy is a quick first diagnostic: a positive gap suggests overconfidence, and a negative gap suggests underconfidence. Now let us implement the Expected Calibration Error calculation.

In[7]:
Code
def expected_calibration_error(y_true, y_prob, y_pred, n_bins=10):
    """
    Calculate Expected Calibration Error.

    Args:
        y_true: True labels (integer indices)
        y_prob: Predicted probabilities for the predicted class
        y_pred: Predicted labels
        n_bins: Number of bins for calibration

    Returns:
        ece: Expected Calibration Error
        bin_accuracies: Accuracy per bin
        bin_confidences: Confidence per bin
        bin_counts: Number of samples per bin
    """
    bin_boundaries = np.linspace(0, 1, n_bins + 1)
    bin_lowers = bin_boundaries[:-1]
    bin_uppers = bin_boundaries[1:]

    ece = 0.0
    bin_accuracies = []
    bin_confidences = []
    bin_counts = []

    for bin_lower, bin_upper in zip(bin_lowers, bin_uppers):
        # Find samples in this bin
        in_bin = (y_prob > bin_lower) & (y_prob <= bin_upper)
        prop_in_bin = np.mean(in_bin)

        if prop_in_bin > 0:
            accuracy_in_bin = np.mean((y_pred == y_true)[in_bin])
            confidence_in_bin = np.mean(y_prob[in_bin])

            ece += np.abs(accuracy_in_bin - confidence_in_bin) * prop_in_bin

            bin_accuracies.append(accuracy_in_bin)
            bin_confidences.append(confidence_in_bin)
            bin_counts.append(np.sum(in_bin))
        else:
            bin_accuracies.append(0.0)
            bin_confidences.append(0.0)
            bin_counts.append(0)

    return (
        ece,
        np.array(bin_accuracies),
        np.array(bin_confidences),
        np.array(bin_counts),
    )


# Calculate ECE for our synthetic model
ece, bin_acc, bin_conf, bin_counts = expected_calibration_error(
    true_labels, confidences, pred_labels, n_bins=10
)
Out[8]:
Console
Expected Calibration Error: 0.0645 (6.45%)

Bin statistics (10 bins):
Bin Range       Count    Accuracy   Confidence   Gap     
------------------------------------------------------------
[0.0, 0.1)0        0.000      0.000        +0.000
[0.1, 0.2)0        0.000      0.000        +0.000
[0.2, 0.3)2        1.000      0.264        -0.736
[0.3, 0.4)45       0.356      0.363        +0.008
[0.4, 0.5)100      0.430      0.452        +0.022
[0.5, 0.6)129      0.465      0.552        +0.087
[0.6, 0.7)145      0.607      0.650        +0.043
[0.7, 0.8)124      0.685      0.748        +0.063
[0.8, 0.9)150      0.733      0.851        +0.118
[0.9, 1.0)305      0.898      0.956        +0.058

Looking at the per-bin statistics, we can read the calibration curve numerically before visualizing it. Bins where the gap is positive indicate overconfidence: the model claims more confidence than its accuracy warrants. Bins where the gap is negative indicate underconfidence. The pattern of gaps across bins reveals whether the miscalibration is uniform (suggesting temperature scaling is appropriate) or localized to specific confidence ranges (suggesting a non-parametric method would be better).

Out[9]:
Visualization
Grouped bar chart comparing confidence and accuracy per bin with gap indicators.
Expected Calibration Error calculation comparing model confidence against empirical accuracy across ten discrete bins. Confidence consistently exceeds accuracy in most bins, with the red gap indicators showing the per-bin contribution to ECE. Bins where the two bars align closely contribute little to the overall miscalibration.

The ECE shows significant miscalibration, with several bins showing large gaps between confidence and accuracy. We visualize this with a reliability diagram.

Out[10]:
Visualization
Calibration plot with confidence on x-axis and accuracy on y-axis, showing points below the diagonal.
Reliability diagram showing systematic overconfidence in the uncalibrated model. Data points falling below the diagonal indicate the model claims higher confidence than its actual accuracy warrants, with points in the high-confidence region showing the largest deviations from ideal calibration.
Out[11]:
Visualization
Histogram showing confidence distribution with vertical reference lines.
Distribution of model confidence scores across predictions, showing strong concentration in high-confidence regions (above 0.8). The vertical lines contrast mean confidence with actual accuracy, revealing the calibration gap and confirming that most predictions carry inflated confidence.

The reliability diagram confirms overconfidence: the model's accuracy is consistently lower than its confidence. The histogram shows that most predictions have high confidence (above 0.7), which is typical for overconfident neural networks. Now we implement temperature scaling to calibrate these probabilities.

In[12]:
Code
def temperature_scaling(logits, y_true, eps=1e-10):
    """
    Find optimal temperature for calibration using negative log-likelihood.
    """
    import numpy as np
    import torch
    import torch.nn.functional as F
    from scipy.optimize import minimize_scalar

    def nll_loss(T):
        # Apply temperature scaling
        scaled_logits = logits / T
        probs = F.softmax(torch.tensor(scaled_logits), dim=1).numpy()

        # Calculate negative log-likelihood
        nll = 0.0
        for i in range(len(y_true)):
            nll -= np.log(probs[i, y_true[i]] + eps)
        return nll / len(y_true)

    # Optimize temperature
    result = minimize_scalar(nll_loss, bounds=(0.1, 10.0), method="bounded")
    optimal_T = result.x

    # Apply optimal temperature
    scaled_logits = logits / optimal_T
    calibrated_probs = F.softmax(torch.tensor(scaled_logits), dim=1).numpy()

    return optimal_T, calibrated_probs


# Apply temperature scaling
optimal_T, cal_probs = temperature_scaling(logits, true_labels)
cal_pred_labels = np.argmax(cal_probs, axis=1)
cal_confidences = np.max(cal_probs, axis=1)
Out[13]:
Console
Optimal Temperature: 1.305
Original Accuracy: 0.678
Calibrated Accuracy: 0.678
Original Mean Confidence: 0.740
Calibrated Mean Confidence: 0.671

The optimal temperature is greater than 1, which softens the distribution and reduces confidence. Notice that accuracy remains unchanged (temperature scaling preserves rankings), but mean confidence has decreased to better match the actual accuracy. The optimal temperature value itself is informative: larger values indicate more severe overconfidence that required more aggressive softening.

We calculate the ECE after calibration and visualize the improvement.

In[14]:
Code
# Calculate ECE after temperature scaling
ece_cal, bin_acc_cal, bin_conf_cal, _ = expected_calibration_error(
    true_labels, cal_confidences, cal_pred_labels, n_bins=10
)
Out[15]:
Console
ECE before calibration: 0.0645
ECE after temperature scaling: 0.0235
Improvement: 0.0410 (63.6% reduction)
Out[16]:
Visualization
Reliability diagram with points scattered below the diagonal.
Uncalibrated reliability diagram showing significant overconfidence. The model's predicted confidence substantially exceeds its actual accuracy, as shown by points falling well below the diagonal reference line.
Reliability diagram with points closer to diagonal.
Reliability diagram after temperature scaling calibration. The calibration curve shifts considerably closer to the diagonal, indicating the adjusted confidence scores now much better match the actual accuracy across all confidence levels.
Out[17]:
Visualization
Histogram showing high confidence values.
Distribution of uncalibrated confidence scores, showing a right-skewed pattern with most predictions having inflated confidence values concentrated above 0.8.
Histogram showing more spread confidence values.
Distribution of calibrated confidence scores after temperature scaling, showing a leftward shift toward more moderate probability values that align better with the model's true accuracy of approximately 70 percent.

Temperature scaling has significantly reduced the ECE by bringing the calibration curve closer to the diagonal. The confidence distribution has also shifted leftward, reducing the overconfidence. Now we implement isotonic regression for comparison.

In[18]:
Code
# For isotonic regression, we need binary outcomes (correct/incorrect)
correct = (pred_labels == true_labels).astype(float)

# Fit isotonic regression
from sklearn.isotonic import IsotonicRegression

iso_reg = IsotonicRegression(out_of_bounds="clip")
iso_reg.fit(confidences, correct)
iso_confidences = iso_reg.predict(confidences)

# Clip to valid probability range
iso_confidences = np.clip(iso_confidences, 0, 1)
Out[19]:
Console
ECE with Temperature Scaling: 0.0235
ECE with Isotonic Regression: 0.0000
Out[20]:
Visualization
Bar chart comparing ECE values for uncalibrated, temperature scaled, and isotonic regression methods.
Comparison of Expected Calibration Error across calibration methods. Both post-hoc techniques substantially reduce the ECE compared to the uncalibrated baseline. Isotonic regression often achieves lower ECE than temperature scaling because it can fit complex non-linear calibration curves, but this comes with a risk of overfitting on small calibration sets.

Isotonic regression often achieves lower ECE than temperature scaling because it is non-parametric and can fit complex calibration curves. However, it requires more data and can overfit on small validation sets. It also does not necessarily preserve the ranking of predictions when applied to out-of-distribution data, unlike temperature scaling, which maintains the argmax ordering.

Calibration in Language Models

The calibration story for modern large language models is more complicated than for classification models, and it comes with unique challenges that standard calibration methods do not fully address. Understanding these challenges is important before applying calibration techniques to LLM outputs in practice.

Token Probabilities vs. Answer Probabilities

A language model produces a probability distribution over the vocabulary at each token position. These token-level probabilities are well-defined and can be measured directly. However, the quantity we usually care about for calibration purposes is not token probability but answer probability: how confident is the model that its multi-token response is correct?

Converting from token probabilities to answer probabilities is non-trivial. One approach is to sum the log probabilities of all tokens in the response and use this as a log-probability for the full answer. This treats the model as having assigned the product of token probabilities to the full sequence. But this approach penalizes longer answers regardless of correctness, confounds fluency (high probability of generating grammatical tokens) with factual accuracy (generating the correct answer), and does not account for the many equivalent phrasings of the same correct answer.

Some evaluations compute the probability of a single correct token (such as "A", "B", "C", or "D" in a multiple-choice question) conditioned on the question and answer prefix. This approach isolates the relevant probability mass but only works for constrained answer formats. For open-ended questions, you typically need to either sample multiple completions and estimate confidence from the distribution, or use prompted techniques like asking the model to express its own confidence verbally.

Verbalized Confidence

One growing line of research asks language models to verbalize their own confidence. Instead of reading the model's internal probability scores, you prompt the model to say "I am 85% confident that..." and then evaluate whether that stated confidence is calibrated. This approach is appealing because it uses the model's language capabilities for uncertainty quantification, does not require access to logits (important for closed-weight models like commercial APIs), and can potentially capture epistemic uncertainty that token probabilities miss.

The research findings are mixed. Models can learn to express calibrated verbal confidence when explicitly trained on confidence calibration tasks, but they do not automatically do so from standard language modeling pre-training. When asked to express confidence without specific training for this skill, models often show systematic verbal overconfidence: they tend to use high-confidence language ("I am certain that...", "Definitely...") more often than their actual accuracy warrants. This verbal overconfidence is separate from the probability-level overconfidence and can be harder to correct because it reflects learned stylistic patterns in the training data rather than intrinsic probability estimates.

Scale and Calibration

Large language models exhibit interesting scaling behavior with respect to calibration. Earlier research suggested that larger models tend to be more overconfident because they fit training data more aggressively. More recent work with models trained with explicit calibration objectives or RLHF finds that model scale does not uniformly worsen calibration: some large models are better calibrated than smaller ones when evaluated carefully on knowledge tasks.

The key factor appears to be what the model is confident about. Large models tend to be well-calibrated on questions they answer correctly most of the time (they correctly recognize that they know the answer) and poorly calibrated on questions where they make systematic errors (they confidently state incorrect facts). This "confident in wrong things" pattern is sometimes called hallucination confidence and represents a particularly dangerous form of miscalibration: the model does not just make errors, it expresses high confidence in those errors.

Selective Prediction and Abstention

One practical response to calibration concerns is selective prediction: rather than always producing an answer, allow the model to abstain from answering when its confidence falls below a threshold. A well-calibrated model with an abstention option can achieve high accuracy on the questions it chooses to answer while avoiding the harmful effects of overconfident errors on difficult questions.

For selective prediction to work correctly, calibration is essential. If the model is overconfident, the threshold for abstention is harder to set correctly: you might accept answers that the model states with 85% confidence, but if that corresponds to 60% accuracy, your selective prediction policy is accepting too many incorrect answers. Conversely, if the model is underconfident, you might be causing unnecessary abstentions on questions the model would have answered correctly.

The coverage-accuracy trade-off is the key metric for selective prediction. Coverage is the fraction of queries the model answers (rather than abstaining), and accuracy is the fraction of answered queries that are correct. A perfectly calibrated model produces an optimal coverage-accuracy curve: by lowering the threshold, you can increase coverage at the cost of lower accuracy, with the trade-off following the empirical accuracy at each confidence level. A miscalibrated model produces a suboptimal curve: you get less accuracy at each coverage level than you would from a well-calibrated model.

Limitations and Impact

Calibration metrics and methods, while powerful, have important limitations that practitioners must understand.

The most fundamental limitation of ECE is its sensitivity to the number of bins and the binning strategy. Equal-width binning can produce misleading ECE estimates if the confidence distribution is imbalanced, which is common in neural networks that tend toward extreme probabilities. Adaptive binning helps but complicates comparison across models and implementations. Different researchers use different bin counts (typically between 10 and 20), making it difficult to compare ECE values across papers without knowing the exact evaluation protocol. A model with ECE of 0.03 reported with 15 bins may not be directly comparable to one with ECE of 0.04 reported with 10 bins.

ECE also only measures calibration, not sharpness. Sharpness is the property that the model's confidence scores are informative: a model that always predicts the marginal class probability (for example, predicting 0.7 confidence for a class with 70% base rate) is perfectly calibrated but useless, because it provides no information about individual predictions beyond the class prior. This highlights the fundamental difference between being well-calibrated and being informative. Ideally, you want a model that is both calibrated (stated probabilities match empirical frequencies) and sharp (predictions vary systematically across examples based on their difficulty). A useful metric combining these properties is the Expected Calibration Error weighted by sharpness, but this is less commonly reported than plain ECE.

Temperature scaling, despite its effectiveness, assumes that miscalibration is uniform across all classes and inputs. It cannot correct cases where the model is overconfident on some classes but underconfident on others, or where calibration changes with input difficulty. More sophisticated methods like vector scaling (learning a temperature per class) or matrix scaling exist but risk overfitting on small calibration datasets. The choice of calibration method is itself a modeling decision that can go wrong, and there is no universally best method across all models and tasks.

Isotonic regression, while flexible, requires careful validation to prevent overfitting and does not generalize well beyond the range of observed confidence values. If the calibration set contains no examples with confidence below 0.2, the isotonic regression function may behave unpredictably when encountering such low-confidence predictions at test time, because the calibration mapping was never trained for that region. In deployment, model inputs can differ from the calibration set in subtle ways, making post-hoc calibration less reliable than might be hoped.

The impact of calibration extends far beyond academic metrics. In high-stakes applications like medical diagnosis or legal advice, a miscalibrated model can be dangerous. If a diagnostic AI reports 99% confidence in a negative result but is wrong 5% of the time in that confidence band, clinicians who rely on that threshold will miss serious conditions. The false sense of security created by overconfident predictions leads to reduced human oversight and missed critical errors. This problem is particularly severe because the harm is invisible: you only notice the miscalibration when you aggregate outcomes over time, not at the point of individual prediction.

In retrieval-augmented generation systems (which we cover in later chapters), calibration affects when the model should defer to retrieved documents versus rely on its parametric knowledge. Poorly calibrated confidence scores can lead to incorrect trust in either the model's internal knowledge or the external retrieval, resulting in either ignoring valuable retrieved information or trusting incorrect parametric memories. A model that knows it is uncertain would ideally defer to retrieved context; a model that is overconfident bypasses the retrieval and relies on potentially outdated or incorrect internal knowledge.

As we move toward more capable language models evaluated on complex benchmarks, calibration becomes important for human-AI collaboration. You need to know when to trust model outputs and when to seek human verification. A model that knows what it knows, and admits what it does not, is far more valuable than one that is merely accurate on average. This metacognitive capability, often called "knowing what you know," represents a critical safety property for deployed AI systems. Well-calibrated models enable appropriate trust calibration in human users, preventing both excessive skepticism of correct predictions and dangerous over-reliance on incorrect ones.

Future work in alignment and safety increasingly focuses on calibration as a prerequisite for trustworthy AI systems. One proposed safety criterion for advanced AI systems is that they should maintain accurate models of their own uncertainty and communicate that uncertainty faithfully to users. Calibration is the measurable operationalization of this principle: a model that is well-calibrated at prediction time has demonstrated one key component of honest epistemic behavior, even if achieving calibration does not guarantee other desired safety properties.

Summary

Calibration measures the alignment between a model's predicted confidence and its actual accuracy. A perfectly calibrated model satisfies the condition that among all predictions where the model claims confidence pp, exactly fraction pp are correct. This is a statistical property that applies to the model's behavior over many predictions, not to individual outputs.

The Expected Calibration Error quantifies miscalibration by binning predictions and measuring the weighted average gap between confidence and accuracy across bins. The weights reflect the fraction of predictions in each bin, so that densely populated confidence regions contribute more to the overall metric. MCE captures worst-case miscalibration in any single bin, which matters more in safety-critical applications. Both metrics depend on the binning strategy, and comparison across implementations requires care.

Neural networks systematically tend toward overconfidence due to the asymmetric gradient of cross-entropy loss, increased model capacity, and interactions with architectural components like batch normalization. Label smoothing during training can partially mitigate this tendency by preventing logits from growing without bound toward infinite confidence.

Temperature scaling provides a simple, effective post-hoc calibration method that divides logits by a learned scalar T>1T > 1 to soften overconfident distributions. It preserves prediction rankings and achieves surprisingly competitive calibration performance given its single parameter. Isotonic regression offers a non-parametric alternative that can capture complex calibration curves but requires more data. Platt scaling and Beta calibration provide intermediate options with two to four parameters. It offers more flexibility than temperature scaling while being less prone to overfitting than isotonic regression.

Reliability diagrams visualize calibration by plotting accuracy against confidence at each binned confidence level, revealing overconfidence (points below the diagonal) or underconfidence (points above). Modern large language models achieve high accuracy but often suffer from overconfidence, making calibration an essential consideration for applications that rely on uncertainty estimates for decision-making, selective prediction, or human-AI collaboration.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about calibration.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026calibrationmachine, author = {Michael Brenndoerfer}, title = {Calibration in Machine Learning: Confidence, Accuracy & ECE}, year = {2026}, url = {https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Calibration in Machine Learning: Confidence, Accuracy & ECE. Retrieved from https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece
MLAAcademic
Michael Brenndoerfer. "Calibration in Machine Learning: Confidence, Accuracy & ECE." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece>.
CHICAGOAcademic
Michael Brenndoerfer. "Calibration in Machine Learning: Confidence, Accuracy & ECE." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Calibration in Machine Learning: Confidence, Accuracy & ECE'. Available at: https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Calibration in Machine Learning: Confidence, Accuracy & ECE. https://mbrenndoerfer.com/writing/calibration-machine-learning-confidence-accuracy-ece

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.