Part of Language AI Handbook
Covers cross-entropy loss for language models. Topics include entropy, KL divergence, perplexity.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Cross-Entropy Loss
Language models learn by predicting the next token in a sequence. But how do we quantify the gap between a model's predictions and reality? We need a mathematical measure that can compare the probability distribution output by our model against the actual observed data. This provides a single scalar value that indicates how well our predictions align with truth. Cross-entropy loss provides the mathematical foundation for this measurement, serving as both the training objective that guides gradient descent through the parameter space and the evaluation metric that tells us how well our model understands the statistical patterns of language.
As we explored in our discussion of Perplexity, language models output probability distributions over possible next tokens. At each position in a sequence, the model produces a vector of probabilities representing its confidence that each vocabulary item will appear next. Cross-entropy gives us a principled way to measure how far these predicted distributions diverge from the true distributions we observe in training data. It connects information theory, probability, and optimization into a single coherent framework that drives modern language AI. This framework allows us to transform the abstract problem of learning language patterns into a concrete optimization task where we can compute gradients and update model weights systematically.
Unlike perplexity, which expresses model uncertainty as an intuitive "effective vocabulary size," cross-entropy operates in the domain of information theory, measuring the average number of bits needed to encode the true data using a code optimized for the model's predicted distribution. This perspective reveals basic limits on compression and provides the optimization target for virtually all modern language models, from the causal objectives discussed in Causal Language Modeling to the masked objectives in Masked Language Modeling. Understanding cross-entropy provides deep insight into why language models work, how they improve during training, and what basic limits constrain their performance.
One reason cross-entropy occupies such a central position in language model training is that it connects naturally to maximum likelihood estimation, the workhorse of classical statistics. When you minimize cross-entropy loss over a dataset, you are simultaneously maximizing the probability that the model assigns to the observed training tokens. This equivalence means cross-entropy is not an arbitrary choice of loss function imposed by convention. It is the theoretically principled answer to the question of how to measure the quality of a probabilistic model. Later in this chapter, we will make this connection precise and see why it justifies using cross-entropy as the objective for training neural language models of any size.
Information Theory Foundations
To understand cross-entropy, we must first grasp entropy and how it quantifies uncertainty. The concept of entropy originates from thermodynamics and statistical mechanics, but Claude Shannon adapted it for information theory in 1948. This provides the mathematical vocabulary for measuring information content. Shannon's framework allows us to quantify how much information is contained in a random event, which directly translates to how surprised we should be when observing different outcomes. For language modeling, this means measuring how unexpected each token is given our current understanding of the language's statistical structure.
Shannon's key insight was that information and uncertainty are two sides of the same coin. An event carries more information precisely because it was unexpected. When a coin lands heads after you knew it would land heads because it is double-headed, you learn nothing. When that same coin lands heads despite being weighted toward tails, you learn something meaningful. This seemingly philosophical distinction has concrete mathematical consequences that turn out to be exactly what we need to train language models.
Shannon Entropy
Entropy measures the uncertainty inherent in a probability distribution. It answers the question: on average, how surprised will we be when we observe an outcome from this distribution? For a discrete distribution over outcomes , entropy is defined as:
where:
- : the entropy of probability distribution , measured in bits (base-2 log) or nats (natural log)
- : a possible outcome from the set of all events
- : the probability of outcome occurring
- : the logarithm function (typically base 2 for bits or natural log for nats)
The logarithm ensures that rare events contribute more to the total uncertainty than common ones. This mathematical property aligns with our intuition: when an unlikely event occurs, we experience more surprise and receive more information than when a common event occurs. When an event is certain (), its contribution is zero because we gain no new information from observing something we already knew would happen. When an event is impossible (), we define by continuity. This reflects that impossible events contribute nothing to the entropy since they never occur.
To see why the logarithm specifically encodes surprise, consider two independent events each with probability . The probability of both occurring together is . If surprise is additive (seeing two independent rare events is twice as surprising as seeing one), then the measure of surprise should satisfy . The only continuous function satisfying this property for all probabilities is the logarithm. So is the only mathematically consistent measure of surprise for an event with probability , and entropy is the expected value of this surprise over all possible outcomes.
Entropy represents the expected information content or "surprise" of observing a random variable. A uniform distribution over outcomes has maximum entropy , while a deterministic distribution has zero entropy. This means that when all outcomes are equally likely, we are maximally uncertain about what will happen next, whereas when one outcome is guaranteed, we have perfect certainty and zero uncertainty. From a compression perspective, entropy is the basic lower bound on how many bits per symbol any lossless compression algorithm can achieve when encoding data drawn from distribution .
Consider a vocabulary of 50,000 tokens. A uniform distribution where every token has probability yields entropy:
This means we need approximately 15.6 bits on average to encode each token if we use an optimal code designed for this uniform distribution. This represents the worst-case scenario for compression, where no token is more predictable than any other. In contrast, if our model is perfect and assigns probability 1 to the correct token and 0 elsewhere, entropy drops to zero, requiring no bits to transmit the message because there is no uncertainty. The difference between these extremes represents the potential for compression that good language models exploit.
Real language is far from uniform. Common words like "the" and "is" appear far more frequently than rare technical terms. Good models learn these frequency differences and assign higher probability to common tokens, reducing average surprise below the maximum. The entropy of real English text has been estimated at 0.6 to 1.2 bits per character depending on context length, far below the 4.7 bits per character that would be required for a uniform distribution over the 26 letters and a space character. This gap is what enables compression algorithms, and it is the gap that language models are trained to exploit.

Kullback-Leibler Divergence
While entropy measures uncertainty within a single distribution, the Kullback-Leibler (KL) divergence measures how one probability distribution diverges from a reference distribution . It quantifies the inefficiency of assuming that the distribution is when the true distribution is :
where:
- : the Kullback-Leibler divergence from distribution to reference distribution
- : the probability of outcome under the true distribution
- : the probability of outcome under the model distribution
- : an outcome from the support of
KL divergence is always non-negative and equals zero if and only if everywhere. This non-negativity follows from Jensen's inequality applied to the concave logarithm function. It represents the extra bits required when using a code optimized for to encode data that follows distribution . This interpretation reveals that KL divergence measures the penalty we pay for using the wrong model: if we build a compression scheme based on our predicted distribution , but reality follows , we will waste bits proportional to the KL divergence.
The asymmetry of KL divergence is worth understanding carefully. is not the same as in general. When we minimize with respect to , we penalize cases where but . This is called the "forward" or "inclusive" KL, and it encourages to cover all regions where has mass. When we minimize (the "reverse" or "exclusive" KL), we penalize cases where but , which encourages to concentrate on modes of . Language model training uses the forward KL form through cross-entropy, which means models are penalized for assigning low probability to tokens that occur in training data, encouraging broad coverage of plausible continuations.
However, KL divergence has an asymmetry that matters for machine learning: it requires knowing the true distribution over all possible outcomes. In practice, we only observe specific outcomes from , not the full distribution. We see individual tokens in our training data, but we do not know the complete probability distribution over all possible next tokens for every context. This limitation prevents us from computing the KL divergence directly during training, necessitating the use of cross-entropy instead.
Defining Cross-Entropy
Cross-entropy combines the entropy of the true distribution with the KL divergence between true and predicted distributions into a single quantity that we can compute from data. For distributions (true) and (model prediction), cross-entropy is:
where:
- : the cross-entropy between the true distribution and the predicted distribution
- : the probability of outcome under the true distribution (empirical distribution from data)
- : the probability of outcome assigned by the model
- : an outcome from the set of possible events
When represents the empirical distribution from training data, is 1 for the actual observed token and 0 for all others. This simplifies cross-entropy dramatically because the sum collapses to a single term involving only the true token:
where:
- : the observed token (the single outcome where )
- : the probability assigned by the model to the true token
This is the negative log-likelihood of the true token under the model's predicted distribution. It penalizes the model proportionally to how surprised it was by the actual observation: if the model assigns high probability to the correct token, the loss is small; if the model assigns low probability, the loss is large, approaching infinity as the predicted probability approaches zero.
The simplification from the full sum to a single term has an important practical consequence. During training, we do not need to compare probability distributions element by element across the entire vocabulary. We only need to look up how much probability the model assigned to the one token that appeared next. This makes gradient computation efficient: backpropagation only needs to update the output neuron corresponding to the true token, though intermediate layers receive gradients from all vocabulary items through the softmax normalization.

Cross-entropy decomposes into entropy plus KL divergence:
where:
- : the cross-entropy between the true distribution and the predicted distribution
- : the entropy of the true distribution (constant with respect to model parameters)
- : the KL divergence from the model distribution to the true distribution
Since is constant with respect to model parameters, minimizing cross-entropy is equivalent to minimizing KL divergence. This decomposition reveals that cross-entropy captures both the inherent uncertainty in the data (the entropy term) and the additional penalty for using a suboptimal model (the KL divergence term). During optimization, the entropy term acts as a constant offset, so gradient descent on cross-entropy effectively minimizes the KL divergence between our model and reality. The lowest achievable cross-entropy on any dataset equals the true entropy of the language, which represents the irreducible uncertainty stemming from inherent ambiguity in text.
Cross-Entropy as Maximum Likelihood Estimation
The connection between cross-entropy and maximum likelihood estimation (MLE) explains the statistical basis for using cross-entropy to train probabilistic language models. Given a dataset of observed tokens , MLE seeks the model parameters that maximize the probability of the observed data:
where:
- : the maximum likelihood estimate of the model parameters
- : the model's probability of the -th observed token given its context
- : the product over all tokens in the dataset
Because products of many small probabilities rapidly underflow to zero numerically, we take the logarithm of both sides. This converts the product into a sum, which is numerically stable:
where:
- : the log-likelihood of the observed tokens
- : the negative log-likelihood, normalized by dataset size
The final expression is exactly the cross-entropy loss averaged over the training set. Minimizing cross-entropy is mathematically identical to performing maximum likelihood estimation. This means we have a strong statistical justification for using cross-entropy: it is the objective that, as grows large, causes our model to converge to the true underlying data distribution (or the closest approximation within our model family). Any other loss function would not have this convergence guarantee in the same form.
The Language Modeling Objective
In autoregressive language modeling, we predict the next token given previous tokens . This sequential factorization allows us to model complex joint distributions over sequences by breaking them down into manageable conditional probabilities. The probability of a sequence factorizes as:
where:
- : the joint probability of the entire sequence
- : the conditional probability of token given all preceding tokens
This factorization is exact: the chain rule of probability guarantees it without any independence assumptions. The cross-entropy loss for a sequence of length is:
where:
- : the average cross-entropy loss over the sequence
- : the length of the target sequence (number of tokens)
- : the token at position in the sequence
- : all tokens preceding position (the context)
- : the model parameters
- : the model's predicted probability of token given previous tokens
This formulation reveals why cross-entropy is called the log-loss: we sum the negative log probabilities of each token. Lower probability assigned to the correct token results in higher loss, with the loss approaching infinity as the predicted probability approaches zero. This creates strong gradients when the model is confidently wrong, pushing the parameters aggressively to correct major errors.
The averaging by (sequence length) ensures that longer sequences do not automatically incur higher loss simply because they contain more tokens. This normalization allows us to compare loss values across sequences of different lengths and to aggregate loss across batches during training. Without this normalization, a long document would always appear to be modeled worse than a short sentence, even if the model's per-token accuracy was identical.
A necessary practical point: in transformer training, the model processes an entire sequence in a single forward pass, computing predicted probability distributions at every position simultaneously. For a sequence of length , this produces losses in parallel. The positions at the beginning of the sequence receive less context than positions near the end, so the model must predict early tokens with limited information. This variability in difficulty across positions is averaged away by the uniform weighting in the loss, which treats all positions equally. Some researchers have proposed weighting later positions more heavily, since they benefit from richer context and thus provide cleaner learning signal, but this remains an active area of investigation.
Cross-Entropy and Perplexity
As established in our discussion of Perplexity, these two metrics are intimately related through a simple mathematical transformation. Understanding this relationship helps you interpret evaluation results and compare models across different papers and frameworks, as they measure the same underlying quantity but present it in different units that emphasize different aspects of model performance.
The Mathematical Relationship
Perplexity is defined as the exponential of cross-entropy:
where:
- : the effective branching factor or exponential of the cross-entropy
- : the cross-entropy between true and predicted distributions
- : the total number of tokens in the evaluation set
- : the model's predicted probability of the -th true token
- : the exponential function (using the same base as the logarithm in cross-entropy)
The exponential uses the same base as the logarithm in the cross-entropy calculation. In natural language processing:
- When using natural logarithm () for cross-entropy, perplexity uses as the base
- When using base-2 logarithm () for cross-entropy, perplexity uses 2 as the base
Both formulations measure the same underlying quantity, just on different scales. Cross-entropy in nats (natural log) or bits () provides an additive measure of uncertainty, while perplexity provides a multiplicative measure that can be interpreted as the effective branching factor at each prediction step. If your model has a perplexity of 100, it is as uncertain as if it were choosing uniformly among 100 equally likely options at each step.
When to Use Each Metric
Cross-entropy offers advantages during training and optimization:
- Differentiability: Cross-entropy is directly differentiable and is the loss function for gradient computation. We can backpropagate through the cross-entropy calculation to update model weights.
- Additivity: Cross-entropy losses add across tokens, making aggregation straightforward. We can sum losses across batches and sequences without complications.
- Numerical stability: Working in log space avoids underflow when multiplying many small probabilities. Summing logs is numerically safer than multiplying probabilities directly.
Perplexity offers advantages for human interpretation:
- Scale intuition: A perplexity of 100 is more meaningful than a cross-entropy of 4.6 nats. Humans can intuitively grasp the difference between perplexity 50 and 100 better than between 3.9 and 4.6 nats.
- Vocabulary-relative performance: Perplexity naturally scales with vocabulary size. This allows rough comparison across models with different vocabularies. A perplexity equal to vocabulary size indicates random guessing. This provides a clear baseline.
- Historical convention: Perplexity remains the standard reporting metric for language model benchmarks. Research papers and leaderboards universally report perplexity, which makes it necessary for comparing with prior work.
In practice, modern frameworks often report both: cross-entropy for training curves and model selection, perplexity for final evaluation and paper results. During development, you might monitor cross-entropy to catch training instabilities early, then report perplexity in your final results to communicate with the broader research community.
Bits Per Character
While cross-entropy measures bits per token, bits per character (BPC) provides a tokenization-agnostic metric for comparing model efficiency. Since different tokenizers produce different sequence lengths for the same text, BPC normalizes by the number of characters in the original text. This normalization is needed for fair comparison because a model using a more aggressive tokenizer (with larger tokens) will naturally achieve lower per-token perplexity than an equivalent model using character-level tokenization, even if both models capture the same underlying linguistic patterns.
Calculating BPC
Given a sequence with characters and cross-entropy loss (in bits), bits per character is:
where:
- : bits per character, measuring compression efficiency independent of tokenization
- : the cross-entropy loss in bits (not nats)
- : the number of tokens in the sequence
- : the number of characters in the original text
For character-level language models where , BPC equals the cross-entropy. For subword tokenizers where (multiple characters per token), we must account for the compression ratio. The multiplication by converts the per-token loss to total bits for the sequence, while division by spreads those bits across the original characters.
BPC represents a basic limit on lossless compression. If a language model achieves 1.0 BPC on English text, it means no compression algorithm can compress English below 1 bit per character on average, assuming the model perfectly captures the true distribution of the language. This connects language modeling to data compression theory, where better language models yield better compression algorithms. In fact, arithmetic coding with a language model as the probability estimator is theoretically optimal: as the model improves, the compressed file size approaches the entropy of the text.
Current state-of-the-art language models achieve approximately 0.7-0.9 BPC on standard English text corpora, approaching but not reaching the estimated entropy of English (typically estimated between 0.6-1.0 BPC depending on the domain and estimation method). This gap suggests there remains room for improvement in language modeling, though diminishing returns may make further advances increasingly difficult.
Character-Level vs Subword Evaluation
When comparing models using different tokenization strategies, BPC provides fairer comparison than per-token metrics:
- Character-level models: Directly report BPC as the loss, since each token corresponds to one character.
- Subword models: Must divide total bits by character count to enable fair comparison with character-level models.
- Byte-level models: Similar to character-level but operating on UTF-8 bytes, where one character may consist of multiple bytes.
This normalization is important when evaluating Byte Pair Encoding or SentencePiece tokenizers against character-level baselines, as the same model architecture can appear to have wildly different per-token perplexities solely due to tokenization differences. Without BPC normalization, you might conclude that a subword model is dramatically better than a character model, when in fact the improvement comes primarily from the tokenizer grouping characters into larger units rather than from better linguistic understanding.
Binary vs Categorical Cross-Entropy
Language modeling uses categorical cross-entropy because we select among many discrete classes (vocabulary tokens). However, understanding binary cross-entropy illuminates special cases and alternative formulations that appear in auxiliary training objectives and discriminative fine-tuning tasks.
Categorical Cross-Entropy
For multi-class classification with classes, where the model outputs a probability distribution and the true class is represented as a one-hot vector :
where:
- : the cross-entropy loss for a single prediction
- : the total number of classes (vocabulary size)
- : the true probability of class (1 for the correct class, 0 otherwise in one-hot encoding)
- : the model's predicted probability for class
Since is one-hot (all zeros except for the true class where ), the sum collapses to a single term:
where:
- : the model's predicted probability for the true class (the single class where )
In language modeling, equals the vocabulary size (often 32,000 to 100,000 tokens). The Softmax Function ensures sums to 1 across all vocabulary items. This formulation handles the extreme multi-class nature of language modeling, where we must discriminate among tens of thousands of possible next tokens at each step.
Binary Cross-Entropy
For binary classification or independent multi-label problems:
where:
- : the binary cross-entropy loss
- : the true binary label (0 or 1)
- : the model's predicted probability that the label is 1
- : the probability that the label is 0
- : the model's predicted probability that the label is 0
Language models rarely use binary cross-entropy directly for next-token prediction, but it appears in specialized objectives like Replaced Token Detection, where the model predicts whether each token is original or replaced. In this case, the task reduces to binary classification for each position independently. The discriminator component of ELECTRA uses this formulation to learn by distinguishing real tokens from synthetically corrupted ones. This provides an alternative to traditional masked language modeling.
Binary cross-entropy also appears in contrastive learning objectives and ranking tasks. When a model must choose between two candidate continuations, the relative log-probabilities determine which is preferred, and binary cross-entropy over these preferences can serve as a training signal. This pattern appears in reinforcement learning from human feedback, where the reward model learns to predict human preferences between pairs of model outputs.
Label Smoothing
A practical modification to cross-entropy prevents overconfidence. Instead of using hard targets (0 or 1), label smoothing distributes a small amount of probability mass across all classes:
where:
- : the smoothing parameter (typically 0.01 to 0.1), representing the amount of probability mass to distribute uniformly
- : the adjusted target probability for the true class (reduced from 1)
- : the uniform smoothing term added to all classes
- : the model's predicted probability for class
This prevents the model from becoming overconfident and improves generalization by so gradients continue to flow even when the model predicts the correct class with high probability. Without smoothing, once a model achieves near-perfect prediction for a training example, gradients vanish and learning stops for that example. With smoothing, the model continues to learn, adjusting its internal representations to better model the underlying distribution rather than merely memorizing training labels. Typical values for range from 0.01 to 0.1, with higher values giving more regularization but potentially slowing convergence.
The effect of label smoothing can be understood from the decomposition perspective. Smoothing replaces the hard empirical distribution (all mass on one token) with a mixture: of the mass stays on the true token, and is spread uniformly. This makes the target distribution softer, requiring the model to assign non-negligible probability to all vocabulary items rather than collapsing its distribution to a single peak. In practice, label smoothing consistently improves downstream task performance and was used in the original Transformer paper, where it improved translation BLEU scores by 0.1 points despite slightly increasing training perplexity.
Gradient Behavior and Why Cross-Entropy Works
Understanding why cross-entropy produces good gradients requires examining what happens when we differentiate through the softmax and cross-entropy combination. This is one of those places where the mathematical elegance of the formulation reveals why certain design choices are not arbitrary but deeply principled.
The Softmax-Cross-Entropy Gradient
Let be the unnormalized logit for class , and let be the predicted probability after normalization. The gradient of the cross-entropy loss with respect to the logit is:
where:
- : the gradient of the loss with respect to logit
- : the model's predicted probability for class
- : the true label for class (1 for the true class, 0 otherwise)
This remarkably simple result comes from the cancellation of terms when differentiating through softmax and logarithm together. For the true class (where ), the gradient is : a negative quantity that pushes the logit upward, increasing the true class probability. For all other classes (where ), the gradient is : a positive quantity that pushes those logits downward, reducing their probability. The gradient magnitude is proportional to the prediction error: when the model is right with high confidence (), gradients are near zero. When the model is wrong with high confidence (the model assigned probability near 1 to the wrong class), gradients are large.
This gradient formula has another important property: it is bounded. No matter how wrong the model is, the gradient of cross-entropy with respect to logits is at most 1.0 in magnitude. Compare this to squared error loss, where gradients grow without bound as predictions diverge from targets. The bounded gradients of cross-entropy make training more stable and reduce the need for gradient clipping in practice. As discussed in Gradient Clipping, managing gradient magnitude is important for stable transformer training, and the naturally bounded logit gradients from cross-entropy help.
Why Cross-Entropy Beats Mean Squared Error for Classification
A natural question arises: why not use mean squared error (MSE) as a classification loss? MSE between predicted probabilities and one-hot targets is:
The gradient of MSE with respect to logits, after accounting for the softmax, involves terms like . When the model is very wrong (say, for the true class), this gradient includes the factor , which nearly vanishes. This is the saturation problem: MSE produces tiny gradients exactly when the model is most wrong and needs correction most urgently. Cross-entropy avoids this problem because its gradient is simply , which remains large when the model is wrong. This is why cross-entropy consistently outperforms MSE for training classifiers, and why you will never see MSE used as the training objective for a language model.
Training Dynamics and Loss Curves
Cross-entropy loss evolves characteristically during language model training. Understanding these patterns helps diagnose training issues, identify when models have converged, and distinguish between healthy optimization and pathological behavior such as overfitting or instability.
Typical Loss Curve Patterns
Training cross-entropy typically follows a three-phase pattern:
-
Rapid initial decrease: The model quickly learns basic statistical patterns: unigram frequencies and simple n-gram statistics. Loss may drop from an initial value near (where is vocabulary size) down to roughly half that value within the first few hundred steps. During this phase, the model discovers that certain tokens (like common punctuation or high-frequency words) appear frequently regardless of context.
-
Gradual improvement: The model learns syntactic patterns, semantic relationships, and longer-range dependencies. Loss decreases slowly but steadily, often following a power law relationship with training compute, as described in Kaplan Scaling Laws. This phase involves learning grammatical rules, semantic categories, and contextual representations that require processing many examples.
-
Plateau: Eventually, loss improvements diminish as the model approaches the entropy of the training data. Further gains require exponentially more compute or data. At this stage, the model has extracted most of the predictable structure from the training distribution, and remaining errors often reflect inherent ambiguity in language or noise in the training data.
The initial loss value provides a useful sanity check during training. For a model with vocabulary size and random initialization, the first prediction should assign roughly uniform probability to each token, yielding initial cross-entropy . For a vocabulary of 50,000 tokens, this is approximately 10.8 nats. If your initial loss differs dramatically from this value, something is wrong with the initialization, the data preprocessing, or the model architecture.
Validation Loss and Overfitting
While training loss typically decreases monotonically (especially with large datasets), validation loss reveals model generalization:
- Healthy training: Training and validation loss decrease together, maintaining a small gap. This indicates the model is learning generalizable patterns rather than memorizing specific training examples.
- Overfitting onset: Validation loss stops decreasing while training loss continues to fall. This signals that the model has begun memorizing training-specific noise or idiosyncrasies rather than learning general linguistic principles.
- Divergence: Validation loss increases while training loss decreases, showing memorization rather than generalization. The model is essentially becoming a lookup table for the training set rather than a compressor of linguistic patterns.
Modern large language models often train for only a single epoch or use repeated data with Causal Language Modeling objectives, making overfitting less apparent than in traditional supervised learning. However, validation loss on held-out corpora remains important for detecting distribution shift or data quality issues. If validation loss suddenly spikes while training loss remains stable, this may indicate that the validation set differs systematically from the training distribution or contains preprocessing errors.
Loss Spikes and Training Instability
A common pattern in large-scale language model training is loss spikes: sudden increases in cross-entropy that may or may not recover. These spikes typically result from:
- Gradient accumulation errors: Numerical precision issues when accumulating gradients over many steps, particularly in mixed-precision training
- Outlier batches: Training examples that are unusually long, contain rare tokens, or have atypical structure
- Learning rate schedule artifacts: Transitions in the learning rate schedule, particularly during warmup or when switching from warmup to decay
- Data quality issues: Corrupted examples or formatting errors in the training corpus that cause the model to receive contradictory gradient signals
When a loss spike occurs, the model often recovers automatically over the next few hundred steps as the optimizer's momentum smooths out the perturbation. However, severe spikes may require restarting from an earlier checkpoint with reduced learning rate. Monitoring the gradient norm alongside cross-entropy loss helps distinguish spikes caused by exploding gradients (requiring gradient clipping) from spikes caused by data issues (requiring data inspection).
Learning Rate Effects
The learning rate schedule strongly affects cross-entropy trajectories:
- High learning rates: Cause loss spikes and instability when gradients explode. The optimization process may overshoot minima in the loss landscape or bounce between valleys without settling.
- Warmup: Gradually increasing learning rates from near-zero prevents early training instability. This technique allows the model to move carefully through the initially chaotic loss landscape before taking larger steps once it has found a reasonable basin of attraction.
- Decay: Cosine or linear decay schedules help loss converge to lower final values than constant rates. As the model approaches convergence, smaller step sizes allow for fine-grained exploration of the local minimum.
The Adam Optimizer and AdamW variants adapt per-parameter learning rates, typically creating smoother loss curves than vanilla stochastic gradient descent. These adaptive methods maintain separate learning rates for each parameter based on the historical magnitude of gradients, letting some parameters to update quickly while others change slowly.
Worked Example
Let's walk through a concrete calculation to solidify these concepts. Consider a tiny language model with a vocabulary of four tokens: {the, cat, sat, mat}. The model predicts the next token given the context "the cat". This simplified scenario allows us to trace through the calculations manually and verify that our mathematical understanding matches our computational implementation.
Token Probability Distribution
Suppose our model outputs the following probabilities for the next token:
| Token | Model Probability | True? |
|---|---|---|
the | 0.1 | No |
cat | 0.05 | No |
sat | 0.7 | Yes |
mat | 0.15 | No |
The true next token is sat.

The true next token is sat. The cross-entropy loss is simply:
where:
- : the cross-entropy loss for this single prediction
- : the model's predicted probability for the true token (
sat)
Converting to bits (dividing by ):
where:
- : the cross-entropy measured in bits
- : the model's predicted probability for the true token
The perplexity is:
where:
- : the exponential function with base
- : the cross-entropy in nats from the previous calculation
Or in base-2:
where:
- : the base of the logarithm (corresponding to bits)
- : the cross-entropy in bits from the previous calculation
This low perplexity indicates the model is quite confident and correct. If the model had predicted uniform probabilities (0.25 each):
where:
- : the cross-entropy loss for uniform prediction
- : the uniform probability for each of the 4 tokens
where:
- : the calculated perplexity value, equal to the vocabulary size for uniform distribution
This matches our intuition: uniform random guessing among 4 items yields perplexity 4. The model with 0.7 probability on the correct answer is effectively as uncertain as if it were choosing among only 1.43 equally likely options. This shows its high confidence.
Sequence-Level Calculation
Now consider a complete sentence: "the cat sat mat" (we omit the final punctuation for simplicity). The model processes this autoregressively, conditioning each prediction on all previous tokens:
- Context:
<bos>, True:the, Prob: 0.4, Loss: - Context:
the, True:cat, Prob: 0.3, Loss: - Context:
the cat, True:sat, Prob: 0.7, Loss: - Context:
the cat sat, True:mat, Prob: 0.5, Loss:
The average cross-entropy is:
where:
- : the average cross-entropy loss over the sequence
- : the individual token losses (negative log probabilities) for each position
- : the number of tokens in the sequence
Notice how token 2 ("cat") has the highest individual loss despite a moderate probability of 0.3. This reflects higher uncertainty at that position: many words could follow "the" in English. Token 3 ("sat") benefits from richer context and receives a much better prediction. This position-by-position view illustrates how the context window gradually reduces uncertainty as the model accumulates information.
The perplexity is:
where:
- : the exponential function with base
- : the average cross-entropy in nats
If the text contains 15 characters (including spaces), the bits per character would be:
where:
- : bits per character, the compression efficiency metric
- : the average cross-entropy loss in nats per token
- : the number of tokens in the sequence
- : the conversion factor from nats to bits ()
- : the total number of characters in the original text
Note the conversion factor: 1 nat = bits. This conversion is necessary because our initial cross-entropy was computed using natural logarithms, but BPC is traditionally reported in bits.
Code Implementation
Let's implement cross-entropy calculations from scratch and using standard libraries. We'll demonstrate the relationship between manual calculation, PyTorch's implementation, and the conversion to perplexity.
import numpy as np
import torch
# Set seed for reproducibility
np.random.seed(42)
torch.manual_seed(42)We'll start with a basic implementation using NumPy to understand the mechanics, then show the optimized PyTorch version used in actual training loops.
def cross_entropy_manual(y_true_idx, y_pred_probs, eps=1e-12):
"""
Manual cross-entropy calculation.
Args:
y_true_idx: Integer index of true class
y_pred_probs: Array of predicted probabilities (must sum to 1)
eps: Small constant to avoid log(0)
Returns:
Cross-entropy in nats (natural log)
"""
# Clip probabilities to avoid log(0)
y_pred_clipped = np.clip(y_pred_probs, eps, 1.0 - eps)
# Get probability assigned to true class
prob_true = y_pred_clipped[y_true_idx]
# Return negative log probability
return -np.log(prob_true)
def perplexity_from_entropy(cross_entropy, base="e"):
"""Convert cross-entropy to perplexity."""
if base == "2":
return 2**cross_entropy
else:
return np.exp(cross_entropy)Now let's test with our worked example from earlier:
Vocabulary: ['the', 'cat', 'sat', 'mat']
True token: 'sat' (index 2)
Model probabilities: {'the': np.float64(0.1), 'cat': np.float64(0.05), 'sat': np.float64(0.7), 'mat': np.float64(0.15)}
Cross-entropy: 0.3567 nats (0.5146 bits)
Perplexity: 1.4286
Uniform distribution cross-entropy: 1.3863 nats
Uniform perplexity: 4.0000 (should equal vocab size 4)The manual calculation confirms our mathematical understanding. Now let's implement sequence-level evaluation using PyTorch, which handles batched operations efficiently:
def evaluate_sequence(model_probs, true_indices):
"""
Calculate cross-entropy and perplexity for a sequence.
Args:
model_probs: Tensor of shape (seq_len, vocab_size) with probability distributions
true_indices: Tensor of shape (seq_len,) with true token indices
Returns:
Dictionary with cross_entropy (nats), bits_per_token, perplexity, and bpc
"""
seq_len = len(true_indices)
# Extract probabilities of true tokens
true_probs = model_probs[torch.arange(seq_len), true_indices]
# Cross-entropy (negative log likelihood averaged over sequence)
ce = -torch.log(true_probs).mean()
# Convert to bits
ce_bits = ce / torch.log(torch.tensor(2.0))
# Perplexity
perp = torch.exp(ce)
return {
"cross_entropy_nats": ce.item(),
"cross_entropy_bits": ce_bits.item(),
"perplexity": perp.item(),
}Let's simulate a sequence evaluation with our example sentence "the cat sat mat":
Sequence evaluation results: Cross-entropy: 0.7925 nats Cross-entropy: 1.1434 bits/token Perplexity: 2.2090 Bits per character: 0.3049 (assuming 15 chars)
PyTorch provides optimized implementations that handle numerical stability through log-softmax and avoid explicit probability calculations when working with logits directly:
import torch.nn.functional as F
# PyTorch's cross_entropy function operates on logits, not probabilities
logits = torch.log(
probs_tensor + 1e-12
) # Convert back to log space for demonstration
# F.cross_entropy expects (N, C) for logits and (N,) for targets
# It combines log_softmax + nll_loss efficiently
pytorch_ce = F.cross_entropy(logits, true_indices)PyTorch cross-entropy: 0.7925 Matches manual calculation: True
The close agreement between PyTorch's optimized implementation and our manual calculation confirms that both approaches correctly compute cross-entropy. The tiny numerical differences (if any) stem from floating-point precision variations in how probabilities are handled. This validation gives us confidence that our manual implementation correctly implements the mathematical definition, while PyTorch's version offers superior numerical stability and computational efficiency for production use.
Visualizing Loss Curves
Let's simulate training curves to demonstrate typical cross-entropy dynamics:
import numpy as np
def simulate_training_curve(
steps=1000, initial_loss=8.0, final_loss=2.5, noise=0.05
):
"""
Simulate realistic training and validation loss curves.
"""
# Training loss: smooth exponential decay with noise
t = np.linspace(0, 1, steps)
train_loss = final_loss + (initial_loss - final_loss) * np.exp(-5 * t)
train_loss += np.random.normal(0, noise, steps) * (
1 - 0.8 * t
) # Decreasing noise
# Validation loss: similar but with gap and potential overfitting at end
val_loss = (
train_loss + 0.3 + 0.2 * np.maximum(0, t - 0.7) ** 2
) # Gap increases slightly at end
val_loss += np.random.normal(0, noise * 1.5, steps)
return train_loss, val_loss
train_loss, val_loss = simulate_training_curve()
steps = np.arange(len(train_loss))
The plot demonstrates how cross-entropy decreases during training while maintaining the exponential relationship with perplexity shown on the secondary axis. Notice how validation loss typically runs 10-30% higher than training loss due to the model not having seen that specific data during optimization.
Comparing Different Model Sizes
Let's examine how cross-entropy scales with model capacity, which reflects the power laws discussed in Power Laws in Deep Learning:
Model Size vs Cross-Entropy (simulated): -------------------------------------------------- 0.1B params: CE = 6.185 nats, Perplexity = 485.3 0.3B params: CE = 5.623 nats, Perplexity = 276.7 0.8B params: CE = 5.301 nats, Perplexity = 200.6 1.3B params: CE = 5.089 nats, Perplexity = 162.3 2.7B params: CE = 4.814 nats, Perplexity = 123.3 6.7B params: CE = 4.493 nats, Perplexity = 89.4 13.0B params: CE = 4.272 nats, Perplexity = 71.7 175.0B params: CE = 3.506 nats, Perplexity = 33.3

This demonstrates the predictable relationship between model scale and cross-entropy: with the simulated exponent above, doubling model parameters yields a constant reduction in loss of about 5% in the relevant regime.
Numerical Stability in Practice
When implementing cross-entropy in real training systems, numerical stability requires careful attention. The naive formulation involves taking logarithms of very small probabilities, which can produce infinities or cause gradient explosions. Modern deep learning frameworks address this through a combination of mathematical reformulations and careful floating-point management.
The Log-Sum-Exp Trick
The standard cross-entropy computation involves first computing softmax probabilities, then taking logarithms. This two-step process is numerically unstable because softmax can produce probabilities that underflow to zero or overflow to infinity for logits with large magnitudes. The solution is the log-sum-exp trick, which computes in a single numerically stable step:
where:
- : the log-probability for class (what we need for cross-entropy)
- : the raw logit for class
- : the log-partition function, computed stably as where
By subtracting the maximum logit before exponentiating, we ensure that at least one term equals , and all other terms are at most 1. This prevents overflow while preserving the relative differences between logits that determine the probability distribution. PyTorch's F.cross_entropy and F.log_softmax implement this automatically.
Mixed Precision Considerations
Modern language model training uses mixed precision (float16 or bfloat16) to reduce memory usage and increase throughput. However, cross-entropy loss computation is often kept in float32 precision even when the forward pass uses float16. This is because float16 has limited range (approximately ) and limited precision (about 3 decimal digits), which can cause loss values to become NaN or Inf during computation. Frameworks like PyTorch's automatic mixed precision (AMP) automatically identify loss computation as a sensitivity point and upcasts to float32 for that portion of the computation, maintaining stability without sacrificing efficiency elsewhere.
Implementation Parameters
When implementing cross-entropy loss in practice, several parameters control numerical stability and behavior. Each choice reflects a tradeoff between computational efficiency, numerical safety, and training quality.
The eps parameter represents a small constant added to probabilities to prevent numerical errors from log(0). Typical values range from to , as shown in our manual implementation where we clip probabilities before taking the logarithm. Without this protection, a model that accidentally assigns zero probability to the true token, either due to numerical underflow or an unfortunate initialization, would produce infinite loss and destroy the training process through exploding gradients. In practice, when operating on logits rather than probabilities (as PyTorch does internally), the log-sum-exp trick makes explicit clipping unnecessary, but the epsilon remains important in custom implementations.
The reduction parameter specifies how to aggregate losses across the batch. The options are 'mean' (average across all elements), 'sum' (total loss), and 'none' (per-element loss without reduction). Our sequence evaluation uses mean reduction to normalize by sequence length. This keeps longer sequences do not dominate the loss disproportionately. When training with gradient accumulation over multiple micro-batches before a weight update, the choice between mean and sum reduction affects whether you need to scale the loss by the number of accumulation steps. Mean reduction is typically safer because it produces loss values that are comparable across different batch sizes. This makes hyperparameter transfer between experiments more reliable.
The from_logits convention (called weight_softmax in some frameworks) indicates whether inputs are raw logits or probabilities. Operating directly on logits, as PyTorch's F.cross_entropy does internally, improves numerical stability by avoiding explicit exponentiation and division in a separate softmax step. When operating on logits, the function applies log-softmax internally in a numerically stable way using the log-sum-exp trick described above.
The label_smoothing parameter takes a float between 0 and 1 that distributes a small amount of probability mass to all classes, preventing overconfidence and improving generalization. Typical values range from 0.01 to 0.1. Higher values provide stronger regularization but may slow convergence and prevent the model from achieving peak accuracy on the training set. The optimal value depends on the task: generation tasks often use smaller values (0.01-0.05) while classification fine-tuning tasks may benefit from larger values (0.05-0.1). Label smoothing is particularly beneficial when the training data contains labeling noise, as it prevents the model from memorizing noisy labels with extreme confidence.
The ignore_index parameter (common in PyTorch) specifies a token index whose loss should be excluded from the computation. In language modeling, this is typically the padding token used to fill sequences to a uniform length within a batch. Without ignoring padding tokens, the model would receive spurious gradient signals pushing it to assign high probability to whatever padding token follows the actual end of each sequence. Setting ignore_index correctly is needed for correct training with variable-length sequences.
Limitations and Impact
Cross-entropy drives modern language model training, but understanding its limitations helps you interpret evaluation results and design better systems. Recognizing what cross-entropy cannot measure is as important as understanding what it can measure.
Semantic Blindness
Cross-entropy is fundamentally a statistical measure that cares only about probability assignment, not semantic content. A model that predicts "The feline rested" instead of "The cat sat" receives the same penalty as predicting a completely unrelated word, even though the first substitution preserves meaning while the second destroys it. Both errors result in identical cross-entropy if the predicted probabilities are the same, despite the first being a reasonable paraphrase and the second being nonsensical.
This limitation becomes most apparent when evaluating generative models. A model that produces a perfect paraphrase of the expected output gets penalized as if it had generated gibberish, because cross-entropy compares against a single reference answer. This is the basic tension in language model evaluation: we train models to maximize probability of specific observed sequences, but we evaluate (and ultimately care about) much broader notions of quality including coherence, factual accuracy, style, and helpfulness.
This limitation motivates the sequence-level evaluation metrics we'll explore in upcoming chapters on BLEU Score and ROUGE Scores, which measure n-gram overlap with reference texts to capture semantic similarity. However, these metrics introduce their own biases, favoring conservative, extractive generation over creative paraphrasing. Neither cross-entropy nor n-gram overlap perfectly captures the fine-grained quality of generated text.
Calibration and Confidence
Minimizing cross-entropy does not guarantee well-calibrated probabilities. A model can achieve low cross-entropy while being overconfident on correct predictions and underconfident on errors. This manifests in generation when sampling strategies like Nucleus Sampling rely on probability rankings that may not reflect true likelihoods. A model might assign 99% probability to its top choice even when that choice is wrong half the time, leading to overconfident generation of hallucinations or factually incorrect text.
Calibration errors often appear as systematic biases tied to training data patterns. If the training corpus contains many confident declarative statements and few expressions of uncertainty, the model may learn to be overconfident in its outputs regardless of the actual uncertainty of a given prediction. This is particularly problematic for factual questions: a model may output false facts with the same high confidence as true facts, because the training objective never required distinguishing confident knowledge from uncertain speculation.
Techniques like temperature scaling and label smoothing address calibration to some degree, but the basic disconnect between cross-entropy optimization and calibrated uncertainty remains a challenge, particularly for safety-necessary applications where probability estimates inform decision-making. A medical diagnosis system or legal analysis tool must provide well-calibrated confidence estimates, not just low cross-entropy on training text.
Context Length Dependencies
Cross-entropy averages uniformly across positions, treating the first token of a document identically to the thousandth. However, language modeling difficulty varies dramatically with position:
- Initial tokens: High entropy, little context to constrain predictions. The model must predict the first word of a document with essentially no prior information.
- Middle tokens: Lower entropy, rich preceding context. The model can rely on established topic and grammatical structure.
- Long-range dependencies: Cross-entropy may underweight rare but necessary long-range dependencies that require maintaining coherence across paragraphs. A document-level consistency error might affect only a few tokens but represents a serious quality failure.
Modern architectures like those discussed in Long Context sections address this through position interpolation and attention mechanisms, but the uniform averaging of cross-entropy loss does not explicitly prioritize maintaining coherence over arbitrary distances. The loss function treats a pronoun resolution error at position 500 as equally important as a common word prediction error at position 5, even though the former may indicate a more serious failure of understanding.
An important practical consequence is that cross-entropy measured on the first 512 tokens of documents may not reflect model quality on long documents. A model that achieves low perplexity on short contexts may still struggle with long-range coherence, narrative consistency, or multi-step reasoning across extended passages. This limitation motivates dedicated evaluation benchmarks that specifically test long-context understanding, rather than relying solely on standard perplexity metrics that average over all positions equally.
Training Data Dependence
Cross-entropy is computed against the specific tokens observed in training data, which means it depends heavily on the particular tokenization, normalization, and formatting choices applied to the corpus. Two models trained on different tokenizations of the same underlying text may achieve very different cross-entropy scores while capturing equivalent linguistic knowledge. A model trained on text with aggressive lowercasing and punctuation removal will achieve lower cross-entropy on similarly preprocessed test data, but may generate outputs that look different from those of a model trained on the original casing and punctuation.
This dependence also means that cross-entropy conflates learning linguistic patterns with learning the specific stylistic conventions of the training corpus. A model trained primarily on formal text will have lower cross-entropy on formal text than on informal text, even if the formal and informal corpora represent the same underlying language at the semantic level. This is one reason why domain adaptation remains important: fine-tuning on in-domain text reduces cross-entropy for that domain by learning domain-specific vocabulary, facts, and stylistic conventions.
The Optimization-Evaluation Gap
During training, we optimize cross-entropy on next-token prediction. During evaluation for generation tasks, we care about sequence quality, coherence, and factual accuracy. This creates a mismatch between the training objective and downstream utility. We train models to predict the next token, but we evaluate them on their ability to generate helpful, harmless, and honest responses.
This gap becomes stark in instruction-following scenarios. A model trained on web text achieves low cross-entropy by predicting how web text continues, but web text rarely contains the kind of helpful, well-structured responses to explicit questions that we want from an assistant. The model learns the statistical patterns of its training distribution, but that distribution may not align well with the target use case. This explains why instruction fine-tuning and RLHF are necessary: they bridge the gap between the statistical objective (minimize cross-entropy on training data) and the behavioral objective (produce responses that humans find helpful and trustworthy).
Techniques like Reinforcement Learning from Human Feedback address this gap by optimizing for human preferences rather than pure likelihood, but cross-entropy remains the foundational pre-training objective that endows models with linguistic knowledge before alignment tuning. Understanding this gap explains why models can achieve excellent cross-entropy scores on benchmarks while still creating unsatisfactory outputs in practice, and why fine-tuning on human preferences often degrades perplexity while improving utility. A model that learns to say "I'm not sure about that" instead of confidently generating a plausible-sounding but incorrect answer will have higher cross-entropy (since the training data contains many confident assertions) but better real-world reliability.
Summary
Cross-entropy loss provides the mathematical foundation for training and evaluating language models, connecting information theory with practical optimization. Key insights to remember:
The core definition: cross-entropy measures the average bits needed to encode data from distribution using a code optimized for distribution :
where:
- : the cross-entropy between the true distribution and the predicted distribution
- : an outcome from the set of possible events
- : the probability of outcome under the true distribution
- : the probability of outcome under the model distribution
The key connections and implications:
- Relationship to perplexity: Perplexity equals the exponential of cross-entropy (), converting additive information measures into multiplicative branching factors that are easier to interpret intuitively.
- Equivalence to MLE: Minimizing cross-entropy maximizes the likelihood of training data, which makes it equivalent to maximum likelihood estimation under the model's parameterization and giving it a strong statistical justification.
- Gradient behavior: The gradient of the combined softmax-cross-entropy with respect to logits simplifies to , a bounded and interpretable signal that is large when the model is wrong and small when it is right.
- Bits per character: Normalizing by character count rather than token count allows fair comparison across different tokenization strategies, with state-of-the-art models achieving approximately 0.7-0.9 BPC on English text.
- Loss curves: Training dynamics typically show rapid initial improvement followed by gradual convergence following power-law relationships with compute and data.
- Numerical stability: Real implementations use the log-sum-exp trick and operate on logits rather than probabilities to avoid overflow and underflow in floating-point computation.
Cross-entropy is the universal training signal for Causal Language Modeling and Masked Language Modeling, but remember that low cross-entropy does not guarantee high-quality generation, factual accuracy, or aligned behavior. In the following chapters, we'll explore complementary metrics like BLEU and ROUGE that better capture generation quality, as well as benchmarks that evaluate capabilities beyond next-token prediction. The limitations of cross-entropy are as important to understand as its strengths: knowing what the training objective cannot measure helps us design better evaluation strategies and recognize when a model's low perplexity score might not translate to practical value.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about cross-entropy loss, from information theory foundations to practical training dynamics.
Cross-Entropy Loss Fundamentals
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!