Cross-Entropy Loss: Information Theory for LLM Training

Michael BrenndoerferFebruary 21, 202655 min read

Part of Language AI Handbook

Covers cross-entropy loss for language models. Topics include entropy, KL divergence, perplexity.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Cross-Entropy Loss

Language models learn by predicting the next token in a sequence. But how do we quantify the gap between a model's predictions and reality? We need a mathematical measure that can compare the probability distribution output by our model against the actual observed data. This provides a single scalar value that indicates how well our predictions align with truth. Cross-entropy loss provides the mathematical foundation for this measurement, serving as both the training objective that guides gradient descent through the parameter space and the evaluation metric that tells us how well our model understands the statistical patterns of language.

As we explored in our discussion of Perplexity, language models output probability distributions over possible next tokens. At each position in a sequence, the model produces a vector of probabilities representing its confidence that each vocabulary item will appear next. Cross-entropy gives us a principled way to measure how far these predicted distributions diverge from the true distributions we observe in training data. It connects information theory, probability, and optimization into a single coherent framework that drives modern language AI. This framework allows us to transform the abstract problem of learning language patterns into a concrete optimization task where we can compute gradients and update model weights systematically.

Unlike perplexity, which expresses model uncertainty as an intuitive "effective vocabulary size," cross-entropy operates in the domain of information theory, measuring the average number of bits needed to encode the true data using a code optimized for the model's predicted distribution. This perspective reveals basic limits on compression and provides the optimization target for virtually all modern language models, from the causal objectives discussed in Causal Language Modeling to the masked objectives in Masked Language Modeling. Understanding cross-entropy provides deep insight into why language models work, how they improve during training, and what basic limits constrain their performance.

One reason cross-entropy occupies such a central position in language model training is that it connects naturally to maximum likelihood estimation, the workhorse of classical statistics. When you minimize cross-entropy loss over a dataset, you are simultaneously maximizing the probability that the model assigns to the observed training tokens. This equivalence means cross-entropy is not an arbitrary choice of loss function imposed by convention. It is the theoretically principled answer to the question of how to measure the quality of a probabilistic model. Later in this chapter, we will make this connection precise and see why it justifies using cross-entropy as the objective for training neural language models of any size.

Information Theory Foundations

To understand cross-entropy, we must first grasp entropy and how it quantifies uncertainty. The concept of entropy originates from thermodynamics and statistical mechanics, but Claude Shannon adapted it for information theory in 1948. This provides the mathematical vocabulary for measuring information content. Shannon's framework allows us to quantify how much information is contained in a random event, which directly translates to how surprised we should be when observing different outcomes. For language modeling, this means measuring how unexpected each token is given our current understanding of the language's statistical structure.

Shannon's key insight was that information and uncertainty are two sides of the same coin. An event carries more information precisely because it was unexpected. When a coin lands heads after you knew it would land heads because it is double-headed, you learn nothing. When that same coin lands heads despite being weighted toward tails, you learn something meaningful. This seemingly philosophical distinction has concrete mathematical consequences that turn out to be exactly what we need to train language models.

Shannon Entropy

Entropy measures the uncertainty inherent in a probability distribution. It answers the question: on average, how surprised will we be when we observe an outcome from this distribution? For a discrete distribution PP over outcomes xx, entropy is defined as:

H(P)=−∑xP(x)log⁡P(x)H(P) = -\sum_{x} P(x) \log P(x)

where:

  • H(P)H(P): the entropy of probability distribution PP, measured in bits (base-2 log) or nats (natural log)
  • xx: a possible outcome from the set of all events
  • P(x)P(x): the probability of outcome xx occurring
  • log⁡\log: the logarithm function (typically base 2 for bits or natural log for nats)

The logarithm ensures that rare events contribute more to the total uncertainty than common ones. This mathematical property aligns with our intuition: when an unlikely event occurs, we experience more surprise and receive more information than when a common event occurs. When an event is certain (P(x)=1P(x) = 1), its contribution is zero because we gain no new information from observing something we already knew would happen. When an event is impossible (P(x)=0P(x) = 0), we define 0⋅log⁡0=00 \cdot \log 0 = 0 by continuity. This reflects that impossible events contribute nothing to the entropy since they never occur.

To see why the logarithm specifically encodes surprise, consider two independent events each with probability pp. The probability of both occurring together is p2p^2. If surprise is additive (seeing two independent rare events is twice as surprising as seeing one), then the measure of surprise should satisfy f(p2)=2f(p)f(p^2) = 2f(p). The only continuous function satisfying this property for all probabilities is the logarithm. So −log⁡P(x)-\log P(x) is the only mathematically consistent measure of surprise for an event with probability P(x)P(x), and entropy is the expected value of this surprise over all possible outcomes.

Entropy as Expected Information

Entropy represents the expected information content or "surprise" of observing a random variable. A uniform distribution over NN outcomes has maximum entropy log⁡N\log N, while a deterministic distribution has zero entropy. This means that when all outcomes are equally likely, we are maximally uncertain about what will happen next, whereas when one outcome is guaranteed, we have perfect certainty and zero uncertainty. From a compression perspective, entropy is the basic lower bound on how many bits per symbol any lossless compression algorithm can achieve when encoding data drawn from distribution PP.

Consider a vocabulary of 50,000 tokens. A uniform distribution where every token has probability 1/50,0001/50,000 yields entropy:

H=−∑i=150000150000log⁡2150000=log⁡250000≈15.6 bits\begin{aligned} H &= -\sum_{i=1}^{50000} \frac{1}{50000} \log_2 \frac{1}{50000} \\ &= \log_2 50000 \\ &\approx 15.6 \text{ bits} \end{aligned}

This means we need approximately 15.6 bits on average to encode each token if we use an optimal code designed for this uniform distribution. This represents the worst-case scenario for compression, where no token is more predictable than any other. In contrast, if our model is perfect and assigns probability 1 to the correct token and 0 elsewhere, entropy drops to zero, requiring no bits to transmit the message because there is no uncertainty. The difference between these extremes represents the potential for compression that good language models exploit.

Real language is far from uniform. Common words like "the" and "is" appear far more frequently than rare technical terms. Good models learn these frequency differences and assign higher probability to common tokens, reducing average surprise below the maximum. The entropy of real English text has been estimated at 0.6 to 1.2 bits per character depending on context length, far below the 4.7 bits per character that would be required for a uniform distribution over the 26 letters and a space character. This gap is what enables compression algorithms, and it is the gap that language models are trained to exploit.

Out[3]:
Visualization
Line chart of binary entropy function peaking at 0.693 nats when p equals 0.5 and approaching zero at both extremes.
Binary entropy function showing how uncertainty varies with probability. Maximum entropy of approximately 0.693 nats occurs at p=0.5, corresponding to the uniform distribution over two outcomes. Entropy approaches zero at both extremes as the distribution becomes more deterministic, illustrating that predictable outcomes carry less information than surprising ones.

Kullback-Leibler Divergence

While entropy measures uncertainty within a single distribution, the Kullback-Leibler (KL) divergence measures how one probability distribution QQ diverges from a reference distribution PP. It quantifies the inefficiency of assuming that the distribution is QQ when the true distribution is PP:

DKL(P∥Q)=∑xP(x)log⁡P(x)Q(x)D_{KL}(P \| Q) = \sum_{x} P(x) \log \frac{P(x)}{Q(x)}

where:

  • DKL(P∥Q)D_{KL}(P \| Q): the Kullback-Leibler divergence from distribution QQ to reference distribution PP
  • P(x)P(x): the probability of outcome xx under the true distribution PP
  • Q(x)Q(x): the probability of outcome xx under the model distribution QQ
  • xx: an outcome from the support of PP

KL divergence is always non-negative and equals zero if and only if P=QP = Q everywhere. This non-negativity follows from Jensen's inequality applied to the concave logarithm function. It represents the extra bits required when using a code optimized for QQ to encode data that follows distribution PP. This interpretation reveals that KL divergence measures the penalty we pay for using the wrong model: if we build a compression scheme based on our predicted distribution QQ, but reality follows PP, we will waste bits proportional to the KL divergence.

The asymmetry of KL divergence is worth understanding carefully. DKL(P∥Q)D_{KL}(P \| Q) is not the same as DKL(Q∥P)D_{KL}(Q \| P) in general. When we minimize DKL(P∥Q)D_{KL}(P \| Q) with respect to QQ, we penalize cases where P(x)>0P(x) > 0 but Q(x)≈0Q(x) \approx 0. This is called the "forward" or "inclusive" KL, and it encourages QQ to cover all regions where PP has mass. When we minimize DKL(Q∥P)D_{KL}(Q \| P) (the "reverse" or "exclusive" KL), we penalize cases where Q(x)>0Q(x) > 0 but P(x)≈0P(x) \approx 0, which encourages QQ to concentrate on modes of PP. Language model training uses the forward KL form through cross-entropy, which means models are penalized for assigning low probability to tokens that occur in training data, encouraging broad coverage of plausible continuations.

However, KL divergence has an asymmetry that matters for machine learning: it requires knowing the true distribution PP over all possible outcomes. In practice, we only observe specific outcomes from PP, not the full distribution. We see individual tokens in our training data, but we do not know the complete probability distribution over all possible next tokens for every context. This limitation prevents us from computing the KL divergence directly during training, necessitating the use of cross-entropy instead.

Defining Cross-Entropy

Cross-entropy combines the entropy of the true distribution with the KL divergence between true and predicted distributions into a single quantity that we can compute from data. For distributions PP (true) and QQ (model prediction), cross-entropy is:

H(P,Q)=−∑xP(x)log⁡Q(x)H(P, Q) = -\sum_{x} P(x) \log Q(x)

where:

  • H(P,Q)H(P, Q): the cross-entropy between the true distribution PP and the predicted distribution QQ
  • P(x)P(x): the probability of outcome xx under the true distribution (empirical distribution from data)
  • Q(x)Q(x): the probability of outcome xx assigned by the model
  • xx: an outcome from the set of possible events

When PP represents the empirical distribution from training data, P(x)P(x) is 1 for the actual observed token and 0 for all others. This simplifies cross-entropy dramatically because the sum collapses to a single term involving only the true token:

H(P,Q)=−∑xP(x)log⁡Q(x)=−log⁡Q(xtrue)\begin{aligned} H(P, Q) &= -\sum_{x} P(x) \log Q(x) \\ &= -\log Q(x_{\text{true}}) \end{aligned}

where:

  • xtruex_{\text{true}}: the observed token (the single outcome where P(x)=1P(x) = 1)
  • Q(xtrue)Q(x_{\text{true}}): the probability assigned by the model to the true token

This is the negative log-likelihood of the true token under the model's predicted distribution. It penalizes the model proportionally to how surprised it was by the actual observation: if the model assigns high probability to the correct token, the loss is small; if the model assigns low probability, the loss is large, approaching infinity as the predicted probability approaches zero.

The simplification from the full sum to a single term has an important practical consequence. During training, we do not need to compare probability distributions element by element across the entire vocabulary. We only need to look up how much probability the model assigned to the one token that appeared next. This makes gradient computation efficient: backpropagation only needs to update the output neuron corresponding to the true token, though intermediate layers receive gradients from all vocabulary items through the softmax normalization.

Out[4]:
Visualization
Line chart showing cross-entropy loss in nats and bits rising steeply as predicted probability approaches zero and falling toward zero as probability approaches one.
Cross-entropy loss as a function of predicted probability for the true class. The loss decreases rapidly as model confidence increases, approaching zero for perfect prediction, and rises steeply to infinity as the predicted probability approaches zero. The steep left tail creates strong gradient signals that push the model to avoid confidently wrong predictions.
Cross-Entropy Decomposition

Cross-entropy decomposes into entropy plus KL divergence:

H(P,Q)=H(P)+DKL(P∥Q)H(P, Q) = H(P) + D_{KL}(P \| Q)

where:

  • H(P,Q)H(P, Q): the cross-entropy between the true distribution PP and the predicted distribution QQ
  • H(P)H(P): the entropy of the true distribution (constant with respect to model parameters)
  • DKL(P∥Q)D_{KL}(P \| Q): the KL divergence from the model distribution to the true distribution

Since H(P)H(P) is constant with respect to model parameters, minimizing cross-entropy is equivalent to minimizing KL divergence. This decomposition reveals that cross-entropy captures both the inherent uncertainty in the data (the entropy term) and the additional penalty for using a suboptimal model (the KL divergence term). During optimization, the entropy term acts as a constant offset, so gradient descent on cross-entropy effectively minimizes the KL divergence between our model and reality. The lowest achievable cross-entropy on any dataset equals the true entropy of the language, which represents the irreducible uncertainty stemming from inherent ambiguity in text.

Cross-Entropy as Maximum Likelihood Estimation

The connection between cross-entropy and maximum likelihood estimation (MLE) explains the statistical basis for using cross-entropy to train probabilistic language models. Given a dataset of NN observed tokens {x1,x2,…,xN}\{x_1, x_2, \ldots, x_N\}, MLE seeks the model parameters θ\theta that maximize the probability of the observed data:

θ^=arg⁡max⁡θ∏i=1NP(xi∣contexti;θ)\hat{\theta} = \arg\max_{\theta} \prod_{i=1}^{N} P(x_i | \text{context}_i; \theta)

where:

  • θ^\hat{\theta}: the maximum likelihood estimate of the model parameters
  • P(xi∣contexti;θ)P(x_i | \text{context}_i; \theta): the model's probability of the ii-th observed token given its context
  • ∏\prod: the product over all tokens in the dataset

Because products of many small probabilities rapidly underflow to zero numerically, we take the logarithm of both sides. This converts the product into a sum, which is numerically stable:

θ^=arg⁡max⁡θ∑i=1Nlog⁡P(xi∣contexti;θ)=arg⁡min⁡θ−1N∑i=1Nlog⁡P(xi∣contexti;θ)\begin{aligned} \hat{\theta} &= \arg\max_{\theta} \sum_{i=1}^{N} \log P(x_i | \text{context}_i; \theta) \\ &= \arg\min_{\theta} -\frac{1}{N} \sum_{i=1}^{N} \log P(x_i | \text{context}_i; \theta) \end{aligned}

where:

  • ∑i=1Nlog⁡P(xi∣contexti;θ)\sum_{i=1}^{N} \log P(x_i | \text{context}_i; \theta): the log-likelihood of the observed tokens
  • −1N∑i=1Nlog⁡P(xi∣contexti;θ)-\frac{1}{N} \sum_{i=1}^{N} \log P(x_i | \text{context}_i; \theta): the negative log-likelihood, normalized by dataset size

The final expression is exactly the cross-entropy loss averaged over the training set. Minimizing cross-entropy is mathematically identical to performing maximum likelihood estimation. This means we have a strong statistical justification for using cross-entropy: it is the objective that, as NN grows large, causes our model to converge to the true underlying data distribution (or the closest approximation within our model family). Any other loss function would not have this convergence guarantee in the same form.

The Language Modeling Objective

In autoregressive language modeling, we predict the next token xtx_t given previous tokens x<tx_{<t}. This sequential factorization allows us to model complex joint distributions over sequences by breaking them down into manageable conditional probabilities. The probability of a sequence factorizes as:

P(x1,x2,…,xT)=∏t=1TP(xt∣x<t)P(x_1, x_2, \ldots, x_T) = \prod_{t=1}^{T} P(x_t | x_{<t})

where:

  • P(x1,x2,…,xT)P(x_1, x_2, \ldots, x_T): the joint probability of the entire sequence
  • P(xt∣x<t)P(x_t | x_{<t}): the conditional probability of token xtx_t given all preceding tokens x<tx_{<t}

This factorization is exact: the chain rule of probability guarantees it without any independence assumptions. The cross-entropy loss for a sequence of length TT is:

L=−1T∑t=1Tlog⁡P(xt∣x<t;θ)\mathcal{L} = -\frac{1}{T} \sum_{t=1}^{T} \log P(x_t | x_{<t}; \theta)

where:

  • L\mathcal{L}: the average cross-entropy loss over the sequence
  • TT: the length of the target sequence (number of tokens)
  • xtx_t: the token at position tt in the sequence
  • x<tx_{<t}: all tokens preceding position tt (the context)
  • θ\theta: the model parameters
  • P(xt∣x<t;θ)P(x_t | x_{<t}; \theta): the model's predicted probability of token xtx_t given previous tokens

This formulation reveals why cross-entropy is called the log-loss: we sum the negative log probabilities of each token. Lower probability assigned to the correct token results in higher loss, with the loss approaching infinity as the predicted probability approaches zero. This creates strong gradients when the model is confidently wrong, pushing the parameters aggressively to correct major errors.

The averaging by TT (sequence length) ensures that longer sequences do not automatically incur higher loss simply because they contain more tokens. This normalization allows us to compare loss values across sequences of different lengths and to aggregate loss across batches during training. Without this normalization, a long document would always appear to be modeled worse than a short sentence, even if the model's per-token accuracy was identical.

A necessary practical point: in transformer training, the model processes an entire sequence in a single forward pass, computing predicted probability distributions at every position simultaneously. For a sequence of length TT, this produces TT losses in parallel. The positions at the beginning of the sequence receive less context than positions near the end, so the model must predict early tokens with limited information. This variability in difficulty across positions is averaged away by the uniform weighting in the loss, which treats all positions equally. Some researchers have proposed weighting later positions more heavily, since they benefit from richer context and thus provide cleaner learning signal, but this remains an active area of investigation.

Cross-Entropy and Perplexity

As established in our discussion of Perplexity, these two metrics are intimately related through a simple mathematical transformation. Understanding this relationship helps you interpret evaluation results and compare models across different papers and frameworks, as they measure the same underlying quantity but present it in different units that emphasize different aspects of model performance.

The Mathematical Relationship

Perplexity is defined as the exponential of cross-entropy:

Perplexity=exp⁡(H(P,Q))=exp⁡(−1N∑i=1Nlog⁡P(xi))\begin{aligned} \text{Perplexity} &= \exp(H(P, Q)) \\ &= \exp\left(-\frac{1}{N} \sum_{i=1}^{N} \log P(x_i)\right) \end{aligned}

where:

  • Perplexity\text{Perplexity}: the effective branching factor or exponential of the cross-entropy
  • H(P,Q)H(P, Q): the cross-entropy between true and predicted distributions
  • NN: the total number of tokens in the evaluation set
  • P(xi)P(x_i): the model's predicted probability of the ii-th true token xix_i
  • exp⁡\exp: the exponential function (using the same base as the logarithm in cross-entropy)

The exponential uses the same base as the logarithm in the cross-entropy calculation. In natural language processing:

  • When using natural logarithm (ln⁡\ln) for cross-entropy, perplexity uses ee as the base
  • When using base-2 logarithm (log⁡2\log_2) for cross-entropy, perplexity uses 2 as the base

Both formulations measure the same underlying quantity, just on different scales. Cross-entropy in nats (natural log) or bits (log⁡2\log_2) provides an additive measure of uncertainty, while perplexity provides a multiplicative measure that can be interpreted as the effective branching factor at each prediction step. If your model has a perplexity of 100, it is as uncertain as if it were choosing uniformly among 100 equally likely options at each step.

When to Use Each Metric

Cross-entropy offers advantages during training and optimization:

  • Differentiability: Cross-entropy is directly differentiable and is the loss function for gradient computation. We can backpropagate through the cross-entropy calculation to update model weights.
  • Additivity: Cross-entropy losses add across tokens, making aggregation straightforward. We can sum losses across batches and sequences without complications.
  • Numerical stability: Working in log space avoids underflow when multiplying many small probabilities. Summing logs is numerically safer than multiplying probabilities directly.

Perplexity offers advantages for human interpretation:

  • Scale intuition: A perplexity of 100 is more meaningful than a cross-entropy of 4.6 nats. Humans can intuitively grasp the difference between perplexity 50 and 100 better than between 3.9 and 4.6 nats.
  • Vocabulary-relative performance: Perplexity naturally scales with vocabulary size. This allows rough comparison across models with different vocabularies. A perplexity equal to vocabulary size indicates random guessing. This provides a clear baseline.
  • Historical convention: Perplexity remains the standard reporting metric for language model benchmarks. Research papers and leaderboards universally report perplexity, which makes it necessary for comparing with prior work.

In practice, modern frameworks often report both: cross-entropy for training curves and model selection, perplexity for final evaluation and paper results. During development, you might monitor cross-entropy to catch training instabilities early, then report perplexity in your final results to communicate with the broader research community.

Bits Per Character

While cross-entropy measures bits per token, bits per character (BPC) provides a tokenization-agnostic metric for comparing model efficiency. Since different tokenizers produce different sequence lengths for the same text, BPC normalizes by the number of characters in the original text. This normalization is needed for fair comparison because a model using a more aggressive tokenizer (with larger tokens) will naturally achieve lower per-token perplexity than an equivalent model using character-level tokenization, even if both models capture the same underlying linguistic patterns.

Calculating BPC

Given a sequence with CC characters and cross-entropy loss HH (in bits), bits per character is:

BPC=H×TC\text{BPC} = \frac{H \times T}{C}

where:

  • BPC\text{BPC}: bits per character, measuring compression efficiency independent of tokenization
  • HH: the cross-entropy loss in bits (not nats)
  • TT: the number of tokens in the sequence
  • CC: the number of characters in the original text

For character-level language models where T=CT = C, BPC equals the cross-entropy. For subword tokenizers where T<CT < C (multiple characters per token), we must account for the compression ratio. The multiplication by TT converts the per-token loss to total bits for the sequence, while division by CC spreads those bits across the original characters.

BPC and Compression

BPC represents a basic limit on lossless compression. If a language model achieves 1.0 BPC on English text, it means no compression algorithm can compress English below 1 bit per character on average, assuming the model perfectly captures the true distribution of the language. This connects language modeling to data compression theory, where better language models yield better compression algorithms. In fact, arithmetic coding with a language model as the probability estimator is theoretically optimal: as the model improves, the compressed file size approaches the entropy of the text.

Current state-of-the-art language models achieve approximately 0.7-0.9 BPC on standard English text corpora, approaching but not reaching the estimated entropy of English (typically estimated between 0.6-1.0 BPC depending on the domain and estimation method). This gap suggests there remains room for improvement in language modeling, though diminishing returns may make further advances increasingly difficult.

Character-Level vs Subword Evaluation

When comparing models using different tokenization strategies, BPC provides fairer comparison than per-token metrics:

  • Character-level models: Directly report BPC as the loss, since each token corresponds to one character.
  • Subword models: Must divide total bits by character count to enable fair comparison with character-level models.
  • Byte-level models: Similar to character-level but operating on UTF-8 bytes, where one character may consist of multiple bytes.

This normalization is important when evaluating Byte Pair Encoding or SentencePiece tokenizers against character-level baselines, as the same model architecture can appear to have wildly different per-token perplexities solely due to tokenization differences. Without BPC normalization, you might conclude that a subword model is dramatically better than a character model, when in fact the improvement comes primarily from the tokenizer grouping characters into larger units rather than from better linguistic understanding.

Binary vs Categorical Cross-Entropy

Language modeling uses categorical cross-entropy because we select among many discrete classes (vocabulary tokens). However, understanding binary cross-entropy illuminates special cases and alternative formulations that appear in auxiliary training objectives and discriminative fine-tuning tasks.

Categorical Cross-Entropy

For multi-class classification with KK classes, where the model outputs a probability distribution y^\hat{y} and the true class is represented as a one-hot vector yy:

L=−∑k=1Kyklog⁡(y^k)\mathcal{L} = -\sum_{k=1}^{K} y_k \log(\hat{y}_k)

where:

  • L\mathcal{L}: the cross-entropy loss for a single prediction
  • KK: the total number of classes (vocabulary size)
  • yky_k: the true probability of class kk (1 for the correct class, 0 otherwise in one-hot encoding)
  • y^k\hat{y}_k: the model's predicted probability for class kk

Since yy is one-hot (all zeros except for the true class where ytrue=1y_{\text{true}} = 1), the sum collapses to a single term:

L=−∑k=1Kyklog⁡(y^k)=−log⁡(y^true)\begin{aligned} \mathcal{L} &= -\sum_{k=1}^{K} y_k \log(\hat{y}_k) \\ &= -\log(\hat{y}_{\text{true}}) \end{aligned}

where:

  • y^true\hat{y}_{\text{true}}: the model's predicted probability for the true class (the single class where yk=1y_k = 1)

In language modeling, KK equals the vocabulary size (often 32,000 to 100,000 tokens). The Softmax Function ensures y^\hat{y} sums to 1 across all vocabulary items. This formulation handles the extreme multi-class nature of language modeling, where we must discriminate among tens of thousands of possible next tokens at each step.

Binary Cross-Entropy

For binary classification or independent multi-label problems:

L=−[ylog⁡(y^)+(1−y)log⁡(1−y^)]\mathcal{L} = -\left[y \log(\hat{y}) + (1-y) \log(1-\hat{y})\right]

where:

  • L\mathcal{L}: the binary cross-entropy loss
  • yy: the true binary label (0 or 1)
  • y^\hat{y}: the model's predicted probability that the label is 1
  • 1−y1-y: the probability that the label is 0
  • 1−y^1-\hat{y}: the model's predicted probability that the label is 0

Language models rarely use binary cross-entropy directly for next-token prediction, but it appears in specialized objectives like Replaced Token Detection, where the model predicts whether each token is original or replaced. In this case, the task reduces to binary classification for each position independently. The discriminator component of ELECTRA uses this formulation to learn by distinguishing real tokens from synthetically corrupted ones. This provides an alternative to traditional masked language modeling.

Binary cross-entropy also appears in contrastive learning objectives and ranking tasks. When a model must choose between two candidate continuations, the relative log-probabilities determine which is preferred, and binary cross-entropy over these preferences can serve as a training signal. This pattern appears in reinforcement learning from human feedback, where the reward model learns to predict human preferences between pairs of model outputs.

Label Smoothing

A practical modification to cross-entropy prevents overconfidence. Instead of using hard targets (0 or 1), label smoothing distributes a small amount of probability mass ϵ\epsilon across all classes:

L=−∑k=1K[(1−ϵ)yk+ϵK]log⁡(y^k)\mathcal{L} = -\sum_{k=1}^{K} \left[(1-\epsilon) y_k + \frac{\epsilon}{K}\right] \log(\hat{y}_k)

where:

  • ϵ\epsilon: the smoothing parameter (typically 0.01 to 0.1), representing the amount of probability mass to distribute uniformly
  • (1−ϵ)yk(1-\epsilon) y_k: the adjusted target probability for the true class (reduced from 1)
  • ϵK\frac{\epsilon}{K}: the uniform smoothing term added to all classes
  • y^k\hat{y}_k: the model's predicted probability for class kk

This prevents the model from becoming overconfident and improves generalization by so gradients continue to flow even when the model predicts the correct class with high probability. Without smoothing, once a model achieves near-perfect prediction for a training example, gradients vanish and learning stops for that example. With smoothing, the model continues to learn, adjusting its internal representations to better model the underlying distribution rather than merely memorizing training labels. Typical values for ϵ\epsilon range from 0.01 to 0.1, with higher values giving more regularization but potentially slowing convergence.

The effect of label smoothing can be understood from the decomposition perspective. Smoothing replaces the hard empirical distribution (all mass on one token) with a mixture: (1−ϵ)(1-\epsilon) of the mass stays on the true token, and ϵ\epsilon is spread uniformly. This makes the target distribution softer, requiring the model to assign non-negligible probability to all vocabulary items rather than collapsing its distribution to a single peak. In practice, label smoothing consistently improves downstream task performance and was used in the original Transformer paper, where it improved translation BLEU scores by 0.1 points despite slightly increasing training perplexity.

Gradient Behavior and Why Cross-Entropy Works

Understanding why cross-entropy produces good gradients requires examining what happens when we differentiate through the softmax and cross-entropy combination. This is one of those places where the mathematical elegance of the formulation reveals why certain design choices are not arbitrary but deeply principled.

The Softmax-Cross-Entropy Gradient

Let zkz_k be the unnormalized logit for class kk, and let y^k=softmax(zk)\hat{y}_k = \text{softmax}(z_k) be the predicted probability after normalization. The gradient of the cross-entropy loss with respect to the logit zjz_j is:

∂L∂zj=y^j−yj\frac{\partial \mathcal{L}}{\partial z_j} = \hat{y}_j - y_j

where:

  • ∂L∂zj\frac{\partial \mathcal{L}}{\partial z_j}: the gradient of the loss with respect to logit zjz_j
  • y^j\hat{y}_j: the model's predicted probability for class jj
  • yjy_j: the true label for class jj (1 for the true class, 0 otherwise)

This remarkably simple result comes from the cancellation of terms when differentiating through softmax and logarithm together. For the true class (where yj=1y_j = 1), the gradient is y^j−1\hat{y}_j - 1: a negative quantity that pushes the logit upward, increasing the true class probability. For all other classes (where yj=0y_j = 0), the gradient is y^j\hat{y}_j: a positive quantity that pushes those logits downward, reducing their probability. The gradient magnitude is proportional to the prediction error: when the model is right with high confidence (y^j≈1\hat{y}_j \approx 1), gradients are near zero. When the model is wrong with high confidence (the model assigned probability near 1 to the wrong class), gradients are large.

This gradient formula has another important property: it is bounded. No matter how wrong the model is, the gradient of cross-entropy with respect to logits is at most 1.0 in magnitude. Compare this to squared error loss, where gradients grow without bound as predictions diverge from targets. The bounded gradients of cross-entropy make training more stable and reduce the need for gradient clipping in practice. As discussed in Gradient Clipping, managing gradient magnitude is important for stable transformer training, and the naturally bounded logit gradients from cross-entropy help.

Why Cross-Entropy Beats Mean Squared Error for Classification

A natural question arises: why not use mean squared error (MSE) as a classification loss? MSE between predicted probabilities and one-hot targets is:

LMSE=1K∑k=1K(yk−y^k)2\mathcal{L}_{\text{MSE}} = \frac{1}{K} \sum_{k=1}^{K} (y_k - \hat{y}_k)^2

The gradient of MSE with respect to logits, after accounting for the softmax, involves terms like (y^j−yj)⋅y^j(1−y^j)(\hat{y}_j - y_j) \cdot \hat{y}_j(1 - \hat{y}_j). When the model is very wrong (say, y^j≈0\hat{y}_j \approx 0 for the true class), this gradient includes the factor y^j(1−y^j)≈0\hat{y}_j(1 - \hat{y}_j) \approx 0, which nearly vanishes. This is the saturation problem: MSE produces tiny gradients exactly when the model is most wrong and needs correction most urgently. Cross-entropy avoids this problem because its gradient is simply y^j−yj\hat{y}_j - y_j, which remains large when the model is wrong. This is why cross-entropy consistently outperforms MSE for training classifiers, and why you will never see MSE used as the training objective for a language model.

Training Dynamics and Loss Curves

Cross-entropy loss evolves characteristically during language model training. Understanding these patterns helps diagnose training issues, identify when models have converged, and distinguish between healthy optimization and pathological behavior such as overfitting or instability.

Typical Loss Curve Patterns

Training cross-entropy typically follows a three-phase pattern:

  1. Rapid initial decrease: The model quickly learns basic statistical patterns: unigram frequencies and simple n-gram statistics. Loss may drop from an initial value near log⁡V\log V (where VV is vocabulary size) down to roughly half that value within the first few hundred steps. During this phase, the model discovers that certain tokens (like common punctuation or high-frequency words) appear frequently regardless of context.

  2. Gradual improvement: The model learns syntactic patterns, semantic relationships, and longer-range dependencies. Loss decreases slowly but steadily, often following a power law relationship with training compute, as described in Kaplan Scaling Laws. This phase involves learning grammatical rules, semantic categories, and contextual representations that require processing many examples.

  3. Plateau: Eventually, loss improvements diminish as the model approaches the entropy of the training data. Further gains require exponentially more compute or data. At this stage, the model has extracted most of the predictable structure from the training distribution, and remaining errors often reflect inherent ambiguity in language or noise in the training data.

The initial loss value provides a useful sanity check during training. For a model with vocabulary size VV and random initialization, the first prediction should assign roughly uniform probability 1/V1/V to each token, yielding initial cross-entropy ≈log⁡V\approx \log V. For a vocabulary of 50,000 tokens, this is approximately 10.8 nats. If your initial loss differs dramatically from this value, something is wrong with the initialization, the data preprocessing, or the model architecture.

Validation Loss and Overfitting

While training loss typically decreases monotonically (especially with large datasets), validation loss reveals model generalization:

  • Healthy training: Training and validation loss decrease together, maintaining a small gap. This indicates the model is learning generalizable patterns rather than memorizing specific training examples.
  • Overfitting onset: Validation loss stops decreasing while training loss continues to fall. This signals that the model has begun memorizing training-specific noise or idiosyncrasies rather than learning general linguistic principles.
  • Divergence: Validation loss increases while training loss decreases, showing memorization rather than generalization. The model is essentially becoming a lookup table for the training set rather than a compressor of linguistic patterns.

Modern large language models often train for only a single epoch or use repeated data with Causal Language Modeling objectives, making overfitting less apparent than in traditional supervised learning. However, validation loss on held-out corpora remains important for detecting distribution shift or data quality issues. If validation loss suddenly spikes while training loss remains stable, this may indicate that the validation set differs systematically from the training distribution or contains preprocessing errors.

Loss Spikes and Training Instability

A common pattern in large-scale language model training is loss spikes: sudden increases in cross-entropy that may or may not recover. These spikes typically result from:

  • Gradient accumulation errors: Numerical precision issues when accumulating gradients over many steps, particularly in mixed-precision training
  • Outlier batches: Training examples that are unusually long, contain rare tokens, or have atypical structure
  • Learning rate schedule artifacts: Transitions in the learning rate schedule, particularly during warmup or when switching from warmup to decay
  • Data quality issues: Corrupted examples or formatting errors in the training corpus that cause the model to receive contradictory gradient signals

When a loss spike occurs, the model often recovers automatically over the next few hundred steps as the optimizer's momentum smooths out the perturbation. However, severe spikes may require restarting from an earlier checkpoint with reduced learning rate. Monitoring the gradient norm alongside cross-entropy loss helps distinguish spikes caused by exploding gradients (requiring gradient clipping) from spikes caused by data issues (requiring data inspection).

Learning Rate Effects

The learning rate schedule strongly affects cross-entropy trajectories:

  • High learning rates: Cause loss spikes and instability when gradients explode. The optimization process may overshoot minima in the loss landscape or bounce between valleys without settling.
  • Warmup: Gradually increasing learning rates from near-zero prevents early training instability. This technique allows the model to move carefully through the initially chaotic loss landscape before taking larger steps once it has found a reasonable basin of attraction.
  • Decay: Cosine or linear decay schedules help loss converge to lower final values than constant rates. As the model approaches convergence, smaller step sizes allow for fine-grained exploration of the local minimum.

The Adam Optimizer and AdamW variants adapt per-parameter learning rates, typically creating smoother loss curves than vanilla stochastic gradient descent. These adaptive methods maintain separate learning rates for each parameter based on the historical magnitude of gradients, letting some parameters to update quickly while others change slowly.

Worked Example

Let's walk through a concrete calculation to solidify these concepts. Consider a tiny language model with a vocabulary of four tokens: {the, cat, sat, mat}. The model predicts the next token given the context "the cat". This simplified scenario allows us to trace through the calculations manually and verify that our mathematical understanding matches our computational implementation.

Token Probability Distribution

Suppose our model outputs the following probabilities for the next token:

Token probability distribution for the context "the cat".
TokenModel ProbabilityTrue?
the0.1No
cat0.05No
sat0.7Yes
mat0.15No

The true next token is sat.

Out[5]:
Visualization
Bar chart showing predicted probabilities for four tokens: the at 0.10, cat at 0.05, sat at 0.70 highlighted in teal, and mat at 0.15, with a coral dashed line at uniform probability 0.25.
Predicted probability distribution over the vocabulary for the context 'the cat'. The model assigns highest probability (0.70) to the correct token 'sat' (shown in teal), with remaining probability spread across 'the' (0.10), 'mat' (0.15), and 'cat' (0.05). The coral dashed line marks uniform probability of 0.25, illustrating how the model concentrates mass on the most likely continuation.

The true next token is sat. The cross-entropy loss is simply:

H=−log⁡(0.7)≈0.357 nats\begin{aligned} H &= -\log(0.7) \\ &\approx 0.357 \text{ nats} \end{aligned}

where:

  • HH: the cross-entropy loss for this single prediction
  • 0.70.7: the model's predicted probability for the true token (sat)

Converting to bits (dividing by ln⁡2\ln 2):

Hbits=−log⁡2(0.7)≈0.515 bits\begin{aligned} H_{\text{bits}} &= -\log_2(0.7) \\ &\approx 0.515 \text{ bits} \end{aligned}

where:

  • HbitsH_{\text{bits}}: the cross-entropy measured in bits
  • 0.70.7: the model's predicted probability for the true token

The perplexity is:

Perplexity=exp⁡(0.357)≈1.43\begin{aligned} \text{Perplexity} &= \exp(0.357) \\ &\approx 1.43 \end{aligned}

where:

  • exp⁡\exp: the exponential function with base ee
  • 0.3570.357: the cross-entropy in nats from the previous calculation

Or in base-2:

Perplexity=20.515≈1.43\begin{aligned} \text{Perplexity} &= 2^{0.515} \\ &\approx 1.43 \end{aligned}

where:

  • 22: the base of the logarithm (corresponding to bits)
  • 0.5150.515: the cross-entropy in bits from the previous calculation

This low perplexity indicates the model is quite confident and correct. If the model had predicted uniform probabilities (0.25 each):

H=−log⁡(0.25)=ln⁡4≈1.386 nats\begin{aligned} H &= -\log(0.25) \\ &= \ln 4 \\ &\approx 1.386 \text{ nats} \end{aligned}

where:

  • HH: the cross-entropy loss for uniform prediction
  • 0.250.25: the uniform probability for each of the 4 tokens
Perplexity=4\text{Perplexity} = 4

where:

  • 44: the calculated perplexity value, equal to the vocabulary size for uniform distribution

This matches our intuition: uniform random guessing among 4 items yields perplexity 4. The model with 0.7 probability on the correct answer is effectively as uncertain as if it were choosing among only 1.43 equally likely options. This shows its high confidence.

Sequence-Level Calculation

Now consider a complete sentence: "the cat sat mat" (we omit the final punctuation for simplicity). The model processes this autoregressively, conditioning each prediction on all previous tokens:

  1. Context: <bos>, True: the, Prob: 0.4, Loss: −ln⁡(0.4)≈0.916-\ln(0.4) \approx 0.916
  2. Context: the, True: cat, Prob: 0.3, Loss: −ln⁡(0.3)≈1.204-\ln(0.3) \approx 1.204
  3. Context: the cat, True: sat, Prob: 0.7, Loss: −ln⁡(0.7)≈0.357-\ln(0.7) \approx 0.357
  4. Context: the cat sat, True: mat, Prob: 0.5, Loss: −ln⁡(0.5)≈0.693-\ln(0.5) \approx 0.693

The average cross-entropy is:

L=0.916+1.204+0.357+0.6934≈0.793 nats\begin{aligned} \mathcal{L} &= \frac{0.916 + 1.204 + 0.357 + 0.693}{4} \\ &\approx 0.793 \text{ nats} \end{aligned}

where:

  • L\mathcal{L}: the average cross-entropy loss over the sequence
  • 0.916,1.204,0.357,0.6930.916, 1.204, 0.357, 0.693: the individual token losses (negative log probabilities) for each position
  • 44: the number of tokens in the sequence

Notice how token 2 ("cat") has the highest individual loss despite a moderate probability of 0.3. This reflects higher uncertainty at that position: many words could follow "the" in English. Token 3 ("sat") benefits from richer context and receives a much better prediction. This position-by-position view illustrates how the context window gradually reduces uncertainty as the model accumulates information.

The perplexity is:

Perplexity=exp⁡(0.793)≈2.21\begin{aligned} \text{Perplexity} &= \exp(0.793) \\ &\approx 2.21 \end{aligned}

where:

  • exp⁡\exp: the exponential function with base ee
  • 0.7930.793: the average cross-entropy in nats

If the text contains 15 characters (including spaces), the bits per character would be:

BPC=0.793×4 nats×1.443 bits/nat15 chars≈0.305 BPC\begin{aligned} \text{BPC} &= \frac{0.793 \times 4 \text{ nats} \times 1.443 \text{ bits/nat}}{15 \text{ chars}} \\ &\approx 0.305 \text{ BPC} \end{aligned}

where:

  • BPC\text{BPC}: bits per character, the compression efficiency metric
  • 0.7930.793: the average cross-entropy loss in nats per token
  • 44: the number of tokens in the sequence
  • 1.4431.443: the conversion factor from nats to bits (1/ln⁡21/\ln 2)
  • 1515: the total number of characters in the original text

Note the conversion factor: 1 nat = 1ln⁡2≈1.443\frac{1}{\ln 2} \approx 1.443 bits. This conversion is necessary because our initial cross-entropy was computed using natural logarithms, but BPC is traditionally reported in bits.

Code Implementation

Let's implement cross-entropy calculations from scratch and using standard libraries. We'll demonstrate the relationship between manual calculation, PyTorch's implementation, and the conversion to perplexity.

In[6]:
Code
import numpy as np
import torch

# Set seed for reproducibility
np.random.seed(42)
torch.manual_seed(42)

We'll start with a basic implementation using NumPy to understand the mechanics, then show the optimized PyTorch version used in actual training loops.

In[7]:
Code
def cross_entropy_manual(y_true_idx, y_pred_probs, eps=1e-12):
    """
    Manual cross-entropy calculation.

    Args:
        y_true_idx: Integer index of true class
        y_pred_probs: Array of predicted probabilities (must sum to 1)
        eps: Small constant to avoid log(0)

    Returns:
        Cross-entropy in nats (natural log)
    """
    # Clip probabilities to avoid log(0)
    y_pred_clipped = np.clip(y_pred_probs, eps, 1.0 - eps)

    # Get probability assigned to true class
    prob_true = y_pred_clipped[y_true_idx]

    # Return negative log probability
    return -np.log(prob_true)


def perplexity_from_entropy(cross_entropy, base="e"):
    """Convert cross-entropy to perplexity."""
    if base == "2":
        return 2**cross_entropy
    else:
        return np.exp(cross_entropy)

Now let's test with our worked example from earlier:

Out[8]:
Console
Vocabulary: ['the', 'cat', 'sat', 'mat']
True token: 'sat' (index 2)
Model probabilities: {'the': np.float64(0.1), 'cat': np.float64(0.05), 'sat': np.float64(0.7), 'mat': np.float64(0.15)}

Cross-entropy: 0.3567 nats (0.5146 bits)
Perplexity: 1.4286

Uniform distribution cross-entropy: 1.3863 nats
Uniform perplexity: 4.0000 (should equal vocab size 4)

The manual calculation confirms our mathematical understanding. Now let's implement sequence-level evaluation using PyTorch, which handles batched operations efficiently:

In[9]:
Code
def evaluate_sequence(model_probs, true_indices):
    """
    Calculate cross-entropy and perplexity for a sequence.

    Args:
        model_probs: Tensor of shape (seq_len, vocab_size) with probability distributions
        true_indices: Tensor of shape (seq_len,) with true token indices

    Returns:
        Dictionary with cross_entropy (nats), bits_per_token, perplexity, and bpc
    """
    seq_len = len(true_indices)

    # Extract probabilities of true tokens
    true_probs = model_probs[torch.arange(seq_len), true_indices]

    # Cross-entropy (negative log likelihood averaged over sequence)
    ce = -torch.log(true_probs).mean()

    # Convert to bits
    ce_bits = ce / torch.log(torch.tensor(2.0))

    # Perplexity
    perp = torch.exp(ce)

    return {
        "cross_entropy_nats": ce.item(),
        "cross_entropy_bits": ce_bits.item(),
        "perplexity": perp.item(),
    }

Let's simulate a sequence evaluation with our example sentence "the cat sat mat":

Out[10]:
Console
Sequence evaluation results:
  Cross-entropy: 0.7925 nats
  Cross-entropy: 1.1434 bits/token
  Perplexity: 2.2090
  Bits per character: 0.3049 (assuming 15 chars)

PyTorch provides optimized implementations that handle numerical stability through log-softmax and avoid explicit probability calculations when working with logits directly:

In[11]:
Code
import torch.nn.functional as F

# PyTorch's cross_entropy function operates on logits, not probabilities
logits = torch.log(
    probs_tensor + 1e-12
)  # Convert back to log space for demonstration

# F.cross_entropy expects (N, C) for logits and (N,) for targets
# It combines log_softmax + nll_loss efficiently
pytorch_ce = F.cross_entropy(logits, true_indices)
Out[12]:
Console
PyTorch cross-entropy: 0.7925
Matches manual calculation: True

The close agreement between PyTorch's optimized implementation and our manual calculation confirms that both approaches correctly compute cross-entropy. The tiny numerical differences (if any) stem from floating-point precision variations in how probabilities are handled. This validation gives us confidence that our manual implementation correctly implements the mathematical definition, while PyTorch's version offers superior numerical stability and computational efficiency for production use.

Visualizing Loss Curves

Let's simulate training curves to demonstrate typical cross-entropy dynamics:

In[13]:
Code
import numpy as np


def simulate_training_curve(
    steps=1000, initial_loss=8.0, final_loss=2.5, noise=0.05
):
    """
    Simulate realistic training and validation loss curves.
    """
    # Training loss: smooth exponential decay with noise
    t = np.linspace(0, 1, steps)
    train_loss = final_loss + (initial_loss - final_loss) * np.exp(-5 * t)
    train_loss += np.random.normal(0, noise, steps) * (
        1 - 0.8 * t
    )  # Decreasing noise

    # Validation loss: similar but with gap and potential overfitting at end
    val_loss = (
        train_loss + 0.3 + 0.2 * np.maximum(0, t - 0.7) ** 2
    )  # Gap increases slightly at end
    val_loss += np.random.normal(0, noise * 1.5, steps)

    return train_loss, val_loss


train_loss, val_loss = simulate_training_curve()
steps = np.arange(len(train_loss))
Out[14]:
Visualization
Line plot of training and validation cross-entropy loss decreasing over 1000 training steps with a secondary perplexity axis on the right.
Simulated training dynamics showing cross-entropy loss decrease over 1000 steps. Training loss (solid) decays smoothly from roughly 8.0 nats toward a plateau near 2.5 nats, while validation loss (dashed) tracks slightly higher with increased variance throughout. The secondary right axis shows the corresponding perplexity values, illustrating how the exponential relationship compresses large perplexity differences at the start of training into a narrow band as the model improves.

The plot demonstrates how cross-entropy decreases during training while maintaining the exponential relationship with perplexity shown on the secondary axis. Notice how validation loss typically runs 10-30% higher than training loss due to the model not having seen that specific data during optimization.

Comparing Different Model Sizes

Let's examine how cross-entropy scales with model capacity, which reflects the power laws discussed in Power Laws in Deep Learning:

Out[15]:
Console
Model Size vs Cross-Entropy (simulated):
--------------------------------------------------
   0.1B params: CE = 6.185 nats, Perplexity = 485.3
   0.3B params: CE = 5.623 nats, Perplexity = 276.7
   0.8B params: CE = 5.301 nats, Perplexity = 200.6
   1.3B params: CE = 5.089 nats, Perplexity = 162.3
   2.7B params: CE = 4.814 nats, Perplexity = 123.3
   6.7B params: CE = 4.493 nats, Perplexity = 89.4
  13.0B params: CE = 4.272 nats, Perplexity = 71.7
 175.0B params: CE = 3.506 nats, Perplexity = 33.3
Out[16]:
Visualization
Scatter plot on a log x-axis showing simulated cross-entropy loss decreasing from about 6.2 nats at 100 million parameters to 3.5 nats at 175 billion parameters along a power-law curve.
Power-law scaling of cross-entropy loss with model parameter count, following Kaplan et al. scaling law predictions. In this simulation, loss decreases predictably from about 6.2 nats at 100 million parameters to 3.5 nats at 175 billion parameters. Each doubling of parameters yields a consistent fractional reduction in cross-entropy, illustrating why scaling remains an effective strategy for improving language model performance.

This demonstrates the predictable relationship between model scale and cross-entropy: with the simulated exponent above, doubling model parameters yields a constant reduction in loss of about 5% in the relevant regime.

Numerical Stability in Practice

When implementing cross-entropy in real training systems, numerical stability requires careful attention. The naive formulation involves taking logarithms of very small probabilities, which can produce infinities or cause gradient explosions. Modern deep learning frameworks address this through a combination of mathematical reformulations and careful floating-point management.

The Log-Sum-Exp Trick

The standard cross-entropy computation involves first computing softmax probabilities, then taking logarithms. This two-step process is numerically unstable because softmax can produce probabilities that underflow to zero or overflow to infinity for logits with large magnitudes. The solution is the log-sum-exp trick, which computes log⁡(softmax(z))\log(\text{softmax}(z)) in a single numerically stable step:

log⁡y^k=zk−log⁡∑j=1Kexp⁡(zj)\log \hat{y}_k = z_k - \log \sum_{j=1}^{K} \exp(z_j)

where:

  • log⁡y^k\log \hat{y}_k: the log-probability for class kk (what we need for cross-entropy)
  • zkz_k: the raw logit for class kk
  • log⁡∑j=1Kexp⁡(zj)\log \sum_{j=1}^{K} \exp(z_j): the log-partition function, computed stably as m+log⁡∑jexp⁡(zj−m)m + \log \sum_{j} \exp(z_j - m) where m=max⁡jzjm = \max_j z_j

By subtracting the maximum logit mm before exponentiating, we ensure that at least one term equals exp⁡(0)=1\exp(0) = 1, and all other terms are at most 1. This prevents overflow while preserving the relative differences between logits that determine the probability distribution. PyTorch's F.cross_entropy and F.log_softmax implement this automatically.

Mixed Precision Considerations

Modern language model training uses mixed precision (float16 or bfloat16) to reduce memory usage and increase throughput. However, cross-entropy loss computation is often kept in float32 precision even when the forward pass uses float16. This is because float16 has limited range (approximately ±65,000\pm 65,000) and limited precision (about 3 decimal digits), which can cause loss values to become NaN or Inf during computation. Frameworks like PyTorch's automatic mixed precision (AMP) automatically identify loss computation as a sensitivity point and upcasts to float32 for that portion of the computation, maintaining stability without sacrificing efficiency elsewhere.

Implementation Parameters

When implementing cross-entropy loss in practice, several parameters control numerical stability and behavior. Each choice reflects a tradeoff between computational efficiency, numerical safety, and training quality.

The eps parameter represents a small constant added to probabilities to prevent numerical errors from log(0). Typical values range from 10−1210^{-12} to 10−710^{-7}, as shown in our manual implementation where we clip probabilities before taking the logarithm. Without this protection, a model that accidentally assigns zero probability to the true token, either due to numerical underflow or an unfortunate initialization, would produce infinite loss and destroy the training process through exploding gradients. In practice, when operating on logits rather than probabilities (as PyTorch does internally), the log-sum-exp trick makes explicit clipping unnecessary, but the epsilon remains important in custom implementations.

The reduction parameter specifies how to aggregate losses across the batch. The options are 'mean' (average across all elements), 'sum' (total loss), and 'none' (per-element loss without reduction). Our sequence evaluation uses mean reduction to normalize by sequence length. This keeps longer sequences do not dominate the loss disproportionately. When training with gradient accumulation over multiple micro-batches before a weight update, the choice between mean and sum reduction affects whether you need to scale the loss by the number of accumulation steps. Mean reduction is typically safer because it produces loss values that are comparable across different batch sizes. This makes hyperparameter transfer between experiments more reliable.

The from_logits convention (called weight_softmax in some frameworks) indicates whether inputs are raw logits or probabilities. Operating directly on logits, as PyTorch's F.cross_entropy does internally, improves numerical stability by avoiding explicit exponentiation and division in a separate softmax step. When operating on logits, the function applies log-softmax internally in a numerically stable way using the log-sum-exp trick described above.

The label_smoothing parameter takes a float between 0 and 1 that distributes a small amount of probability mass to all classes, preventing overconfidence and improving generalization. Typical values range from 0.01 to 0.1. Higher values provide stronger regularization but may slow convergence and prevent the model from achieving peak accuracy on the training set. The optimal value depends on the task: generation tasks often use smaller values (0.01-0.05) while classification fine-tuning tasks may benefit from larger values (0.05-0.1). Label smoothing is particularly beneficial when the training data contains labeling noise, as it prevents the model from memorizing noisy labels with extreme confidence.

The ignore_index parameter (common in PyTorch) specifies a token index whose loss should be excluded from the computation. In language modeling, this is typically the padding token used to fill sequences to a uniform length within a batch. Without ignoring padding tokens, the model would receive spurious gradient signals pushing it to assign high probability to whatever padding token follows the actual end of each sequence. Setting ignore_index correctly is needed for correct training with variable-length sequences.

Limitations and Impact

Cross-entropy drives modern language model training, but understanding its limitations helps you interpret evaluation results and design better systems. Recognizing what cross-entropy cannot measure is as important as understanding what it can measure.

Semantic Blindness

Cross-entropy is fundamentally a statistical measure that cares only about probability assignment, not semantic content. A model that predicts "The feline rested" instead of "The cat sat" receives the same penalty as predicting a completely unrelated word, even though the first substitution preserves meaning while the second destroys it. Both errors result in identical cross-entropy if the predicted probabilities are the same, despite the first being a reasonable paraphrase and the second being nonsensical.

This limitation becomes most apparent when evaluating generative models. A model that produces a perfect paraphrase of the expected output gets penalized as if it had generated gibberish, because cross-entropy compares against a single reference answer. This is the basic tension in language model evaluation: we train models to maximize probability of specific observed sequences, but we evaluate (and ultimately care about) much broader notions of quality including coherence, factual accuracy, style, and helpfulness.

This limitation motivates the sequence-level evaluation metrics we'll explore in upcoming chapters on BLEU Score and ROUGE Scores, which measure n-gram overlap with reference texts to capture semantic similarity. However, these metrics introduce their own biases, favoring conservative, extractive generation over creative paraphrasing. Neither cross-entropy nor n-gram overlap perfectly captures the fine-grained quality of generated text.

Calibration and Confidence

Minimizing cross-entropy does not guarantee well-calibrated probabilities. A model can achieve low cross-entropy while being overconfident on correct predictions and underconfident on errors. This manifests in generation when sampling strategies like Nucleus Sampling rely on probability rankings that may not reflect true likelihoods. A model might assign 99% probability to its top choice even when that choice is wrong half the time, leading to overconfident generation of hallucinations or factually incorrect text.

Calibration errors often appear as systematic biases tied to training data patterns. If the training corpus contains many confident declarative statements and few expressions of uncertainty, the model may learn to be overconfident in its outputs regardless of the actual uncertainty of a given prediction. This is particularly problematic for factual questions: a model may output false facts with the same high confidence as true facts, because the training objective never required distinguishing confident knowledge from uncertain speculation.

Techniques like temperature scaling and label smoothing address calibration to some degree, but the basic disconnect between cross-entropy optimization and calibrated uncertainty remains a challenge, particularly for safety-necessary applications where probability estimates inform decision-making. A medical diagnosis system or legal analysis tool must provide well-calibrated confidence estimates, not just low cross-entropy on training text.

Context Length Dependencies

Cross-entropy averages uniformly across positions, treating the first token of a document identically to the thousandth. However, language modeling difficulty varies dramatically with position:

  • Initial tokens: High entropy, little context to constrain predictions. The model must predict the first word of a document with essentially no prior information.
  • Middle tokens: Lower entropy, rich preceding context. The model can rely on established topic and grammatical structure.
  • Long-range dependencies: Cross-entropy may underweight rare but necessary long-range dependencies that require maintaining coherence across paragraphs. A document-level consistency error might affect only a few tokens but represents a serious quality failure.

Modern architectures like those discussed in Long Context sections address this through position interpolation and attention mechanisms, but the uniform averaging of cross-entropy loss does not explicitly prioritize maintaining coherence over arbitrary distances. The loss function treats a pronoun resolution error at position 500 as equally important as a common word prediction error at position 5, even though the former may indicate a more serious failure of understanding.

An important practical consequence is that cross-entropy measured on the first 512 tokens of documents may not reflect model quality on long documents. A model that achieves low perplexity on short contexts may still struggle with long-range coherence, narrative consistency, or multi-step reasoning across extended passages. This limitation motivates dedicated evaluation benchmarks that specifically test long-context understanding, rather than relying solely on standard perplexity metrics that average over all positions equally.

Training Data Dependence

Cross-entropy is computed against the specific tokens observed in training data, which means it depends heavily on the particular tokenization, normalization, and formatting choices applied to the corpus. Two models trained on different tokenizations of the same underlying text may achieve very different cross-entropy scores while capturing equivalent linguistic knowledge. A model trained on text with aggressive lowercasing and punctuation removal will achieve lower cross-entropy on similarly preprocessed test data, but may generate outputs that look different from those of a model trained on the original casing and punctuation.

This dependence also means that cross-entropy conflates learning linguistic patterns with learning the specific stylistic conventions of the training corpus. A model trained primarily on formal text will have lower cross-entropy on formal text than on informal text, even if the formal and informal corpora represent the same underlying language at the semantic level. This is one reason why domain adaptation remains important: fine-tuning on in-domain text reduces cross-entropy for that domain by learning domain-specific vocabulary, facts, and stylistic conventions.

The Optimization-Evaluation Gap

During training, we optimize cross-entropy on next-token prediction. During evaluation for generation tasks, we care about sequence quality, coherence, and factual accuracy. This creates a mismatch between the training objective and downstream utility. We train models to predict the next token, but we evaluate them on their ability to generate helpful, harmless, and honest responses.

This gap becomes stark in instruction-following scenarios. A model trained on web text achieves low cross-entropy by predicting how web text continues, but web text rarely contains the kind of helpful, well-structured responses to explicit questions that we want from an assistant. The model learns the statistical patterns of its training distribution, but that distribution may not align well with the target use case. This explains why instruction fine-tuning and RLHF are necessary: they bridge the gap between the statistical objective (minimize cross-entropy on training data) and the behavioral objective (produce responses that humans find helpful and trustworthy).

Techniques like Reinforcement Learning from Human Feedback address this gap by optimizing for human preferences rather than pure likelihood, but cross-entropy remains the foundational pre-training objective that endows models with linguistic knowledge before alignment tuning. Understanding this gap explains why models can achieve excellent cross-entropy scores on benchmarks while still creating unsatisfactory outputs in practice, and why fine-tuning on human preferences often degrades perplexity while improving utility. A model that learns to say "I'm not sure about that" instead of confidently generating a plausible-sounding but incorrect answer will have higher cross-entropy (since the training data contains many confident assertions) but better real-world reliability.

Summary

Cross-entropy loss provides the mathematical foundation for training and evaluating language models, connecting information theory with practical optimization. Key insights to remember:

The core definition: cross-entropy measures the average bits needed to encode data from distribution PP using a code optimized for distribution QQ:

H(P,Q)=−∑xP(x)log⁡Q(x)H(P, Q) = -\sum_{x} P(x) \log Q(x)

where:

  • H(P,Q)H(P, Q): the cross-entropy between the true distribution PP and the predicted distribution QQ
  • xx: an outcome from the set of possible events
  • P(x)P(x): the probability of outcome xx under the true distribution PP
  • Q(x)Q(x): the probability of outcome xx under the model distribution QQ

The key connections and implications:

  • Relationship to perplexity: Perplexity equals the exponential of cross-entropy (exp⁡(H)\exp(H)), converting additive information measures into multiplicative branching factors that are easier to interpret intuitively.
  • Equivalence to MLE: Minimizing cross-entropy maximizes the likelihood of training data, which makes it equivalent to maximum likelihood estimation under the model's parameterization and giving it a strong statistical justification.
  • Gradient behavior: The gradient of the combined softmax-cross-entropy with respect to logits simplifies to y^j−yj\hat{y}_j - y_j, a bounded and interpretable signal that is large when the model is wrong and small when it is right.
  • Bits per character: Normalizing by character count rather than token count allows fair comparison across different tokenization strategies, with state-of-the-art models achieving approximately 0.7-0.9 BPC on English text.
  • Loss curves: Training dynamics typically show rapid initial improvement followed by gradual convergence following power-law relationships with compute and data.
  • Numerical stability: Real implementations use the log-sum-exp trick and operate on logits rather than probabilities to avoid overflow and underflow in floating-point computation.

Cross-entropy is the universal training signal for Causal Language Modeling and Masked Language Modeling, but remember that low cross-entropy does not guarantee high-quality generation, factual accuracy, or aligned behavior. In the following chapters, we'll explore complementary metrics like BLEU and ROUGE that better capture generation quality, as well as benchmarks that evaluate capabilities beyond next-token prediction. The limitations of cross-entropy are as important to understand as its strengths: knowing what the training objective cannot measure helps us design better evaluation strategies and recognize when a model's low perplexity score might not translate to practical value.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about cross-entropy loss, from information theory foundations to practical training dynamics.

Cross-Entropy Loss Fundamentals

Question 1 of 80 of 8 completed
What does cross-entropy loss fundamentally measure?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026crossentropy, author = {Michael Brenndoerfer}, title = {Cross-Entropy Loss: Information Theory for LLM Training}, year = {2026}, url = {https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Cross-Entropy Loss: Information Theory for LLM Training. Retrieved from https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory
MLAAcademic
Michael Brenndoerfer. "Cross-Entropy Loss: Information Theory for LLM Training." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory>.
CHICAGOAcademic
Michael Brenndoerfer. "Cross-Entropy Loss: Information Theory for LLM Training." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Cross-Entropy Loss: Information Theory for LLM Training'. Available at: https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Cross-Entropy Loss: Information Theory for LLM Training. https://mbrenndoerfer.com/writing/cross-entropy-loss-language-models-information-theory

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.