Part of Language AI Handbook
Covers the mathematical framework for speculative decoding, including the exact acceptance criterion, rejection sampling logic.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Speculative Decoding Math
In the previous chapter, we introduced speculative decoding as a technique for accelerating autoregressive generation by having a small draft model propose multiple tokens that a larger target model verifies in parallel. The key insight was that verification is faster than sequential generation because it allows parallel processing. However, a necessary question remains: how do we decide which draft tokens to accept?
The answer involves carefully designed acceptance criteria that guarantee mathematical correctness. We need a rejection sampling scheme that preserves the target model's output distribution exactly, not approximately. Speculative decoding is mathematically equivalent to standard sampling, creating outputs that are statistically indistinguishable from what you would have generated by running only the large model. This chapter develops the mathematical framework for speculative decoding, including the acceptance criterion, expected speedup analysis, draft quality effects, and optimal draft length selection.
To appreciate why the mathematics matters here, consider what would go wrong with naive approaches. One tempting shortcut is to accept draft tokens whenever both models agree and reject them otherwise. But this biases output toward the intersection of the two distributions, changing the character of the generated text in ways that depend on how similar the models are. Another approach might be to always accept draft tokens and only run the target model occasionally for "spot checking." This would be fast, but it fundamentally changes the output distribution to match the draft model rather than the target. Neither approach achieves the goal of preserving exact target model behavior while gaining speed.
Think of the speculative decoding problem as analogous to quality control in manufacturing. A fast but imperfect inspector (the draft model) checks items on a production line and marks some as passing. A slower, authoritative inspector (the target model) can then review a batch of flagged items simultaneously rather than examining each one independently. The key challenge is designing the fast inspector's acceptance rules so that the batch of passed items, after the authoritative review, has exactly the same quality distribution as if the authoritative inspector had checked each item individually. You want speed without sacrificing quality guarantees.
The mathematical framework we develop draws on classical statistics, specifically rejection sampling, and connects the performance of speculative decoding to the distributional similarity between draft and target models. We will derive precise formulas for expected speedup, understand how draft model quality affects performance, and see how to choose the optimal number of draft tokens for any hardware configuration. This analysis provides both theoretical understanding and practical guidance for deploying speculative decoding in real systems.
Speculative decoding was independently proposed by Chen et al. (2023) in "Accelerating Large Language Model Decoding with Speculative Sampling" and by Leviathan et al. (2023) in "Fast Inference from Transformers via Speculative Decoding." Both papers arrived at essentially the same acceptance criterion based on rejection sampling, though with slightly different presentations. The concurrent discovery reflects how the problem naturally suggests the same mathematical solution: once you frame token verification as a sampling correction problem, rejection sampling emerges as the canonical approach. The technique has since been incorporated into Hugging Face Transformers and vLLM as well as cloud-provider inference APIs. This shows both its practical value and the correctness of the underlying mathematics.
The Acceptance Criterion
The basic challenge in speculative decoding is accepting draft tokens in a way that produces the exact same distribution as sampling directly from the target model. If we simply accept tokens when the draft and target models agree, we would bias the output toward the intersection of their distributions. Instead, we need a principled rejection sampling approach.
The acceptance criterion is the heart of speculative decoding. Everything else, the expected speedup, the optimal draft length, the effect of model quality, all flows from this single decision rule. Getting this criterion right means speculative decoding is as correct as running only the target model. Getting it wrong means the output distribution changes subtly, potentially introducing biases that are hard to detect but real in their effect on generated text quality.
To understand why this problem is non-trivial, consider a small vocabulary example. Suppose our vocabulary has just four tokens and the target model assigns probabilities while the draft model assigns . The models agree that the first token is most likely, but they disagree on the others. If we simply accept the draft token whenever it happens to match what the target would have chosen, we need to think carefully about what distribution that produces. The answer depends on complex interactions between the two distributions that are not immediately obvious without the mathematical framework we are about to develop.
Rejection sampling provides exactly the right tool for this problem. The classical version of rejection sampling, developed by John von Neumann in the 1950s, allows you to sample from a target distribution by filtering samples from a proposal distribution. The core idea is to accept proposal samples with a probability that is proportional to how much the target distribution "wants" that sample relative to how often the proposal generates it. Applied to token acceptance, this gives us a precise mathematical rule.
Setting Up the Problem
To understand the acceptance criterion, we must first appreciate the subtle problem it solves. When a draft model proposes a token, we face a basic question: how do we decide whether to keep that token while ensuring our final output looks exactly as if we had sampled directly from the target model? The decision must maintain a precise mathematical relationship between what we accept and what the target model would have produced on its own.
Think of the two distributions as two people independently ranking a list of candidates. The target model's ranking is authoritative. The draft model's ranking is a fast approximation. We want to use the fast approximation as much as possible, but we need to correct for cases where the approximation's preferences differ from the authoritative ranking. The acceptance criterion is precisely this correction mechanism.
The formal setup is clean. Let denote the target model's probability distribution over the next token, and let denote the draft model's distribution. These are distributions over the same vocabulary , so . When the draft model proposes token (sampled from ), we need an acceptance probability such that the final output, considering both accepted draft tokens and resampled tokens on rejection, follows exactly .
The key insight is that rejection sampling provides an acceptance probability that automatically adjusts for the ratio of target to draft probabilities. If we accept with probability:
where:
- : acceptance probability for token
- : target model probability of token
- : draft model probability of token
then tokens where are always accepted, while tokens where are accepted with probability proportional to how much the target model favors them relative to the draft model.
To build intuition for this formula, consider what happens in two contrasting scenarios. In the first scenario, suppose the target model assigns probability 0.3 to token while the draft model only assigns probability 0.1. The ratio exceeds 1, so we set and always accept. This makes sense: the draft model is underproposing this token relative to what the target wants, so we should accept it whenever it appears. In the second scenario, suppose the target model assigns probability 0.1 to token while the draft model assigns probability 0.3. Now the ratio , so we accept with probability 1/3. The draft model is overproposing this token, so we need to reject some instances to bring its frequency down to what the target model expects.
The minimum function ensures acceptance probabilities stay in . When , the ratio exceeds 1, but we cap it at 1 since probabilities cannot exceed 1. When , the ratio is less than 1 and represents the correct acceptance probability. The symmetry of this rule is elegant: tokens are accepted or rejected based purely on whether the draft model over- or underrepresents them relative to the target.
Why This Works
Let's verify that this acceptance criterion produces the correct distribution. The verification proceeds through careful probability calculations that track what happens when we combine sampling from the draft model with our acceptance decision. When we sample and accept with probability , the probability of accepting a specific token is:
where:
- : probability of generating and accepting token
- : probability of drafting token
- : probability of accepting drafted token
- : probability of token under the target distribution
This derivation reveals that the probability of accepting any particular token is the minimum of the two distributions at that point. This creates a kind of "clipping" effect where we keep only the overlapping probability mass between the draft and target distributions. Think of it as finding the common ground between two people's preferences: we only commit to choices where both parties have at least some agreement.
The total acceptance probability across all tokens is:
where:
- : probability that a drafted token is accepted
- : sum over the entire vocabulary
- : draft model probability of token
- : target model probability of token
This sum has a beautiful geometric interpretation. If you plot both distributions as histograms over the vocabulary, equals the total area of overlap between the two histograms. When the distributions are identical, every token is accepted (total overlap equals 1). When they share no common support, no tokens are accepted (overlap equals 0). The more similar the two models, the more frequently we can reuse draft tokens without needing to resample.
Conditional on acceptance, the distribution over accepted tokens is:
where:
- : probability of token given it was accepted
- : normalization sum over all possible tokens
- : draft model probability
- : target model probability
This does not yet equal . The distribution is corrected by handling rejections appropriately.
The Residual Distribution
When we reject a draft token, we don't simply redraft. Instead, we sample from a carefully constructed residual distribution that "fills in" the probability mass that rejection sampling missed. This residual distribution is the mathematical key that makes speculative decoding exact rather than approximate.
To understand why we need a residual distribution, consider what happens with pure rejection sampling. When we accept with probability , we capture all the probability mass where the draft distribution meets or exceeds the target. But what about tokens where ? For these tokens, the draft model underestimates the target probability, and our rejection sampling only captures the portion up to . If we simply re-ran the draft model on rejection, we would keep trying to sample from and would never correctly generate the tokens that favors more than .
Think of the residual distribution as a "correction fund." The acceptance mechanism captures the portion of that already covers. The residual distribution covers what missed. Together, they add up to the full target distribution . The residual samples only from tokens where , specifically in proportion to how large that gap is, which is exactly what we need to account for the "missing mass."
Define the residual distribution as:
where:
- : residual distribution probability for token
- : target model probability of token
- : draft model probability of token
- : the positive part of the difference between target and draft probabilities (zero when draft exceeds target)
- : normalizing constant,
The residual distribution captures the probability mass where , which is exactly the mass that acceptance sampling misses. Geometrically, if you imagine the target distribution as a histogram and the draft distribution as another histogram overlaid on top, the residual distribution represents the portions of the target histogram that "stick out above" the draft histogram. By sampling from on rejection, we ensure the overall distribution matches .
Notice that the normalizing constant equals , which is the total rejection probability. To see this, note that . The probability mass captured by the residual distribution exactly equals the probability mass that rejection sampling misses. This keeps everything adds up correctly and the final distribution is properly normalized.

Complete Algorithm for One Token
The full acceptance procedure for a single token position works as follows. This algorithm combines the acceptance criterion with the residual distribution to guarantee exact sampling from the target model:
- Sample draft token
- Sample uniform
- If , accept
- Otherwise, sample
This procedure guarantees that the output follows distribution exactly. To see why, consider that there are two mutually exclusive ways to generate a token . The first path is through acceptance: we draft with probability and accept it with probability , contributing to the total probability. The second path is through rejection and resampling: with probability we reject and then sample from the residual with probability , contributing to the total probability. Adding these two contributions gives , exactly as desired.
The key insight is that the two paths together cover the entire target distribution without double-counting. The acceptance path handles the overlap region (where both models assign positive probability), while the residual path handles the regions where the target model exceeds the draft model. The algebraic identity is the mathematical guarantee that these two contributions sum to the target probability exactly.
The acceptance criterion with residual resampling produces outputs that are statistically indistinguishable from sampling directly from the target model. This is not an approximation; it is mathematically exact. The proof follows from the decomposition , which shows that the acceptance path and residual path together reconstruct the full target distribution.
Extending to Multiple Tokens
In practice, the draft model proposes tokens at once. We verify them sequentially from left to right, accepting tokens until the first rejection. This sequential verification is needed because language models produce conditional distributions: the probability of each token depends on all preceding tokens.
The sequential structure of verification mirrors the sequential nature of language itself. Each word in a sentence depends on everything that came before it. When the draft model proposes "the cat sat on the mat," the probability of "sat" depends on having seen "the cat" first, and the probability of "on" depends on having seen "the cat sat" first. If we reject "cat" and replace it with something else, then all subsequent positions, "sat," "on," "the," "mat," become invalid because they were conditioned on "cat." This is why we must verify left to right and stop at the first rejection.
Let be the draft sequence. For position , conditioning on the prefix :
where:
- : acceptance probability for the -th token
- : the -th token in the draft sequence
- : target model conditional probability given all preceding tokens
- : draft model conditional probability given all preceding tokens
If position rejects, we sample from the residual distribution and discard positions . This maintains correctness because each accepted token was drawn from the correct conditional distribution. The key insight is that we cannot "skip" a rejection and continue verifying later tokens, because those tokens were conditioned on the rejected token. Once we sample a different token from the residual distribution, the entire subsequent sequence becomes invalid and must be regenerated.
This left-to-right verification creates an important efficiency consideration. Even though we verify all positions in a single parallel forward pass of the target model, we can only use tokens up to the first rejection. However, because we always sample from the residual distribution upon rejection, we are guaranteed to produce at least one valid token per iteration, even if all draft tokens are rejected. In fact, when all tokens are accepted, we get tokens, since the target model's forward pass also provides the distribution for the next position beyond the draft sequence. This bonus token is a small but consistent benefit: we always get at least one more token than we pay for in draft generation.
Worked Example: Manual Token Verification
Before diving into expected speedup analysis, let's work through a complete numerical example of the acceptance criterion to build concrete intuition. We will use a tiny vocabulary of five tokens and trace through the entire verification process step by step.
Suppose our vocabulary is and the current token position has the following distributions from each model:
| Token | (target) | (draft) | ||
|---|---|---|---|---|
| A | 0.40 | 0.20 | 2.00 | 1.00 |
| B | 0.30 | 0.50 | 0.60 | 0.60 |
| C | 0.15 | 0.15 | 1.00 | 1.00 |
| D | 0.10 | 0.10 | 1.00 | 1.00 |
| E | 0.05 | 0.05 | 1.00 | 1.00 |
The draft model overestimates token B's probability (0.50 vs 0.30) and underestimates token A's probability (0.20 vs 0.40). All other tokens match exactly.
Scenario 1: Draft proposes token A. We compute . We draw and since , we always accept. Token A is used as the next token.
Scenario 2: Draft proposes token B. We compute . We draw . Suppose , so we reject. Now we must sample from the residual distribution.
To compute the residual distribution, we calculate for each token:
- :
- :
- :
- :
- :
The unnormalized residual concentrates all mass on token A with value 0.20. The normalizing constant is , so . When the draft proposes B and we reject, we always sample A from the residual.
Verification that the combined procedure matches . Let's check token A's total probability:
The probability of creating A through acceptance: draft proposes A with probability and we always accept, contributing .
The probability of creating A through rejection: draft proposes B with probability and we reject with probability , then sample A from residual with probability 1.0, contributing .
Total probability for A: . Correct.
The probability of creating B: draft proposes B with probability and we accept with probability , contributing . Correct.
Tokens C, D, E: the draft proposes each with the same probability as the target, so they are always accepted and match the target distribution exactly.
This worked example demonstrates concretely why the residual distribution is necessary. When we reject B (because the draft overproposed it), we do not simply resample from the draft. We sample from the corrected residual that redirects all the "saved" probability mass to tokens the draft underproposed. In this case, every rejection of B leads to an acceptance of A, which is exactly what we need to bring A's total probability from 0.20 (what the draft alone would give) up to 0.40 (what the target requires).
Expected Speedup Analysis
Understanding the expected speedup from speculative decoding requires analyzing how many tokens we expect to accept from each draft sequence and how this translates to wall-clock improvements. This analysis provides both theoretical insights and practical guidance for deploying speculative decoding systems.
The speedup analysis connects the abstract acceptance probability to concrete performance numbers. You might wonder why we need a full mathematical treatment rather than just running experiments. The answer is that experiments can only test specific configurations, while the mathematical framework reveals the basic tradeoffs. It tells us how speedup depends on draft model quality, draft length, and the relative speed of the two models. Armed with this formula, you can predict performance for new model pairs without running extensive benchmarks, and you can optimize configuration choices analytically.
Think of the speedup formula as a budget analysis. Each speculative decoding iteration has a cost (running the draft model multiple times and the target model once) and a benefit (generating multiple tokens in one iteration). The ratio of benefit to cost is the speedup. The mathematical framework makes this ratio precise and reveals how each parameter affects the balance.
Token Acceptance Probability
Let denote the average acceptance probability for a single token. Under simplifying assumptions where each position has the same acceptance rate, we can derive a clean expression for this probability. While real-world acceptance rates vary by position and context, this i.i.d. assumption provides valuable analytical insights and gives formulas that match empirical measurements reasonably well when using the observed acceptance rate as the input.
The expected acceptance probability is computed by averaging over all tokens that the draft model might propose, weighting each token's acceptance probability by how often the draft model proposes it:
where:
- : expected acceptance probability
- : expectation over tokens proposed by the draft model
- : draft model probability of token
- : target model probability of token
- : total probability mass overlap between the two distributions
This quantity measures the overlap between the draft and target distributions. When the distributions are identical, . When they are completely disjoint, . In practice, typically falls between 0.6 and 0.9 for well-matched draft and target model pairs. Because equals the distribution overlap, it directly measures draft model quality from first principles and connects acceptance rates to the statistical notion of distributional similarity.
Expected Accepted Tokens
Given a draft of length , the number of accepted tokens follows a truncated geometric distribution. This distribution arises because we accept tokens sequentially until the first rejection, but we stop at position even if no rejection has occurred. Under the i.i.d. assumption with uniform acceptance probability , the probability that the first tokens are all accepted is , and the probability of a rejection at position is .
Let be the number of tokens we generate (accepted drafts plus one from rejection sampling or the th position). The expected value is:
where:
- : expected number of accepted tokens (plus one verification)
- : draft sequence length
- : probability of accepting a single token
- : position of the first rejection
- : probability that the first rejection occurs at position
- : contribution from the case where all tokens are accepted
The first term accounts for rejecting at position (we keep tokens including the resampled one). The second term accounts for accepting all drafts (we get tokens including the bonus verification token).
We can simplify this sum by observing that is 1 plus the count of accepted draft tokens. Since the probability of accepting the first tokens is for all (under the simplified i.i.d. assumption), we can use the linearity of expectation:
where:
- : closed-form expected number of tokens generated per iteration
- : acceptance probability
- : draft length
- : sum of the geometric series
This closed-form expression shows that expected tokens per iteration grow as a geometric series in the acceptance probability. When is close to 1, almost all draft tokens are accepted, and approaches . When is close to 0, most drafts are rejected immediately, and approaches 1 (we generate just one token from the residual distribution). The geometric series structure is key: each additional draft token contributes to the expected yield, which decays exponentially with position, creating natural diminishing returns at longer draft lengths.
![Line chart showing expected tokens E[N] on the y-axis versus acceptance probability alpha on the x-axis for k values of 1, 3, 5, and 10; curves rise sharply toward their respective k+1 asymptotes as alpha approaches 1.](https://assets.mbrenndoerfer.com/_optimized/notebooks/12_speculative_decoding_math_files/expected-tokens-vs-alpha-1920w.webp)
The Speedup Formula
Let be the cost ratio, defined as the time for one target model forward pass divided by the time for one draft model forward pass. Typically since the target model is much larger. For example, if the target model has 70 billion parameters and the draft model has 7 billion parameters, we might expect to be roughly 10, depending on hardware and batching configurations. In memory-bandwidth-limited regimes (which describes most autoregressive inference), scales roughly with the ratio of model parameter counts.
Without speculative decoding, generating tokens requires target model passes, taking time . This is the baseline we are trying to beat. Each token requires a full forward pass through the large model, with no opportunity for parallelism across tokens in standard autoregressive generation.
With speculative decoding, each iteration requires:
- One draft model pass generating tokens, taking time
- One target model pass verifying all positions, taking time
The iteration produces tokens on average and takes time .
The speedup ratio compares how fast we generate tokens with speculative decoding versus standard autoregressive generation. We derive this by computing the time per token under both approaches and taking their ratio. Standard generation takes per token. Speculative decoding takes per token. The speedup is:
where:
- : speedup ratio (values above 1 mean speculative decoding is faster)
- : time for standard autoregressive generation per token
- : draft length
- : cost ratio ()
- : acceptance probability (used in )
- : cost of one speculative iteration in units of target model passes
- : average tokens produced per iteration
This formula captures the needed tradeoff in speculative decoding. The numerator measures benefit, specifically how many tokens we expect to generate per iteration. The denominator measures cost, specifically how much computation that iteration requires. Speedup greater than 1 means speculative decoding is faster than standard generation. The formula also shows that speedup depends on all three of , , and in ways that interact: high is more valuable when is large, and large is only worth it when is large enough.
When the draft model is very fast (), this simplifies to:
where:
- : theoretical maximum speedup with a zero-cost draft model
- : acceptance probability
- : draft length
This limiting case represents the best possible speedup when the draft model adds negligible overhead. In practice, draft models incur costs, so the actual speedup is lower than this bound. However, this formula provides a useful upper limit for what speculative decoding can achieve. When and , the theoretical maximum speedup is , meaning even with a free draft model, we cannot exceed 3.7 times the speed of standard generation in this configuration.

Draft Quality Effects
The acceptance probability directly depends on how well the draft model approximates the target model. Understanding this relationship helps us choose appropriate draft models and predict performance. The choice of draft model is a necessary practical decision in deploying speculative decoding, and the mathematical connection between distributional similarity and acceptance rate gives us a principled framework for that choice.
Think of the draft model as a stand-in actor rehearsing the lines of the lead performer. The more similar their performance style, the less the director (target model) needs to intervene with corrections. If the stand-in naturally speaks in the same rhythm and vocabulary as the lead, most rehearsal takes place without correction. If their styles differ substantially, constant intervention is needed and the rehearsal provides little benefit over having the lead perform every scene directly.
The mathematical connection between draft quality and speedup is direct and quantifiable. We do not need to guess how much a better draft model will improve throughput. The formula with tells us exactly how changes in translate to changes in speedup. A draft model that improves from 0.7 to 0.8 with and increases from about 2.84 to 3.28 tokens, translating to a speedup increase from roughly to , a 16% improvement from a single quality metric increase.
Measuring Draft Quality
Draft quality can be quantified by the expected acceptance probability:
where:
- : measure of draft quality (higher is better)
- : draft distribution
- : target distribution
This quantity has a clear interpretation: it is the probability mass that the two distributions share. A high-quality draft model assigns similar probabilities to most tokens as the target model, resulting in large overlap. A poor draft model might assign high probability to tokens the target model dislikes, or vice versa, resulting in small overlap.
This is related to the total variation distance between distributions, one of the most basic measures of dissimilarity between probability distributions. Total variation distance measures the maximum probability that an observer could assign to an event using one distribution minus the probability they would assign using the other. It is a tight bound on how distinguishable two distributions are:
where:
- : total variation distance between target and draft distributions
- : target distribution probability for token
- : draft distribution probability for token
- : absolute difference in probability for token
- : acceptance probability (equals )
The relationship is remarkably clean. It says that the acceptance probability is exactly one minus the total variation distance between the draft and target distributions. Higher draft quality means smaller total variation distance, which means higher acceptance probability. This connection to total variation distance indicates that the acceptance probability captures the most basic notion of distributional similarity: the maximum probability that an adversary could distinguish between samples from the two distributions.
In practical terms, this connection is useful for intuition. Total variation distance of 0.2 means acceptance probability of 0.8. It also means that if you drew a sample from either distribution at random, an optimal observer could identify which distribution it came from with probability at most 0.2 better than chance. This is the sense in which the draft and target models are "close" when acceptance rates are high.
Quality-Size Tradeoffs
Larger draft models tend to produce distributions closer to the target, increasing . However, larger models are slower, decreasing . The optimal draft model balances these factors, and finding this balance requires understanding the tradeoffs involved. This is one of the key engineering decisions in deploying speculative decoding at scale.
Consider three scenarios that illustrate how the quality-speed tradeoff plays out in practice:
-
Very small draft model: Fast inference ( large) but poor distribution match ( small). We accept few tokens per iteration, limiting speedup. As a concrete example, imagine using a tiny 125M parameter model as the draft for a 70B parameter target. The draft model runs 100 times faster than the target, but it might only match the target's distribution 40% of the time. Most iterations produce just one or two tokens before rejection. The high cannot compensate for the poor .
-
Medium draft model: Moderate speed and moderate acceptance rate. Often the sweet spot for practical applications. A 7B draft model for a 70B target might run 10 times faster while achieving 75% acceptance rates. With and , this yields a speedup of roughly . This combination often yields the highest speedups in practice precisely because neither dimension of the tradeoff dominates.
-
Large draft model: Good distribution match ( close to 1) but slow ( small). The overhead of running the draft model eats into speedup gains. A 30B draft for a 70B target might achieve 90% acceptance rates, but if it only runs 2-3 times faster than the target, the iteration cost is too high to achieve meaningful speedup. With and , the speedup is , lower than the medium model despite better acceptance rates.
The counterintuitive conclusion is that you do not always want the best possible draft model. A draft model that is somewhat worse but much faster can outperform a nearly perfect but slower draft model. The mathematics gives you the tools to find the optimal point for your specific hardware configuration.
Empirical Observations
In practice, acceptance rates typically range from 0.6 to 0.9 depending on several factors that interact in fine-grained ways. Using a smaller model from the same training family (for example, LLaMA-7B as a draft for LLaMA-70B) yields higher acceptance rates than using a distantly related model, because models trained on the same data with similar objectives develop similar probability distributions. Task difficulty matters too: code completion and factual text have more predictable continuations, leading to high acceptance rates, while creative writing or open-ended reasoning sees lower rates as both models explore different possibilities.
Temperature settings also have a significant effect on acceptance rates. At very low temperatures where sampling becomes nearly deterministic, both draft and target models tend to select the same highest-probability token, resulting in acceptance rates approaching 1. At higher temperatures where sampling explores more of the distribution, the models are more likely to disagree, reducing acceptance rates. This creates an interesting interaction: the temperatures that make output most diverse and creative are precisely the temperatures where speculative decoding helps least, while the temperatures that produce most predictable output are where speculative decoding provides maximum benefit.
These factors interact in interesting ways. For instance, using a very low temperature on a strongly matched draft-target pair can push acceptance rates above 0.95, making even aggressive draft lengths ( or more) productive. Monitoring acceptance rates per task type and adjusting draft lengths accordingly can significantly improve practical throughput.
Optimal Draft Length
Choosing the draft length involves balancing the cost of generating more drafts against the diminishing probability of accepting them all. This optimization problem has both theoretical and practical importance. In theory, it determines the maximum achievable speedup for a given model pair and hardware configuration. In practice, it guides the single most important configuration choice when deploying speculative decoding.
The optimal draft length is not fixed. It depends on the acceptance probability (which varies by task and context), the cost ratio (which depends on hardware and model sizes), and potentially on dynamic factors like recent acceptance rate history. Understanding these dependencies allows you to make good static choices and design adaptive strategies that respond to changing conditions.
The Optimization Problem
We want to find that maximizes speedup. Substituting our formula for , the optimization problem is:
where:
- : optimal draft length
- : expected tokens generated with draft length
- : cost ratio
Taking the derivative with respect to and setting to zero is analytically intractable because is an integer, but we can analyze the continuous relaxation and then round to find the optimal integer value. The continuous analysis reveals the structure of the problem and shows why an interior optimum always exists for finite .
Diminishing Returns
The expected tokens grows sublinearly in . This can be seen from the geometric series formula:
where:
- : expected tokens for draft length
- : asymptotic limit of expected tokens as grows without bound
- : acceptance probability
The growth rate decreases exponentially because decays exponentially. Meanwhile, the cost grows linearly in . This tension creates an interior optimum. Intuitively, each additional draft token has a diminishing marginal benefit (it is only useful if all previous tokens were accepted, which becomes increasingly unlikely) but a constant marginal cost (one more draft model forward pass). At some point, the marginal cost exceeds the marginal benefit, and we should stop drafting.
Think of it as a classic diminishing returns problem. The first draft token you add is very valuable because there is a high probability it gets accepted (). The second draft token is somewhat less valuable because it only contributes when the first was also accepted ( probability of usefulness). The tenth draft token only contributes when all nine preceding tokens were accepted ( probability), which for is only about 10.7%. At some point, the additional draft tokens contribute so rarely that running the draft model to generate them costs more than it saves.
Approximate Optimal
For large (fast draft model), we can derive an approximate optimal draft length. When , the optimal is infinite since there's no cost to drafting more tokens. For finite , the optimal satisfies the first-order condition from setting the derivative of the speedup function to zero. Treating as continuous and applying the quotient rule:
where:
- : acceptance probability
- : draft length
- : cost ratio
- : derivative of the numerator term (positive since for )
- : derivative of the denominator term
- : squared denominator from the quotient rule
Setting the numerator to zero gives:
where:
- : acceptance probability
- : draft length
- : cost ratio
- : marginal gain from extending the draft by one position
- Right side: normalized marginal cost of extension
This equation balances marginal benefit against marginal cost. The left side represents the expected additional tokens from extending the draft by one position (how often does the th token get accepted, and how much does that help). The right side represents the cost of that extension in terms of reduced speedup per existing token. While this equation cannot be solved in closed form, it can be solved numerically for any specific values of and .
For typical values (, ), optimal ranges from 4 to 8 tokens. This explains why practical systems typically use draft lengths in this range. The formula predicts that using is close to optimal for many common configurations, which is why it has become a common default setting in speculative decoding implementations.
Adaptive Draft Length
In practice, the optimal varies based on context. Different parts of a response have different acceptance rates: the beginning of a sentence is often more predictable than the middle, and common phrases have much higher acceptance rates than rare ones. Systems that use a fixed throughout generation leave performance on the table because they cannot adjust to these variations.
Practical adaptive strategies include:
- Fixed : Simple to implement, works well when acceptance rates are stable. A good default for production systems where simplicity matters.
- Adaptive : Adjust based on recent acceptance rates. If the last several iterations had high acceptance rates, increase ; if rates were low, decrease . This can improve average throughput by 10-20% over fixed in practice.
- Tree-based drafting: Generate multiple candidate continuations as a tree rather than a single sequence, increasing effective acceptance probability. When the draft model is uncertain about the best next token, it can propose several alternatives, and the target model accepts the best match. This approach trades increased memory usage for higher effective acceptance rates.
The adaptive approach monitors acceptance rates during generation and increases when rates are high (showing the draft model is performing well on the current content) or decreases when rates are low (showing more challenging content where longer drafts would be wasteful). Implementing this correctly requires careful bookkeeping to avoid introducing biases from the adaptation process itself, but when done correctly, it maintains the exact-sampling guarantee while improving average throughput.
Code Implementation
Let's implement the speculative decoding mathematics to see these concepts in action. The implementation serves two purposes: it verifies our mathematical derivations empirically, and it provides a clean reference implementation that shows how the theoretical concepts translate into actual code.
import numpy as np
rng = np.random.default_rng(42)First, we'll implement the acceptance criterion for a single token:
from typing import Tuple
import numpy as np
def acceptance_probability(p: np.ndarray, q: np.ndarray) -> float:
"""
Calculate the expected acceptance probability alpha.
Args:
p: Target model distribution over vocabulary
q: Draft model distribution over vocabulary
Returns:
Expected acceptance probability
"""
return np.sum(np.minimum(p, q))
def sample_with_rejection(
p: np.ndarray, q: np.ndarray, draft_token: int
) -> Tuple[int, bool]:
"""
Apply acceptance criterion to a draft token.
Args:
p: Target model distribution
q: Draft model distribution
draft_token: Token proposed by draft model
Returns:
Tuple of (final_token, was_accepted)
"""
# Calculate acceptance probability for this specific token
accept_prob = min(1.0, p[draft_token] / q[draft_token])
# Sample uniform and compare
u = rng.uniform()
if u < accept_prob:
return draft_token, True
else:
# Sample from residual distribution
residual = np.maximum(0, p - q)
residual = residual / residual.sum() # Normalize
new_token = rng.choice(len(p), p=residual)
return new_token, FalseLet's verify that this produces the correct distribution:
# Create example distributions
vocab_size = 10
p = np.array([0.3, 0.25, 0.15, 0.1, 0.08, 0.05, 0.03, 0.02, 0.01, 0.01])
q = np.array([0.2, 0.2, 0.2, 0.15, 0.1, 0.05, 0.04, 0.03, 0.02, 0.01])
# Run many samples to verify distribution
n_samples = 100000
final_tokens = []
for _ in range(n_samples):
# Draft model proposes a token
draft_token = rng.choice(vocab_size, p=q)
# Apply acceptance criterion
final_token, _ = sample_with_rejection(p, q, draft_token)
final_tokens.append(final_token)
# Calculate empirical distribution
empirical_dist = np.bincount(final_tokens, minlength=vocab_size) / n_samplesTarget distribution p(x): [0.3 0.25 0.15 0.1 0.08 0.05 0.03 0.02 0.01 0.01] Draft distribution q(x): [0.2 0.2 0.2 0.15 0.1 0.05 0.04 0.03 0.02 0.01] Empirical distribution: [0.298 0.25 0.152 0.101 0.08 0.049 0.03 0.02 0.01 0.01 ] Max deviation from target: 0.0018
The empirical distribution closely matches the target, confirming our acceptance criterion preserves the correct distribution. With 100,000 samples, the maximum deviation from target probabilities is well within statistical noise bounds. This provides strong empirical evidence that the theoretical guarantee holds in practice.
Now let's calculate expected speedup for different configurations:
def expected_tokens(alpha: float, k: int) -> float:
"""Calculate expected number of tokens per iteration."""
return (1 - alpha ** (k + 1)) / (1 - alpha)
def speedup(alpha: float, k: int, c: float) -> float:
"""
Calculate speedup ratio.
Args:
alpha: Acceptance probability
k: Draft length
c: Cost ratio (target time / draft time)
Returns:
Speedup factor
"""
expected_n = expected_tokens(alpha, k)
iteration_cost = 1 + k / c
return expected_n / iteration_costLet's visualize how speedup varies with draft length for different acceptance rates:
k_values = np.arange(1, 16)
c = 20 # Target model is 20x slower than draft
alphas = [0.5, 0.6, 0.7, 0.8, 0.9]
# Calculate speedups for each alpha
speedup_results = {}
for alpha in alphas:
speedup_results[alpha] = [speedup(alpha, k, c) for k in k_values]
The plot shows that optimal draft length increases with acceptance probability. In this configuration, peaks near 13 draft tokens, while peaks near 3. The exact optimum changes with the draft-to-target cost ratio, so the broader lesson is to lengthen drafts only while the additional accepted-token yield exceeds their generation cost.
Let's find the optimal for different scenarios:
def find_optimal_k(
alpha: float, c: float, max_k: int = 20
) -> Tuple[int, float]:
"""Find the optimal draft length and corresponding speedup."""
best_k = 1
best_speedup = speedup(alpha, 1, c)
for k in range(2, max_k + 1):
s = speedup(alpha, k, c)
if s > best_speedup:
best_speedup = s
best_k = k
return best_k, best_speedupconfigs = []
for alpha in [0.6, 0.7, 0.8, 0.9]:
for c in [10, 20, 50]:
opt_k, opt_speedup = find_optimal_k(alpha, c)
configs.append((alpha, c, opt_k, opt_speedup))Optimal draft length for different configurations: -------------------------------------------------- α c Optimal k Speedup -------------------------------------------------- 0.6 10 3 1.67x 0.6 20 4 1.92x 0.6 50 6 2.17x 0.7 10 4 1.98x 0.7 20 6 2.35x 0.7 50 8 2.76x 0.8 10 6 2.47x 0.8 20 8 3.09x 0.8 50 11 3.82x 0.9 10 10 3.43x 0.9 20 13 4.67x 0.9 50 19 6.37x
The results confirm our analysis: higher acceptance rates and faster draft models (larger ) support longer draft sequences and achieve greater speedup. Notice that the optimal increases with both and , which matches our theoretical understanding of the diminishing returns structure.
Finally, let's visualize the acceptance rate as a function of distribution similarity:
# Generate random distribution pairs with varying divergence
alphas_sim = []
kl_divs = []
n_pairs = 200
for _ in range(n_pairs):
# Create base distribution
base = rng.dirichlet(np.ones(20))
# Add noise to create draft distribution with varying divergence
noise_scale = rng.uniform(0, 2)
noise = rng.dirichlet(np.ones(20) * (1 / (noise_scale + 0.1)))
mix = rng.uniform(0.3, 1.0)
q_sim = mix * base + (1 - mix) * noise
q_sim = q_sim / q_sim.sum()
p_sim = base
# Calculate alpha and KL divergence
alpha_sim = np.sum(np.minimum(p_sim, q_sim))
# KL divergence (with smoothing to avoid log(0))
eps = 1e-10
kl = np.sum(p_sim * np.log((p_sim + eps) / (q_sim + eps)))
alphas_sim.append(alpha_sim)
kl_divs.append(kl)
The relationship between distribution divergence and acceptance probability is clear: as the draft model's distribution diverges from the target, acceptance rates drop substantially. This shows the importance of choosing draft models that approximate the target well.
Key Parameters
The key parameters governing speculative decoding performance are:
- k (Draft length): The number of tokens proposed by the draft model in each step. Longer drafts amortize the target model pass cost over more tokens but face diminishing returns as acceptance probability decays geometrically.
- c (Cost ratio): The ratio , representing how much faster the draft model is compared to the target model. Larger allows longer drafts and greater theoretical speedup.
- alpha (Acceptance probability): The probability that a draft token is accepted by the target model. Equal to the probability mass overlap between draft and target distributions, and equal to .
Understanding how these three parameters interact is the key to successfully configuring and optimizing speculative decoding deployments. The speedup formula ties them together into a single performance prediction that you can optimize.
Limitations and Practical Considerations
While the mathematics of speculative decoding provides strong theoretical guarantees, several practical challenges affect real-world performance. Understanding these limitations is important for setting realistic expectations and for designing systems that work well despite them.
The acceptance criterion assumes we can efficiently compute both and for all tokens. In practice, this means running the full forward pass of both models, even though we only need the probability of specific draft tokens. Some implementations optimize this by caching intermediate states, but the full softmax computation remains a bottleneck. The residual distribution sampling also requires computing the full target distribution, which makes it difficult to apply speculative decoding with certain memory-efficient generation techniques that avoid materializing full vocabulary distributions. Systems that use top-k or nucleus sampling at inference time must be careful to ensure that the acceptance criterion is applied to the full distribution, not the truncated one, or the mathematical guarantees break down.
The i.i.d. assumption in our speedup analysis rarely holds in practice. Acceptance rates vary significantly based on context: function names and common phrases have high acceptance rates, while creative or technical content sees more rejections. This variance means actual speedups may differ from theoretical predictions, and systems should monitor acceptance rates to adapt draft lengths dynamically. When the current generation context has low acceptance rates (perhaps because the model is generating a rare technical term or exploring a creative tangent), the fixed- analysis overestimates speedup, and adaptive systems that reduce in these situations will outperform static configurations. Additionally, the tree-structured extensions to speculative decoding (generating multiple candidate continuations) can improve effective acceptance rates but add significant implementation complexity and memory overhead.
Memory constraints present another practical challenge. In single-GPU deployments, both the draft and target models must fit in GPU memory simultaneously. For a 70B target model already consuming nearly all available VRAM, adding even a 7B draft model may require quantization, CPU offloading, or other compromises that affect performance. Multi-GPU deployments solve this by placing the models on separate GPUs, but then the verification step requires communicating logit tensors between devices, adding latency that reduces the effective cost ratio . Production systems at scale must carefully characterize these communication overheads and include them in the cost model used to select optimal .
Hardware considerations also affect practical speedup. The theoretical model assumes draft and target passes don't interfere with each other's memory or compute. On GPUs with limited memory bandwidth, loading both models' weights can create contention. Some deployments use separate GPUs for draft and target models, while others time-multiplex on a single device. Continuous batching systems, which we will explore in the next chapter, add another layer of complexity to speculative decoding integration because batch composition changes between iterations, and the draft and target models may disagree more when generating for diverse batch members simultaneously than when generating for a single request.
There is also a subtle correctness concern when speculative decoding interacts with stateful generation features. Key-value (KV) caches must be managed carefully to ensure that the cached states reflect the accepted tokens, not all the tokens that were proposed. When tokens are rejected and the sequence diverges from the draft, the KV cache must be rolled back or updated accordingly. Implementation bugs in KV cache management can silently corrupt the output distribution without triggering obvious errors, making them particularly dangerous in production systems.
Summary
This chapter developed the mathematical foundations of speculative decoding. This provides tools to understand and optimize this important inference acceleration technique.
The acceptance criterion uses rejection sampling with a residual distribution to guarantee that speculative decoding produces outputs identical in distribution to standard autoregressive sampling. The acceptance probability ensures we accept tokens where the draft model underestimates the target probability while probabilistically rejecting overestimated tokens. The residual distribution on rejection corrects for the missing probability mass, completing the guarantee through the algebraic identity .
The connection to total variation distance, , provides a basic measure of draft model quality rooted in classical statistics. This tells us that acceptance rates are not arbitrary metrics but reflect the most basic notion of how distinguishable two probability distributions are.
Expected speedup depends on the acceptance probability , draft length , and cost ratio . The formula captures the tradeoff between generating more draft tokens and the diminishing probability of accepting them all. The geometric series structure of is responsible for the diminishing returns that create an interior optimum for draft length.
Optimal draft length balances these factors. Higher acceptance rates support longer drafts, with practical optima typically ranging from 4 to 10 tokens depending on model quality and hardware configuration. The first-order condition for optimality equates the marginal benefit of extending the draft with the marginal cost, and this balance shifts predictably with changes in and .
Draft model quality directly impacts speedup through the acceptance probability . Smaller models within the same family often provide the best balance of speed and distribution similarity, and the optimal draft model is not necessarily the one with the highest acceptance rate but the one that maximizes the speedup formula when both quality and speed are accounted for.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about speculative decoding math.
Speculative Decoding Mathematics Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!