Part of Language AI Handbook
Explains how nucleus sampling dynamically selects tokens based on cumulative probability, solving top-k limitations for coherent and creative text generation.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Nucleus Sampling
When GPT models generate text, they face a fundamental challenge: how do you sample from a probability distribution over thousands of possible next tokens in a way that stays coherent while retaining variety? We've seen how temperature scaling adjusts the sharpness of the distribution and how top-k sampling restricts choices to the k most likely tokens. But top-k has a flaw: the number of reasonable next tokens varies dramatically depending on context. Sometimes only one or two tokens make sense; other times, dozens are equally valid. A fixed k cannot adapt to this variation, so it inevitably fails in one direction or the other.
Nucleus sampling, introduced by Holtzman et al. in their 2020 paper "The Curious Case of Neural Text Degeneration", solves this problem by shifting the question entirely. Instead of asking "how many tokens should I consider?", it asks "how much probability mass should I keep?". The algorithm dynamically selects the smallest set of tokens whose cumulative probability exceeds a threshold , a value you set once and reuse across all generation steps. This adaptive approach captures the "nucleus" of the probability mass at each step: all tokens the model considers plausible, while excluding the unreliable tail of unlikely candidates.
Think of the nucleus as the model's zone of reasonable options. In high-confidence situations, that zone is small. In open-ended situations, it is large. Nucleus sampling finds that zone automatically without you having to specify its size in advance. This lets it adapt better than top-k when generating over long sequences with varying levels of model certainty.
The paper itself coined the phrase "nucleus" to describe this idea: the high-density core of the probability distribution, surrounded by a diffuse tail of low-probability tokens. Just as an atomic nucleus contains most of the atom's mass in a tiny volume, the probability nucleus contains most of the probability mass in a small number of tokens. Everything outside the nucleus is statistically noise, and sampling from it produces incoherent text.
Understanding nucleus sampling deeply requires understanding why the tail is harmful. A language model trained on vast amounts of text assigns non-zero probability to millions of possible continuations, many of which are syntactically or semantically wrong in the current context. These wrong tokens sit in the tail. When you include the tail in your sampling pool, you occasionally pick one of these tokens, causing the generated text to veer off course. Once one wrong token is chosen, subsequent predictions are made with a corrupted context, compounding the error. Nucleus sampling prevents this degradation by excluding the tail before sampling even begins.
In this chapter, we'll build up nucleus sampling from first principles, trace through a concrete worked example, implement it in PyTorch, and explore how it interacts with temperature scaling. We'll also look at its limitations and the practical guidance you need to use it effectively in real systems.
The Problem with Fixed-k Sampling
Top-k sampling works by selecting the tokens with the highest probabilities and redistributing probability mass among them. This works well when the model's confidence is consistent, but language is anything but consistent.
Consider two scenarios:
-
High certainty: The model predicts "The capital of France is ___" and assigns 95% probability to "Paris". Here, sampling from the top 50 tokens includes 49 tokens that collectively share only 5% of the probability mass, many of which would produce nonsensical completions.
-
Low certainty: The model predicts "I had a wonderful ___" where "day", "time", "experience", "meal", "trip", and dozens of other tokens are all reasonable. With top-k=10, we might exclude perfectly valid continuations.
The core issue is that is a hyperparameter that cannot adapt to the context. What we really want is to keep all tokens that represent "reasonable" choices, and the most principled way to define "reasonable" is through probability mass.
This problem is more subtle than it first appears. When you set k=50 and hope for the best, you're making an implicit assumption: that every generation step involves roughly the same number of plausible continuations. That assumption is almost always wrong. A story generation task might have high-certainty moments when a character's name is being repeated and low-certainty moments when the next plot event is being chosen. A code generation task might have near-deterministic steps when closing a parenthesis and uncertain steps when choosing between algorithmic approaches. A fixed handles none of these variations well.
There is also an asymmetry in how top-k fails. When is too large, you include harmful tail tokens and generation quality degrades. When is too small, you exclude legitimate options and generation becomes repetitive or stilted. These failure modes pull in opposite directions, and neither extreme is acceptable for open-ended generation. Practitioners often discover this by trying multiple values of and noticing that the right value changes between task types, prompt styles, and even between sentences in the same passage.
The deeper philosophical issue is that top-k conflates two separate things: the threshold for "plausibility" and the count of plausible items. In most generation contexts, you care about plausibility as measured by probability, not about having exactly options. When the model's distribution is uniform across 50 tokens, all 50 deserve consideration. When the distribution puts 95% on one token, you probably want just that one. Nucleus sampling honors this by letting the distribution itself determine how many tokens to include.
The Top-p Formulation
The insight behind nucleus sampling is simple: instead of asking "how many tokens should I consider?", we ask "how much probability mass should I capture?" This shift in perspective leads to an adaptive algorithm that naturally handles both high-confidence and uncertain predictions.
From Intuition to Definition
Think about what we really want when sampling the next token. We want to include all tokens that have a reasonable chance of being correct, and exclude all tokens that are essentially noise. But "reasonable" depends on context. After "The sky is", the token "blue" might have 40% probability, meaning we should consider other options. After "2 + 2 =", the token "4" might have 99% probability, meaning alternatives are probably mistakes.
The key insight is that probability itself tells us what's reasonable. If we capture 90% of the total probability mass, we've included essentially all the tokens the model considers plausible. Everything in the remaining 10% is, by definition, something the model thinks is unlikely. We're not guessing how many tokens to include; we're reading the distribution and respecting what it says.
Notice that this reframing changes what the hyperparameter means. With top-k, you are saying "always consider exactly options." With nucleus sampling, you are saying "always capture at least fraction of the probability mass." The second statement is much easier to reason about. Setting means "I want to include the options in the 90% probability core." That is a semantically meaningful statement that stays valid regardless of whether the model is confident or uncertain.
This leads us to define the nucleus as the smallest set of highest-probability tokens that together account for at least fraction of the total probability. Formally, given a probability distribution over vocabulary , the nucleus is the minimal set such that:
where:
- : the nucleus, the minimal set of highest-probability tokens we'll sample from
- : the probability assigned to token by the model
- : the cumulative probability threshold, typically set between 0.9 and 0.95
The nucleus is the minimal set of highest-probability tokens whose cumulative probability mass meets or exceeds the threshold . Tokens outside the nucleus are discarded, and the remaining probabilities are renormalized.
The Algorithm Step by Step
How do we find this minimal set? The algorithm is straightforward once you see the logic:
-
Sort tokens by probability in descending order. This puts the most likely tokens first.
-
Walk through the sorted list, accumulating probabilities. Keep adding tokens until your running sum reaches or exceeds .
-
Stop as soon as you cross the threshold. The tokens you've accumulated form the nucleus. Discard everything else.
-
Renormalize so that the remaining probabilities sum to 1, giving you a valid distribution to sample from.
Let's express this mathematically. After sorting, we have tokens ordered so that , where is the vocabulary size. We find the smallest such that:
where:
- : the token with the -th highest probability
- : the number of tokens in the nucleus (determined dynamically based on the distribution)
- : the total vocabulary size
The nucleus is then , containing exactly the top tokens needed to reach the probability threshold. The key property of this formulation is that emerges from the distribution itself. A peaked distribution yields small ; a flat distribution yields large .
In practice, this means the nucleus can be as small as 1 token (when the model is extremely confident) or as large as thousands of tokens (when the model is almost completely uncertain across its vocabulary). Top-k cannot span this range with a single setting; nucleus sampling spans it automatically.
Why Renormalization Matters
After truncating to the nucleus, the probabilities no longer sum to 1. If our nucleus contains 90% of the probability mass, the probabilities inside it sum to 0.90, not 1.0. To sample correctly, we need to renormalize.
The renormalized probability for each token in the nucleus is simply its original probability divided by the total probability mass in the nucleus:
where:
- : the renormalized probability used for sampling
- : the original probability of token
- : the sum of original probabilities over all tokens in the nucleus, which is the normalizing constant
This renormalization preserves the relative ordering of tokens. If "nice" was twice as likely as "good" before, it remains twice as likely after. We're simply scaling everything up so the probabilities form a valid distribution that sums to 1.
The renormalization step also has a subtle effect on quality: it redistributes the probability mass that was assigned to tail tokens back into the nucleus. This slightly increases the probability of every nucleus token, which is exactly what we want. When the model was "uncertain" about whether to pick from the tail, we're saying: no, allocate that residual probability to the plausible options instead.
The Unreliable Tail
To appreciate why nucleus sampling matters, it helps to understand what exactly lives in the tail of a language model's output distribution. The tail contains tokens that are much less likely than the leading candidates. In the current context, a well-calibrated model should assign these tokens near-zero probability. The fact that they receive small but non-zero probabilities reflects the imperfection of current models.
Holtzman et al. published "The Curious Case of Neural Text Degeneration" in 2020, and the paper's central finding was striking. They showed empirically that neural language models, despite achieving low perplexity, could generate text that degraded rapidly in quality. The culprit was the unreliable tail: over long sequences, sampling from the full distribution repeatedly landed on low-probability tokens, and each such error compounded the next. They coined the term "neural text degeneration" and proposed nucleus sampling as the cure. The paper immediately influenced how practitioners thought about generation quality and established the standard that nucleus sampling should be the default for open-ended generation.
Language models learn to assign small probabilities to nearly everything in the vocabulary, because neural networks with softmax outputs are rarely completely certain. Even when "Paris" has 95% probability, the remaining 5% is spread across thousands of tokens: "Pari", "paris", "PARIS", "Lyon", and also nonsense like "the", "xyz", "purple", and every other token in the vocabulary. The tail is mostly nonsense, with occasional legitimate alternatives mixed in.
The problem is that when you have 50,000 tokens each with 0.001% probability, sampling from all of them occasionally picks something terrible. In practice, over a 200-token generation with a large vocabulary, you might accidentally sample from the unreliable tail several times. Each such accident corrupts the context for subsequent tokens, often in ways that are hard to recover from. A sentence that starts generating well can suddenly produce a strange word, and subsequent tokens try to make sense of the strange word, leading to increasingly incoherent text.
Top-k sampling with a large mitigates this by excluding the very bottom of the tail, but it still includes tokens with probabilities like 0.01% when . Nucleus sampling is more principled: by defining the nucleus as 90% of the total mass, it excludes everything below a contextually-determined probability floor. In high-certainty contexts, that floor is high (only a few tokens qualify). In uncertain contexts, the floor is lower (many tokens qualify), but the floor always reflects the actual distribution shape.
In practice, the safest way to think about the tail is this: any token with probability low enough to be excluded by nucleus sampling is a token the model is, in effect, saying should not be generated here. Forcing it into the sampling pool overrides the model's own judgment. Nucleus sampling respects what the model learned.
A Worked Example
Abstract formulas come alive with concrete numbers. Let's trace through nucleus sampling step by step with a realistic example.
Setting Up the Problem
Suppose a language model predicts the next token after "The weather is" and produces the following probability distribution:
| Token | Probability |
|---|---|
| nice | 0.35 |
| good | 0.25 |
| bad | 0.15 |
| great | 0.10 |
| terrible | 0.05 |
| wonderful | 0.04 |
| cold | 0.03 |
| hot | 0.02 |
| (other tokens) | 0.01 |
This distribution is already sorted by probability. The model strongly favors positive weather descriptions, with "nice" and "good" together accounting for 60% of the mass.
Finding the Nucleus
With , we walk through the tokens from highest to lowest probability, keeping a running sum:
| Step | Token | Probability | Cumulative Sum | In Nucleus? |
|---|---|---|---|---|
| 1 | nice | 0.35 | 0.35 | ✓ |
| 2 | good | 0.25 | 0.60 | ✓ |
| 3 | bad | 0.15 | 0.75 | ✓ |
| 4 | great | 0.10 | 0.85 | ✓ |
| 5 | terrible | 0.05 | 0.90 | ✓ (threshold reached) |
| 6 | wonderful | 0.04 | . | ✗ (excluded) |
| 7 | cold | 0.03 | . | ✗ (excluded) |
At step 5, our cumulative sum hits exactly 0.90, meeting the threshold. We stop here. The nucleus is {"nice", "good", "bad", "great", "terrible"}, containing 5 tokens.
Notice what happened: we didn't need to decide in advance how many tokens to include. The distribution itself determined that 5 tokens were needed to capture 90% of the probability mass. If we had used top-k=5, we would get the same result here, but only by coincidence. Change the distribution slightly and top-k=5 would fail; nucleus sampling at would adapt.
Notice also what got excluded: "wonderful", "cold", and "hot" are all perfectly reasonable weather descriptions in English. The model just assigned them less probability in this context, so they fall below the threshold. This is the right call. The model's distribution reflects its training, and if it thinks "wonderful" is less likely than "terrible" here, we should respect that judgment.
Renormalizing for Sampling
The five tokens in our nucleus have probabilities summing to 0.90. To sample from a valid probability distribution, we divide each by this sum:
| Token | Original | Calculation | Renormalized |
|---|---|---|---|
| nice | 0.35 | 0.389 | |
| good | 0.25 | 0.278 | |
| bad | 0.15 | 0.167 | |
| great | 0.10 | 0.111 | |
| terrible | 0.05 | 0.056 |
You can verify: (the small discrepancy is rounding). We now have a valid probability distribution over just five tokens.


The left panel shows how cumulative probability grows as we add tokens. The curve rises steeply at first (the top tokens contribute most of the mass) then flattens. We stop when we cross . The right panel shows how renormalization scales up each probability proportionally, preserving relative rankings while creating a valid distribution.
Comparing to Top-k
In this particular case, top-k=5 would include the same tokens. But consider what happens if the model were 90% confident in a single token. Say "Paris" has 0.90 probability and everything else shares the remaining 0.10.
With nucleus sampling at , we'd include just "Paris" (0.90 meets the threshold immediately). With top-k=5, we'd include "Paris" plus four tokens that collectively have only 10% probability. Those four tokens are noise that nucleus sampling correctly excludes.
This is the adaptive behavior that makes nucleus sampling effective: it contracts when the model is confident and expands when the model is uncertain.
How Different Values of p Change the Result
Let's revisit our weather example to see how the nucleus changes as we vary . With , we only need "nice" and "good" (cumulative sum: 0.60). With , we need to add "bad" as well. With , we add "great". With , we add "terrible". With , we bring in "wonderful" as well.
Each threshold draws the line at a different point in the distribution. The choice of is ultimately a statement about how cautious you want to be: how much of the distribution's mass must you capture before you feel comfortable sampling? A of 0.9 says "I need the 90% core before sampling." A of 0.99 says "I need nearly everything." The first is conservative and produces coherent text; the second approaches pure sampling and produces more varied but potentially less coherent text.
Implementation
Now that we understand the algorithm conceptually, let's implement it. We'll build nucleus sampling from scratch, verify it works correctly, then see how to use the production-ready version in Hugging Face's transformers library.
Building the Core Algorithm
The implementation follows our algorithm directly. We'll work with logits (the raw model outputs before softmax), apply optional temperature scaling, convert to probabilities, then perform the nucleus truncation.
import torch
import torch.nn.functional as F
def nucleus_sample(
logits: torch.Tensor, p: float = 0.9, temperature: float = 1.0
) -> int:
"""
Sample from the nucleus (top-p) of a probability distribution.
Args:
logits: Raw logits from the model (before softmax)
p: Cumulative probability threshold
temperature: Temperature scaling factor
Returns:
Sampled token index
"""
# Apply temperature scaling
scaled_logits = logits / temperature
# Convert to probabilities
probs = F.softmax(scaled_logits, dim=-1)
# Sort probabilities in descending order
sorted_probs, sorted_indices = torch.sort(probs, descending=True)
# Compute cumulative probabilities
cumulative_probs = torch.cumsum(sorted_probs, dim=-1)
# Find the cutoff index where cumulative probability exceeds p
# We shift by 1 to include the token that crosses the threshold
sorted_indices_to_remove = cumulative_probs > p
sorted_indices_to_remove[1:] = sorted_indices_to_remove[:-1].clone()
sorted_indices_to_remove[0] = False
# Zero out probabilities for tokens outside the nucleus
sorted_probs[sorted_indices_to_remove] = 0.0
# Renormalize
sorted_probs = sorted_probs / sorted_probs.sum()
# Sample from the filtered distribution
sampled_idx = torch.multinomial(sorted_probs, num_samples=1)
# Map back to original vocabulary indices
return sorted_indices[sampled_idx].item()One subtlety deserves attention: the shift operation in the cutoff logic. When we compute cumulative_probs > p, the first position where this is True is the token that crosses the threshold. But we want to include that token in the nucleus, not exclude it. The shift ensures we remove tokens after the one that crosses the threshold, keeping the nucleus at exactly the right size.
This shift is critical for correctness. Without it, you would exclude the token that pushes cumulative probability over , leaving you with a nucleus that covers slightly less than of the mass. The resulting behavior would be slightly more conservative than intended. With the shift, you always include exactly enough tokens to meet the threshold.
In production code, you would also want to handle edge cases: what if (keep the entire vocabulary)? What if all probabilities are equal (the nucleus becomes the entire vocabulary)? What if the first token alone exceeds (the nucleus contains exactly one token)? The implementation above handles all these cases correctly because the shift logic works properly at the boundaries.
Verifying Our Implementation
Let's confirm our implementation produces the expected behavior using the "weather" example:
# Create a probability distribution matching our example
token_names = [
"nice",
"good",
"bad",
"great",
"terrible",
"wonderful",
"cold",
"hot",
"other",
]
probs = torch.tensor([0.35, 0.25, 0.15, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01])
# Convert to logits (inverse of softmax, approximately)
logits = torch.log(probs)
# Sample 1000 times to see the empirical distribution
samples = []
for _ in range(1000):
idx = nucleus_sample(logits, p=0.9, temperature=1.0)
samples.append(token_names[idx])Sample distribution (1000 samples with p=0.9): nice: 37.0% good: 28.1% bad: 18.0% great: 11.9% terrible: 5.0% wonderful: 0.0%
The samples concentrate on the five nucleus tokens, with "wonderful" appearing rarely or never. The empirical frequencies should approximate our renormalized probabilities: "nice" around 39%, "good" around 28%, and so on. Running 1,000 samples provides enough trials to see the distribution clearly, even if individual runs vary due to randomness.
Visualizing Adaptive Nucleus Size
The real insight comes from seeing how nucleus size varies with . Let's visualize this:

With our example distribution, requires only 1 token (just "nice" at 35%), while needs 7 tokens. The relationship is non-linear because probability mass concentrates in the top tokens. Moving from to adds just one token, but moving from to adds two because the next tail probabilities are small. Each additional token contributes less marginal probability.
The Adaptive Advantage: Peaked vs. Flat Distributions
The advantage of nucleus sampling becomes clear when we compare it to top-k across different distributional shapes. Consider two scenarios:


These visualizations show why nucleus sampling outperforms top-k. When the model is confident (peaked distribution), nucleus sampling automatically tightens to just 2 tokens, while top-k=5 wastefully includes 3 near-zero tokens. When the model is uncertain (flat distribution), nucleus sampling expands to 8 tokens to capture the spread probability mass, while top-k=5 arbitrarily excludes 3 reasonable options.
Top-k makes the same decision regardless of context. Nucleus sampling reads the distribution and responds appropriately.
The left panel shows the peaked case: even though top-k=5 is marked (gray dashed line), nucleus sampling determines that only 2 tokens are needed. The right panel shows the flat case: nucleus sampling expands past the top-k=5 mark because the 6th, 7th, and 8th tokens still carry meaningful probability mass. This is exactly the behavior we want, and it happens automatically from a single shared hyperparameter .
Production Use with Hugging Face Transformers
In practice, you won't implement nucleus sampling yourself. The transformers library provides it out of the box. Here's how to generate text with GPT-2 using top-p:
from transformers import GPT2LMHeadModel, GPT2Tokenizer
# Load model and tokenizer
tokenizer = GPT2Tokenizer.from_pretrained("gpt2")
model = GPT2LMHeadModel.from_pretrained("gpt2")
model.eval()
# Set pad token to avoid warnings
tokenizer.pad_token = tokenizer.eos_token
# Encode prompt
prompt = "The future of artificial intelligence is"
input_ids = tokenizer.encode(prompt, return_tensors="pt")The generate() method accepts top_p directly. Set do_sample=True to enable sampling, and top_k=0 to disable top-k filtering so nucleus sampling operates alone:
# Generate with nucleus sampling (top_p)
with torch.no_grad():
outputs = model.generate(
input_ids,
max_new_tokens=50,
do_sample=True, # Enable sampling
top_p=0.9, # Nucleus probability threshold
top_k=0, # Disable top-k (set to 0 or very large)
temperature=1.0, # No temperature scaling
pad_token_id=tokenizer.eos_token_id,
num_return_sequences=3, # Generate 3 different completions
)
generated_texts = [
tokenizer.decode(output, skip_special_tokens=True) for output in outputs
]Generated completions with nucleus sampling (p=0.9): [1] The future of artificial intelligence is a pivotal one. The basic characteristics of robots have not been fully developed for billions of years, so the consequences of human error could evolve in much the same way. Over the past 70 years, we have discovered dozens of glitches in the implementation of artificial [2] The future of artificial intelligence is uncertain. For one thing, what happens when a computer only learns about features the human eye can't see? As AI develops, each individual person will be deprived of what little information their past activities might enable. As these generations experience more time on Earth [3] The future of artificial intelligence is likely to become more complex." What the future of AI will look like can be seen in more detail below:
Each completion takes a different path while remaining coherent. At every generation step, nucleus sampling includes only tokens the model considers plausible, allowing for creativity without introducing obvious errors.
Choosing the Right p Value
The probability threshold controls the trade-off between creativity and coherence, and choosing it well matters more than it might initially seem. Unlike top-k's , which is difficult to reason about in the abstract, has a clear semantic meaning: it is the fraction of the probability distribution you want to sample from. Higher means you're including more of the tail, which produces more varied but potentially less coherent text. Lower means you're sampling from a more concentrated distribution, producing coherent but possibly repetitive or predictable text.
Here are practical guidelines for selecting :
-
to : General-purpose creative writing. This is the most common setting and works well for open-ended writing tasks, including story and dialogue generation. The nucleus is large enough for variety but excludes the unreliable tail.
-
to : More focused generation. Useful when you want coherent output with moderate diversity, such as paraphrasing or generating multiple options for the same intent.
-
: Very focused generation. Approaches greedy decoding behavior. Use for tasks requiring high precision, like code completion or factual responses.
-
: Very permissive. Rarely used alone because it includes many low-probability tokens. Can produce surprising but occasionally incoherent outputs.
The default of has held up surprisingly well across tasks and model sizes. Holtzman et al.'s original paper used this value in their experiments, and subsequent work has largely confirmed it as a useful starting point. The intuition is that 90% of the probability mass captures virtually all the plausible options while leaving out the long tail of noise.
In practice, the right can depend on the domain and the model's calibration. A model that is well-calibrated for a specific domain, such as a fine-tuned legal document generator, may benefit from lower because the distribution is already focused on domain-appropriate vocabulary. A general-purpose model generating creative fiction may benefit from higher to encourage novelty. The best approach is to generate several passages at candidate values of and evaluate them qualitatively, since automated metrics often fail to capture the trade-off between diversity and quality.
The value of interacts with the model's overall uncertainty, which changes over the course of a sequence. The beginning of a generation tends to have more uncertainty (the model hasn't established context yet), while later tokens tend to be more constrained by prior choices. This means the effective nucleus size varies naturally across a sequence even with a fixed . In the early tokens, might yield a nucleus of 50 tokens; in later tokens constrained by prior context, it might yield a nucleus of just 3. This dynamic adaptation is one of nucleus sampling's key strengths.
Let's visualize how affects generation diversity:

Combining Nucleus Sampling with Temperature
Nucleus sampling and temperature scaling address different aspects of the probability distribution. Temperature reshapes the distribution, making it more peaked or more uniform. Nucleus sampling then truncates based on the reshaped distribution. Understanding how they interact helps you use both controls more precisely.
Temperature scaling modifies the logits before the softmax step. Dividing logits by a temperature sharpens the distribution, pushing more mass to the top tokens. Dividing by flattens it, spreading mass more evenly across the vocabulary. The nucleus threshold then determines how much of this reshaped distribution to keep.
Think of temperature as the "shape dial" and as the "inclusion dial". Temperature decides how concentrated the distribution is; decides what fraction of that distribution to retain. When you lower temperature, the model becomes more confident, so the nucleus naturally shrinks (even with the same , fewer tokens are needed to reach the threshold). When you raise temperature, the model spreads probability more widely, so the nucleus expands.
This interaction gives you two independent dimensions of control over generation behavior. You can hold fixed at 0.9 and vary temperature to change how deterministic the generation feels. Alternatively, you can hold temperature fixed and vary to change how cautiously you sample from the model's distribution. In practice, most systems expose both parameters and let users tune them independently, which provides the flexibility needed across diverse generation tasks.



The visualization above shows how temperature and nucleus size interact. At low temperature (T=0.5), probability concentrates heavily in the top token, so the nucleus contains just 1-2 tokens. At high temperature (T=1.5), probability spreads more evenly, requiring 5+ tokens to capture 90% of the mass. This interaction lets you fine-tune generation behavior: temperature controls the shape of the distribution, while the nucleus threshold determines how much of that reshaped distribution you sample from.
Using both together gives you fine-grained control:
- Low temperature + high p: Coherent output with some variety. The temperature concentrates probability mass, while high p still allows sampling from multiple tokens.
- High temperature + low p: Unusual but still controlled. Temperature spreads probability, but low p keeps only the tokens that remain relatively high after spreading.
- High temperature + high p: Maximum creativity. Use with caution as outputs can become incoherent.
# Generate with combined temperature and nucleus sampling
with torch.no_grad():
# Lower temperature for more focused but still diverse output
focused_outputs = model.generate(
input_ids,
max_new_tokens=50,
do_sample=True,
top_p=0.9,
temperature=0.7, # Slightly lower temperature
pad_token_id=tokenizer.eos_token_id,
num_return_sequences=2,
)
# Higher temperature for more creative output
creative_outputs = model.generate(
input_ids,
max_new_tokens=50,
do_sample=True,
top_p=0.95,
temperature=1.2, # Higher temperature
pad_token_id=tokenizer.eos_token_id,
num_return_sequences=2,
)Focused (temp=0.7, p=0.9): The future of artificial intelligence is in the hands of the future, but it may not be that way at the moment." This is a quote from the paper "A new class of artificial intelligence is emerging that will be able to outperform human intelligence in tasks such as the The future of artificial intelligence is now in the hands of a group of engineers who have created a tool that will make it easier to predict the future. Advertisement The new software, called "Big Data," will be called "Machine Learning," and is expected to be Creative (temp=1.2, p=0.95): The future of artificial intelligence is the world, in which every computer is an entirely different thing – and at least one of them will have an uncanny ability to pick up patterns. One of these "intelligence-enhanced" machines known as the "AI" will emerge to help humanity The future of artificial intelligence is looking brighter. It does it from a computational neuroscience and not through any of the usual "mechanics" techniques. But then there's machine learning, which is doing this sort of stuff on a large scale. It gives you an idea why something
The focused outputs with lower temperature tend to stay closer to common phrasings and high-probability continuations. The creative outputs with higher temperature explore more unusual word choices and directions, though they may occasionally produce less coherent passages. Finding the right balance depends on your application: conversational AI typically benefits from moderate settings, while creative writing tools can push toward higher values.
Nucleus Sampling vs. Top-k: When to Use Which
Both methods truncate the vocabulary, but they do so with different philosophies:
| Aspect | Top-k | Nucleus (Top-p) |
|---|---|---|
| Truncation criterion | Fixed count | Probability mass |
| Adapts to context | No | Yes |
| Parameter intuition | "Keep this many options" | "Keep this much probability" |
| Risk of over-truncation | Yes (when few tokens dominate) | No |
| Risk of under-truncation | Yes (when probability spreads) | No |
| Compute cost | Slightly lower (no cumsum) | Slightly higher |
In practice, nucleus sampling is preferred for creative text generation because of its adaptive behavior. Top-k remains useful when you want explicit control over the number of options or when slight computational savings matter at scale.
Some systems use both: first apply top-k as a coarse filter (e.g., k=100), then apply nucleus sampling within that set. This combines the efficiency of top-k with the adaptivity of top-p. The top-k pass quickly eliminates the very bottom of the vocabulary without needing a cumulative sum, and the subsequent nucleus pass makes the fine-grained probability-based cutoff within the reduced candidate set. For very large vocabularies (100,000+ tokens), this two-pass approach can reduce computation noticeably.
Another consideration is reproducibility. Top-k is deterministic in its selection set: the same distribution always produces the same candidates. Nucleus sampling is also deterministic in selection but can include different numbers of tokens across calls with the same , since the nucleus size is data-dependent. This flexibility is usually an advantage, but for debugging purposes, it helps to know that the nucleus size will vary across steps.
Decoding Strategies
Nucleus sampling is one of several strategies for decoding from a language model's probability distribution, and understanding how it fits among the alternatives helps you choose the right tool for each task.
Greedy decoding always selects the single highest-probability token at each step. It is fast and deterministic, but it is prone to repetition and produces text that feels mechanical. The problem with greedy decoding is that locally optimal choices are not globally optimal: choosing the most probable token at step can close off better paths at steps and beyond.
Beam search addresses this by maintaining multiple candidate sequences simultaneously, expanding each by one token at a time and keeping the top- sequences by cumulative log-probability. This produces more globally coherent completions than greedy decoding, but it tends to generate safe, generic text that lacks the diversity and naturalness of human writing. Beam search also has failure modes in long-form generation, where it produces degenerate repetitive sequences.
Pure sampling draws a token from the full probability distribution without any truncation. This maximizes diversity but regularly samples from the tail, producing incoherent text over longer generations.
Temperature scaling combined with truncation (top-k or top-p) occupies a middle ground. The truncation eliminates the harmful tail while preserving diversity within the plausible region. Nucleus sampling represents the principled version of this middle ground because it defines "plausible" in terms of probability mass rather than token count.
More recently, researchers have proposed contrastive search, typical sampling, and other variants that address different failure modes. Contrastive search penalizes repetition by comparing generated text against the context. Typical sampling targets the "typical" probability region of the distribution rather than the highest-probability tokens. These methods are complementary to nucleus sampling and can sometimes be combined with it.
For most practical applications, nucleus sampling at combined with moderate temperature remains the standard starting point. The alternatives are worth exploring when you have specific failure modes to address.
Limitations and Impact
Nucleus sampling represented a significant advance in text generation quality, but it comes with limitations worth understanding in depth.
The choice of remains a hyperparameter that requires tuning. While works well in many settings, different tasks and domains may benefit from different thresholds. There's no universal value that works optimally across all contexts, and the "right" can even vary within a single generation as the model moves through different parts of a sequence. For instance, the beginning of a story might benefit from higher to establish creative premises, while later passages benefit from lower to maintain consistency with established facts and character voices. This within-sequence variability is a fundamental challenge that a single fixed cannot fully address.
Nucleus sampling also does not address all failure modes of autoregressive generation. Repetition loops, factual errors, and incoherent long-range structure can still occur. The method operates on local token probabilities at each step and has no mechanism for enforcing global coherence or factual accuracy. If the model assigns high probability to factually incorrect tokens, nucleus sampling will include them in the candidate pool. The method filters the distribution's tail but does not correct the distribution's shape for domains the model doesn't handle well. Modern systems often combine nucleus sampling with additional techniques like repetition penalties, contrastive search, or post-hoc filtering to address these remaining failure modes.
The cumulative probability computation adds modest overhead compared to simpler methods like pure sampling or greedy decoding. For most applications this is negligible, but in latency-critical systems processing millions of requests, it can add up. Optimized implementations pre-sort once and reuse the sorted indices, and hardware-level sorting on GPUs makes the overhead minimal in practice. Libraries like vLLM and TensorRT-LLM include optimized nucleus sampling implementations designed for high-throughput serving.
Another limitation is that nucleus sampling does not account for semantic quality directly. It filters by probability, not by whether a token is coherent with the meaning the model is trying to convey. A grammatically valid but semantically strange continuation might have a high probability in the model's distribution. Nucleus sampling will include it in the nucleus and might sample it. The method's effectiveness depends entirely on how well the model's probability distribution aligns with human judgments of quality, which is itself an imperfect alignment for current models.
Despite these limitations, nucleus sampling's impact on practical text generation has been substantial. It became the default sampling method in many popular language model interfaces and APIs. The insight that truncation should adapt to distributional shape rather than use a fixed count influenced subsequent work on decoding strategies. Nucleus sampling also gave practitioners a principled way to balance coherence and creativity that was previously achieved only through extensive trial-and-error with temperature and top-k. When OpenAI released the GPT-3 API in 2020, nucleus sampling via the top_p parameter was immediately available and quickly became the recommended approach for generation tasks. Its simplicity, single hyperparameter, and strong empirical performance made it an enduring standard.
Key Parameters
When using nucleus sampling in Hugging Face's generate() method, the following parameters control the sampling behavior:
-
top_p(float, 0.0 to 1.0): The cumulative probability threshold. Only tokens whose cumulative probability mass is within this threshold are considered. Higher values include more tokens, increasing diversity. Typical range: 0.9 to 0.95 for creative tasks, 0.7 to 0.9 for more focused generation. -
temperature(float, > 0): Scales the logits before applying softmax. Values below 1.0 sharpen the distribution (more deterministic), values above 1.0 flatten it (more random). Applied before nucleus truncation. -
top_k(int): When using nucleus sampling alone, set to 0 to disable top-k filtering. Can be combined with top-p for hybrid filtering. -
do_sample(bool): Must be set toTrueto enable any sampling strategy. WhenFalse, the model uses greedy decoding. -
num_return_sequences(int): Number of independent completions to generate. Useful for generating multiple diverse options.
Summary
Nucleus sampling addresses a fundamental limitation of top-k sampling by adapting the number of candidate tokens to the probability distribution at each step. Rather than asking "how many tokens should I consider?", nucleus sampling asks "how much probability mass should I keep?", and this reframing produces more consistent generation quality across varying contexts.
The core innovation is the shift from a count-based criterion to a probability-mass-based criterion. This single change enables the nucleus to contract when the model is confident and expand when the model is uncertain, tracking the actual structure of the distribution rather than imposing a fixed structure on it. Combined with temperature scaling, nucleus sampling gives practitioners two orthogonal controls over generation behavior: one for shaping the distribution and one for determining how much of it to retain.
The key ideas are:
- Cumulative probability threshold: The nucleus is the smallest token set whose probabilities sum to at least
- Adaptive truncation: High-confidence predictions yield small nuclei; uncertain predictions yield large nuclei
- The unreliable tail: Tokens below the nucleus threshold are statistically noise that degrade generation quality when sampled
- Typical values: to for general creative generation
- Combines with temperature: Temperature reshapes the distribution; nucleus sampling truncates it
- Practical default: Nucleus sampling is the most common choice for open-ended generation in modern systems
When implementing text generation, nucleus sampling should be your default starting point. Adjust based on your tolerance for creativity versus coherence, and combine with temperature scaling for finer control over the output distribution. Start at with temperature , then tune from there: lower or lower temperature for more conservative outputs, higher values for more creative ones.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about nucleus sampling and top-p text generation.
Nucleus Sampling Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!