Top-k Sampling: Controlling Language Model Text Generation

Michael BrenndoerferUpdated July 29, 202552 min read

Part of Language AI Handbook

Top-k sampling limits each generation step to the k most probable tokens. Covers probability renormalization, diversity, failure modes, and tuning.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Top-k Sampling

When a language model generates text, it produces a probability distribution over tens of thousands of possible next tokens at each step. The question of how to turn that distribution into an actual token choice is called the decoding strategy, and it turns out this choice matters enormously for the quality and character of the resulting text.

Greedy decoding always picks the single most probable token, which sounds reasonable but in practice produces repetitive, bland text. If the most likely word at every step is always selected, the model converges on safe, common phrasing rather than engaging with the richness of language. You might ask a language model to write a story and receive paragraph after paragraph of "The man walked to the store. He bought some things. He went home." Grammatically correct, statistically probable, and deeply uninteresting.

Temperature scaling offers one way out. By dividing logits by a temperature parameter before computing softmax, we can flatten the distribution, giving lower-ranked tokens a better chance. But temperature has a important weakness: it raises the probability of every token, including the nonsensical ones buried in the long tail. A large vocabulary contains tokens like unusual punctuation sequences, rare Unicode characters, and word fragments that should never appear in coherent text. Raising temperature to improve diversity simultaneously makes these garbage tokens more likely to sneak through.

Top-k sampling offers a different solution: rather than redistributing probability across all tokens, first eliminate the ones that cannot possibly be good choices, then sample from what remains. By keeping only the kk most probable tokens and discarding everything else, we define a bounded sampling space that excludes the incoherent long tail while preserving meaningful variation among plausible candidates.

The beauty of this approach lies in its simplicity. You are not changing the model's beliefs about what is likely, only restricting which possibilities you are willing to consider. The model still says "Paris" is thirty times more likely than "café" for completing "The capital of France is", and that ratio is preserved after truncation. You are acting like a careful editor who says "I'll accept anything the model finds plausible, but I refuse to publish something the model itself considers highly unlikely."

This chapter explores top-k sampling in depth. We examine the mathematical foundation of truncated sampling, implement the algorithm from scratch and verify it behaves correctly, study how to choose kk for different contexts and tasks, see how temperature and top-k interact when combined, and analyze the basic limitations that motivated the development of more sophisticated alternatives like nucleus sampling. By the end, you will have both the conceptual understanding to reason about decoding strategy design and the practical implementation knowledge to apply top-k in real generation pipelines.

Historical Context

Top-k sampling was popularized in the context of neural language models around 2018, though truncation-based sampling ideas appeared earlier in the literature. Angela Fan, Mike Lewis, and Yann Dauphin formalized and studied top-k sampling extensively in their 2018 paper "Hierarchical Neural Story Generation," which used k=10k=10 for creative story generation with the WritingPrompts dataset. Their work demonstrated that truncating the sampling distribution to a fixed-size candidate set produced substantially more coherent and creative text than either greedy decoding or pure temperature sampling. The approach quickly became standard in the language model community and was adopted as a default in systems like GPT-2 and its successors. Nucleus (top-p) sampling was introduced in 2019 by Holtzman et al. as a more adaptive alternative, but top-k remains widely used today due to its computational simplicity and predictable behavior.

The Problem with Full Vocabulary Sampling

Before diving into top-k, it is worth spending time understanding exactly why sampling from the full vocabulary distribution causes problems. The issue is not immediately obvious, because any individual low-probability token should be selected very rarely. The problem only becomes clear when you think statistically about what happens over many generation steps.

A trained language model assigns a non-zero probability to every token in its vocabulary, even tokens that would be semantically absurd given the context. This is an unavoidable consequence of training: the model is trained on a finite corpus and learns a smooth probability distribution over the vocabulary. The smoothness prevents it from assigning exactly zero probability to any token, because that would create sharp discontinuities that generalize poorly. The result is a long tail of tokens with tiny but non-zero probabilities.

Think of full vocabulary sampling as a lottery where most tickets have a small but real chance of winning. In a vocabulary of 50,000 tokens, even if each "bad" token has probability 0.001%, collectively thousands of such tokens might represent 10-20% of the total probability mass. When you sample hundreds of tokens to generate a paragraph, the cumulative probability of at least one bad token slipping through becomes substantial. Just one incoherent token can derail the generation, because subsequent tokens are conditioned on the previous context, and an incoherent token shifts the distribution toward other incoherent tokens.

The key insight is that the model's probability distribution is shaped like a power law: a small number of tokens hold most of the probability mass, and the rest is distributed across a very long tail. The top 50 tokens might hold 95% of the probability, while the remaining 49,950 tokens share the other 5%. When we sample, we almost always get one of those top 50 tokens, but occasionally we get something from the long tail. That occasional bad draw is enough to degrade generation quality significantly, especially in longer sequences where errors can compound.

In[5]:
Code
def get_next_token_distribution(model, tokenizer, prompt):
    """Get the probability distribution over next tokens."""
    inputs = tokenizer(prompt, return_tensors="pt")

    with torch.no_grad():
        outputs = model(inputs["input_ids"])
        logits = outputs.logits[0, -1, :]
        probs = torch.softmax(logits, dim=0)

    return probs
Out[6]:
Console
Prompt: 'The capital of France is'

Vocabulary size: 50,257
Tokens with P > 1%: 14
Tokens with P > 0.1%: 109

Top 5 predictions:
  1. ' the': 0.0846
  2. ' now': 0.0479
  3. ' a': 0.0462
  4. ' France': 0.0324
  5. ' Paris': 0.0322

Notice how only a handful of tokens have probability above 1%, and only a few dozen have probability above 0.1%. The remaining 50,000+ tokens collectively hold a small but non-trivial share of the probability mass. This observation motivates truncation: if we only sample from the top few hundred tokens, we capture essentially all the probability that matters while guaranteeing we never produce true garbage.

Out[7]:
Visualization
Log-scale bar chart showing token probabilities ranked by likelihood, with steep dropoff after top tokens.
Next-token probability distribution for ''The capital of France is''. The distribution is highly skewed: a few tokens dominate while thousands have negligible but non-zero probability. Sampling from the full distribution occasionally selects from this long tail, creating incoherent text.

The figure illustrates why the long tail is problematic. After the first few tokens, probability drops rapidly, but thousands of tokens retain some small probability. Summed together, this tail can be significant. If we sample proportionally, low-probability tokens occasionally win, creating outputs like "The capital of France is ????????" or "The capital of France is asdf". Even rarer, a sequence of low-probability choices can cascade into completely incoherent text that never recovers. Top-k sampling is our circuit breaker: once we identify the drop-off point in the probability ranking, we simply refuse to sample below it.

How Top-k Sampling Works

Now that we understand the problem, let us build up the solution piece by piece. The core insight behind top-k sampling is deceptively simple: if most of the probability mass concentrates in a handful of tokens, why not just sample from those and ignore the rest?

The elegance of this solution is that it makes almost no assumptions about the shape of the distribution. It does not try to model why probability falls off or where the "natural" cutoff should be. It simply imposes a hard count-based boundary and accepts the consequences. This makes top-k both easy to understand and easy to implement, which explains why it became so widely adopted despite its simplicity.

Think of top-k sampling as a casting director who says: "I will only consider actors in the top kk of my ranking. Everyone else is not being called back, regardless of how close they were to the boundary." The director does not re-evaluate the top actors based on who was excluded; they simply rescale their assessments of the remaining candidates. Someone ranked 3rd out of 50,000 is still ranked 3rd out of the top kk, and their relative standing compared to other top-kk candidates is unchanged.

This analogy also highlights an important subtlety: the excluded tokens receive zero probability, not some small residual probability. This is a hard cut, not a soft fade. The benefit is that we guarantee no sampling from the tail. The cost is that tokens just outside the boundary (ranked k+1k+1 and k+2k+2) receive zero probability even if they were only slightly less probable than the token at rank kk. Whether this boundary effect matters in practice depends on how steeply the probability drops off near rank kk.

The Intuition: Focus on What Matters

Think about how you complete sentences. Given "The capital of France is", you do not mentally consider every possible word in English. You immediately focus on a small set of plausible continuations: "Paris", maybe "a", "the", or "known". Your mental probability distribution is not uniform across all words, and neither is the model's. Top-k sampling formalizes this intuition by explicitly restricting attention to the most likely candidates.

The approach works in five steps:

  1. Compute probabilities: The model outputs a distribution over all vocabulary tokens
  2. Rank by likelihood: Sort tokens from most to least probable
  3. Keep the top k: Select only the kk highest-probability tokens
  4. Zero out the rest: Set all other token probabilities to exactly zero
  5. Renormalize and sample: Rescale the remaining probabilities so they sum to 1, then sample
Top-k Sampling

Top-k sampling restricts the sampling space to the kk most probable tokens. After zeroing out tokens outside the top-kk, the remaining probabilities are renormalized before sampling.

From Intuition to Mathematics

Let us formalize this process. At generation step ii, the model has seen context x<ix_{<i} (all previous tokens) and outputs a probability P(xi∣x<i)P(x_i | x_{<i}) for each possible next token xix_i. This distribution spans the entire vocabulary, but we only want to sample from the top performers.

First, we need to identify which tokens make the cut. We define VkV_k as the set containing the kk tokens with highest probability. Formally, Vk={v1,v2,…,vk}V_k = \{v_1, v_2, \ldots, v_k\} where P(v1∣x<i)≥P(v2∣x<i)≥…≥P(vk∣x<i)P(v_1 | x_{<i}) \geq P(v_2 | x_{<i}) \geq \ldots \geq P(v_k | x_{<i}) and every token outside VkV_k has probability no greater than the smallest element of VkV_k. If the vocabulary has 50,000 tokens and k=50k=50, then VkV_k contains exactly the 50 most likely tokens for this specific context.

Now we need to handle a subtle but important issue: after removing tokens from consideration, the probabilities of the remaining tokens no longer sum to 1. A probability distribution must sum to 1 for sampling to work correctly. If the top-kk tokens collectively hold 85% of the probability mass, the removed tokens held the other 15%, and our truncated distribution sums to only 0.85. The solution is renormalization: we divide each remaining probability by the total mass of the kept tokens, scaling everything up so the distribution is valid again.

To compute the renormalization constant, we sum the probabilities of all top-kk tokens:

Zk=∑x∈VkP(x∣x<i)Z_k = \sum_{x \in V_k} P(x | x_{<i})

where:

  • VkV_k: the set of kk tokens with highest probability
  • P(x∣x<i)P(x | x_{<i}): the model's original probability for token xx given context x<ix_{<i}
  • ZkZ_k: the total probability mass held by the top-kk tokens, which acts as a normalizing constant

This gives us the formal definition of top-k sampling. Given the original distribution and the normalization constant, the truncated distribution assigns probability:

Pk(xi∣x<i)={P(xi∣x<i)Zkif xi∈Vk0otherwiseP_k(x_i | x_{<i}) = \begin{cases} \frac{P(x_i | x_{<i})}{Z_k} & \text{if } x_i \in V_k \\ 0 & \text{otherwise} \end{cases}

where:

  • xix_i: a candidate token at position ii
  • x<ix_{<i}: the context (all tokens before position ii)
  • P(xi∣x<i)P(x_i | x_{<i}): the original probability the model assigns to token xix_i given the context
  • VkV_k: the set of kk tokens with highest probability under P(⋅∣x<i)P(\cdot | x_{<i})
  • Zk=∑x∈VkP(x∣x<i)Z_k = \sum_{x \in V_k} P(x | x_{<i}): the normalization constant, which equals the sum of probabilities of the top-kk tokens
  • Pk(xi∣x<i)P_k(x_i | x_{<i}): the renormalized probability used for sampling

Why does this formula make sense? Notice that for tokens inside VkV_k, dividing by ZkZ_k (which is less than 1) makes each probability larger than the original. This is correct: we are redistributing the probability that was held by the excluded tokens and giving it proportionally to the included tokens. For tokens outside VkV_k, the formula assigns zero. The resulting distribution sums to exactly 1 because ∑x∈VkPk(x∣x<i)=∑x∈VkP(x∣x<i)/Zk=Zk/Zk=1\sum_{x \in V_k} P_k(x | x_{<i}) = \sum_{x \in V_k} P(x | x_{<i}) / Z_k = Z_k / Z_k = 1.

Why Renormalization Preserves Relative Probabilities

A necessary property of this construction is that renormalization preserves the relative likelihood of tokens within the top-kk. This matters because the model's ranking captures important information about which tokens are more semantically appropriate than others, and we want to honor that information.

Consider two tokens: "Paris" with original probability 0.30 and "the" with probability 0.15. In the original distribution, "Paris" is exactly twice as likely as "the".

After truncation, suppose the top-kk tokens have cumulative probability Zk=0.90Z_k = 0.90. The renormalized probabilities become:

  • "Paris": 0.30/0.90=0.3330.30 / 0.90 = 0.333
  • "the": 0.15/0.90=0.1670.15 / 0.90 = 0.167

Notice that 0.333/0.167=20.333 / 0.167 = 2. The ratio is preserved! This happens because we divide both probabilities by the same constant ZkZ_k. Renormalization scales all probabilities equally, maintaining their relative ordering and ratios.

This property matters because it means top-k sampling respects the model's preferences among plausible tokens. We are not arbitrarily reweighting tokens; we are simply restricting which tokens can be sampled while honoring the model's ranking within that restricted set. If you trust the model's learned distribution over the top candidates, renormalized top-k sampling samples from that distribution faithfully.

The key insight is that top-k sampling changes the support of the distribution (which tokens have non-zero probability) without changing the shape of the distribution over the supported tokens. The relative probabilities among kept tokens are identical to what they would be under the original distribution, up to the constant scaling factor 1/Zk1/Z_k.

Visualizing the Truncation

The following visualization shows this process concretely. On the left, we see the original distribution over the top 20 tokens. On the right, we see the truncated distribution after keeping only k=10k=10 tokens and renormalizing.

Out[8]:
Visualization
Horizontal bar chart showing top 20 token probabilities. The top 10 bars are blue (kept) and the bottom 10 are gray (truncated).
Original distribution (top 20 tokens shown). Blue bars indicate tokens kept in top-10, gray bars are truncated.
Horizontal bar chart showing the 10 kept tokens with renormalized probabilities that sum to 1.0.
Top-10 truncated distribution after renormalization. Probabilities now sum to 1.

Worked Example: Step-by-Step Numerical Trace

To make the algorithm fully concrete, let us trace through a complete top-k sampling step with small, explicit numbers. This example uses a toy vocabulary of 8 tokens so we can inspect every calculation by hand.

Suppose we have just processed the context "The sky is" and our language model outputs the following raw logits over 8 vocabulary tokens:

TokenLogit
blue3.2
clear2.8
dark1.5
gray1.1
red0.4
green-0.3
loud-1.2
empty-2.5

Step 1: Convert logits to probabilities using softmax. For each logit ziz_i, compute ezie^{z_i}, then divide by the sum of all exponentials.

The exponentials are:

  • blue: e3.2=24.53e^{3.2} = 24.53
  • clear: e2.8=16.44e^{2.8} = 16.44
  • dark: e1.5=4.48e^{1.5} = 4.48
  • gray: e1.1=3.00e^{1.1} = 3.00
  • red: e0.4=1.49e^{0.4} = 1.49
  • green: e−0.3=0.74e^{-0.3} = 0.74
  • loud: e−1.2=0.30e^{-1.2} = 0.30
  • empty: e−2.5=0.08e^{-2.5} = 0.08

The sum of all exponentials is 24.53+16.44+4.48+3.00+1.49+0.74+0.30+0.08=51.0624.53 + 16.44 + 4.48 + 3.00 + 1.49 + 0.74 + 0.30 + 0.08 = 51.06.

Dividing each exponential by 51.06 gives the probability distribution:

TokenProbability
blue0.480
clear0.322
dark0.088
gray0.059
red0.029
green0.014
loud0.006
empty0.002

Step 2: Select the top k=4k=4 tokens. Ranking by probability, the top 4 are: blue (0.480), clear (0.322), dark (0.088), gray (0.059). The remaining tokens (red, green, loud, empty) are excluded.

Step 3: Compute the normalization constant. We sum the probabilities of the kept tokens:

Z4=0.480+0.322+0.088+0.059=0.949Z_4 = 0.480 + 0.322 + 0.088 + 0.059 = 0.949

Step 4: Renormalize. Divide each kept token's probability by Z4=0.949Z_4 = 0.949:

TokenOriginal ProbabilityRenormalized Probability
blue0.4800.480 / 0.949 = 0.506
clear0.3220.322 / 0.949 = 0.339
dark0.0880.088 / 0.949 = 0.093
gray0.0590.059 / 0.949 = 0.062
Sum0.9491.000

Step 5: Sample. We now draw one token from this renormalized distribution. With probability 0.506 we select "blue"; with probability 0.339, "clear"; with probability 0.093, "dark"; and with probability 0.062, "gray". The tokens "red", "green", "loud", and "empty" can never be selected.

Verification of ratio preservation. Let us check that the ratio of "blue" to "clear" is preserved. Originally: 0.480/0.322=1.4910.480 / 0.322 = 1.491. After renormalization: 0.506/0.339=1.4930.506 / 0.339 = 1.493. The tiny difference is rounding error; the ratio is preserved because both are divided by the same Z4Z_4. The model's belief that "blue" is about 1.5 times more likely than "clear" is fully honored.

This numerical trace shows every operation explicitly: softmax converts logits to probabilities, top-kk selection identifies the candidates, summation computes the normalization constant, division produces a valid probability distribution, and multinomial sampling draws the final token.

Implementing Top-k Sampling

With the mathematics established, let us translate the algorithm into code. The implementation is straightforward, but walking through it step by step reveals how each line corresponds to a piece of the formula.

The implementation relies on two key PyTorch operations. First, torch.topk(tensor, k) efficiently finds the kk largest values in a tensor and returns both those values and their original positions (indices) in the tensor. We need both the values (to compute probabilities) and the indices (to identify which vocabulary tokens were selected). Second, torch.multinomial(probs, num_samples=1) draws a random sample from a discrete probability distribution, returning the index of the sampled element.

Think of the implementation as a pipeline: raw logits flow in one end, get temperature-scaled, get truncated to the top kk, get converted to probabilities via softmax, and then a single sample emerges from the other end. Each stage is a simple operation, and the pipeline is easy to read and debug.

Building the Core Sampler

Our implementation needs to accomplish four things:

  1. Apply temperature scaling to the logits
  2. Find the kk highest-scoring tokens
  3. Convert those scores to a valid probability distribution
  4. Sample one token from that distribution

The key insight is that we do not need to explicitly zero out the non-top-k tokens. Instead, we can compute softmax over only the top-k logits, which automatically gives us a properly normalized distribution over just those tokens. This is equivalent to the renormalization step in our formula and is numerically cleaner than setting probabilities to zero and then renormalizing.

In[9]:
Code
def top_k_sample(logits, k, temperature=1.0):
    """
    Sample from top-k truncated distribution.

    Args:
        logits: Raw logits from model, shape (vocab_size,)
        k: Number of top tokens to keep
        temperature: Temperature for scaling before truncation

    Returns:
        Sampled token index
    """
    # Apply temperature scaling
    scaled_logits = logits / temperature

    # Get top-k values and indices
    top_k_values, top_k_indices = torch.topk(scaled_logits, k)

    # Convert to probabilities
    top_k_probs = torch.softmax(top_k_values, dim=0)

    # Sample from truncated distribution
    sample_idx = torch.multinomial(top_k_probs, num_samples=1)

    # Map back to original vocabulary index
    return top_k_indices[sample_idx].item()

Let's trace through each step:

  • Temperature scaling divides each logit by the temperature, controlling how peaked or flat the distribution becomes before truncation. Dividing by a value less than 1 sharpens the distribution; dividing by a value greater than 1 flattens it.
  • torch.topk efficiently finds the kk largest values and their positions in the vocabulary, returning both the values and their indices. The indices tell us which vocabulary tokens these are.
  • Softmax over top-k values computes ezi/∑jezje^{z_i} / \sum_j e^{z_j} using only the kept tokens, which is equivalent to the renormalization step in our formula. We get a valid probability distribution automatically.
  • torch.multinomial samples according to the probability distribution, returning an index into our top-k list (a number from 0 to k−1k-1, not the vocabulary index).
  • Index mapping converts the sampled position (0 to k−1k-1) back to the actual vocabulary token ID using top_k_indices[sample_idx].

Seeing It in Action

Let's sample multiple times from the same context to see the variety top-k produces:

Out[10]:
Console
Prompt: 'The capital of France is'

Sampling with k=5 (10 samples):
  Sample 1: ' the'
  Sample 2: ' France'
  Sample 3: ' the'
  Sample 4: ' now'
  Sample 5: ' the'
  Sample 6: ' the'
  Sample 7: ' the'
  Sample 8: ' a'
  Sample 9: ' now'
  Sample 10: ' a'

Notice how all samples are reasonable completions. With k=5k=5, we are sampling from only the five most likely tokens, which for this factual prompt are all sensible choices. The variation comes from the probabilistic sampling, but every option is plausible. If we had used k=1k=1, we would always get the same token; if we had used k=50000k=50000, we might occasionally get something bizarre.

The output also demonstrates one of the practical advantages of top-k sampling for interactive applications: you can run the same prompt multiple times and get different, interesting responses, unlike greedy decoding which always gives you the same deterministic output.

From Single Tokens to Full Text Generation

A single sample is interesting, but the real power of top-k sampling emerges when we generate entire sequences. Each token we generate becomes part of the context for the next token, creating a chain of sampling decisions. The model sees the prompt plus all previously generated tokens, creating a new distribution, from which we sample again. This auto-regressive loop continues until we hit a stopping condition.

The generation loop follows a simple pattern: get logits, sample a token, append it to the sequence, repeat. At each step, top-k ensures we only consider reasonable continuations. The accumulated context shapes which tokens are plausible at each new step, so the generated text remains coherent even as it explores different directions.

In[11]:
Code
def generate_with_top_k(
    model, tokenizer, prompt, max_tokens=30, k=50, temperature=1.0
):
    """Generate text using top-k sampling."""
    input_ids = tokenizer(prompt, return_tensors="pt")["input_ids"]
    generated_ids = input_ids.clone()

    for _ in range(max_tokens):
        with torch.no_grad():
            outputs = model(generated_ids)
            next_token_logits = outputs.logits[0, -1, :]

        # Sample using top-k
        next_token_id = top_k_sample(
            next_token_logits, k=k, temperature=temperature
        )

        # Append and continue
        generated_ids = torch.cat(
            [generated_ids, torch.tensor([[next_token_id]])], dim=1
        )

        # Stop if EOS token
        if next_token_id == tokenizer.eos_token_id:
            break

    return tokenizer.decode(generated_ids[0], skip_special_tokens=True)

The Effect of Different k Values

Now we can see how kk shapes the character of generated text. Smaller kk means stricter filtering, keeping only the most confident predictions. Larger kk allows more exploration of the probability space, creating text with more varied word choices and sentence structures.

The relationship between kk and text quality is not monotonic. At k=1k=1 we get greedy decoding: deterministic but repetitive. As kk increases toward 10-20, we gain diversity while maintaining high quality. As kk grows past 100-200, we start including tokens that are implausible for the context, and quality begins to degrade. The sweet spot depends on the task and the model, but values in the range 40-100 work well for most general text generation applications.

Out[12]:
Console
Prompt: 'Artificial intelligence will'

Generated continuations:
--------------------------------------------------
k=10: Artificial intelligence will also be more complex, and more diverse, than ever before. For example, artificial intelligence, which is already a key technology for the world's financial
k=50: Artificial intelligence will also create even more exciting challenges for our nation's workers, and many of the benefits will come from improving our ability to build our own factories and reduce
k=200: Artificial intelligence will become so advanced that, like our pets, children, pets will be doing almost any form of AI-related activity. Companies are taking our money.

The outputs reveal distinct personalities. With k=10k=10, the text tends toward common, expected phrasings. With k=50k=50, there's more room for varied word choices while maintaining fluency. With k=200k=200, the model has significant freedom, though even here, we've eliminated the large majority of the vocabulary's long tail.

Out[13]:
Visualization
Heatmap showing sampling counts across token ranks for different k values.
Empirical sampling frequency across 100 samples for different k values. With k=5, nearly all samples come from the top 2-3 tokens. As k increases, the distribution spreads more evenly across the kept tokens, though higher-probability tokens still dominate.

The heatmap reveals how kk shapes sampling behavior empirically. With k=5k=5, the dark cells concentrate in the leftmost columns, meaning the top-ranked tokens receive nearly all samples. As kk increases, the distribution spreads rightward, but higher-ranked tokens still dominate because they have higher probability even after renormalization. Increasing kk does not magically make all kept tokens equally likely; it simply allows lower-ranked tokens to have a non-zero chance of being selected.

Choosing the Right Value of k

The choice of kk significantly affects generation quality, and there is no universally optimal value. The best kk depends on the model, the task, the context, and even the individual prompt. Understanding the principles behind kk selection helps you make informed decisions rather than guessing.

The basic tension is between quality and diversity. Lower kk values restrict sampling to the model's most confident predictions, creating text that is safe and coherent but potentially monotonous. Higher kk values allow the model to explore less certain territory, creating more varied text but risking occasional incoherence. This is not a tradeoff you can resolve once and forget; it lives at the heart of every generation decision.

One useful way to think about kk selection is in terms of coverage: how much of the probability mass does the top-kk set hold? For a factual prompt where the model is highly confident, the top 5 tokens might hold 90% of the probability mass. Choosing k=50k=50 in this case would include 45 tokens that collectively hold only 10% of the mass. You are sampling mostly from the same small set of dominant tokens, with occasional low-probability diversions that may not add value. For a creative writing prompt where probability is spread more evenly, the top 5 tokens might hold only 30% of the mass, and k=50k=50 would be needed to capture a meaningful range of plausible continuations.

Think of coverage-based reasoning as calibrating kk to the distribution's shape. When the distribution is sharp (high confidence), a small kk captures most of the mass. When the distribution is flat (low confidence), you need a larger kk to stay in the region where the model finds things plausible. This is exactly the intuition behind nucleus (top-p) sampling, which adaptively selects kk based on cumulative probability. But even for fixed-kk sampling, understanding coverage helps you choose kk wisely for a given application.

Out[14]:
Visualization
Line plot showing diversity increasing and coherence decreasing as k increases, with optimal zone highlighted.
Trade-off between diversity and coherence as k varies. Lower k values produce more focused text but risk repetition. Higher k values increase diversity but may introduce occasional off-topic tokens. The shaded region represents typical values used in practice (k=30-100).

Guidelines for Selecting k

Consider these factors when choosing kk:

  • Task type: Creative writing benefits from higher kk (50-100) for variety. Factual responses work better with lower kk (10-30) for precision.

  • Context confidence: When the model is highly confident (one token dominates), even large kk won't introduce much diversity since top tokens have most of the mass anyway.

  • Generation length: Longer generations may need lower kk to prevent drift. Errors accumulate over many steps, and each improbable token can push the generation off course.

  • User expectations: Interactive applications often use kk=40-50 as a reasonable default that balances fluency with variety.

In[15]:
Code
def analyze_top_k_coverage(probs, k_values):
    """Analyze what fraction of probability mass top-k captures."""
    sorted_probs, _ = torch.sort(probs, descending=True)

    coverage = {}
    for k in k_values:
        coverage[k] = sorted_probs[:k].sum().item()

    return coverage
Out[16]:
Console
Top-k Probability Coverage by Context Type:
============================================================

Prompt: 'The capital of France is'
  k=  5: 24.3% of probability mass
  k= 10: 35.9% of probability mass
  k= 20: 45.7% of probability mass
  k= 50: 54.9% of probability mass
  k=100: 62.4% of probability mass

Prompt: 'Once upon a time in a land far'
  k=  5: 68.3% of probability mass
  k= 10: 79.8% of probability mass
  k= 20: 86.2% of probability mass
  k= 50: 92.8% of probability mass
  k=100: 95.4% of probability mass

Prompt: '2 + 2 ='
  k=  5: 40.8% of probability mass
  k= 10: 52.5% of probability mass
  k= 20: 62.0% of probability mass
  k= 50: 70.5% of probability mass
  k=100: 76.0% of probability mass

The coverage analysis reveals an important pattern: when the model is confident, a small kk captures most of the probability mass. "2 + 2 =" concentrates probability in very few tokens, so k=10k=10 might capture 99%+ of the mass. Open-ended prompts like "Once upon a time" spread probability more evenly, requiring larger kk to maintain diverse sampling. Looking at coverage percentages is one of the most practical guides to choosing kk for a specific application.

Out[17]:
Visualization
Line plot showing cumulative probability curves for three prompts, showing different coverage rates.
Cumulative probability coverage as k increases for three different contexts. Confident predictions (blue) reach near-complete coverage with tiny k values. Open-ended contexts (green) require larger k to capture the same probability mass. The dashed lines show the 90% and 99% coverage thresholds.

This visualization makes the intuition concrete. For confident predictions like "2 + 2 =", the curve shoots up almost vertically, reaching 99% coverage with just a handful of tokens. Open-ended prompts show a gentler slope, requiring larger kk to achieve the same coverage. This explains why a fixed kk works differently across contexts: the same kk value can mean very different things depending on how the probability mass is distributed. A value of k=50k=50 might cover 99.9% of the mass for a factual prompt and only 60% for a creative one.

Combining Top-k with Temperature

Top-k sampling and temperature scaling are two different tools that address the same underlying problem from different angles. Temperature reshapes the entire distribution before any truncation occurs, while top-k determines the size of the sampling space after reshaping. Used together, they give you two independent dimensions of control over generation quality.

Temperature scaling, which we covered in the previous chapter, works by dividing all logits by a temperature parameter TT before applying softmax. When T<1T < 1, the distribution becomes more peaked, exaggerating the difference between high and low logits. When T>1T > 1, the distribution flattens, compressing those differences. The key point is that temperature operates on the full distribution and affects every token's probability.

Top-k, in contrast, operates after temperature has been applied. It looks at the shaped distribution and says: "From this reshaped distribution, I will only consider the top kk candidates." The combination means you first decide what shape of distribution you want (via temperature), then decide how many candidates to sample from (via kk). These are separate controls: you can have low temperature with high kk, high temperature with low kk, or any combination.

Think of the combination as a two-stage filter. Temperature is the first stage: it determines how spread out or concentrated the probability distribution looks. Top-k is the second stage: it draws a boundary around the highest-probability region of that distribution and refuses to sample outside it. Together, they give you fine-grained control that neither parameter alone provides.

Out[18]:
Visualization
Horizontal bar chart of top-20 tokens at temperature 0.5 showing concentrated probability mass in the top few tokens.
Low temperature (T=0.5) concentrates probability in top tokens, making k=20 effectively sample from fewer options.
Horizontal bar chart of top-20 tokens at temperature 1.5 showing more evenly spread probability across all kept tokens.
High temperature (T=1.5) flattens the distribution, spreading probability more evenly across all 20 kept tokens.

The figure shows how temperature reshapes the distribution within the top-kk candidates. At low temperature, the top token dominates even among the kept tokens, meaning the effective sampling space is much smaller than kk would suggest. At high temperature, probability spreads more evenly across all kk tokens, giving each candidate a meaningful chance.

Temperature-First vs Top-k-First

The order of operations matters. Standard practice applies temperature first, then top-k:

  1. Temperature scaling: Divide logits by temperature
  2. Top-k selection: Keep the kk highest values
  3. Softmax: Convert to probabilities
  4. Sample: Draw from the truncated distribution
In[19]:
Code
def top_k_sample_with_temperature(logits, k, temperature=1.0):
    """
    Top-k sampling with temperature applied first.

    This is the standard order: temperature reshapes the distribution,
    then top-k truncates, then we sample.
    """
    # Step 1: Temperature scaling
    scaled_logits = logits / temperature

    # Step 2: Get top-k
    top_k_vals, top_k_idx = torch.topk(scaled_logits, k)

    # Step 3: Softmax over top-k only
    top_k_probs = torch.softmax(top_k_vals, dim=0)

    # Step 4: Sample
    sample_idx = torch.multinomial(top_k_probs, num_samples=1)

    return top_k_idx[sample_idx].item()

Why temperature first? Temperature affects which tokens qualify as "top-k." When you apply high temperature, the differences between logits are compressed, which can change which tokens make the cut near the boundary. A token at rank 11 with the original distribution might move to rank 9 after high-temperature scaling because its neighbors' scores were squeezed more. By applying temperature before the top-k cut, you ensure that the set of candidates you are sampling from reflects the distribution you want to sample from, not the original distribution.

Applying temperature after top-k selection would mean truncating based on the original distribution, then reshaping the truncated distribution's probabilities. This gives slightly different results and is less theoretically clean. The standard order, temperature first, is conceptually cleaner: it says "here is the distribution I want, now take the top kk of it."

Out[20]:
Console
Prompt: 'The best programming language is'

Top-10 tokens at different temperatures:
--------------------------------------------------
T=0.5: [' Java', ' Python', ' the', ' a', ' not', ' C', ' one', ' JavaScript', ' Haskell', ' probably']
T=1.0: [' Java', ' Python', ' the', ' a', ' not', ' C', ' one', ' JavaScript', ' Haskell', ' probably']
T=2.0: [' Java', ' Python', ' the', ' a', ' not', ' C', ' one', ' JavaScript', ' Haskell', ' probably']

The output shows that the top-kk tokens remain mostly stable across temperatures, with the highest-ranked tokens appearing in all lists. However, tokens near the boundary (rank 8-10) may swap in and out as temperature changes, which can subtly affect sampling behavior. At very low temperature, the ranking becomes almost deterministic because logit differences are amplified. At high temperature, tokens near the boundary are nearly interchangeable, so the specific set included in the top-kk becomes somewhat arbitrary.

Practical Considerations

Moving from theory to production, several practical aspects affect how top-k sampling performs in real applications. This section covers computational efficiency, edge case handling, and batched generation, all of which matter when deploying language models at scale.

Top-k sampling is designed to add minimal computational overhead. The bottleneck in language model inference is the forward pass through the transformer, which involves matrix multiplications across the model's full parameter count. The decoding strategy is applied after the forward pass to the logit vector, which is a much simpler operation. Nevertheless, understanding the computational profile of top-k helps you reason about latency budgets and whether more complex decoding strategies are worth their overhead.

The implementation details also matter for numerical stability. Language model logits can be very large in magnitude, especially for well-trained models where the model is highly confident about certain tokens. Very large logits can cause numerical overflow when exponentiating, which is why we should always apply temperature scaling and work in logit space as long as possible before converting to probabilities.

Computational Efficiency

Top-k sampling adds minimal overhead to generation. The torch.topk operation is efficient, running in O(nlog⁡k)O(n \log k) time, where nn is the vocabulary size and kk is the number of tokens to select. Since k≪nk \ll n (typically kk is 50-100 while nn is 50,000+), this is essentially linear in vocabulary size.

In[21]:
Code
import time


def benchmark_top_k(
    vocab_size=50257, k_values=[10, 50, 100, 500], n_trials=1000
):
    """Benchmark top-k selection speed."""
    logits = torch.randn(vocab_size)

    results = {}
    for k in k_values:
        start = time.perf_counter()
        for _ in range(n_trials):
            _ = torch.topk(logits, k)
        elapsed = time.perf_counter() - start
        results[k] = elapsed / n_trials * 1000  # Convert to ms

    return results
Out[22]:
Console
Top-k selection time (per call):
  k= 10: 0.0923 ms
  k= 50: 0.1005 ms
  k=100: 0.1082 ms
  k=500: 0.1835 ms

The timings confirm that top-k selection takes only fractions of a millisecond, even for the largest kk values. Since a typical GPT-2 forward pass takes 10-50ms (and larger models take 100ms+), the top-k operation adds less than 1% overhead. This makes top-k sampling practical for production use without any performance concerns.

Handling Edge Cases

Several edge cases require attention in production implementations. These are situations where the naive implementation would produce incorrect or undefined behavior, and a reliable system needs to handle them gracefully.

The first edge case is when kk exceeds the vocabulary size. If someone passes k=100000k=100000 to a model with a 50,000-token vocabulary, the algorithm should gracefully fall back to sampling from the entire vocabulary rather than crashing. The second edge case is temperature at or near zero. A temperature of exactly zero causes division by zero, and temperatures very close to zero produce extremely large logits that can overflow in exponential calculations. The correct behavior at zero temperature is deterministic greedy selection of the most probable token.

The key edge cases to handle are:

  • k larger than vocabulary: If k≥k \geq vocabulary size, top-k degenerates to full sampling
  • All zero logits: Rare but possible with certain inputs; results in uniform sampling
  • Very small k: k=1k=1 is equivalent to greedy decoding
In[23]:
Code
def robust_top_k_sample(logits, k, temperature=1.0, min_tokens=1):
    """
    Robust top-k sampling with edge case handling.
    """
    vocab_size = logits.size(0)

    # Clamp k to valid range
    k = max(min_tokens, min(k, vocab_size))

    # Handle temperature
    if temperature <= 0:
        # Greedy selection
        return logits.argmax().item()

    scaled_logits = logits / temperature

    # Check for numerical issues
    if torch.isnan(scaled_logits).any() or torch.isinf(scaled_logits).any():
        # Fall back to argmax
        return logits.argmax().item()

    top_k_vals, top_k_idx = torch.topk(scaled_logits, k)
    top_k_probs = torch.softmax(top_k_vals, dim=0)

    # Handle case where all probabilities become 0 or nan
    if top_k_probs.sum() == 0 or torch.isnan(top_k_probs).any():
        # Uniform sampling over top-k
        sample_idx = torch.randint(0, k, (1,))
    else:
        sample_idx = torch.multinomial(top_k_probs, num_samples=1)

    return top_k_idx[sample_idx].item()

The reliable version above handles all the problematic cases: kk larger than vocabulary, temperature at or below zero, NaN or infinite logits from numerical instability, and probability collapse after softmax. Each fallback behavior is chosen to be sensible: when things go wrong, default to the most deterministic option (argmax) rather than creating garbage.

Batched Generation

For efficiency with multiple sequences, top-k can be applied in parallel across a batch. This is important for production systems that need to serve multiple users simultaneously or generate multiple candidate completions for evaluation.

The key difference from the single-sequence case is the dimension along which we apply torch.topk. For a batch of shape (batch_size, vocab_size), we want to take the top-k independently for each row, which is accomplished by passing dim=-1 to torch.topk.

In[24]:
Code
def batched_top_k_sample(logits, k, temperature=1.0):
    """
    Top-k sampling for batched logits.

    Args:
        logits: Shape (batch_size, vocab_size)
        k: Number of top tokens
        temperature: Temperature for scaling

    Returns:
        Tensor of sampled token indices, shape (batch_size,)
    """
    batch_size = logits.size(0)

    scaled_logits = logits / temperature
    top_k_vals, top_k_idx = torch.topk(scaled_logits, k, dim=-1)
    top_k_probs = torch.softmax(top_k_vals, dim=-1)

    # Sample one token per sequence
    sample_indices = torch.multinomial(top_k_probs, num_samples=1).squeeze(-1)

    # Gather the actual token IDs
    selected_tokens = top_k_idx[torch.arange(batch_size), sample_indices]

    return selected_tokens
Out[25]:
Console
Batch size: 4
Sampled token IDs: [29552, 19987, 48649, 40351]

The batched implementation processes all four sequences in a single operation, returning one sampled token ID per sequence. This vectorized approach is needed for efficient inference when generating multiple sequences in parallel, as it avoids the overhead of looping through sequences individually. On GPU hardware, the parallelization benefit is substantial: processing a batch of 32 sequences takes roughly the same wall-clock time as processing a single sequence.

Limitations of Top-k Sampling

While top-k is widely used and practically effective, it has notable limitations that have driven research into more sophisticated decoding strategies. Understanding these limitations is important both for knowing when top-k might fail and for appreciating the motivation behind alternatives like nucleus sampling, beam search, and minimum Bayes risk decoding.

The core tension in top-k sampling is that it uses a fixed count to define the sampling boundary, but the appropriate boundary is fundamentally probabilistic and context-dependent. The model itself does not think in terms of how many tokens should be reasonable; it thinks in terms of probabilities. By imposing a count-based cutoff, we are translating a probabilistic judgment into a structural one, and this translation inevitably loses information.

A second important limitation is that top-k sampling is memoryless at the selection level: it makes the same kind of decision at every generation step, regardless of what has been generated so far. Some researchers have argued that good text generation requires the system to track the "narrative direction" of the text and bias future sampling accordingly. Standard top-k has no such mechanism; it looks only at the current context and applies the same fixed-kk rule every time.

Fixed k Ignores Context

The basic limitation is that kk is fixed regardless of context. Consider two scenarios:

  • High-confidence context: "The president of the United States in 2020 was Donald", where very few tokens are reasonable next.
  • Open-ended context: "I think the best way to", where many tokens could reasonably follow.

A fixed kk treats both identically. If k=50k=50, the first case includes many implausible tokens, since the model might only need k=5k=5 to cover all reasonable options. In the second case, k=50k=50 might not be enough to capture the full range of plausible continuations, artificially restricting diversity when the model sees many good options.

This context-blindness of fixed-kk means that the effective quality of the sampling boundary varies wildly across prompts. For some contexts, k=50k=50 is overly generous and includes garbage. For others, it is overly restrictive and excludes good options. You cannot tune a single kk to be optimal across all possible contexts, because the contexts themselves vary in how concentrated or dispersed their distributions are.

Out[26]:
Visualization
Bar chart of token probabilities for a high-confidence prompt. A red dashed line marks k=50; the top few tokens hold nearly all probability mass.
High confidence context: k=50 includes many unnecessary low-probability tokens when the model is already certain.
Bar chart of token probabilities for a low-confidence prompt. A red dashed line marks k=50; probability is spread more evenly across the top tokens.
Low confidence context: k=50 may exclude viable options when probability is spread more evenly.

Quality-Diversity Trade-off

No single kk value works optimally across all situations. Lower kk improves quality but reduces diversity. Higher kk increases diversity but risks including low-quality tokens. This creates a trade-off when the goal is text that is both high quality and interestingly varied.

The challenge is amplified by the fact that quality and diversity are difficult to measure automatically. Human judgment is the gold standard, but it is expensive and slow. Automatic metrics like perplexity measure how "expected" the text is by the model, which correlates with quality but not with diversity. Metrics like distinct-n (the fraction of distinct n-grams in generated text) measure diversity but not quality. There is no single number that captures the balance well, which makes empirical kk tuning difficult.

Exposure Bias from Hard Truncation

A more subtle limitation is the hard boundary effect. A token ranked kk and a token ranked k+1k+1 might have nearly identical probabilities, but one is included and the other is completely excluded. This sharp boundary creates a kind of "exposure bias" where the model never has to deal with its own slightly-below-cutoff predictions during generation. During training, the model saw all tokens including low-probability ones. During inference with top-k, it never generates from that region of its distribution. This mismatch between training and inference can lead to subtle degradation in long-form generation.

Nucleus sampling (top-p), introduced by Holtzman et al. in 2019, addresses the fixed-kk limitation by instead choosing the smallest set of tokens whose cumulative probability exceeds a threshold pp. This allows the sampling set to shrink when the distribution is peaked and expand when it is flat, automatically adapting to each context. The next chapter explores this approach in detail, including the mathematical relationship between top-k and top-p and how to choose pp analogously to how you choose kk.

Repetition and Degeneration

A third practical limitation of top-k sampling (shared with most sampling methods) is its tendency to produce repetitive text in certain regimes. When the context strongly predicts specific tokens, those tokens dominate the top-kk set and keep getting sampled. The model then sees those tokens in the context, which further increases their probability at the next step, creating a self-reinforcing loop.

This repetition problem is mitigated by repetition penalties, which downweight the logits of tokens that have already appeared in the generated sequence. The penalty is applied before top-k selection, reducing the probability of repeated tokens and encouraging more diverse output. Systems like Hugging Face's generate() support a repetition_penalty parameter for this purpose. In practice, combining top-k with a mild repetition penalty (around 1.1 to 1.3) substantially improves the quality of long-form text generation.

Comparison with Other Decoding Methods

Let us compare top-k sampling with other decoding strategies to understand when each is most appropriate. The space of decoding strategies is broader than just top-k and temperature; understanding the full range helps you choose the right tool for each task.

Greedy decoding is the simplest strategy: always take the most probable token. It is deterministic and fast but produces repetitive, predictable text because it always commits to the single best local choice, which often leads to globally suboptimal sequences. Beam search improves on greedy by maintaining multiple candidate sequences simultaneously, but it tends to produce generic, averaged-out text that lacks the quality and diversity of human writing.

Temperature sampling is conceptually simple: reshape the distribution and sample from all tokens. It introduces diversity but does not eliminate the long-tail problem. Top-k addresses the long tail by truncation but uses a fixed count that does not adapt to distribution shape. Nucleus sampling adapts the truncation threshold to probability mass, addressing top-k's context-blindness. Each approach makes a different trade-off.

In[27]:
Code
def generate_comparison(model, tokenizer, prompt, max_tokens=15):
    """Generate using different decoding strategies."""
    results = {}

    # Greedy: k=1 (always take top token)
    results["greedy"] = generate_with_top_k(
        model, tokenizer, prompt, max_tokens=max_tokens, k=1, temperature=1.0
    )

    # Temperature sampling (T=0.8) with large k to approximate full-vocab sampling
    results["temperature_0.8"] = generate_with_top_k(
        model, tokenizer, prompt, max_tokens=max_tokens, k=200, temperature=0.8
    )

    # Top-k (k=50)
    results["top_k_50"] = generate_with_top_k(
        model, tokenizer, prompt, max_tokens=max_tokens, k=50, temperature=1.0
    )

    # Top-k + Temperature (k=50, T=0.8)
    results["top_k_50_temp_0.8"] = generate_with_top_k(
        model, tokenizer, prompt, max_tokens=max_tokens, k=50, temperature=0.8
    )

    return results
Out[28]:
Console
Prompt: 'The future of renewable energy'
============================================================

greedy:
  The future of renewable energy is in the hands of the people.

"We need to be

temperature_0.8:
  The future of renewable energy is a reality as it provides a source of clean energy for solar technologies.

top_k_50:
  The future of renewable energy and how governments will respond to climate change is still not well understood," said

top_k_50_temp_0.8:
  The future of renewable energy?

The European Commission will be watching where it goes from here.

The comparison reveals characteristic differences. Greedy decoding produces focused but potentially repetitive text. Pure temperature sampling adds variety but can drift. Top-k restricts the sampling space while preserving diversity. Combining top-k with temperature offers the most control, letting you to tune both how wide the candidate set is and how the probability is distributed within that set.

Out[29]:
Visualization
Bar chart of a greedy decoding distribution with only the highest-probability token highlighted.
Greedy decoding retains only the highest-probability token, producing deterministic output.
Bar chart of a temperature-scaled distribution with all ten tokens highlighted and probabilities flattened.
Temperature sampling reshapes the full distribution before sampling from every token.
Out[30]:
Visualization
Bar chart of a top-k distribution with the five highest-probability tokens highlighted.
Top-k sampling retains a fixed number of highest-probability tokens before sampling.
Bar chart of a nucleus distribution with the tokens inside the cumulative-probability boundary highlighted.
Nucleus sampling retains the smallest token set reaching a cumulative probability threshold.

Key Parameters

When using top-k sampling for text generation, the following parameters have the greatest impact on output quality:

  • k (top_k): The number of highest-probability tokens to keep. Values of 40-100 work well for most applications. Lower values (10-30) produce more focused, deterministic output suited for factual content. Higher values (100-200) allow more creative diversity but risk occasional incoherent tokens.

  • temperature: Scales the logits before computing probabilities. Values below 1.0 sharpen the distribution, making the top tokens more dominant. Values above 1.0 flatten the distribution, spreading probability more evenly. Common values range from 0.7 to 1.2. Temperature is typically applied before top-k selection.

  • do_sample: Boolean flag in Hugging Face's generate() method. Must be True to enable any sampling strategy including top-k. When False, the model uses greedy decoding regardless of other parameters.

  • max_new_tokens: Limits the number of tokens to generate. Longer generations may accumulate sampling noise, so lower kk values can help maintain coherence over extended outputs.

  • repetition_penalty: An additional parameter (not part of the core top-k algorithm, but commonly combined with it) that downweights the logits of previously generated tokens. Values slightly above 1.0, such as 1.1 or 1.2, reduce repetition without drastically altering the distribution.

Summary

Top-k sampling provides a practical solution to the long-tail problem in language model decoding. By truncating the vocabulary to the kk most likely tokens and renormalizing, it eliminates improbable tokens while preserving meaningful diversity among plausible choices.

Key takeaways from this chapter:

  • Truncation principle: Top-k zeros out all tokens outside the kk highest-probability candidates, then renormalizes the remaining probabilities. This prevents sampling from the incoherent long tail while letting variety among reasonable options.

  • Temperature interaction: Temperature scaling reshapes the distribution before truncation. Low temperature concentrates probability in fewer tokens; high temperature spreads it more evenly. Applying temperature first affects which tokens make the top-kk cut.

  • Parameter selection: Common values range from kk=40 to kk=100. Lower kk produces more focused text; higher kk allows more diversity. The optimal choice depends on task, context confidence, and user preferences.

  • Fixed-k limitation: A constant kk applies regardless of context, including too many tokens when the model is confident and potentially too few when it's uncertain. This motivates adaptive approaches like nucleus sampling.

  • Computational efficiency: Top-k selection adds negligible overhead, with O(nlog⁡k)O(n \log k) time complexity where nn is vocabulary size and kk is the number of tokens kept. This is fast compared to the model forward pass, which makes it practical for production use.

  • Edge cases: Production implementations should handle edge cases including kk larger than vocabulary size, near-zero temperature, and numerical instability in logits. A reliable fallback to greedy decoding ensures reliable behavior across all inputs.

Top-k sampling strikes a useful balance between greedy decoding (deterministic but repetitive) and pure sampling (diverse but sometimes incoherent). While nucleus sampling offers more adaptive truncation, top-k remains widely used due to its simplicity and interpretability. The next chapter explores nucleus sampling, which addresses the fixed-k limitation by adapting the threshold to each distribution's shape rather than maintaining a fixed count.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about top-k sampling and how it controls language model text generation.

Top-k Sampling Quiz

Question 1 of 80 of 8 completed
What problem does top-k sampling solve in language model text generation?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025topk, author = {Michael Brenndoerfer}, title = {Top-k Sampling: Controlling Language Model Text Generation}, year = {2025}, url = {https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Top-k Sampling: Controlling Language Model Text Generation. Retrieved from https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation
MLAAcademic
Michael Brenndoerfer. "Top-k Sampling: Controlling Language Model Text Generation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation>.
CHICAGOAcademic
Michael Brenndoerfer. "Top-k Sampling: Controlling Language Model Text Generation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Top-k Sampling: Controlling Language Model Text Generation'. Available at: https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Top-k Sampling: Controlling Language Model Text Generation. https://mbrenndoerfer.com/writing/top-k-sampling-language-model-text-generation

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.