Repetition Penalties: Preventing Loops in LLM Generation

Michael BrenndoerferUpdated August 1, 202556 min read

Part of Language AI Handbook

Repetition, frequency, and presence penalties alter token probabilities to reduce loops. Covers n-gram blocking, tuning, and effects on generated text.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Repetition Penalties

Language models have a peculiar tendency to get stuck in loops. Ask GPT to write a story, and without intervention, you might get output like "The cat sat on the mat. The cat sat on the mat. The cat sat on the mat..." This repetitive behavior emerges from a fundamental property of autoregressive generation: at each step, the model selects tokens that maximize likelihood given the context, and if a pattern worked once, the same pattern often has high probability again.

Think of it like a river carving a channel: once water flows a particular path, the channel deepens and that path becomes increasingly easy to follow. In language models, each generated token strengthens the patterns already in the context, gradually making those exact patterns more attractive at future steps. A model generating a list might produce one item, then another that uses the same structure, then another, then another, until it is trapped forever in a loop it cannot escape on its own. The channel has been carved so deep that the water cannot climb out.

Repetition is not always undesirable. Certain phrases naturally repeat in language: "again and again", "more and more", or the rhythmic repetition in poetry and song lyrics. Technical writing legitimately repeats specialized terms: a machine learning paper must say "gradient descent" many times. Legal documents repeat precise phrasings deliberately. But uncontrolled repetition signals a failure mode where the model has collapsed into a degenerate loop, producing text that no human would write intentionally. The challenge is to suppress pathological repetition while preserving the natural repetition that language requires.

This chapter explores four techniques to prevent such loops while preserving natural language patterns. Each operates at a different level of granularity and with a different mathematical mechanism. Understanding all four gives you a toolkit to match the intervention to the specific failure mode you are encountering.

We'll examine three probabilistic approaches that modify the probability distribution during generation. Repetition penalty scales down the logits of previously used tokens. Frequency penalty applies increasingly strong penalties based on how many times each token has appeared. Presence penalty applies a flat penalty to any token that has appeared at all. We'll also explore n-gram blocking, a deterministic approach that prevents exact phrase repetition. Each technique offers different trade-offs between preventing loops and preserving natural repetition.

Historical Context

Repetition in neural language models became a serious research concern around 2015-2017 as recurrent neural network (RNN) language models were being applied to open-ended generation tasks. Early sequence-to-sequence models for machine translation and summarization would sometimes produce highly repetitive outputs. The repetition penalty approach was formalized by Keskar et al. in the 2019 CTRL paper ("CTRL: A Conditional Transformer Language Model for Controllable Generation"), which introduced it as a straightforward logit-scaling technique. OpenAI popularized the frequency and presence penalty framing through their GPT-3 API in 2020, giving practitioners fine-grained control over different repetition patterns. N-gram blocking was widely adopted in summarization research, including BART (Lewis et al., 2020), where the no_repeat_ngram_size parameter became a standard configuration option. Today all major language model APIs and frameworks expose some variant of these mechanisms.

Why Models Repeat Themselves

Before diving into solutions, we should understand why repetition happens. The answer lies in how autoregressive models generate text and how they are trained. This is not a superficial bug but a deep consequence of the architecture and training objective.

During training, language models learn to predict the next token given all previous tokens. They're optimized to assign high probability to tokens that frequently follow specific contexts in the training data. When a particular phrase or structure appears often in training text, the model learns to give it high probability in similar contexts. The model is, in a sense, a compressed summary of what patterns appeared in the training corpus, and repetitive patterns show up strongly in that corpus.

The problem manifests during generation. Suppose the model generates "The results show that" and then produces "the model performs well." Now "The results show that the model performs well" becomes part of the context. If the training data contained similar patterns of presenting multiple results, the phrase "The results show that" might again have high probability, leading to another similar sentence. Each repetition reinforces the pattern in the context, making further repetition even more likely. The model is not doing anything wrong by its own internal logic: it is faithfully predicting what typically follows such contexts. The training data likely contained many documents that followed this structure, so the model learned to associate such contexts with such continuations.

There is also a more subtle contributor: maximum likelihood training creates pressure toward safe, high-probability tokens. During training, the model is penalized for predicting low-probability tokens and rewarded for predicting high-probability ones. This creates a systematic bias toward tokens the model is confident about, and confidence often correlates with familiarity. A token that just appeared is familiar: its embedding has just been activated, similar patterns have just been processed. It is, in a sense, "top of mind" for the model, making it an appealing choice at the next step.

The main point is that repetition is not random noise but a systematic attractor in the probability distribution. Once a phrase is in context, it raises the probability of similar phrases, which when generated raise the probability of even more similar phrases. Left unchecked, this positive feedback loop converges to a fixed point where the model endlessly repeats the same content.

In[5]:
Code
def generate_without_penalty(model, tokenizer, prompt, max_tokens=100):
    """Generate text without any repetition penalty."""
    input_ids = tokenizer.encode(prompt, return_tensors="pt")

    with torch.no_grad():
        output = model.generate(
            input_ids,
            max_new_tokens=max_tokens,
            do_sample=True,
            temperature=0.8,
            top_p=0.9,
            repetition_penalty=1.0,  # No penalty
            pad_token_id=tokenizer.eos_token_id,
        )

    return tokenizer.decode(output[0], skip_special_tokens=True)

The output may contain repetitive phrases or sentence structures. While sampling strategies like temperature and nucleus sampling add randomness, they don't specifically target repetition. A token that appeared recently still has the same probability as before, so the model can easily select it again. Temperature and sampling truncation change the shape of the distribution but do not specifically push previously used tokens toward lower probability. We need dedicated mechanisms for that.

The Feedback Loop

Repetition often starts subtly and escalates. The model might repeat a word, then a phrase, then entire sentences. This acceleration happens because each repetition adds more of the same pattern to the context, which the model uses to predict the next token. The context becomes increasingly dominated by the repeated content, making continuation of that content more and more likely.

Consider a model generating a list. After producing "First, we need to consider..." it might generate "Second, we need to consider..." and then "Third, we need to consider...". So far, this is reasonable parallel structure. But without intervention, the model might continue with "Fourth, we need to consider..." long after the list should have ended, trapped in a pattern that keeps reinforcing itself. The parallel structure pattern is strong enough in the training data that the model keeps predicting more of it.

The feedback loop has a characteristic shape. Early in generation, repetition probability is roughly at baseline. After a repeated phrase first appears, its probability edges up slightly. The second repetition raises it further. By the third or fourth repetition, the probability can be so high that almost any intervention short of hard blocking is insufficient. This is why the problem is called a "degenerate loop": the model has reached a stable fixed point from which it cannot easily escape through normal sampling variation. Understanding this escalation pattern helps explain why even moderate penalties are effective: they flatten the early growth of repetition probability before the feedback loop has a chance to establish itself.

The feedback loop also helps explain why beam search, a common decoding strategy, is particularly vulnerable to repetition. Beam search maintains multiple high-probability hypotheses and keeps expanding the most likely ones. Since repetition temporarily maximizes local probability, beam search selects into loops very aggressively. This is why n-gram blocking is most commonly associated with beam search in summarization: without hard blocking, beam search almost always collapses into repetitive output.

The Repetition Penalty

The most direct approach to preventing repetition modifies the logits for tokens that have already appeared in the generated sequence. The repetition penalty, introduced by Keskar et al. (2019) in the CTRL paper, divides the logits of previously seen tokens by a penalty factor, making them less likely to be selected.

Think of the repetition penalty as a "familiarity tax." Every token that has appeared before faces a handicap when competing for the next position. Fresh tokens, those not yet used, compete on their own merits. Used tokens start at a disadvantage proportional to the strength of the penalty. The effect is to tilt the playing field toward novelty without eliminating previously used tokens entirely.

This approach has an appealing simplicity. It requires only a single parameter, the penalty factor θ\theta, which has a natural interpretation: at θ=1.0\theta = 1.0 there is no effect, and higher values impose stronger penalties. Implementation requires only tracking which tokens have appeared, a set that is trivially maintained as generation proceeds. The computational overhead is minimal: one division or multiplication per token in the used-token set per generation step.

Repetition Penalty

The repetition penalty modifies logits for tokens that appear in the existing context. For each token tt that has already been generated, its logit is divided by the penalty factor θ\theta if positive, or multiplied by θ\theta if negative. This reduces the probability of repeating tokens without eliminating them entirely.

Mathematical Formulation

To understand how repetition penalty works, we need to trace the path from model output to token selection. When a language model predicts the next token, it doesn't output probabilities directly. Instead, it produces logits: raw, unnormalized scores for each token in the vocabulary. These logits then pass through the softmax function to become probabilities, and finally, we sample from that probability distribution.

This pipeline gives us a natural intervention point. If we want to reduce the probability of certain tokens, we can modify their logits before softmax. The question is: how should we modify them?

The Core Insight

Our goal is straightforward: make previously used tokens less likely to appear again. Since softmax converts logits to probabilities, reducing a logit reduces its corresponding probability. But there's a subtlety. Logits can be positive or negative, and the operation that reduces a positive number (division) increases a negative number (makes it closer to zero). We need an operation that pushes logits in the "less likely" direction regardless of their sign.

The solution is a piecewise function that applies different operations based on the sign of the logit. Let ziz_i be the logit for token ii and let GG be the set of tokens that have already appeared in the generated sequence. The repetition penalty modifies logits as follows:

zi′={zi/θif i∈G and zi>0zi⋅θif i∈G and zi≤0ziif i∉Gz'_i = \begin{cases} z_i / \theta & \text{if } i \in G \text{ and } z_i > 0 \\ z_i \cdot \theta & \text{if } i \in G \text{ and } z_i \leq 0 \\ z_i & \text{if } i \notin G \end{cases}

where:

  • ziz_i: the original logit for token ii, the raw score from the model before softmax
  • zi′z'_i: the modified logit after applying the penalty
  • θ\theta: the repetition penalty factor, typically in the range [1.0,2.0][1.0, 2.0]
  • GG: the set of token indices that have appeared in the context
  • i∈Gi \in G: notation meaning "token ii is a member of set GG" (i.e., token ii has appeared before)

The formula reads: for each token, check if it has appeared before. If it has and its logit is positive, divide by the penalty factor. If it has and its logit is zero or negative, multiply by the penalty factor. If it hasn't appeared, leave it unchanged.

Why does this formula make sense? Notice that in both the positive and negative cases, the modification moves the logit away from large positive values. For a positive logit, division reduces the magnitude (e.g., 4.0 becomes 2.0). For a negative logit, multiplication increases the magnitude in the negative direction (e.g., -2.0 becomes -3.0). Both operations reduce the value in the mathematical sense (the result is smaller or more negative), and because softmax is an increasing function of each logit, a smaller logit always maps to a smaller probability. The penalty is thus correctly reducing probability in both cases.

Why the Asymmetric Treatment?

This asymmetry might seem arbitrary, but it follows directly from how softmax works. Recall that softmax converts a logit ziz_i to a probability:

P(i)=ezi∑jezjP(i) = \frac{e^{z_i}}{\sum_j e^{z_j}}

The probability depends on ezie^{z_i}, the exponential of the logit. The exponential function is strictly increasing: larger inputs produce larger outputs. So to reduce P(i)P(i), we must reduce ziz_i.

where:

  • P(i)P(i): the probability assigned to token ii
  • ezie^{z_i}: the exponential of token ii's logit, which must be positive
  • ∑jezj\sum_j e^{z_j}: the sum of exponentials across all tokens in the vocabulary, serving as the normalizing constant

For a positive logit like zi=2.0z_i = 2.0, dividing by θ=1.5\theta = 1.5 gives zi′=1.33z'_i = 1.33. Since 1.33<2.01.33 < 2.0, we've reduced the logit and thus reduced the probability. Division works as intended.

For a negative logit like zi=−2.0z_i = -2.0, the same division gives zi′=−2.0/1.5=−1.33z'_i = -2.0 / 1.5 = -1.33. But −1.33>−2.0-1.33 > -2.0 (it's closer to zero), so we've increased the logit and the probability. Division fails here.

The fix is multiplication. Multiplying −2.0-2.0 by θ=1.5\theta = 1.5 gives zi′=−3.0z'_i = -3.0. Since −3.0<−2.0-3.0 < -2.0, we've pushed the logit further negative, reducing the probability. Both operations, division for positive logits and multiplication for negative ones, consistently reduce probability.

Why does this asymmetry arise in the first place? The root cause is that division by a number greater than 1 is not a "reduction" operation for negative numbers. Division by 1.5 takes -2.0 to -1.33, which is closer to 0 and thus larger. Multiplication by the same factor takes -2.0 to -3.0, which is further from 0 and thus smaller. The piecewise formula selects the operation that achieves the desired effect in each case.

The Neutral Case

When θ=1.0\theta = 1.0, the penalty has no effect. Dividing by 1 or multiplying by 1 leaves any number unchanged. This gives us a natural "off switch" for the penalty and a baseline for comparison. As θ\theta increases above 1.0, the penalty strengthens, pushing repeated tokens further toward improbability.

Out[6]:
Visualization
Line plot showing logit transformations for positive values under different penalty strengths.
Positive logits are divided by the penalty factor, pulling values toward zero and reducing probability.
Line plot showing logit transformations for negative values under different penalty strengths.
Negative logits are multiplied by the penalty factor, pushing values further negative and reducing probability.

The visualization shows the key insight: for positive logits, division pulls values toward zero (reducing probability), while for negative logits, multiplication pushes values further negative (also reducing probability). Both transformations achieve the same goal through different arithmetic operations. The dashed line at θ=1.0\theta = 1.0 shows the identity case where the penalty has no effect, serving as a useful reference for understanding how each higher θ\theta value moves tokens further from their original probability.

Implementation

Let's implement the repetition penalty from scratch to see exactly how it works:

In[7]:
Code
def apply_repetition_penalty(logits, generated_ids, penalty):
    """
    Apply repetition penalty to logits for tokens that have been generated.

    Args:
        logits: Tensor of shape (vocab_size,) containing raw logits
        generated_ids: Set of token IDs that have been generated
        penalty: Penalty factor (1.0 = no penalty, >1.0 = reduce repetition)

    Returns:
        Modified logits with penalty applied to generated tokens
    """
    modified_logits = logits.clone()

    for token_id in generated_ids:
        if modified_logits[token_id] > 0:
            modified_logits[token_id] = modified_logits[token_id] / penalty
        else:
            modified_logits[token_id] = modified_logits[token_id] * penalty

    return modified_logits

The function iterates through each token that has appeared in the context and applies the appropriate modification based on the sign of its logit. This is conceptually simple but reveals an important property: the penalty treats all repeated tokens equally, whether they appeared once or a hundred times. "The" after one use faces the same penalty as "the" after fifty uses. This equality is a design choice with consequences we will revisit when comparing to frequency penalty.

In[8]:
Code
# Create example logits and demonstrate the effect
vocab_size = 10
example_logits = torch.tensor(
    [2.5, -1.0, 3.2, 0.5, -0.3, 1.8, -2.0, 0.0, 1.2, -0.5]
)
token_names = [
    "the",
    "cat",
    "sat",
    "on",
    "mat",
    "dog",
    "ran",
    "and",
    "big",
    "small",
]

# Suppose tokens 0, 2, 3 have been generated (the, sat, on)
generated_tokens = {0, 2, 3}
penalty = 1.5

modified_logits = apply_repetition_penalty(
    example_logits, generated_tokens, penalty
)
Out[9]:
Console
Effect of repetition penalty (θ = 1.5):
------------------------------------------------------------
Token         Logit   Modified    P(orig)     P(mod)
------------------------------------------------------------
the            2.50       1.67      0.241      0.194 *
cat           -1.00      -1.00      0.007      0.013  
sat            3.20       2.13      0.485      0.309 *
on             0.50       0.33      0.033      0.051 *
mat           -0.30      -0.30      0.015      0.027  
dog            1.80       1.80      0.120      0.221  
ran           -2.00      -2.00      0.003      0.005  
and            0.00       0.00      0.020      0.037  
big            1.20       1.20      0.066      0.121  
small         -0.50      -0.50      0.012      0.022  
------------------------------------------------------------
* = token appeared in context (penalty applied)

The table shows how the repetition penalty redistributes probability mass. Tokens that appeared in the context (marked with *) have their probabilities reduced, while tokens that haven't appeared receive proportionally higher probabilities. Notice that "sat" (token 2), which had the highest original logit, is no longer the most likely token after applying the penalty. The probability mass that was taken from the penalized tokens flows to unpenalized tokens, effectively boosting the chances of less familiar choices.

Visualizing the Effect

Out[10]:
Visualization
Bar chart showing original token probabilities.
Original probability distribution before applying any penalty.
Bar chart showing penalized token probabilities with decreased probability for repeated tokens.
Distribution after repetition penalty (θ = 1.5). Red bars show penalized tokens with reduced probability.

Choosing the Penalty Value

The repetition penalty parameter requires careful tuning. Too low, and repetition persists. Too high, and the model avoids necessary repetition, producing awkward text that never uses "the" or "a" more than once. Finding the right value is partly a matter of matching the penalty to the type of content being generated.

The challenge is that function words like "the", "a", "is", "and" legitimately appear hundreds of times in any long document. A penalty strong enough to suppress pathological repetition of content words will, if applied uniformly, also heavily penalize these necessary function words. The repetition penalty applies to all tokens equally regardless of their type, which means tuning it always involves a trade-off between suppressing content repetition and preserving grammatical fluency. This limitation motivates the frequency and presence penalties discussed later, which can be calibrated to allow some repetition (of function words) while punishing excessive repetition (of content words).

Common values and their effects:

  • θ=1.0\theta = 1.0: No penalty applied. Baseline behavior with full repetition potential.
  • θ=1.1\theta = 1.1 to 1.21.2: Light penalty. Gently discourages repetition while allowing natural patterns. Good starting point for most applications.
  • θ=1.3\theta = 1.3 to 1.51.5: Moderate penalty. Noticeably reduces repetition. May affect fluency for texts requiring repeated terms (technical writing, legal documents).
  • θ>1.5\theta > 1.5: Strong penalty. Significantly suppresses any repeated token. Can produce stilted or unnatural text, especially for longer outputs.
In[11]:
Code
def generate_with_penalty(model, tokenizer, prompt, penalty, max_tokens=80):
    """Generate text with specified repetition penalty."""
    input_ids = tokenizer.encode(prompt, return_tensors="pt")

    with torch.no_grad():
        output = model.generate(
            input_ids,
            max_new_tokens=max_tokens,
            do_sample=True,
            temperature=0.8,
            top_p=0.9,
            repetition_penalty=penalty,
            pad_token_id=tokenizer.eos_token_id,
        )

    return tokenizer.decode(output[0], skip_special_tokens=True)

As the penalty increases, you'll notice the text becomes more varied in vocabulary but may also become less coherent. The model is forced to find alternative words even when repetition would be natural. At high penalty values, the output can read as if someone deliberately avoided using the same word twice, which is not how fluent writing works. The sweet spot depends on the specific task: creative writing benefits from more diversity, while technical writing where precision requires consistent terminology benefits from milder penalties.

Out[12]:
Visualization
Heatmap showing probability reduction percentages for different combinations of original logit values and penalty strengths.
Probability reduction as a function of original logit and penalty strength. Higher original logits experience larger absolute probability reductions, but the relative reduction is more uniform.

The heatmap reveals an important pattern: the probability reduction depends on both the penalty strength and the token's original standing. Tokens with moderate positive logits (around 1-3) experience the largest percentage reductions because they start with meaningful probability that can be substantially reduced. Very high logits still dominate even after penalty, while very low logits have little probability to lose. This means that the most problematic tokens, the ones the model is most confident about and most likely to repeat, are also the most resistant to moderate penalties. Extremely confident predictions require strong penalties to dislodge, while less confident predictions can be steered with mild ones.

Frequency Penalty

While repetition penalty treats all repeated tokens identically, frequency penalty scales with occurrence count. A token used once receives a small penalty; a token used ten times receives ten times that penalty. This graduated approach allows some repetition while strongly discouraging excessive use of any single token.

Think of frequency penalty as a progressive tax on word usage. The first time you use a word, you pay a small surcharge. The second time, you pay twice that surcharge. By the tenth use, the cumulative surcharge is substantial enough to make the word expensive to use again. Function words that legitimately appear throughout a document incur some cost, but content words that are being pathologically overused incur enormous cumulative cost. The graduated nature of the penalty naturally adapts to the difference between normal and pathological repetition.

This approach is particularly well suited to long-form generation tasks where you want the model to explore a topic broadly rather than drilling into one narrow aspect. When writing a blog post, you want to cover multiple subtopics and use diverse vocabulary. Frequency penalty naturally pushes the model in this direction by making it progressively more costly to return to the same words and phrases that have already appeared.

Frequency Penalty

Frequency penalty subtracts a value from each token's logit proportional to how many times that token has appeared in the generated text. The formula is zi′=zi−α⋅ciz'_i = z_i - \alpha \cdot c_i, where α\alpha is the penalty strength and cic_i is the number of times token ii has been generated.

Mathematical Formulation

Repetition penalty has a limitation: it treats a token that appeared once the same as one that appeared fifty times. Both receive the same penalty. But intuitively, we might want to allow occasional repetition (the word "the" naturally appears multiple times in most texts) while strongly discouraging tokens that have become overused. This calls for a penalty that accumulates with each occurrence.

From Binary to Graduated Penalties

Frequency penalty introduces a simple but powerful shift: instead of asking "has this token appeared?", we ask "how many times has this token appeared?" The penalty then scales proportionally with the answer.

The mechanism is additive rather than multiplicative. For each occurrence of a token, we subtract a fixed amount from its logit. Let cic_i denote the number of times token ii has appeared in the generated sequence. The frequency penalty modifies logits as:

zi′=zi−α⋅ciz'_i = z_i - \alpha \cdot c_i

where:

  • ziz_i: the original logit for token ii
  • zi′z'_i: the modified logit after applying the penalty
  • α\alpha: the frequency penalty coefficient, typically in the range [0,2][0, 2], controlling how strongly each occurrence is penalized
  • cic_i: the count of token ii in the generated sequence (0 if the token has not appeared)

The formula subtracts α\alpha once for each time the token has appeared. If a token appeared three times and α=0.5\alpha = 0.5, we subtract 0.5×3=1.50.5 \times 3 = 1.5 from its logit. Note that if ci=0c_i = 0, the formula applies no penalty at all, since subtracting zero has no effect. New tokens are left untouched.

Why does this formula make sense? Each appearance of a token is a "vote" that the model has already expressed its preference for that token. The frequency penalty accumulates these votes into a debt that the token must overcome to be chosen again. The more it has been chosen in the past, the heavier the debt, and the harder it must compete against fresher alternatives. This creates a natural tendency toward exploration: once the model has heavily used a token, it becomes progressively less attractive, and the model naturally turns to other options.

Why Additive Works

Unlike repetition penalty, frequency penalty doesn't need to handle positive and negative logits differently. Subtraction always reduces a number, regardless of its sign. Subtracting from a positive logit makes it smaller (or even negative). Subtracting from a negative logit makes it more negative. Both operations reduce the logit and thus reduce the probability after softmax.

This simplicity is a feature. The formula is easy to implement, easy to understand, and produces predictable behavior: each occurrence adds the same "cost" to using that token again. The mathematical mechanism is also more transparent: we know exactly how much the logit has been reduced by looking at the count, whereas the multiplicative mechanism of repetition penalty interacts with the original logit value in a more complex way.

Linear Scaling in Action

The linearity of frequency penalty creates a distinctive pattern. Early occurrences of a token receive mild penalties, allowing natural repetition of common words. But as a token accumulates, the penalty grows relentlessly. Consider α=0.5\alpha = 0.5:

Frequency penalty scaling with α = 0.5. The penalty grows linearly with occurrence count.
OccurrencesPenalty Applied
00.0
10.5
21.0
52.5
105.0

By the tenth occurrence, the token's logit has been reduced by 5.0, a substantial penalty that makes it far less competitive against tokens that haven't been overused. This graduated pressure naturally prevents any single token from dominating the output without harsh early constraints. Notice that a token appearing twice is only twice as penalized as one appearing once, which is quite lenient: the real impact accumulates at higher counts where pathological repetition is more likely to be occurring.

Implementation

In[13]:
Code
def apply_frequency_penalty(logits, token_counts, alpha):
    """
    Apply frequency penalty based on token occurrence counts.

    Args:
        logits: Tensor of shape (vocab_size,) containing raw logits
        token_counts: Dict mapping token IDs to their occurrence counts
        alpha: Frequency penalty coefficient

    Returns:
        Modified logits with frequency-based penalties applied
    """
    modified_logits = logits.clone()

    for token_id, count in token_counts.items():
        modified_logits[token_id] = modified_logits[token_id] - alpha * count

    return modified_logits

Let's see how frequency penalty differs from repetition penalty:

In[14]:
Code
# Create a scenario where tokens have different occurrence counts
example_logits = torch.tensor([2.0, 2.0, 2.0, 2.0, 2.0])
token_names = ["data", "model", "the", "learning", "result"]

# Token counts: "the" appeared 5 times, "model" appeared 3 times, "data" once
token_counts = {0: 1, 1: 3, 2: 5}  # data: 1, model: 3, the: 5

alpha = 0.5  # Frequency penalty coefficient
modified_logits = apply_frequency_penalty(example_logits, token_counts, alpha)
Out[15]:
Console
Frequency penalty effect (α = 0.5):
-----------------------------------------------------------------
Token         Count    Logit   Modified    P(orig)     P(mod)
-----------------------------------------------------------------
data              1     2.00       1.50      0.200      0.208
model             3     2.00       0.50      0.200      0.077
the               5     2.00      -0.50      0.200      0.028
learning          0     2.00       2.00      0.200      0.343
result            0     2.00       2.00      0.200      0.343
-----------------------------------------------------------------

The key difference is apparent: tokens with higher counts receive proportionally stronger penalties. "The", which appeared 5 times, has its logit reduced by 2.5 (5 × 0.5), while "data", which appeared once, only loses 0.5 from its logit. This graduated penalty is particularly effective for suppressing overused words while allowing reasonable repetition of less common terms. Notice that "learning" and "result", which have not appeared at all, receive no penalty and maintain their original probability relative to each other.

Visualizing Frequency Penalty

Out[16]:
Visualization
Line plot showing linear growth of frequency penalty with token count.
Frequency penalty grows linearly with token occurrence count for different α values.
Line plot showing probability decay as occurrence count increases.
Probability decay as a token appears more frequently (α = 0.5).

Presence Penalty

Presence penalty takes an even simpler approach: penalize any token that has appeared, regardless of how many times. This binary penalty treats "appeared once" and "appeared twenty times" identically. At first glance this might seem crude compared to the graduated precision of frequency penalty, but it captures a different goal.

Think of presence penalty as an entry tax rather than a per-use tax. The first time a token appears in the generated text, it pays a fixed fee. Subsequent appearances pay nothing extra. This structure is ideal when your primary concern is vocabulary breadth rather than repetition frequency. You want the model to explore as many different words and concepts as possible, with each new token representing a new idea being introduced. Once a concept is on the table, having it mentioned again is less concerning than having entirely new concepts go unmentioned.

This makes presence penalty particularly valuable for brainstorming tasks, where the goal is to generate as many distinct ideas as possible in a limited number of tokens. It also works well for summarization, where you want to ensure coverage of multiple aspects of the source document rather than repeatedly circling around one or two salient points. In both cases, the question is "have I introduced this yet?" rather than "how many times have I used this?", which maps directly to the binary structure of presence penalty.

Presence Penalty

Presence penalty applies a fixed penalty to the logit of any token that has appeared in the generated text, regardless of frequency. The formula is zi′=zi−β⋅1[i∈G]z'_i = z_i - \beta \cdot \mathbb{1}[i \in G], where β\beta is the penalty strength and 1[i∈G]\mathbb{1}[i \in G] is an indicator that equals 1 if token ii has appeared and 0 otherwise.

Mathematical Formulation

We've now seen two approaches: repetition penalty treats all repeated tokens equally (ignoring count), while frequency penalty scales with count. Presence penalty takes the opposite extreme from frequency penalty: it also ignores count, but uses the simpler additive mechanism.

The Binary Question

Presence penalty asks the simplest possible question about each token: "Have you appeared before, yes or no?" If yes, apply a fixed penalty. If no, leave the logit unchanged. The number of previous occurrences is irrelevant.

To express this mathematically, we use an indicator function, a standard notation for encoding yes/no conditions as 1/0 values:

zi′=zi−β⋅1[i∈G]z'_i = z_i - \beta \cdot \mathbb{1}[i \in G]

where:

  • ziz_i: the original logit for token ii
  • zi′z'_i: the modified logit after applying the penalty
  • β\beta: the presence penalty coefficient, controlling how strongly any prior appearance is penalized
  • 1[i∈G]\mathbb{1}[i \in G]: the indicator function, which equals 1 if token ii has appeared (is in set GG), and 0 otherwise
  • GG: the set of tokens that have appeared in the generated sequence

Reading the Indicator Function

The notation 1[i∈G]\mathbb{1}[i \in G] might look intimidating, but it simply acts like a switch. The condition inside the brackets, "i∈Gi \in G" (token ii is in set GG), is evaluated as true or false. The indicator function converts this to a number:

  • If the condition is true (token has appeared): 1[true]=1\mathbb{1}[\text{true}] = 1
  • If the condition is false (token has not appeared): 1[false]=0\mathbb{1}[\text{false}] = 0

Substituting back into the formula:

  • For a token that has appeared: zi′=zi−β⋅1=zi−βz'_i = z_i - \beta \cdot 1 = z_i - \beta
  • For a token that hasn't appeared: zi′=zi−β⋅0=ziz'_i = z_i - \beta \cdot 0 = z_i

The indicator function elegantly handles both cases in a single equation. In practice, you can implement this without any indicator function machinery: simply iterate over the set of appeared tokens and subtract β\beta from each. The indicator notation is useful for mathematical exposition but translates to a simple loop in code.

Why does this formula make sense? Because the logit space is linear in its effect on probability (through the exponentiation in softmax), a fixed subtraction of β\beta from a logit translates to dividing the unnormalized probability by eβe^\beta. So presence penalty is equivalent to multiplying all appeared tokens' unnormalized probabilities by a constant factor e−βe^{-\beta}, which is less than 1 for any positive β\beta. This is a clean, uniform discounting of all previously seen tokens.

Presence vs. Frequency: A Comparison

The contrast with frequency penalty is instructive. Frequency penalty applies α\alpha for each occurrence, so a token appearing 10 times receives 10 times the penalty of one appearing once. Presence penalty applies β\beta exactly once regardless of occurrence count. Whether a token appeared 1 time or 100 times, it receives the same penalty β\beta.

This makes presence penalty particularly effective for encouraging vocabulary diversity. Once a word has been used, it faces a fixed "tax" on appearing again. The model is pushed to explore alternatives, to find synonyms, to vary its phrasing. It is a blunt instrument compared to frequency penalty's graduated pressure, but for tasks like brainstorming or generating diverse lists, that bluntness is exactly what is needed. The model cannot "amortize" the penalty by spreading uses over time, the way it can with frequency penalty at low counts. The penalty is immediate and permanent from the first use.

Out[17]:
Visualization
Line plot comparing flat presence penalty curve against linearly increasing frequency penalty curves.
Presence penalty vs frequency penalty as occurrence count increases. Presence penalty (dashed) applies a constant penalty regardless of count, while frequency penalty (solid) grows linearly.

The contrast is stark. Presence penalty (dashed lines) jumps to its full value after the first occurrence and stays flat. Frequency penalty (solid lines) starts at zero and grows steadily. For tokens that appear many times, frequency penalty eventually dominates; for tokens that appear just once or twice, presence penalty applies stronger immediate pressure. If you want a token to be discouraged immediately upon first use, presence penalty is the sharper tool. If you want the penalty to grow with overuse rather than applying immediately, frequency penalty is better suited.

Implementation

In[18]:
Code
def apply_presence_penalty(logits, generated_tokens, beta):
    """
    Apply presence penalty to tokens that have appeared at least once.

    Args:
        logits: Tensor of shape (vocab_size,) containing raw logits
        generated_tokens: Set of token IDs that have appeared
        beta: Presence penalty coefficient

    Returns:
        Modified logits with presence-based penalties applied
    """
    modified_logits = logits.clone()

    for token_id in generated_tokens:
        modified_logits[token_id] = modified_logits[token_id] - beta

    return modified_logits

Notice how much simpler this implementation is compared to frequency penalty. We do not need to maintain a count dictionary, only a set of token IDs that have appeared. For long generation runs with large vocabularies, this can be a meaningful efficiency advantage, particularly when the set of appeared tokens is substantially smaller than the full vocabulary. The set can be maintained with O(1) average insertion and lookup time, making the penalty application efficient even for very long sequences.

Comparing All Three Penalties

Let's see how the three penalties differ when applied to the same scenario:

In[19]:
Code
# Setup: tokens with varying occurrence counts
logits = torch.tensor([2.0, 2.0, 2.0, 2.0, 2.0, 2.0])
names = ["novel", "data", "the", "model", "is", "good"]

# Counts: novel=0, data=1, the=5, model=3, is=2, good=0
counts = {1: 1, 2: 5, 3: 3, 4: 2}  # data, the, model, is
appeared = {1, 2, 3, 4}  # Set of tokens that appeared

# Apply each penalty type
rep_penalty = 1.5
freq_alpha = 0.3
pres_beta = 1.0

logits_rep = apply_repetition_penalty(logits, appeared, rep_penalty)
logits_freq = apply_frequency_penalty(logits, counts, freq_alpha)
logits_pres = apply_presence_penalty(logits, appeared, pres_beta)
Out[20]:
Console
Comparison of penalty types:
================================================================================
Token     Count   Original   Rep(θ=1.5)  Freq(α=0.3)    Pres(β=1)
================================================================================
novel         0      0.167        0.247        0.255        0.288
data          1      0.167        0.127        0.189        0.106
the           5      0.167        0.127        0.057        0.106
model         3      0.167        0.127        0.104        0.106
is            2      0.167        0.127        0.140        0.106
good          0      0.167        0.247        0.255        0.288
================================================================================
Out[21]:
Visualization
Grouped bar chart comparing original, repetition, frequency, and presence penalty probabilities across six tokens.
Comparison of three penalty types on the same token distribution. Frequency penalty scales with occurrence count, heavily penalizing 'the' (5 occurrences) while lightly penalizing 'data' (1 occurrence).

The visualization reveals each penalty's character:

  • Repetition penalty (blue) uniformly reduces all appeared tokens, regardless of count. "data" (×1) and "the" (×5) are penalized identically, which is the key limitation of this approach.
  • Frequency penalty (red) creates graduated reductions: "the" (×5) drops dramatically while "data" (×1) drops minimally. This is the most fine-grained of the three approaches.
  • Presence penalty (green) applies equal reduction to all appeared tokens, similar to repetition penalty but using additive rather than multiplicative adjustment. The difference from repetition penalty is visible in tokens with very negative logits, but for the equal-logit scenario here, the two approaches look similar.

N-gram Blocking

The penalties we've explored so far are probabilistic: they reduce the likelihood of repetition without preventing it entirely. N-gram blocking takes a deterministic approach by absolutely forbidding the repetition of specific token sequences. This is a qualitatively different intervention: rather than adjusting the shape of the probability distribution, it removes certain tokens from consideration entirely.

Think of n-gram blocking as a constraint rather than a preference. The probabilistic penalties say "you should probably not repeat this token." N-gram blocking says "you are not allowed to repeat this exact sequence." The difference matters when the model is highly confident: a very strong preference can override a probabilistic penalty, but it cannot override an absolute constraint.

This guarantee comes at a cost. Absolute constraints remove tokens from consideration without evaluating whether their repetition would be problematic in context. A model blocked from repeating "the quick brown" cannot use that phrase again even if it naturally continues the same thought. The question is whether the benefit of the guarantee outweighs the loss of flexibility in specific use cases.

N-gram Blocking

N-gram blocking prevents the model from generating any n-gram (sequence of n consecutive tokens) that has already appeared in the generated text. When the model would complete a forbidden n-gram, that token's probability is set to zero, forcing selection of a different continuation.

How N-gram Blocking Works

The penalties we've explored, repetition and frequency alongside presence penalties, all work by adjusting probabilities. They make repetition less likely but don't prevent it entirely. Sometimes a token's original probability is so high that even after penalization, it remains the most likely choice. For applications where exact repetition would be clearly wrong (legal documents, safety-critical outputs), we need a stronger guarantee.

N-gram blocking provides that guarantee through a fundamentally different mechanism: instead of adjusting probabilities, it eliminates certain tokens from consideration entirely by setting their probability to zero.

What Is an N-gram?

An n-gram is simply a contiguous sequence of n tokens. The terminology comes from computational linguistics:

  • A bigram (2-gram) is a sequence of 2 tokens, like ["the", "cat"]
  • A trigram (3-gram) is a sequence of 3 tokens, like ["sat", "on", "the"]
  • A 4-gram is a sequence of 4 tokens, like ["the", "quick", "brown", "fox"]

N-gram blocking maintains a record of all n-grams that have appeared in the generated text. Before each token is sampled, the algorithm checks: "If I generate this token, will it complete an n-gram that already exists?" If so, that token is forbidden.

The Blocking Mechanism

Consider 3-gram blocking with the sequence "machine learning is powerful. Machine learning is". The current context ends with "Machine learning". If the model generates "is", it would create the trigram ["Machine", "learning", "is"], which already appeared earlier. N-gram blocking detects this and sets the probability of "is" to zero, forcing the model to choose a different continuation.

This is absolute prevention, not probabilistic discouragement. No matter how strongly the model wants to generate "is", it cannot. The token is masked out before sampling occurs. The model must find an alternative that does not complete any previously seen n-gram of length n. This forces a novel local phrase structure and keeps exact phrases from being copied from earlier in the generation.

In[22]:
Code
def get_ngrams(token_ids, n):
    """
    Extract all n-grams from a sequence of token IDs.

    Args:
        token_ids: List of token IDs
        n: Size of n-grams to extract

    Returns:
        Set of n-gram tuples
    """
    ngrams = set()
    for i in range(len(token_ids) - n + 1):
        ngram = tuple(token_ids[i : i + n])
        ngrams.add(ngram)
    return ngrams


def get_banned_tokens(token_ids, n):
    """
    Get tokens that would create a repeated n-gram if generated next.

    Args:
        token_ids: List of previously generated token IDs
        n: Size of n-grams to block

    Returns:
        Set of token IDs that should not be generated
    """
    if len(token_ids) < n - 1:
        return set()

    # Get existing n-grams
    existing_ngrams = get_ngrams(token_ids, n)

    # Get the (n-1)-gram that would be completed by the next token
    context = tuple(token_ids[-(n - 1) :])

    # Find tokens that would complete a repeated n-gram
    banned = set()
    for ngram in existing_ngrams:
        if ngram[:-1] == context:
            banned.add(ngram[-1])

    return banned

Let's trace through an example:

In[23]:
Code
# Simulate a generation scenario
sentence = "The quick brown fox jumps over the lazy dog and the quick brown"
tokens = tokenizer.encode(sentence)
token_strings = [tokenizer.decode([t]) for t in tokens]
Out[24]:
Console
Sentence: 'The quick brown fox jumps over the lazy dog and the quick brown'

Tokens: ['The', ' quick', ' brown', ' fox', ' jumps', ' over', ' the', ' lazy', ' dog', ' and', ' the', ' quick', ' brown']

Existing 3-grams:
  ['the', 'quick', 'brown']
  ['the', 'lazy', 'dog']
  ['and', 'the', 'quick']
  ['The', 'quick', 'brown']
  ['over', 'the', 'lazy']
  ['quick', 'brown', 'fox']
  ['dog', 'and', 'the']
  ['brown', 'fox', 'jumps']
  ['lazy', 'dog', 'and']
  ['jumps', 'over', 'the']
  ['fox', 'jumps', 'over']

Context ends with: [' quick', ' brown']
Banned next tokens: [' fox']

The output shows all trigrams extracted from the sentence. Since the context ends with "quick brown" and the trigram ["quick", "brown", "fox"] already exists, generating "fox" next would create an exact repetition. N-gram blocking identifies this and adds "fox" to the banned token list, forcing the model to choose a different continuation.

Implementation in Generation

In[25]:
Code
def generate_with_ngram_blocking(model, tokenizer, prompt, n, max_tokens=50):
    """
    Generate text with n-gram blocking to prevent exact repetition.

    Args:
        model: Language model
        tokenizer: Tokenizer
        prompt: Input prompt string
        n: Size of n-grams to block (e.g., 3 for trigram blocking)
        max_tokens: Maximum tokens to generate

    Returns:
        Generated text string
    """
    input_ids = tokenizer.encode(prompt, return_tensors="pt")
    generated_ids = input_ids[0].tolist()

    for _ in range(max_tokens):
        with torch.no_grad():
            outputs = model(torch.tensor([generated_ids]))
            logits = outputs.logits[0, -1, :]

        # Get banned tokens
        banned_tokens = get_banned_tokens(generated_ids, n)

        # Mask out banned tokens
        for token_id in banned_tokens:
            logits[token_id] = float("-inf")

        # Apply temperature and sample
        probs = F.softmax(logits / 0.8, dim=0)
        next_token = torch.multinomial(probs, num_samples=1).item()

        generated_ids.append(next_token)

        if next_token == tokenizer.eos_token_id:
            break

    return tokenizer.decode(generated_ids, skip_special_tokens=True)

With 2-gram blocking, the text avoids repeating any two consecutive tokens, which can make the output feel choppy since common phrases like "of the" can only appear once. With 3-gram blocking, the constraint relaxes slightly, allowing more natural flow while still preventing obvious repetition. At 4-gram blocking, only longer repeated phrases are blocked, preserving most natural language patterns while catching more egregious loops.

The practical range of useful n values is 2-5. Below 2 makes no sense (you cannot block unigrams with this approach), and above 5 the blocking is so permissive that it rarely activates. In Hugging Face's generate() function, the default for summarization tasks is often 3 or 4, which catches most copy-paste repetition without over-constraining the output.

Trade-offs of N-gram Blocking

N-gram blocking guarantees no exact phrase repetition but has notable limitations that make it unsuitable as a universal solution. Understanding these limitations helps you decide when the guarantee is worth the constraints.

Rigidity: N-gram blocking cannot distinguish between undesirable repetition and natural language patterns. Phrases like "on the other hand" or "in addition to" might be blocked after first use, even when appropriate. The algorithm has no understanding of semantics: it treats all repeated n-grams as equally problematic. A repeated cliche is blocked just as readily as a repeated argument, even though the latter repetition might be deliberate and rhetorically appropriate.

Local focus: The blocking only prevents exact token-level matches. "The cat sat" and "A cat sat" are different trigrams, so both could appear despite semantic similarity. Two sentences that say the same thing in slightly different words will both be generated without any penalty, even though the semantic repetition might be just as problematic as token-level repetition. This gap between token-level and semantic-level repetition is a fundamental limitation of all the techniques in this chapter: they operate on tokens, not meaning.

Brittleness: The blocking can force awkward workarounds when the model needs to repeat a phrase. If the phrase "neural network" is blocked after first use, the model must avoid it for the rest of the generation, even when it is the most precise term available. The forced alternatives may be less clear or less accurate. In technical writing, where precision of terminology matters, this rigidity can reduce quality significantly.

Smaller n values (2 or 3) are more restrictive but may hurt fluency. Larger values (4 or 5) allow more natural repetition while still preventing obvious loops. As a rule of thumb, 3-gram blocking is the most common choice for general summarization and generation, while 4-gram or 5-gram blocking is better for tasks where natural repetition of short phrases is expected.

Combining Penalties with Sampling Strategies

In practice, repetition penalties work alongside temperature, top-k, and nucleus sampling. The interactions between these components matter for understanding and tuning generation quality. Each technique operates at a different stage of the logit-to-token pipeline, and applying them in the correct order is essential for the intended behavior.

The typical processing order is well-defined and follows a natural logic:

  1. Get raw logits from the model's final layer
  2. Apply repetition/frequency/presence penalties to adjust logit values for previously used tokens
  3. Apply temperature scaling to control distribution sharpness (dividing logits by temperature)
  4. Apply top-k or nucleus truncation to limit candidates to the most promising tokens
  5. Sample from the resulting distribution

This order is not arbitrary. Repetition penalties should be applied before temperature because the penalty magnitudes are calibrated for logit-scale values. If you apply temperature first (which rescales all logits), the same penalty value would have a different effective strength depending on the temperature used. Applying penalties first ensures consistent behavior regardless of temperature.

Similarly, sampling truncation (top-k or nucleus) should come after penalties because the penalties change which tokens are most competitive. If you truncated first, then applied penalties, you might inadvertently keep penalized tokens in the candidate set while excluding better alternatives. Applying penalties before truncation ensures the truncation operates on the post-penalty distribution, correctly identifying the best candidates after repetition is accounted for.

In[26]:
Code
def generate_with_combined_strategies(
    model,
    tokenizer,
    prompt,
    repetition_penalty=1.0,
    frequency_penalty=0.0,
    presence_penalty=0.0,
    temperature=1.0,
    top_p=0.9,
    max_tokens=50,
):
    """
    Generate text combining multiple penalty types with sampling strategies.
    """
    input_ids = tokenizer.encode(prompt, return_tensors="pt")
    generated_ids = input_ids[0].tolist()
    prompt_length = len(generated_ids)
    token_counts = {}

    for _ in range(max_tokens):
        with torch.no_grad():
            outputs = model(torch.tensor([generated_ids]))
            logits = outputs.logits[0, -1, :].clone()

        # Track tokens generated so far (excluding prompt)
        generated_so_far = generated_ids[prompt_length:]
        generated_set = set(generated_so_far)

        # Count occurrences
        for token_id in generated_so_far:
            token_counts[token_id] = token_counts.get(token_id, 0) + 1

        # Apply repetition penalty
        if repetition_penalty != 1.0:
            for token_id in generated_set:
                if logits[token_id] > 0:
                    logits[token_id] = logits[token_id] / repetition_penalty
                else:
                    logits[token_id] = logits[token_id] * repetition_penalty

        # Apply frequency penalty
        if frequency_penalty > 0:
            for token_id, count in token_counts.items():
                logits[token_id] = logits[token_id] - frequency_penalty * count

        # Apply presence penalty
        if presence_penalty > 0:
            for token_id in generated_set:
                logits[token_id] = logits[token_id] - presence_penalty

        # Apply temperature
        logits = logits / temperature

        # Apply nucleus sampling
        sorted_logits, sorted_indices = torch.sort(logits, descending=True)
        cumulative_probs = torch.cumsum(F.softmax(sorted_logits, dim=0), dim=0)
        sorted_indices_to_remove = cumulative_probs > top_p
        sorted_indices_to_remove[1:] = sorted_indices_to_remove[:-1].clone()
        sorted_indices_to_remove[0] = False
        indices_to_remove = sorted_indices_to_remove.scatter(
            0, sorted_indices, sorted_indices_to_remove
        )
        logits[indices_to_remove] = float("-inf")

        # Sample
        probs = F.softmax(logits, dim=0)
        next_token = torch.multinomial(probs, num_samples=1).item()
        generated_ids.append(next_token)

        if next_token == tokenizer.eos_token_id:
            break

    return tokenizer.decode(generated_ids, skip_special_tokens=True)
Out[27]:
Console
Prompt: 'Here are some ideas for improving productivity:'

======================================================================

No penalties:
Here are some ideas for improving productivity:

Show Me The Details

First off, let's talk about how much time I spend on Twitter. I know you're not the only one who's tired of being told that you shouldn't spend more time on Twitter than you should. I know that I want to see my tweets read more and more frequently.

This is something
----------------------------------------------------------------------

Repetition only (θ=1.2):
Here are some ideas for improving productivity:
"Don't make a bunch of extra widgets just to get stuff done. Put them on top and do it well." – Gary Dewey, YouTube Engineer (not long ago). Stop making some arbitrary split-screen widget with no taskbar or multitasking effects that makes your life feel like shit for hours at the least! New features: V
----------------------------------------------------------------------

Frequency only (α=0.5):
Here are some ideas for improving productivity:

Show Me The Details: Show me what you need to get done at work. Make it very clear where the project is being held, and how productive that piece of information will be for a change in your life or career! Focus on one thing per task group; focus more often than not (with less distraction) On each separate activity specific
----------------------------------------------------------------------

Presence only (β=0.8):
Here are some ideas for improving productivity:

Show Me The Details: Show me what you need to get done at work. Make it very clear where the project is being done, where you're being done and where you're going. Show people the value that you have created. Show people the benefits of your work. Show them how to spend their money.

This all comes
----------------------------------------------------------------------

Combined (θ=1.1, α=0.3):
Here are some ideas for improving productivity:
"Don't make a bunch of extra widgets just to get stuff done. Put them on top and do it well." – Gary Dewey, YouTube Engineer (not long ago). Stop making some arbitrary split-screen widget with no taskbar or multitasking effects that makes your life feel like shit for hours at the least! New features: V
----------------------------------------------------------------------

Comparing the outputs reveals how each penalty type shapes generation. Without penalties, the model may fall into repetitive patterns. Repetition penalty alone provides uniform discouragement of repeated tokens. Frequency penalty creates graduated pressure that builds as tokens accumulate, while presence penalty encourages the model to continuously introduce new vocabulary. The combined approach balances these effects, often producing the most natural-sounding output by gently discouraging repetition without forcing unnatural word choices.

You can also combine penalties with n-gram blocking for a hybrid approach: the probabilistic penalties soften the probability distribution to make repetition less attractive, while n-gram blocking provides a hard guarantee against the most egregious exact repetitions. This combination is particularly effective for tasks like abstractive summarization, where you want rich vocabulary diversity but also need to ensure the summary does not copy entire sentences from the source document.

Worked Example: Tracing a Penalty Step by Step

To solidify understanding, let's trace through a concrete numerical example that shows exactly how frequency penalty modifies the generation decision at one step. We'll use a toy vocabulary of 5 tokens and walk through every calculation.

Suppose we are generating text and have just produced the sequence: "the cat sat on the mat and the cat". The current context window contains these 8 tokens. We track their counts:

  • "the": appeared 3 times
  • "cat": appeared 2 times
  • "sat": appeared 1 time
  • "on": appeared 1 time
  • "mat": appeared 1 time
  • "and": appeared 1 time

The model produces the following raw logits for the next token (we use a tiny vocabulary for clarity):

TokenRaw Logit
"the"3.2
"cat"2.8
"sat"2.1
"again"1.5
"purred"1.0

Without any penalty, softmax would strongly favor "the" and "cat". Let's apply frequency penalty with α=0.5\alpha = 0.5.

Step 1: Compute modified logits. For each token, subtract α×ci\alpha \times c_i:

  • "the": 3.2−0.5×3=3.2−1.5=1.73.2 - 0.5 \times 3 = 3.2 - 1.5 = 1.7
  • "cat": 2.8−0.5×2=2.8−1.0=1.82.8 - 0.5 \times 2 = 2.8 - 1.0 = 1.8
  • "sat": 2.1−0.5×1=2.1−0.5=1.62.1 - 0.5 \times 1 = 2.1 - 0.5 = 1.6
  • "again": 1.5−0.5×0=1.51.5 - 0.5 \times 0 = 1.5 (no penalty: first use)
  • "purred": 1.0−0.5×0=1.01.0 - 0.5 \times 0 = 1.0 (no penalty: first use)

Step 2: Apply softmax to modified logits. We compute ezi′e^{z'_i} for each:

  • "the": e1.7≈5.47e^{1.7} \approx 5.47
  • "cat": e1.8≈6.05e^{1.8} \approx 6.05
  • "sat": e1.6≈4.95e^{1.6} \approx 4.95
  • "again": e1.5≈4.48e^{1.5} \approx 4.48
  • "purred": e1.0≈2.72e^{1.0} \approx 2.72

Sum: 5.47+6.05+4.95+4.48+2.72=23.675.47 + 6.05 + 4.95 + 4.48 + 2.72 = 23.67

Probabilities: "the" = 0.231, "cat" = 0.256, "sat" = 0.209, "again" = 0.189, "purred" = 0.115

Step 3: Compare to unpenalized probabilities. With the original logits, the top choices were overwhelmingly "the" and "cat". After frequency penalty, "cat" is still favored but by a much smaller margin, and "again" and "purred" have become viable competitors. The penalty has not eliminated the model's preference but has substantially leveled the playing field, giving novel tokens a real chance at selection.

The key insight from this example is that frequency penalty is most effective when the penalized tokens have modest logit advantages over alternatives. If "the" had a logit of 10.0 rather than 3.2, the same α=0.5\alpha = 0.5 with count 3 would reduce it to 8.5, which would still completely dominate. This is why very confident model predictions can resist even strong penalties, and why n-gram blocking is sometimes necessary as a fallback guarantee.

When to Use Each Penalty

The choice between penalty types depends on your generation task. Each technique has a domain where it performs best, and using the wrong technique can introduce problems while failing to solve the ones you have.

Repetition Penalty works well as a general-purpose solution. It's the most widely implemented (available in Hugging Face's generate() method) and handles most cases adequately. Use it when you want a simple, effective baseline for reducing repetition, particularly when you're not sure which penalty type is most appropriate. The single-parameter interface makes it easy to tune and reason about. Typical starting value: θ=1.1\theta = 1.1 to 1.21.2.

Frequency Penalty excels when you want to allow some repetition while preventing excessive use. It's particularly useful for:

  • Long-form content where occasional word repetition is acceptable but word overuse is not
  • Creative writing where you want natural variation without harsh constraints
  • Technical writing where certain terms must repeat but shouldn't dominate

Presence Penalty promotes vocabulary diversity and topic breadth. Consider it for:

  • Brainstorming and idea generation where you want many distinct concepts
  • Summarization where you want to cover multiple points without redundancy
  • Conversational agents that should avoid fixating on specific words

N-gram Blocking provides guaranteed protection against exact repetition. Use it when:

  • Exact phrase repetition would be clearly wrong (e.g., safety-critical applications)
  • You need deterministic behavior rather than probabilistic reduction
  • Other penalties aren't sufficiently preventing loops

For many applications, combining moderate repetition penalty (θ≈1.1\theta \approx 1.1-1.21.2) with nucleus sampling provides a good balance between preventing loops and maintaining fluent output.

Key Parameters

When implementing repetition penalties in generation, these parameters control behavior:

  • repetition_penalty (float, typically 1.0-2.0): Multiplicative penalty applied to logits of repeated tokens. A value of 1.0 applies no penalty. Values around 1.1-1.3 provide gentle repetition reduction. Values above 1.5 strongly discourage any repetition.

  • frequency_penalty (float, typically 0-2.0): Additive penalty proportional to token occurrence count. Higher values increasingly penalize frequently used tokens. OpenAI's API uses this parameter with typical values of 0-1.0.

  • presence_penalty (float, typically 0-2.0): Flat additive penalty for tokens that have appeared at all. Encourages vocabulary diversity. Also used in OpenAI's API with typical values of 0-1.0.

  • no_repeat_ngram_size (int): Size of n-grams to block from repeating. In Hugging Face's generate(), setting this to 2 prevents any bigram repetition, 3 prevents trigram repetition, etc. Set to 0 to disable.

  • encoder_no_repeat_ngram_size (int): For encoder-decoder models, prevents generating n-grams that appear in the encoder input. Useful for abstractive summarization to avoid copying source phrases.

Limitations

Repetition penalties, for all their utility, have several important limitations that practitioners need to understand before deploying them in production systems.

The most fundamental limitation is that all these techniques operate on token identity, not semantic content. They can prevent the model from repeating the exact token sequence ["neural", "network"] but cannot prevent semantic repetition: sentences that convey the same information using different words. A model avoiding "neural network" might say "the artificial intelligence system" instead, producing essentially identical meaning with no token overlap. If your actual concern is semantic redundancy rather than surface-level repetition, these token-level techniques address only a symptom of the problem, not the underlying cause. Semantic-level repetition detection would require computing embeddings and measuring semantic similarity, which is computationally much more expensive.

The tuning challenge is also significant. The right parameter values depend heavily on the task, the model, the prompt length, and the expected output length. There are no universal values that work well across all settings. A θ=1.5\theta = 1.5 that prevents pathological loops in story generation might produce stilted, awkward text in a domain requiring technical precision. Frequency penalty values that work well for a 200-token response may over-penalize function words in a 2000-token essay. This means practitioners typically need to run empirical evaluations with human raters to find good settings for each deployment context, which requires significant time and resources.

There is also an inherent tension between repetition prevention and coherence. Language models achieve coherence partly through consistent use of terminology and consistent patterns of expression. A model that never repeats a content word may feel incoherent, jumping from synonym to synonym in ways that are hard to follow. Repetition and coherence are not fully separable: some amount of repetition is needed for text to feel like a unified whole rather than a collection of unrelated sentences. Strong penalties that eliminate most repetition can make text feel fragmented, even if it is technically non-repetitive. Finding the balance requires accepting that you cannot fully optimize both goals simultaneously.

Finally, these methods are reactive rather than proactive. They respond to repetition after it begins to occur, rather than preventing the model from entering repetition-prone states in the first place. More sophisticated approaches to this problem include training with diversity-promoting objectives, using contrastive decoding to compare outputs with a smaller "amateur" model, or training with reinforcement learning from human feedback where repetitive outputs receive lower rewards. These methods address the root cause rather than patching the symptom at inference time, but they require access to the training process, which is not always available.

Summary

Repetition is a common failure mode in language model generation, emerging from the autoregressive process where patterns reinforce themselves through context accumulation. The feedback loop at the heart of this failure mode is well-understood: each generated token becomes part of the context used to predict the next token, and high-probability patterns in the training data become increasingly attractive attractors that the model can spiral into. Several techniques address this problem by modifying token probabilities based on prior usage.

Key takeaways:

  • Repetition penalty divides logits of previously generated tokens by a penalty factor, uniformly reducing their probability regardless of occurrence count. Simple and effective for general use.

  • Frequency penalty subtracts from logits proportionally to occurrence count, creating graduated penalties that scale with overuse. Allows natural repetition while strongly penalizing excessive use.

  • Presence penalty applies a flat reduction to any token that has appeared, promoting vocabulary diversity. Effective for brainstorming and ensuring broad coverage.

  • N-gram blocking deterministically prevents exact phrase repetition by masking tokens that would complete a previously seen n-gram. Provides guaranteed protection but can reduce fluency.

  • Order matters: Apply penalties before temperature and sampling truncation. The modified logits then flow through the standard sampling pipeline.

  • Context window: Consider whether penalties apply to the full context (including prompt) or only generated tokens. Most implementations penalize tokens anywhere in the sequence, which can affect prompt-relevant words. Some frameworks allow specifying that only the generated portion should be tracked.

  • Limitations: All four techniques operate on token identity rather than semantic content, cannot prevent semantic repetition through paraphrase, and require empirical tuning for each deployment context.

The optimal configuration depends heavily on your use case. Start with a moderate repetition penalty (θ=1.1\theta = 1.1-1.21.2), observe the output quality, and adjust based on whether you see too much repetition or unnaturally varied vocabulary. For production systems, A/B testing different configurations against human evaluations provides the most reliable guidance. When probabilistic penalties are insufficient, n-gram blocking provides a hard guarantee at the cost of reduced flexibility.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about repetition penalties in language model generation.

Repetition Penalties Quiz

Question 1 of 80 of 8 completed
Why does repetition penalty apply different operations to positive and negative logits?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025repetitionpenalties, author = {Michael Brenndoerfer}, title = {Repetition Penalties: Preventing Loops in LLM Generation}, year = {2025}, url = {https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Repetition Penalties: Preventing Loops in LLM Generation. Retrieved from https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation
MLAAcademic
Michael Brenndoerfer. "Repetition Penalties: Preventing Loops in LLM Generation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation>.
CHICAGOAcademic
Michael Brenndoerfer. "Repetition Penalties: Preventing Loops in LLM Generation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Repetition Penalties: Preventing Loops in LLM Generation'. Available at: https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Repetition Penalties: Preventing Loops in LLM Generation. https://mbrenndoerfer.com/writing/repetition-penalties-language-model-generation

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.