LLM Watermarking: Schemes, Detection, and Robustness

Michael BrenndoerferMarch 13, 202665 min read

Part of Language AI Handbook

Explains how token-level watermarking embeds hidden statistical signals into LLM outputs, enabling cryptographically verifiable attribution and AI provenance.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Watermarking

When an LLM generates text, it leaves no fingerprint. The output looks like any other human-written paragraph. You cannot tell whether a news article, a homework essay, or a product review came from a language model just by reading it. This opacity creates a serious problem: if AI-generated text can circulate undetected, it becomes impossible to enforce academic integrity policies, identify AI-driven misinformation campaigns, audit model outputs for safety, or attribute text to its source.

Watermarking addresses this by embedding a hidden statistical signal into generated text at inference time. The signal is invisible to readers but detectable by anyone who knows the key. Unlike post-hoc classifiers that try to guess whether text was machine-written based on stylistic patterns, watermarking schemes work proactively. The LLM's decoding process is modified to produce text that carries a verifiable, cryptographically secured signature.

The core idea is elegant: if you can influence which tokens the model selects during generation, you can bias those choices in ways that are statistically detectable but imperceptible to human readers. A detector with the right key can compute a test statistic over a piece of text and determine, with high statistical confidence, whether the signal is present. The signal strength accumulates with text length, so longer passages become increasingly certain attributions while shorter ones remain ambiguous.

What makes this different from simply adding metadata to a file? Metadata is trivially stripped. A watermark, by contrast, is woven into the fabric of the text itself, residing in the specific word choices rather than in any attached wrapper. You cannot remove it without changing the words, and changing the words costs something: either effort, quality, or both.

This chapter covers the main families of watermarking schemes, the statistical machinery used for detection, the robustness challenges that attackers exploit, and the metrics used to evaluate watermark quality. Watermarking applies information theory and hypothesis testing to language model decoding. The upcoming chapter on Model Cards connects back to watermarking in the context of documenting model capabilities and responsible disclosure of watermarking policies.

Why Watermarking Matters

The case for watermarking comes from a simple asymmetry: generating convincing AI text is cheap and fast, but detecting it after the fact is hard. Post-hoc detection classifiers have inherent limitations. They require training data that spans the current generation of models, they degrade when models are updated or fine-tuned, and they produce false positives that can harm innocent people. A student accused of using ChatGPT based on a classifier's output may have written every word themselves.

The core issue with classifier-based detection is that it is an inductive problem. The classifier learns a boundary between "looks like human text" and "looks like AI text" from examples, but that boundary shifts whenever models improve. GPT-2 text was once reliably detectable; GPT-4 text is much harder to distinguish. A watermark-based system sidesteps this arms race entirely. The question "does this text contain the signal we embedded?" has a precise statistical answer regardless of how fluent the model has become.

Watermarking also gives you a different kind of certainty. A classifier says "I am 87% confident this was AI-written," a probabilistic assessment over a learned representation. A watermark detector says "under the null hypothesis that no watermark was present, the probability of observing this many green tokens is less than 0.001," a statement grounded in combinatorics and probability theory. The confidence is calibratable: you can set an exact false positive rate and hold it.

The practical applications extend well beyond academic integrity:

  • Provenance tracking: Identifying which model generated a piece of text, and when. Different users or time periods can receive different keys, creating a chain of custody.
  • Safety auditing: Verifying that a deployed model is using the watermarked decoding path rather than an unmodified one. If outputs from a production system lack the watermark, the operator can detect tampering.
  • Accountability: Creating an auditable record of model outputs used in high-stakes decisions, such as medical summaries or legal briefs.
  • Copyright and training data: Distinguishing model-generated content from human-written content in datasets, preventing the feedback loop where AI-generated text is used to train future models.
  • Regulatory compliance: Emerging AI regulations in several jurisdictions are beginning to require disclosure of AI-generated content in certain contexts. Watermarking provides a technical mechanism for such disclosure.

Watermarking is not a complete solution to AI safety or misuse, and we will discuss its limitations carefully throughout this chapter. But it is a powerful primitive that enables a class of accountability mechanisms that would otherwise be technically impossible.

Watermarking Schemes

The field has converged on two main families of watermarking schemes: token-level schemes that operate on the sampling distribution at each decoding step, and distortion-free schemes that modify how sampling randomness is used without changing the marginal output distribution. Token-level schemes have dominated research because they offer strong statistical guarantees and integrate cleanly into the decoding pipeline. Both families share the requirement that the watermark key must be kept private for the scheme to remain secure.

Before diving into the mechanics, it helps to think about what a watermarking scheme must accomplish simultaneously:

  • Detectability: The signal must be strong enough that a detector can reliably confirm its presence in a reasonable amount of text.
  • Imperceptibility: The text must still read like natural language. A watermark that causes the model to produce awkward or ungrammatical sentences would be both obvious and counterproductive.
  • Security: An attacker who does not know the key should not be able to remove the watermark without degrading text quality significantly, nor should they be able to forge a watermark to falsely attribute text.

These three requirements exist in tension, and much of the research in this area examines the tradeoffs between them.

Token-Level Watermarking

The key insight in token-level watermarking is that language model decoding involves choosing tokens from a probability distribution at each step. If you can systematically bias those choices, even slightly, you introduce a pattern that accumulates over many tokens and becomes statistically detectable.

The most influential scheme was introduced by Kirchenbauer et al. (2023), often called the KGW watermark. It works by partitioning the vocabulary into two sets at each token position: a "green" set and a "red" set. The partition depends on the previous token, using a hash function keyed to a secret key. During generation, the logits for green-list tokens are increased by a fixed additive amount δ\delta, making the model more likely to select them. Tokens from the red list are not penalized but compete against boosted green tokens.

To understand why this works, consider what happens to the softmax distribution when you add δ\delta to a subset of logits. If token ii has logit lil_i, then adding δ\delta to green tokens changes the selection probability. For green token gg versus red token rr:

P(green token g)P(red token r)=elg+δelr=elg−lr+δ\frac{P(\text{green token } g)}{P(\text{red token } r)} = \frac{e^{l_g + \delta}}{e^{l_r}} = e^{l_g - l_r + \delta}

Even when the red token has a higher base logit (lr>lgl_r > l_g), if δ\delta is large enough, the green token wins. This is the boost mechanism: you are not forcing specific tokens but tilting the odds.

More precisely, before sampling token tt, the scheme:

  1. Hashes the previous token t−1t-1 (or a window of previous tokens) together with the secret key to obtain a pseudo-random partition
  2. Marks a fraction γ\gamma of the vocabulary as green
  3. Adds δ\delta to the logit of every green-list token
  4. Samples normally from the modified distribution

The detection procedure counts how many green tokens appear in the candidate text, computes a z-score under the null hypothesis that the text was generated without watermarking, and rejects the null if the score exceeds a threshold.

The partition changes at every step because it depends on the current context. This means you cannot predict future green lists from past ones without the key, which prevents attackers from simply avoiding green tokens. If you change token t−1t-1 to avoid a red-token situation at position tt, you simultaneously change the partition at position t+1t+1. The dependency chain makes targeted removal difficult.

Green List / Red List Partition

At each decoding step, the KGW scheme uses a hash of the previous token and a secret key to split the vocabulary into a green list (fraction γ\gamma, typically 0.5) and a red list (fraction 1−γ1 - \gamma). During generation, green-list logits are boosted by δ\delta. During detection, a z-test counts how many tokens in the candidate text are green under the same partition, using the same key to reproduce the identical partition for each position.

The KGW Detection Statistic

The statistical heart of the KGW scheme is a hypothesis test. Let TT be the number of tokens in the candidate text, and let ∣s∣G|s|_G denote the count of green tokens, where "green" is determined by the same hash function and key used during generation.

Under the null hypothesis H0H_0 (text was not watermarked), each token is independently green with probability γ\gamma. Why? Because without the logit boost, the token choice is determined solely by the model's learned distribution, which has no reason to match any particular hash-derived split. The expected number of green tokens is γT\gamma T, and by the central limit theorem, the count is approximately normally distributed with mean γT\gamma T and variance T⋅γ(1−γ)T \cdot \gamma(1 - \gamma).

The z-score is computed as:

z=∣s∣G−γTT⋅γ(1−γ)z = \frac{|s|_G - \gamma T}{\sqrt{T \cdot \gamma(1 - \gamma)}}

where:

  • ∣s∣G|s|_G: the number of tokens in the text that fall in the green list under the key-derived partition
  • TT: total number of tokens in the text
  • γ\gamma: the fraction of the vocabulary assigned to the green list (typically 0.5)
  • T⋅γ(1−γ)T \cdot \gamma(1 - \gamma): the variance of a binomial distribution with TT trials and success probability γ\gamma

This is a standard z-test for a proportion. The numerator measures how far the observed green count is from what you would expect by chance. The denominator normalizes by the standard deviation, putting the statistic on a scale where values above roughly 2 are unlikely under the null.

If the text was watermarked, green tokens appear at a rate higher than γ\gamma because their logits were boosted by δ\delta during generation. A large positive zz score indicates the presence of the watermark. The detector rejects H0H_0 and concludes the text is watermarked when zz exceeds a threshold zαz_\alpha, where α\alpha is the desired false positive rate.

For a one-sided test at significance level α=0.01\alpha = 0.01, the threshold is zα≈2.33z_\alpha \approx 2.33. For α=0.001\alpha = 0.001, the threshold rises to zα≈3.09z_\alpha \approx 3.09. These correspond to the upper tail of the standard normal distribution: a z-score of 2.33 means the observed green count is in the top 1% of what you would expect from unwatermarked text.

The intuition for why the signal accumulates is essentially the same as why large samples give more precise estimates. Each token is a single Bernoulli trial: was it green? With only 20 tokens, random fluctuations dominate. With 200 tokens, the law of large numbers ensures that the observed green rate converges to its true mean, and the true mean under the watermarked distribution is higher than γ\gamma by an amount that depends on δ\delta.

Soft vs. Hard Watermarking

The KGW scheme has two variants that differ in how they handle high-entropy versus low-entropy positions.

Hard watermarking strictly forbids red tokens by setting their logits to −∞-\infty. This maximizes detection power because every generated token is guaranteed to be green. The z-score grows at the maximum rate, and detection is reliable even with fewer tokens. The problem is quality degradation. When the model would naturally continue with a word like "not" or "however" but these happen to be on the red list for the current context, the model is forced to produce something awkward. At low-entropy positions, where one or two tokens strongly dominate the distribution, the forced exclusion of red tokens can produce grammatically incorrect or semantically inconsistent text.

Soft watermarking adds δ\delta to green-list logits without zeroing out red tokens. Quality loss is much smaller because the model can still choose red tokens if they dominate the distribution strongly enough. Consider a position where the model assigns logit 8.0 to "not" (a red token) and logit 4.0 to its best green alternative. With δ=2.0\delta = 2.0, the best green token now has logit 6.0, but "not" still wins at 8.0. The sentence continues naturally. Detection power is lower than hard watermarking because not every generated token is green, but it degrades gracefully and remains reliable at moderate text lengths.

The tradeoff between quality and detection power is fundamental and cannot be fully eliminated. Higher δ\delta makes the watermark easier to detect but also more perceptible as unnatural text because the model is more strongly biased away from its most natural continuations. Kirchenbauer et al. showed that soft watermarking with δ≈2.0\delta \approx 2.0 and γ=0.5\gamma = 0.5 achieves a useful balance: text quality remains close to the unwatermarked baseline as measured by perplexity and human evaluation, while detection succeeds reliably with T≥200T \geq 200 tokens at meaningful significance levels.

The choice between hard and soft watermarking is not merely technical. Hard watermarking might be appropriate for internal auditing where text quality is not consumer-facing. Soft watermarking is necessary for any deployment where the watermarked model interacts directly with users who would notice quality degradation.

Unigram and Context-Dependent Watermarks

The context-dependent partitioning of the KGW scheme has both strengths and weaknesses. The strength is that the partition is unpredictable without the key. An attacker observing the generated text cannot infer which tokens are green for future positions because the partition changes with each new token. The weakness is that the hash of consecutive tokens creates correlations between adjacent decisions, complicating the statistical analysis and making it harder to reason about detection power theoretically.

A simpler alternative is the unigram watermark, which uses a single fixed partition of the vocabulary that does not change across positions. The same green list applies at every token position throughout a document. Detection is still a z-test on the green token count, but the partition is fixed rather than context-dependent. This simplification has two consequences: the statistical analysis is cleaner because token choices are more nearly independent, and the scheme is more vulnerable to attack because an attacker who discovers which words are on the green list can avoid them entirely. Since the partition does not change, discovering even part of the green list gives the attacker useful information for targeted removal.

A more sophisticated approach, proposed by Christ et al. (2023), generates a pseudo-random binary signal rt∈{0,1}r_t \in \{0, 1\} for each position tt using a pseudo-random function keyed to the secret key and the current position index. The bits determine which tokens to prefer without partitioning the full vocabulary. This scheme offers provably optimal detection under certain assumptions because it achieves the maximum possible information content per token while remaining undetectable to a party without the key. The theoretical guarantees come at the cost of more complex detection algorithms, but the framework has influenced subsequent work.

Distortion-Free Watermarking

A significant limitation of logit-bias schemes is that they always distort the output distribution. The generated text is not drawn from the true model distribution. This matters for several reasons. First, calibration: watermarked models are slightly less faithful to the input prompt because they systematically prefer certain tokens even when those tokens are not the most natural continuation. Second, detectability: the distortion is in principle measurable by anyone who knows enough about the model's natural distribution, even without knowing the watermark key. Third, downstream use: if a watermarked model's outputs are used as training data for another model, the biased distribution introduces systematic errors.

Kuditipudi et al. (2023) introduced a family of distortion-free watermarks that preserve the original sampling distribution in expectation. The key idea is to use a pseudo-random coupling between the model's probability distribution and a secret signal, rather than modifying the logits. Because the signal is embedded through the randomness of the sampling process rather than through biased logits, the marginal distribution over outputs is unchanged.

To understand the mechanism, recall how token sampling normally works. The model produces a probability distribution PP over the vocabulary. You sample a token by drawing a uniform random variable u∼Uniform(0,1)u \sim \text{Uniform}(0, 1) and applying the inverse CDF of PP: you find the token tt such that ∑t′<tP(t′)≤u<∑t′≤tP(t′)\sum_{t' < t} P(t') \leq u < \sum_{t' \leq t} P(t'). Different values of uu produce different tokens according to their probabilities.

The distortion-free watermark replaces the uniformly random uu with a value derived from the secret key: ut=F(k,t)u_t = F(k, t) where FF is a pseudo-random function. The output distribution is still PP because a pseudo-random utu_t is statistically indistinguishable from a truly uniform one if FF is a good pseudo-random function. But now the specific token chosen at each position is determined by the key in a way that a detector can verify.

Detection works by testing whether the observed tokens are consistent with having been chosen using the key-derived randomness. Concretely, for each token tt in the candidate text, the detector computes the rank of tt under the probability distribution PP given the preceding context. Under the watermark, each token should have a specific expected rank determined by utu_t. Under the null hypothesis, ranks are uniformly distributed. A rank-based test statistic exploits this asymmetry to identify watermarked text.

The tradeoff is detection efficiency. Because distortion-free schemes do not boost any signals above the natural distribution, each token provides less evidence about the watermark's presence than it would under a logit-bias scheme. Empirically, distortion-free schemes typically require more tokens to achieve the same detection confidence. The theoretical guarantee of zero distribution shift is useful, but practical performance at short text lengths is lower than logit-bias approaches. For long documents, the difference diminishes.

Statistical Detection

Detection is a hypothesis testing problem. The detector receives a piece of text and must decide: was this generated by the watermarked model or not? The quality of the detector is measured by its false positive rate and true positive rate.

The challenge is that the same test statistic must work for text about any topic, written in any style, at any level of formality. A watermark scheme that only works on formal prose would be easy to evade by switching styles. The robustness of the statistical test across diverse text types is an important practical concern.

The Hypothesis Test Framework

Define:

  • H0H_0: the text was not watermarked (null hypothesis)
  • H1H_1: the text was generated by the watermarked model (alternative)

The detector computes a test statistic SS from the text and compares it to a threshold τ\tau. If S>τS > \tau, the detector rejects H0H_0 and claims the text is watermarked.

The two error types are:

  • False positive (Type I error): Declaring human-written text as watermarked. This harms innocent people and destroys trust in the detector. If a university relies on a watermark detector to identify academic dishonesty, a false positive could result in a student being wrongly disciplined.
  • False negative (Type II error): Failing to detect watermarked text. This lets attackers evade detection. Depending on the application, false negatives may be more or less acceptable than false positives.

The FPR is controlled by setting τ\tau at the appropriate quantile of the null distribution. If SS under H0H_0 follows a standard normal distribution, then τ=2.33\tau = 2.33 gives FPR =0.01= 0.01, meaning 1 in 100 human-written texts will be falsely flagged. The null distribution does not depend on the text's content, only on the assumption that token choices are independent of the key-derived partition. This assumption holds when the text was written without any knowledge of the key.

The TPR depends on how far the watermarked distribution has shifted relative to the null. More signal (higher δ\delta, more tokens) means higher TPR at the same FPR. The formal tool for quantifying this tradeoff is the Receiver Operating Characteristic (ROC) curve, which plots TPR against FPR across all threshold values. A strong watermark pushes the curve toward the upper-left corner, indicating high TPR at low FPR. A weak watermark, insufficient token count, or heavy attack places the curve close to the diagonal, indicating near-chance performance.

A useful single-number summary is the area under the ROC curve (AUC). AUC =1.0= 1.0 means perfect separation between watermarked and unwatermarked text; AUC =0.5= 0.5 means the detector is no better than random. In practice, AUC values above 0.95 at text lengths above 100 tokens are achievable with soft watermarking at δ=2.0\delta = 2.0.

Minimum Text Length for Detection

Watermark detection is a function of the number of tokens TT. With very few tokens (say T=20T = 20), the green token count has high variance under both hypotheses, and the z-score is noisy. Detection becomes reliable only when enough evidence accumulates.

The statistical argument is straightforward. Under the watermarked distribution, the expected green count is μ1T\mu_1 T where μ1>γ\mu_1 > \gamma (because green logits are boosted). Under the null, it is γT\gamma T. The z-score is approximately:

z≈(μ1−γ)Tγ(1−γ)z \approx \frac{(\mu_1 - \gamma) \sqrt{T}}{\sqrt{\gamma(1-\gamma)}}

This grows as T\sqrt{T}. To achieve a threshold z-score of zα=2.33z_\alpha = 2.33 at significance level α=0.01\alpha = 0.01, you need:

T≥zα2⋅γ(1−γ)(μ1−γ)2T \geq \frac{z_\alpha^2 \cdot \gamma(1-\gamma)}{(\mu_1 - \gamma)^2}

A higher logit boost δ\delta increases μ1−γ\mu_1 - \gamma, reducing the required TT quadratically. Doubling δ\delta approximately quarters the required token count for reliable detection.

The minimum text length required for reliable detection also depends on:

  • γ\gamma: the green list fraction. Asymmetric partitions (γ≠0.5\gamma \neq 0.5) reduce variance under H1H_1 but may degrade quality more noticeably for tokens in the disadvantaged list.
  • Target FPR/TPR: More stringent requirements demand more tokens. Going from FPR =0.01= 0.01 to FPR =0.001= 0.001 increases the required z-score from 2.33 to 3.09, requiring roughly 75% more tokens.

As a rough heuristic, soft watermarking with δ=2.0\delta = 2.0 and γ=0.5\gamma = 0.5 achieves TPR ≈0.99\approx 0.99 at FPR =0.01= 0.01 with about T=200T = 200 tokens. For shorter texts, power drops significantly. This is one of the scheme's principal limitations in practice.

The chart below illustrates how the z-score distributions under H0H_0 and H1H_1 shift apart as the logit boost δ\delta increases. The overlap between the two distributions corresponds to the unavoidable tradeoff between false positive rate and false negative rate.

Out[4]:
Visualization
Two overlapping normal curves showing minimal separation between H0 and H1 at delta=0.5.
Z-score distributions under H0 (no watermark, orange) and H1 (watermarked, blue) with logit boost delta=0.5. The two distributions nearly overlap, meaning the detector struggles to separate watermarked from non-watermarked text at this weak boost level.
Two normal curves showing moderate separation between H0 and H1 at delta=1.0.
Z-score distributions with delta=1.0. The watermarked distribution (blue) shifts noticeably rightward, creating a visible gap above the detection threshold (dotted line at z=2.33) and substantially increasing true positive rate.
Two well-separated normal curves showing strong detection power at delta=2.0.
Z-score distributions with delta=2.0. At this boost level, the watermarked distribution has shifted far enough that most probability mass lies above the detection threshold, enabling reliable detection with high TPR and low FNR.

When δ=0.5\delta = 0.5, the two distributions overlap almost completely, and detection is barely better than chance at most thresholds. At δ=1.0\delta = 1.0, a clear gap opens, and the fraction of the H1H_1 distribution to the right of the threshold grows substantially. At δ=2.0\delta = 2.0, the watermarked distribution has shifted so far right that most of its mass lies above the detection threshold, yielding high TPR. The cost is that the logit boost at this level is strong enough to measurably distort the output distribution, a tradeoff that practitioners must account for.

Multiple Testing Considerations

In practice, a watermark detector might be applied to millions of texts. With FPR =0.01= 0.01 and one million texts, approximately 10,000 innocent texts will be falsely flagged. This is a multiple testing problem, and failing to account for it leads to a flood of false accusations.

Correcting for multiple testing with a Bonferroni correction tightens the per-test significance threshold. For 1 million tests and a family-wise error rate of 0.05, the per-test threshold becomes:

αcorrected=0.051,000,000=5×10−8\alpha_\text{corrected} = \frac{0.05}{1{,}000{,}000} = 5 \times 10^{-8}

This corresponds to a z-score of about 5.5 rather than 2.33. Achieving that z-score reliably requires significantly longer text than the 200-token rule of thumb at the standard threshold.

The Bonferroni correction is conservative. False discovery rate (FDR) methods like Benjamini-Hochberg are less stringent while controlling the expected proportion of false positives among all detected positives. For a watermarking deployment where the proportion of watermarked texts is substantial (say, 10% of submissions), FDR control may be more appropriate than family-wise error rate control. The choice depends on the application's tolerance for false positives versus the statistical power needed.

Practical deployments must balance detection sensitivity against the false positive burden imposed by the scale of deployment. A school that processes 1,000 submissions per term can afford a tighter threshold than a platform that evaluates 100 million posts per day.

Aggregating Evidence Across Documents

When the goal is to determine whether an AI system (rather than a specific document) is generating watermarked text, evidence can be pooled across multiple short texts. This is especially valuable when individual documents are too short for reliable individual detection: a social media post of 30 words falls well below the minimum length, but if you observe 1,000 such posts from the same account, you can test whether they collectively show the watermark signal.

A meta-analysis approach tests whether the distribution of z-scores across a corpus is consistent with the null or shows systematic enrichment. Fisher's method combines independent p-values:

F=−2∑i=1nln⁡(pi)F = -2 \sum_{i=1}^{n} \ln(p_i)

where:

  • pip_i: the p-value from the z-test on the ii-th document
  • nn: the number of documents in the corpus
  • FF: follows a chi-squared distribution with 2n2n degrees of freedom under H0H_0

Under H0H_0, each pip_i is uniformly distributed on [0,1][0, 1], so −ln⁡(pi)∼Exponential(1)-\ln(p_i) \sim \text{Exponential}(1) and FF follows a χ2(2n)\chi^2(2n) distribution. If the texts are watermarked, p-values will be systematically smaller than uniform, making FF larger than expected and yielding a significant combined test.

This allows detection of watermarked content even when individual documents are too short for reliable individual detection. The cost is that you can only attribute the collection to the watermarked model, not any specific document. This is often sufficient for the use case: you want to know whether an account is primarily using a specific AI system, not whether any single post was AI-generated.

Stouffer's z-score method offers an alternative that is more interpretable in terms of the watermark z-score framework:

Zcombined=∑i=1nzinZ_\text{combined} = \frac{\sum_{i=1}^{n} z_i}{\sqrt{n}}

where ziz_i is the watermark z-score for document ii. Under H0H_0, each ziz_i has mean 0 and variance 1, so Zcombined∼N(0,1)Z_\text{combined} \sim N(0, 1) under H0H_0. Under H1H_1, the mean z-score is positive, and the combined statistic grows as n\sqrt{n}. This has the appealing interpretation that you are accumulating token evidence across documents as if they were a single long text.

Watermark Robustness

A watermark that dissolves under mild editing is not useful. Any practical deployment must assume that adversaries will attempt to remove or spoof the watermark. Robustness research characterizes what attacks are possible and how schemes can be designed to resist them.

The tension here is fundamental. The watermark signal lives in token choices, and any operation that changes tokens has the potential to reduce the signal. But many legitimate text transformations change tokens: editing for clarity, correcting errors, translating to another language, summarizing, quoting, or continuing text all involve token changes. A robust watermark must survive these benign transformations while remaining detectable, which is an inherently difficult requirement.

Attack Taxonomy

Attacks on watermarks fall into three categories, each with different goals and requirements for the attacker.

Removal attacks attempt to eliminate the watermark signal while preserving semantic content. The attacker has the watermarked text and wants to produce a clean version that passes no watermark test. Examples include:

  • Paraphrase attacks: Using a second language model (or human) to rewrite the text in different words. If the watermark is tied to specific token choices, paraphrasing that changes those tokens may remove the signal. The effectiveness depends on how aggressively the paraphraser rewrites and whether synonym substitutions happen to preserve green token rates by chance.
  • Word substitution: Replacing individual words with synonyms using a thesaurus or masked language model. Each substitution may flip a token from green to red or vice versa. Targeted substitution, where only red-list words are replaced with green-list alternatives, would theoretically increase the z-score (an insertion attack) rather than decrease it, so naive word substitution reduces the signal randomly rather than systematically.
  • Translation attacks: Translating to a different language and back. The round-trip typically changes many tokens while preserving meaning, potentially removing enough of the green-token signal to defeat detection. This is one of the more effective practical attacks because translation systems are widely available and produce fluent text.
  • Cropping and insertion: Removing tokens from the end of the text (reducing TT and thus statistical power) or inserting red-list tokens to dilute the green-token fraction. Inserting 50 randomly chosen tokens into a 100-token watermarked text can reduce the z-score significantly, depending on how many of the inserted tokens happen to be green.

Forgery attacks attempt to produce watermarked text without knowing the key. If an attacker can generate text that passes the watermark test, they can falsely attribute text to the watermarking entity, creating a false provenance claim. Forgery is generally much harder than removal because it requires knowing the green list partition, which is derived from the secret key. Without the key, an attacker cannot systematically choose tokens that will score highly under the detector. They could try to brute-force the key space, but cryptographic hash functions make this computationally infeasible.

Spoofing attacks target the detector's behavior rather than any specific watermark. An attacker who knows the detection algorithm (but not the key) might generate a large volume of texts and submit them for detection, attempting to overwhelm the system or to learn information about the key from which texts pass and which fail. Against an oracle detector, an attacker can conduct a binary search over the vocabulary: submit a text, observe whether it passes, modify one token, repeat. Over many queries, this allows the attacker to infer which tokens are green in specific contexts.

Robustness of Token-Level Schemes

The KGW scheme has been analyzed extensively for robustness. Paraphrase attacks are the most effective known removal attack. Because paraphrasing changes specific words, each substituted token may be on the red list in the new context, reducing the green-token count. With enough substitutions, the z-score drops below the detection threshold.

The key variable is how many tokens must be changed to remove the watermark. If a text has 200 tokens and needs to maintain a z-score above 2.33, it needs approximately γT+2.33Tγ(1−γ)≈123\gamma T + 2.33 \sqrt{T \gamma (1-\gamma)} \approx 123 green tokens out of 200 (with γ=0.5\gamma = 0.5). An attacker who substitutes 30% of tokens (60 tokens) reduces the expected green count by roughly 30 tokens, potentially pushing the z-score below the threshold.

The quality-robustness tradeoff is important here. Each token substitution changes the text slightly. If the substitutions preserve meaning well (synonyms, reorderings), the quality cost is low. If they degrade meaning significantly, the attacker has succeeded only by making the text worse. The question is whether a sufficiently capable paraphraser can substitute enough tokens to defeat the watermark at acceptable quality cost. Kirchenbauer et al. showed empirically that naive paraphrase attacks using T5 or GPT-3.5 as the paraphraser could reduce TPR substantially at typical detection thresholds, suggesting that soft watermarks are not fully robust against capable paraphrasers. Newer, more capable paraphrase models increase this vulnerability.

Improving Robustness: Context Windows and Robust Statistics

Several techniques improve robustness against token-level modifications, each with its own tradeoffs.

Context window extension: Instead of hashing only the previous token to determine the green list, hash the previous kk tokens. This creates longer-range dependencies in the watermark signal. An attacker who substitutes one token must also verify that the change does not cascade: since each green list depends on the preceding kk tokens, changing token at position tt alters the green lists at positions t+1t+1 through t+kt+k, potentially undoing other carefully chosen substitutions. The cost is computational: detection requires re-computing the green list for each position based on a window of preceding tokens, and detecting the watermark after arbitrary edits requires alignment between the original and edited sequences.

Robust statistics: Replacing the z-test on raw green counts with a more robust test statistic. For example, a test based on the fraction of runs of consecutive green tokens is more resistant to isolated red-token insertions than a simple count. If an attacker inserts single red tokens between otherwise green sequences, the run-based test can partially discount these interruptions. The tradeoff is reduced detection power on unattacked texts, because the robust statistic is less sensitive than the optimal z-test under the clean watermarked distribution.

Edit distance-aware detection: Detecting the watermark not in the exact sequence of tokens but in the closest matching alignment between the candidate text and the expected green-list pattern. This requires solving an alignment problem similar to dynamic programming sequence alignment, which is computationally more expensive than a simple token scan but enables detection after arbitrary insertions and deletions. The statistic is based on the alignment score rather than a simple count, making it robust to the cropping and insertion attacks that defeat naive detection.

Synonym-resilient watermarking: Embedding the watermark at the semantic level by using multiple green-list tokens that are near-synonyms. If "happy," "glad," and "pleased" are all on the green list for a given context, then synonym substitution between these words does not remove the signal. Designing such coherent green lists requires additional structure beyond random partitioning, but the robustness gains can be significant against synonym-based attacks.

Semantic Invariance

One direction for achieving robustness is to embed the watermark in semantic content rather than specific token choices. A semantically invariant watermark produces a signal that survives paraphrasing because it is tied to concepts or propositions, not surface forms.

Implementing this is challenging because it requires the model to have fine-grained semantic control over its output at a level that current decoding-time watermarks do not. Some proposals encode bits into the choice of which facts to include or emphasize rather than which words to use. For example, a model might be instructed to include certain supporting details in one variant of an explanation and different details in another, with the pattern of included details encoding the watermark. These approaches are more resistant to paraphrase because paraphrasing changes words but typically preserves the factual content of a passage.

The practical challenges are substantial. Semantic watermarks require a sophisticated encoding and decoding scheme that operates at the discourse level rather than the token level. The signal capacity per sentence is much lower because there are fewer discrete choices about what facts to include than about which tokens to use. Detection requires understanding the semantic content of the text, which is a harder problem than counting green tokens. These approaches remain largely in the research prototype stage.

Spoofing and Key Security

The security of watermarking schemes relies on the secret key remaining private. If an attacker learns the key, they can generate text that passes the watermark test (forgery) or identify the green list and avoid those tokens entirely (targeted removal). The security model is analogous to symmetric cryptography: the scheme is secure as long as the key is not disclosed.

Multi-key schemes address this by using different keys for different users or time periods. An enterprise deploying a watermarked LLM might issue different keys to different departments, so that if one department's key is leaked, only texts generated with that key are compromised. Key rotation limits the exposure window: even if an attacker infers the key from a corpus of watermarked texts generated during a specific period, future texts use a different key.

The more fundamental vulnerability is oracle access: an attacker who can submit texts to the detector and observe whether they pass can, in principle, reconstruct which tokens are green by binary search. Against this oracle attack, context-dependent watermarks are more resistant than fixed-partition unigram watermarks, because the attacker's query tells them the partition only at specific context strings. If the attacker changes the preceding token, the partition changes completely. A context-dependent scheme requires the attacker to conduct a separate oracle attack for each context, making full key reconstruction computationally expensive even with unlimited query access.

Rate limiting and query monitoring are operational mitigations. If the detector logs queries and flags accounts that submit suspiciously many queries, the oracle attack becomes detectable. The attacker must balance information gain per query against the risk of detection, which limits how much of the key space they can probe.

Watermark Evaluation

Evaluating a watermarking scheme requires balancing three properties that are often in tension: detection reliability, text quality, and robustness to attack. A scheme that achieves near-perfect detection at the cost of garbled text is useless in practice. A scheme that preserves perfect quality but degrades after a single paraphrase is also insufficient. Good evaluation methodology quantifies all three dimensions and reports the Pareto frontier of achievable tradeoffs.

Detection Reliability Metrics

The core metrics for detection quality come from the hypothesis testing framework.

True positive rate (TPR) at a fixed FPR: The fraction of watermarked texts correctly identified. Reporting TPR at a standard FPR (such as 0.01 or 0.001) allows comparison across schemes, because different schemes optimize the TPR/FPR tradeoff differently. A scheme with high TPR at FPR 0.01 might have lower TPR than another scheme at FPR 0.001, depending on the shape of the ROC curve.

Area under the ROC curve (AUC): A threshold-free summary of detection performance. AUC =1.0= 1.0 means perfect separation; AUC =0.5= 0.5 means chance performance. AUC is useful for comparing schemes across all operating points, but it can obscure important differences in the tails. Two schemes with the same AUC might have very different performance at very low FPR values, which is the region that matters most for high-stakes deployments.

Average token count for reliable detection: How many tokens are needed to achieve, say, TPR =0.95= 0.95 at FPR =0.01= 0.01. Lower is better, because it extends the scheme's utility to shorter texts. This metric is especially important for comparison between soft and hard watermarking and between logit-bias and distortion-free schemes.

These metrics should be reported separately for different text types, including long-form prose, dialogue, code, non-English text, and mixed-language content, because detection power varies with domain. Code, for example, has a highly constrained vocabulary in many contexts (keywords, common identifier names), which affects how often the natural token choice happens to be on the green list.

Text Quality Metrics

The watermark must not noticeably degrade the quality of generated text. Quality is measured at multiple levels, and different metrics capture different aspects of degradation.

Automatic metrics:

Perplexity measures the watermarked text's unexpectedness under a reference language model. A text with higher perplexity contains more unusual token choices, which may signal watermark-induced distortion. Perplexity is sensitive to the strength of the logit boost: at δ=2.0\delta = 2.0, perplexity increases measurably compared to the unwatermarked baseline, while at δ=0.5\delta = 0.5, the increase may be negligible. The reference model for perplexity evaluation should ideally be different from the watermarked model to avoid measuring the original model's own preferences.

BLEU and ROUGE scores measure n-gram overlap between watermarked and unwatermarked outputs on the same prompts. If the watermarked model is forced to use unusual tokens, overlap with the natural output will be lower. These metrics are blunt because high-quality paraphrases can have low n-gram overlap with the original while conveying the same meaning. They are useful as a quick sanity check but should not be the primary quality measure.

Semantic similarity using sentence embeddings captures meaning preservation better than n-gram overlap. By embedding both the watermarked and unwatermarked outputs in a high-dimensional semantic space and measuring cosine similarity, you can assess whether the watermark forced the model to change the meaning of its outputs. A watermark that changes words but not meaning should show high semantic similarity despite low BLEU.

Human evaluation:

Automatic metrics do not always correlate with human perception. Human judges assess fluency (does the text read naturally?), coherence (does it stay on topic and make logical sense?), and relevance (does it faithfully address the prompt?). A watermark that increases perplexity by 10% may still produce text that humans find perfectly natural. Conversely, a small perplexity increase concentrated on a few unnatural word choices may be immediately obvious to human readers.

Human evaluation is the gold standard but expensive to run at scale. Practical evaluation combines automatic screening to identify conditions where watermarks may cause visible degradation, followed by targeted human evaluation of those cases.

Robustness Metrics

Robustness is measured by the TPR of the detector after various attack budgets are applied.

Paraphrase robustness: TPR after applying a paraphrase model, where the budget is constrained so that the paraphrased text maintains semantic similarity above some threshold (typically 0.8 cosine similarity with the original). This measures how much watermark signal survives fluent paraphrasing. A scheme is considered robust to paraphrase if TPR remains above 0.8 at FPR 0.01 after paraphrasing with a constraint of semantic similarity 0.8.

Word substitution robustness: TPR after randomly substituting a fraction pp of tokens. This is a controlled experiment that characterizes how many token substitutions the scheme can tolerate. Results should be reported at multiple substitution rates (10%, 20%, 30%, 50%) to reveal the robustness profile rather than just the failure point.

Cropping robustness: TPR when only a prefix or suffix of the text is evaluated. This matters because attackers may crop text to reduce TT and thus statistical power, or because legitimate users may quote only part of a watermarked document.

The tradeoff between robustness and detection power is a Pareto frontier. A scheme with high detection power at T=100T = 100 tokens may have low robustness to paraphrase; a scheme with high paraphrase robustness may require T=500T = 500 tokens for reliable detection. Evaluation must characterize this frontier rather than reporting a single operating point.

Evaluating Distortion-Free Watermarks

Distortion-free watermarks require a different evaluation approach because they do not change the output distribution. Perplexity-based quality metrics will show no degradation by construction. The relevant quality metric is instead semantic fidelity: do the watermarked outputs carry the same meaning and satisfy the same preferences as the unwatermarked outputs? Since the marginal distribution is unchanged but specific samples differ, this is a subtle question. Two texts drawn from the same distribution may differ considerably in quality due to sampling variance.

Detection metrics are the same (TPR/FPR) but typically show lower efficiency (higher required TT) compared to logit-bias schemes, showing the weaker signal per token that comes from not modifying the output distribution. Robustness comparisons between distortion-free and logit-bias schemes must account for this difference in baseline detection efficiency: it is not fair to compare robustness at the same text length if the two schemes have very different detection power at that length.

Implementation

Let's implement the KGW watermarking scheme from scratch to see exactly how the generation and detection pipelines work. We will use a simplified vocabulary and a toy language model to illustrate the mechanics, then apply it to a realistic scenario.

Setup and Vocabulary

We begin by setting up the vocabulary and hash function for partitioning:

In[5]:
Code
import numpy as np

# Seed for reproducibility
rng = np.random.default_rng(42)

# Simplified vocabulary (actual LLMs have 50k+ tokens)
vocab = [
    "the",
    "cat",
    "sat",
    "on",
    "mat",
    "a",
    "dog",
    "ran",
    "fast",
    "slow",
    "jumped",
    "over",
    "fence",
    "small",
    "large",
    "red",
    "blue",
    "green",
    "happy",
    "sad",
    "sun",
    "moon",
    "star",
    "tree",
    "house",
    "water",
    "fire",
    "earth",
    "wind",
    "rain",
    "snow",
    "cold",
    "warm",
    "hot",
    "good",
    "bad",
    "big",
    "little",
    "old",
    "new",
    "black",
    "white",
    "left",
    "right",
    "up",
    "down",
    "in",
    "out",
    "yes",
    "no",
]
vocab_size = len(vocab)
word_to_id = {w: i for i, w in enumerate(vocab)}
id_to_word = {i: w for w, i in word_to_id.items()}

# Secret watermark key
SECRET_KEY = "language-ai-handbook-watermark-key-2024"
GAMMA = 0.5  # fraction of vocab in green list
DELTA = 2.0  # logit boost for green tokens
Out[6]:
Console
Vocabulary size: 50
Green list fraction (gamma): 0.5
Logit boost (delta): 2.0
Expected green list size: 25 tokens

With a vocabulary of 50 words and a green fraction of 0.5, roughly 25 tokens will be on the green list at any given step. The partition changes at each step based on the previous token, so the same word can be green at one point in a sequence and red at another.

Green List Generation

The key function maps a context token to a green list using a cryptographic hash:

In[7]:
Code
def get_green_list(
    prev_token_id: int, secret_key: str, vocab_size: int, gamma: float
) -> set:
    """
    Derive the green list for the current position using a hash of
    the previous token and the secret key. Returns the set of token IDs
    that are on the green list.
    """
    # Create a deterministic seed from the context and key
    seed_str = f"{secret_key}:{prev_token_id}"
    hash_bytes = hashlib.sha256(seed_str.encode()).digest()
    seed = int.from_bytes(hash_bytes[:4], "big")

    # Use the seed to deterministically shuffle the vocabulary
    local_rng = np.random.default_rng(seed)
    shuffled = local_rng.permutation(vocab_size)

    # The first gamma fraction of the shuffled vocabulary is the green list
    green_size = int(gamma * vocab_size)
    return set(shuffled[:green_size].tolist())
Out[8]:
Console
Green list after 'cat' (first 10): ['a', 'bad', 'cat', 'cold', 'dog', 'earth', 'green', 'house', 'large', 'left']
Green list after 'dog' (first 10): ['a', 'black', 'cat', 'down', 'earth', 'fast', 'fire', 'good', 'green', 'little']
Overlap between the two lists: 15 tokens (60%)

Token 'mat' in green list after 'cat': False
Token 'mat' in green list after 'dog': False

The overlap between green lists for different context tokens is approximately 50%, consistent with random partitioning. The same token can be green after one context word and red after another, which is why the watermark cannot be removed by simply avoiding specific words. An attacker who observes that "mat" is green after "cat" has learned nothing about whether "mat" will be green after "dog" or "tree."

Watermarked Text Generation

The generation function modifies the logits before sampling. The modification is minimal: we add δ\delta to the logit of each green-list token and leave red-list logits unchanged. Sampling then proceeds normally from the modified distribution.

In[9]:
Code
def generate_watermarked(
    prompt_ids: list[int],
    base_logits_fn,
    vocab_size: int,
    secret_key: str,
    gamma: float,
    delta: float,
    n_tokens: int,
    seed: int = 0,
) -> list[int]:
    """
    Generate tokens using the KGW watermarking scheme.
    At each step, boost the logits of green-list tokens by delta.
    """
    gen_rng = np.random.default_rng(seed)
    tokens = list(prompt_ids)

    for _ in range(n_tokens):
        prev_token_id = tokens[-1]

        # Get base logits from the language model
        logits = base_logits_fn(tokens, gen_rng)

        # Get green list for the current context
        green_list = get_green_list(
            prev_token_id, secret_key, vocab_size, gamma
        )

        # Boost green-list logits by delta (soft watermarking)
        for token_id in green_list:
            logits[token_id] += delta

        # Convert logits to probabilities via softmax
        logits_shifted = logits - logits.max()
        probs = np.exp(logits_shifted) / np.exp(logits_shifted).sum()

        # Sample next token
        next_token = gen_rng.choice(vocab_size, p=probs)
        tokens.append(int(next_token))

    return tokens[len(prompt_ids) :]  # Return only generated tokens


def generate_unwatermarked(
    prompt_ids: list[int],
    base_logits_fn,
    vocab_size: int,
    n_tokens: int,
    seed: int = 0,
) -> list[int]:
    """Generate tokens without watermarking, for comparison."""
    gen_rng = np.random.default_rng(seed)
    tokens = list(prompt_ids)

    for _ in range(n_tokens):
        logits = base_logits_fn(tokens, gen_rng)
        logits_shifted = logits - logits.max()
        probs = np.exp(logits_shifted) / np.exp(logits_shifted).sum()
        next_token = gen_rng.choice(vocab_size, p=probs)
        tokens.append(int(next_token))

    return tokens[len(prompt_ids) :]


def dummy_lm_logits(
    token_ids: list[int], rng: np.random.Generator
) -> np.ndarray:
    """
    A simple uniform language model. In practice, this would be the actual
    transformer's output logits. Here we use uniform base logits plus a
    small random perturbation to simulate natural variation.
    """
    base = np.zeros(vocab_size)
    # Small random perturbation to avoid perfectly uniform outputs
    noise = rng.normal(0, 0.5, vocab_size)
    return base + noise
Out[10]:
Console
Watermarked   : new up good in right little mat white star warm in little over old over left hap...
Unwatermarked : right up bad in black old mat left fire rain down old green bad over new sun bad...

Tokens generated: 100

Both sequences look like random strings of words because our dummy language model is nearly uniform. In practice, a real language model would produce coherent sentences whose token choices are heavily constrained by grammar and semantics. The key point is that the watermarked sequence's token choices are systematically biased toward the green list at each position, an invisible pattern that accumulates into a detectable signal.

The Detection Procedure

The detector needs only the candidate text tokens, the secret key, and the parameters γ\gamma and threshold. It recomputes the green list at each position (using the same hash function and key used during generation), counts how many tokens fall on the green list, and runs the z-test.

In[11]:
Code
def detect_watermark(
    token_ids: list[int],
    secret_key: str,
    vocab_size: int,
    gamma: float,
    z_threshold: float = 2.33,  # Corresponds to FPR ~0.01
) -> dict:
    """
    Run the KGW watermark detector on a sequence of token IDs.
    Returns the z-score, p-value, and detection decision.
    """
    from scipy import stats

    T = len(token_ids)
    if T < 2:
        return {
            "z_score": 0.0,
            "p_value": 1.0,
            "is_watermarked": False,
            "n_tokens": T,
        }

    # Count green tokens
    green_count = 0
    for pos in range(1, T):
        prev_token_id = token_ids[pos - 1]
        token_id = token_ids[pos]
        green_list = get_green_list(
            prev_token_id, secret_key, vocab_size, gamma
        )
        if token_id in green_list:
            green_count += 1

    n_tested = T - 1  # First token has no previous context
    expected_green = gamma * n_tested
    std_green = np.sqrt(n_tested * gamma * (1 - gamma))

    z_score = (green_count - expected_green) / std_green
    p_value = 1 - stats.norm.cdf(z_score)
    is_watermarked = z_score > z_threshold

    return {
        "z_score": z_score,
        "green_count": green_count,
        "expected_green": expected_green,
        "p_value": p_value,
        "is_watermarked": is_watermarked,
        "n_tokens": n_tested,
    }
Out[12]:
Console
Detection results (z-threshold = 2.33, FPR ~0.01):

Watermarked text:
  Green tokens : 85 / 101 (expected 50.5)
  Z-score      : 6.866
  p-value      : 0.000000
  Decision     : WATERMARKED

Unwatermarked text:
  Green tokens : 50 / 101 (expected 50.5)
  Z-score      : -0.100
  p-value      : 0.539631
  Decision     : NOT WATERMARKED

The watermarked text shows a z-score substantially above the threshold, while the unwatermarked text's z-score is close to zero, consistent with the null hypothesis. The p-value for the watermarked text is very small, indicating strong statistical evidence for the presence of the watermark. Notice that the detection function has no access to the generation function or the original model: it operates only on the token sequence and the key. This is what makes the scheme practical for real-world deployment, where the detector may run on a different server from the generator.

Visualizing Detection Power vs. Text Length

The detection power increases with the number of tokens. Let's visualize how the z-score and detection decision evolve as more text is processed:

In[13]:
Code
def simulate_detection_power(
    n_tokens_range: list[int],
    n_trials: int,
    vocab_size: int,
    secret_key: str,
    gamma: float,
    delta: float,
    z_threshold: float,
) -> dict:
    """
    For each text length, generate n_trials watermarked and unwatermarked
    texts and compute TPR and FPR.
    """
    tpr_list = []
    fpr_list = []
    z_wm_means = []
    z_uw_means = []

    prompt = [0]  # Single starting token (token ID 0)

    for n_tokens in n_tokens_range:
        wm_detections = 0
        uw_detections = 0
        z_wm_vals = []
        z_uw_vals = []

        for trial in range(n_trials):
            wm_gen = generate_watermarked(
                prompt,
                dummy_lm_logits,
                vocab_size,
                secret_key,
                gamma,
                delta,
                n_tokens,
                seed=trial * 100,
            )
            uw_gen = generate_unwatermarked(
                prompt, dummy_lm_logits, vocab_size, n_tokens, seed=trial * 100
            )

            wm_result = detect_watermark(
                prompt + wm_gen, secret_key, vocab_size, gamma, z_threshold
            )
            uw_result = detect_watermark(
                prompt + uw_gen, secret_key, vocab_size, gamma, z_threshold
            )

            wm_detections += int(wm_result["is_watermarked"])
            uw_detections += int(uw_result["is_watermarked"])
            z_wm_vals.append(wm_result["z_score"])
            z_uw_vals.append(uw_result["z_score"])

        tpr_list.append(wm_detections / n_trials)
        fpr_list.append(uw_detections / n_trials)
        z_wm_means.append(np.mean(z_wm_vals))
        z_uw_means.append(np.mean(z_uw_vals))

    return {
        "n_tokens_range": n_tokens_range,
        "tpr": tpr_list,
        "fpr": fpr_list,
        "z_wm_means": z_wm_means,
        "z_uw_means": z_uw_means,
    }
Out[15]:
Console
Detection power vs. text length:
  Tokens |      TPR |      FPR |    E[z|WM] |   E[z|no WM]
----------------------------------------------------------
      10 |    0.610 |    0.010 |       2.33 |         0.07
      20 |    0.900 |    0.005 |       3.33 |         0.06
      30 |    1.000 |    0.010 |       4.15 |         0.10
      50 |    1.000 |    0.005 |       5.40 |         0.10
      75 |    1.000 |    0.000 |       6.63 |         0.10
     100 |    1.000 |    0.010 |       7.62 |         0.06
     150 |    1.000 |    0.010 |       9.34 |         0.03
     200 |    1.000 |    0.020 |      10.79 |         0.03

The table shows how TPR grows from near-chance at 10 tokens toward reliable detection at 100 or more tokens, while FPR stays close to the theoretical value of 0.01. This empirical confirmation that the test is well-calibrated under the null (FPR tracks 0.01 across all lengths) is important: it means the hypothesis testing framework's assumptions are satisfied in this implementation.

Visualization: Detection Power and Z-Score Distributions

Out[16]:
Visualization
Line chart showing watermarked TPR rising toward 1.0 with text length while unwatermarked FPR remains near the 0.01 target.
Watermark detection rates as a function of token count (200 trials per length). TPR (solid) rises steeply between 20 and 100 tokens, approaching reliable detection above 100 tokens, while FPR (dashed) stays near the nominal 0.01 threshold throughout.
Line chart showing the watermarked mean z-score rising above the 2.33 threshold with text length while the unwatermarked mean stays near zero.
Mean z-score for watermarked (solid) and unwatermarked (dashed) texts as a function of token count. The watermarked z-score grows steadily with sequence length as evidence accumulates, while the unwatermarked score stays near zero, confirming the test's calibration.

The TPR curve rises steeply between 20 and 100 tokens, reaching reliable detection in the 100 to 150 token range for this toy model. The FPR line stays near the 0.01 theoretical value throughout, confirming that the hypothesis test is well-calibrated. The z-score plot tells the same story from a different angle: the watermarked mean z-score grows proportionally to T\sqrt{T}, while the unwatermarked mean stays near zero as expected.

Simulating a Paraphrase Attack

To understand robustness, let's simulate an attacker who replaces a random fraction of tokens. In practice, a paraphrase model would select semantically appropriate replacements, but random substitution gives us a controlled experiment that quantifies the signal decay as a function of substitution rate.

In[17]:
Code
def paraphrase_attack(
    token_ids: list[int],
    substitution_rate: float,
    vocab_size: int,
    seed: int = 0,
) -> list[int]:
    """
    Simulate a paraphrase attack by randomly substituting a fraction of tokens.
    In a real attack, a paraphrase model would be used to substitute tokens
    with semantically similar alternatives. Here we use random substitution
    as a worst-case upper bound on attack effectiveness.
    """
    attack_rng = np.random.default_rng(seed)
    attacked = list(token_ids)
    for i in range(len(attacked)):
        if attack_rng.random() < substitution_rate:
            attacked[i] = int(attack_rng.integers(0, vocab_size))
    return attacked
Out[19]:
Console
Paraphrase attack robustness (100 tokens, 300 trials):
   Substitution Rate |      TPR
----------------------------------
                0.0% |    1.000
               10.0% |    1.000
               20.0% |    0.993
               30.0% |    0.940
               40.0% |    0.710
               50.0% |    0.393
Out[20]:
Visualization
Line chart showing TPR decreasing from near 1.0 to about 0.4 as substitution rate increases from 0% to 50%, with the nominal 0.01 FPR baseline shown below.
TPR decay under random token substitution attack for 100-token watermarked texts (300 trials per rate). As substitution rate increases, the green token signal is progressively diluted and TPR falls from near-perfect detection to about 0.4 at 50% substitution. The nominal FPR threshold (dotted line) shows that detection remains above chance but becomes substantially less reliable. Real paraphrase attacks would be somewhat less effective at low rates because synonym substitutions often preserve green-list membership.

Even with 30% of tokens randomly substituted, detection often remains reliable. At 50% substitution, TPR falls sharply to around 0.4: still well above the false positive baseline, but no longer reliable enough for confident attribution. A real paraphrase attack using a language model would be less effective at low substitution rates because synonym substitutions often preserve green-list membership by chance. Two synonyms for the same concept tend to appear in similar contexts, and the hash function assigns them the same green list for any given preceding token. However, at high substitution rates, even semantically aware paraphrasers become effective.

Comparing Watermarked and Unwatermarked Token Distributions

Out[22]:
Visualization
Overlapping density histograms with unwatermarked green-token rates centered near 0.5 and watermarked rates concentrated near 0.85, plus a vertical expected-rate reference at 0.5.
Distribution of green token rates across 500 generated texts (50 tokens each). Watermarked texts (blue) show a higher green token rate centered above 0.5 due to the logit boost, while unwatermarked texts (orange) center near the expected rate of 0.5. The separation between the two distributions is the detectable signal. A larger delta or more tokens per document would shift the blue distribution further right, increasing detection power.

The separation between the two distributions is the signal that the z-test exploits. The unwatermarked distribution is centered near γ=0.5\gamma = 0.5, exactly as the null hypothesis predicts. The watermarked distribution is shifted to the right because the logit boost increases the probability of choosing green tokens. A higher δ\delta would shift the watermarked distribution further right, increasing detection power at the cost of more noticeable distortion in text quality.

Key Parameters

The main parameters of the KGW watermarking scheme are:

  • γ\gamma (gamma): The fraction of the vocabulary assigned to the green list, typically 0.5. Smaller values make detection harder because the expected green count under both hypotheses is lower, reducing the signal-to-noise ratio. They may also cause more quality degradation at positions where the model's best token is always on the smaller green list.
  • δ\delta (delta): The additive logit boost for green tokens, typically 2.0. Higher values increase detection power because each token provides more evidence, but they distort the output distribution more noticeably. The relationship between δ\delta and perplexity increase is roughly log-linear for moderate values.
  • z_thresholdz\_\text{threshold}: The detection threshold corresponding to a chosen FPR. Use 2.33 for FPR 0.01, 3.09 for FPR 0.001, and 5.5 for Bonferroni-corrected FPR at one million tests. Higher thresholds reduce false positives but require more tokens for reliable detection.
  • Context window: The number of previous tokens used to hash the green list, default 1. Larger windows increase robustness against targeted substitution but complicate detection after text edits and increase computational cost.

Limitations and Impact

Watermarking is a promising primitive, but it comes with significant practical constraints that limit how and where it can be deployed. Understanding these constraints is essential for setting appropriate expectations and designing systems that use watermarking appropriately.

The Fundamental Robustness Limit

The most fundamental limitation is that no watermarking scheme can be both high-quality and provably robust against paraphrase. This is not a failure of current schemes but a theoretical impossibility result. If the watermarked text contains the same semantic content as the original, a sufficiently capable paraphraser can produce text that conveys the same meaning with different tokens. Since the watermark resides in the tokens, not the meaning, a perfect paraphraser would remove it with zero quality loss.

To see why this is fundamental rather than a gap to be closed by better engineering, consider the information theory perspective. A watermark encodes information (the presence of the watermarked generation) into specific token choices. But the meaning of the text is encoded in the joint distribution of tokens, not in any specific realization. There are exponentially many token sequences that express the same meaning. A perfect paraphraser explores this space and finds one that does not carry the watermark signal. As long as meaning is separable from specific token choices, a sufficiently capable semantic manipulator can defeat any token-level watermark.

This concern already applies in practice. Modern LLMs are increasingly capable paraphrasers. A student with access to a frontier model can paraphrase AI-generated text enough to defeat current watermarking schemes while preserving the argument structure and even much of the phrasing. The robustness of any scheme is ultimately bounded by the capability of available paraphrasers, and that capability is increasing rapidly.

Multi-Model Deployment Complexity

Watermarking requires modifying the sampling procedure at inference time. Every deployment of the model must run the watermarked decoding path, which means the watermarking logic must be integrated into every inference server, edge deployment, and API endpoint. In practice, organizations deploy many fine-tuned variants, distilled models, and quantized versions of their base models. Ensuring that all variants carry the watermark is operationally complex, requiring policy enforcement throughout the deployment pipeline.

Fine-tuning on non-watermarked data can also destroy the watermark signal in some schemes. If the fine-tuned model's logit distributions shift significantly from the base model's, the correlation between logit-boosted tokens and the key-derived green lists may break down, reducing TPR. Distillation is even more problematic: a student model trained on watermarked teacher outputs may produce similar text patterns without the statistical signal, because distillation optimizes for output distribution matching, not for reproducing the watermark's bias.

The Open-Weights Problem

Watermarking is most effective when the model weights are private and the sampling procedure is controlled by the operator. For open-weights models, anyone can download the weights and run the model with their own sampling code, trivially bypassing the watermark injection step. This means watermarking cannot solve the provenance problem for publicly released models, which is precisely the category of models most used for unmonitored generation. A student who downloads Llama and removes the watermarking code from the generation script faces no technical barrier.

This creates an uncomfortable asymmetry: watermarking is easiest to enforce for the large commercial API providers who have the most operational control, but these providers also have the strongest incentives to build trust through other means. Open-source models, which arguably pose the greater regulatory challenge, are the hardest to watermark reliably.

Text Length Dependency

Most watermarking schemes require a minimum number of tokens for reliable detection, typically 200 or more for soft schemes. Short texts (a sentence, a social media post, a code completion, a product description) cannot be reliably attributed. This is a significant limitation because many high-value uses of LLMs produce short outputs. Targeted misinformation, for instance, often takes the form of short, pithy statements that spread virally. A 30-word disinformation headline falls well below the minimum length for any current watermarking scheme. Multi-bit watermarks (which encode a full message rather than a binary flag) face even worse token requirements because they must transmit more bits of information through the same token-choice channel.

False Positives in High-Stakes Decisions

Even a very low false positive rate becomes a burden at scale. If watermark detection is used in academic integrity enforcement at a university with 10,000 students submitting one essay each per year, a FPR of 0.001 implies roughly 10 false accusations per year. False positives in this context carry serious consequences for students who did nothing wrong, including academic sanctions that can affect careers. A FPR that seems negligible in a system context (0.1% of submissions) translates to many real people harmed.

This argues for extremely conservative thresholds in high-stakes settings, which in turn requires longer texts for reliable detection. The practical consequence is that watermarking may not be usable for short-form submissions (a 300-word essay) where the stakes are high and the text is too short for reliable detection at a Bonferroni-corrected threshold. Human review of borderline cases is essential, but this reduces the automation benefit and introduces its own biases.

Domain and Language Dependence

Detection power varies significantly across domains and languages. Code generation presents a particular challenge: many code tokens are highly constrained by syntax (keywords, mandatory operators), leaving little flexibility for the logit boost to influence choices. A Python function body that must start with def, use standard indentation, and call specific APIs has few degrees of freedom for watermarking. Detection power in constrained domains is substantially lower than in free-form prose.

Non-English languages present a related challenge. Watermarking schemes are typically developed and evaluated on English text. The distribution of logit values, the frequency of near-tie token choices, and the correlation structure of consecutive tokens all vary by language. A scheme calibrated to achieve FPR 0.01 on English text may have a different effective FPR on Chinese, Arabic, or synthetic code, potentially leading to incorrect detection decisions. Careful evaluation across languages and domains is necessary before deployment in multilingual systems.

Impact on the Field

Despite these limitations, watermarking has opened several important research directions and has changed how the field thinks about AI attribution.

The formalization of the problem as hypothesis testing gave the field a rigorous language for comparing schemes. Before this framework, discussions of watermarking tended to be informal and qualitative. The z-test framework made it possible to state precise claims about FPR, TPR, and minimum text length, enabling apples-to-apples comparison between competing schemes. This formalization is itself a contribution, independent of any specific scheme.

The distortion-free watermarking work resolved an apparent impossibility result and showed that the quality-robustness tradeoff is not fundamental. Before Kuditipudi et al. (2023), it was widely assumed that any detectable watermark must distort the output distribution. The distortion-free scheme demonstrated that this assumption was wrong, opening a new design space. The insight that sampling randomness itself can carry information, independently of the sampled values, is a fundamental result that extends beyond watermarking.

The robustness analysis work helped clarify what watermarking can and cannot guarantee, preventing overconfident claims. Early presentations of watermarking sometimes implied that it was a complete solution to AI attribution. The robustness research showed that this was not so, and redirected attention toward the hardest problems: robustness against paraphrase, detection at short text lengths, and deployment in open-weights settings.

Watermarking also sparked related work on other attribution methods. Provenance via fine-tuning fingerprints embeds information in a model's behavior rather than its outputs, so that the model's origin is recoverable from its response to specific probes. Cryptographic commitments allow a model to prove it generated text without revealing the content, using zero-knowledge protocols. Membership inference approaches test whether specific texts were included in training data. These approaches address different threat models and complement watermarking in a complete attribution system. The watermarking literature's careful problem formulation has benefited all of these adjacent areas.

The field continues to evolve rapidly. Multi-bit watermarks that embed a message rather than just a flag are an active research area, with applications to content tracking and digital rights management. Watermarks that survive multiple rounds of AI-assisted editing are under investigation, with approaches ranging from more reliable hash functions to semantic-level encoding. The policy dimension is also developing: questions about when watermarking should be mandatory, who should hold the keys, what legal status watermark-based attributions carry, and how to handle international deployments where key custody crosses jurisdictions are actively debated in regulatory and academic forums.

Summary

Watermarking embeds a hidden statistical signal into LLM outputs at inference time by modifying the decoding process. The main families of schemes are token-level logit-bias methods (the KGW scheme and its variants) and distortion-free schemes that preserve the output distribution while still enabling detection through the controlled use of sampling randomness.

The core mechanism of the KGW scheme is vocabulary partitioning. At each generation step, a hash of the previous token and a secret key splits the vocabulary into a green list and a red list. Green-list logits are boosted by δ\delta, biasing token selection toward green tokens without forcing them. Because the partition changes with each context token, the signal is unpredictable to anyone without the key.

Detection is a hypothesis test. The KGW detector counts the fraction of green tokens in the candidate text and computes a z-score. The null hypothesis holds that each token is independently green with probability γ\gamma. A high z-score, above the chosen threshold corresponding to the desired FPR, provides statistical evidence that the text was generated with the watermark active. Detection power increases with text length, growing proportionally to T\sqrt{T}, and reliable detection typically requires 100-200 tokens at standard significance levels.

Robustness is the central challenge. Paraphrase attacks, word substitution, and translation can reduce TPR significantly by replacing green-list tokens with red-list alternatives. Context-dependent partitioning, extended hash windows, and robust test statistics improve resistance but do not eliminate it. The theoretical limit is that a sufficiently capable paraphraser can remove any token-level watermark while preserving meaning.

Evaluation covers three dimensions: detection reliability (TPR/FPR and AUC), text quality (perplexity, semantic similarity, human judgments), and robustness (TPR after attack). These dimensions trade off against each other, and the right operating point depends on the deployment context. Distortion-free schemes require different quality metrics because they do not change the output distribution, but they also provide weaker detection signals per token.

Watermarking works best when model weights are private, outputs are long, and the stakes of false attribution are moderate. It is less effective for open-weights models, short outputs, or settings where false positives carry severe consequences. The scheme cannot guarantee robustness against a capable paraphraser, and text length requirements make it inapplicable to the shortest forms of AI-generated content. Used within its appropriate scope, and combined with human review at the margins, watermarking provides a powerful accountability primitive that would otherwise be technically impossible.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about LLM watermarking.

LLM Watermarking Quiz

Question 1 of 80 of 8 completed
In the KGW watermarking scheme, what determines which tokens are assigned to the green list at each decoding step?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026llmwatermarking, author = {Michael Brenndoerfer}, title = {LLM Watermarking: Schemes, Detection, and Robustness}, year = {2026}, url = {https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). LLM Watermarking: Schemes, Detection, and Robustness. Retrieved from https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness
MLAAcademic
Michael Brenndoerfer. "LLM Watermarking: Schemes, Detection, and Robustness." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness>.
CHICAGOAcademic
Michael Brenndoerfer. "LLM Watermarking: Schemes, Detection, and Robustness." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness.
HARVARDAcademic
Michael Brenndoerfer (2026) 'LLM Watermarking: Schemes, Detection, and Robustness'. Available at: https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). LLM Watermarking: Schemes, Detection, and Robustness. https://mbrenndoerfer.com/writing/watermarking-llm-schemes-statistical-detection-robustness

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.