Part of Language AI Handbook
Explains how span corruption works in T5, including span selection strategies, geometric distributions, sentinel tokens.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Span Corruption
Masked language modeling (MLM) trains models to predict individual tokens hidden behind [MASK] placeholders. But what if we masked entire phrases or multi-word expressions instead? Span corruption takes this idea further: rather than masking isolated tokens, it corrupts contiguous spans of text and asks the model to reconstruct them. This simple shift fundamentally changes what the model learns during pretraining.
T5 (Text-to-Text Transfer Transformer) popularized span corruption as its primary pretraining objective. By treating the corrupted input as source text and the original spans as target text, T5 frames pretraining as a sequence-to-sequence problem. This approach produces models that excel at generation tasks while remaining competitive on understanding benchmarks.
Think of the difference between patching a handful of isolated holes in a wall versus repairing entire sections that have been stripped away. Patching individual holes is a targeted, local task. Restoring whole sections requires understanding the structure of the wall, the relationship between surrounding elements, and how to produce a coherent surface. Span corruption gives the model the harder, more structurally demanding task, which builds richer internal representations.
The key insight behind span corruption is that natural language is not a bag of independent tokens. Words cluster into phrases, phrases into clauses, and clauses into sentences. Noun phrases, verb phrases, and named entities are coherent linguistic units whose meaning is distributed across multiple tokens. By masking and reconstructing these multi-token units, the model is forced to internalize phrase-level structure and produce text that is locally coherent, not just statistically plausible token by token.
In practice, span corruption also produces computational benefits that make it attractive for large-scale pretraining. Because each corrupted span collapses into a single sentinel token, the input sequence shrinks. Attention computation, which scales quadratically with sequence length, becomes cheaper. The model sees nearly the same amount of text but processes it in a more compact, information-dense form. These savings compound at the scale of hundreds of billions of pretraining tokens.
In this chapter, you'll learn how span corruption works, why it outperforms token-level masking for certain applications, and how to implement it from scratch. We'll cover span selection strategies, the mathematics of span length distributions, sentinel token design, and the computational advantages that make this approach attractive at scale.
T5 was introduced by Raffel et al. at Google in 2019 in the paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer." The central thesis was that every NLP task, whether classification, translation, summarization, or question answering, could be cast as a sequence-to-sequence problem where both inputs and outputs are text strings. This unification allowed a single model architecture and training procedure to handle an enormous variety of tasks without task-specific heads or loss functions. Span corruption was the pretraining objective that made this text-to-text framing work at scale: by treating corrupted spans as the source and original spans as the target, pretraining itself became a sequence-to-sequence task. T5 was pretrained on the Colossal Clean Crawled Corpus (C4), a 750 GB filtered web text corpus, and the resulting models became foundational to a generation of encoder-decoder architectures.
From Token Masking to Span Corruption
Standard MLM randomly selects 15% of tokens and replaces them with [MASK]. The model predicts each masked token independently, using the surrounding context. This works well but has limitations.
To understand why, consider what exactly the model is learning. When we mask "sat" in "the cat sat on the mat," the model must predict a verb that fits syntactically and semantically in that position. This is a useful signal, and BERT showed that even individual token predictions, aggregated across billions of examples, produce powerful contextual representations. But each prediction is a local, one-shot decision. The model produces a single token and moves on.
Consider the phrase "the cat sat on the mat." If MLM masks "sat" and "mat" independently, the model makes two separate predictions. It never learns that these words might be related or that predicting multi-word expressions requires different reasoning than predicting single tokens. There is no need for the model to produce a coherent continuation across multiple positions at once.
Span corruption addresses this by masking contiguous sequences. Instead of the [MASK] sat on the [MASK], we might see the [X] on [Y] where [X] replaces "cat sat" and [Y] replaces "the mat." The model must now generate complete phrases, not isolated words. This is qualitatively different: the model must decide on the first token of a span, then condition each subsequent token on what it has already generated within that span, while staying anchored to the surrounding uncorrupted context.
Notice that this also changes the topology of the training signal. With MLM, every mask position is a separate, independent prediction. The loss is a sum of per-token cross-entropies, each depending only on the context around that individual mask. With span corruption, the target is a multi-token sequence that must be generated autoregressively. The loss for each token within a span depends on all previously generated tokens in that span, creating intra-span dependencies that MLM cannot capture at all.
A pretraining objective that replaces contiguous spans of tokens with single sentinel tokens. The model learns to reconstruct the original spans from the corrupted input, treating the task as sequence-to-sequence generation.
This shift has several consequences. First, the input sequence becomes shorter because multiple tokens collapse into single sentinels. Second, the target sequence contains multiple tokens per sentinel, requiring the model to generate coherently. Third, the model learns about phrase-level patterns and dependencies that token-level masking might miss. Fourth, because the decoder must produce multi-token outputs for each sentinel, the model is training on a task that closely resembles real-world text generation: given context, produce a fluent, coherent continuation.
Span corruption retains MLM's overall fraction of corrupted tokens (15%) and still trains the model to fill in missing information from context. What changes is the granularity and nature of the reconstruction challenge, from isolated token prediction to phrase generation. This shift improves generation capabilities without requiring a different architecture or training infrastructure.
Span Selection Strategies
Now that we understand what span corruption does conceptually, we need to answer a practical question: how do we decide which parts of a sequence to corrupt? This leads us to two interconnected design choices that significantly affect what the model learns.
Think of span corruption as a controlled demolition. We want to remove enough of the building (the text) to create a meaningful reconstruction challenge, but not so much that the remaining structure provides no clues about what was removed. This balance requires careful decisions about two quantities: how much total material to remove (the corruption rate) and how to divide that material into chunks (the span length distribution).
The two parameters interact. If we fix the corruption rate at 15% and increase the average span length, we need fewer, longer spans. The model encounters fewer reconstruction tasks per sequence but each one is harder. If we reduce the mean span length, we get more, shorter spans. The model encounters more reconstruction tasks but each one is easier and more local. The right balance depends on what downstream capabilities matter most.
Corruption Rate: How Much to Remove
The corruption rate specifies what fraction of the original tokens should end up inside corrupted spans. T5 adopts , matching BERT's masking rate, but the mechanics differ substantially from token-level masking.
In standard MLM, a 15% corruption rate means we independently flip a coin for each token, with 15% probability of masking. The result: exactly 15% of tokens become [MASK], scattered throughout the sequence. With span corruption, we instead select contiguous regions totaling approximately 15% of the sequence. The key difference is that these tokens are grouped, not isolated.
This grouping creates an interesting constraint. If we want to corrupt a fixed fraction of tokens, the number of spans we create depends on how long each span is. Suppose we have a sequence of tokens and want to corrupt fraction of them. If each span contains an average of tokens, simple arithmetic tells us how many spans we need:
where:
- : the total number of tokens in the input sequence
- : the corruption rate (fraction of tokens to corrupt, e.g., 0.15 for 15%)
- : the average span length (e.g., 3 tokens per span)
The reasoning is direct: we need to "cover" tokens with our spans. If each span covers tokens on average, we need spans to achieve the target coverage.
Let's make this concrete. For a 512-token sequence with and :
- Tokens to corrupt: tokens
- Spans needed: spans
- Input length after corruption: Each span gets replaced by one sentinel token, so the corrupted input contains tokens
This calculation reveals an important property: span corruption naturally compresses the input sequence. We remove 77 tokens but add only 26 sentinels, achieving a net reduction of 51 tokens (about 10% shorter). This compression has computational benefits we'll explore later.
The T5 paper explored corruption rates between 10% and 50%. The team found that performance changed little in the 10%-25% range, with 15% consistently performing well across tasks. Rates above 30% begin to remove too much context, making reconstruction harder but also noisier as a training signal. Very high rates (40%-50%) can hurt downstream performance because the model spends most of pretraining on a task so hard that it cannot receive a clear learning signal from the context.
Span Length Distribution: The Shape of Uncertainty
The corruption rate tells us how much to remove; the span length distribution tells us how to partition that removal into individual spans. This choice significantly affects what patterns the model learns.
Consider the extremes. If all spans had length 1, span corruption would collapse to token-level masking: we'd have 77 isolated masks, each predicting a single token independently. If all spans had length 77, we'd have one massive gap containing over half the sequence. This provides almost no useful training signal.
The sweet spot lies somewhere between. We want a mix of span lengths that exposes the model to diverse reconstruction challenges: some single-token predictions (maintaining fine-grained language modeling ability), some short phrases (learning local coherence), and occasional longer spans (forcing the model to reason about larger structures).
T5 uses a geometric distribution to achieve this mix. The geometric distribution is the discrete analog of the exponential distribution, modeling a natural process: imagine flipping a biased coin at each position within a span, continuing until we get "heads" (stop) instead of "tails" (continue). The span length equals how many positions we visited before stopping.
Mathematically, if the probability of stopping at each step is , the probability of a span having exactly tokens is:
where:
- : the probability of sampling a span of exactly tokens
- : the span length (a positive integer: 1, 2, 3, ...)
- : the stopping probability at each step, controlling the distribution's shape
- : the probability of "continuing" for steps before finally "stopping"
This formula captures a simple generative process: to get a span of length , we need consecutive "continue" decisions (each with probability ) followed by one "stop" decision (with probability ).
The geometric distribution has a convenient property: its mean equals . This gives us a direct way to control the average span length. If we want spans to average tokens, we simply set:
For T5's choice of , this gives . At each step within a span, there's a 1-in-3 chance of stopping. The resulting distribution strongly favors shorter spans while maintaining a "heavy tail" of longer ones:
- (one-third of spans are single tokens)
- (about one-fifth are two tokens)
- (around one-sixth are three tokens)
- (one-tenth are four tokens)
- (about one-fifth contain five or more tokens)
This distribution creates a varied training curriculum. The model frequently encounters single-token predictions, maintaining its vocabulary knowledge. It regularly sees two and three-token spans, learning common phrases and local syntax. And it occasionally faces longer spans requiring compositional reasoning.
Why is the geometric distribution a natural choice here, rather than some other distribution with a similar mean? The geometric distribution has the "memoryless" property: the probability of stopping at the next step does not depend on how far we have already traveled within a span. This means there is no "target length" that spans tend toward. Each step within a span is a fresh, independent decision, which mirrors the structure of natural language to some extent: clauses do not have a predetermined length, and there is no linguistic pressure to end a phrase at exactly step three versus step four. The memoryless property keeps the sampling process simple and avoids introducing artificial length preferences.
Let's visualize this distribution by sampling 10,000 span lengths and plotting the empirical frequencies. This will confirm our theoretical calculations and reveal the characteristic exponential decay.
import matplotlib.pyplot as plt # noqa: F401
import numpy as np
# Geometric distribution parameters
mean_span_length = 3
p = 1 / mean_span_length
# Sample 10000 span lengths
span_lengths = np.random.geometric(p, size=10000)
# Calculate distribution statistics
unique, counts = np.unique(span_lengths, return_counts=True)
probabilities = counts / len(span_lengths)
The empirical distribution matches our theoretical expectations. Notice the characteristic exponential decay: each bar is roughly two-thirds the height of the previous one (since ). The red dashed line marks the mean at 3, which falls slightly to the right of the mode (1). This reflects the distribution's right skew.
Why does this shape work well for pretraining? The geometric distribution provides a natural curriculum:
- High-frequency short spans maintain the model's ability to predict individual tokens accurately, preserving fine-grained vocabulary knowledge
- Medium-length spans teach local coherence and common phrases, the bread-and-butter of fluent text generation
- The heavy tail occasionally challenges the model with longer reconstructions, forcing it to reason about how syntax and entities contribute to discourse structure
This is more effective than a uniform distribution, which would waste equal capacity on trivially short spans and overwhelmingly long ones. The geometric decay concentrates learning where it's most useful while still providing exposure to diverse span lengths.
Another way to understand the distribution is through the cumulative perspective: what fraction of spans have length at most ? This cumulative distribution function (CDF) answers practical questions like "what percentage of spans are single tokens?" or "how often do we see spans longer than 5 tokens?"

The CDF reveals that roughly half of all spans have length 2 or less, and about 80% have length 4 or less. Only the remaining 20% challenge the model with longer reconstructions. This steep rise followed by a gradual tail perfectly balances frequent short-span practice with occasional longer challenges.
Sentinel Tokens
Sentinel tokens serve as placeholders for corrupted spans. Unlike BERT's single [MASK] token, span corruption uses multiple distinct sentinels, typically denoted <extra_id_0>, <extra_id_1>, and so on.
Why use different sentinels for each span? Consider reconstructing two spans from the corrupted input. If both used the same [MASK] token, the model couldn't distinguish which output corresponds to which span. Distinct sentinels create a clear mapping between input positions and output targets. The sentinel index directly encodes the span's position in the sequence, giving the decoder an unambiguous signal about which reconstruction task it is working on at each moment.
Special placeholder tokens that replace corrupted spans in the input. Each span receives a unique sentinel (e.g., <extra_id_0>, <extra_id_1>), enabling the model to map reconstructed spans back to their original positions.
The target sequence concatenates all corrupted spans, each preceded by its corresponding sentinel:
Original: "The quick brown fox jumps over the lazy dog"
Corrupted Input: "The <extra_id_0> fox <extra_id_1> lazy dog"
Target: "<extra_id_0> quick brown <extra_id_1> jumps over the"
This format enables autoregressive generation of the target. The model sees <extra_id_0> and generates "quick brown" before the next sentinel signals a new span.
The target sequence design reflects a deliberate choice about when to stop generating for each span. When the decoder encounters a new sentinel token (<extra_id_1>) while generating after <extra_id_0>, it knows the current span has ended and a new one is beginning. This boundary signaling is implicit: the model learns through training that sentinel tokens delimit span boundaries, not through any special stopping logic in the architecture itself. The loss computed during training naturally teaches the model to emit a sentinel when the span is complete, because that is what the target sequence shows.
T5's vocabulary reserves 100 sentinel tokens (<extra_id_0> through <extra_id_99>). This typically suffices since even long sequences rarely contain more than 50-60 spans with a 15% corruption rate and average span length of 3. The 100-sentinel budget provides headroom for unusual cases: very long sequences, higher corruption rates during ablations, or fine-tuning settings where practitioners experiment with different hyperparameters. Reserving these tokens in the vocabulary does carry a small cost: 100 vocabulary slots that could otherwise encode additional subword units are dedicated to sentinels. In practice, given vocabularies of 32,000 tokens, this overhead is negligible.
It is worth understanding what happens if a sequence somehow requires more spans than available sentinels. The implementation simply caps the number of corrupted spans at the sentinel count. In practice with T5's standard hyperparameters, a 512-token sequence with 15% corruption and mean span length 3 produces around 26 spans, far below the 100-sentinel limit. The limit only becomes active if practitioners push corruption rates very high or apply span corruption to unusually long sequences.
Implementing Span Corruption
Now that we understand the mathematics behind span selection, let's translate these concepts into working code. We'll build the algorithm incrementally, starting from the core span selection logic and gradually assembling the complete corruption pipeline. By the end, you'll have a clear picture of how each formula we derived connects to the implementation.
Selecting Span Boundaries
Our first task is to decide which token positions fall within corrupted spans. We need an algorithm that:
- Determines how many spans to create (using our formula: )
- Samples each span's length from a geometric distribution
- Places spans at random positions without overlap
We'll represent the result as a corruption mask: a boolean array where True indicates that token should be part of some corrupted span.
def select_corruption_mask(
seq_length, corruption_rate=0.15, mean_span_length=3
):
"""
Create a boolean mask indicating which positions to corrupt.
Spans are selected to achieve approximately the target corruption rate.
"""
# Calculate expected number of spans
num_tokens_to_corrupt = int(seq_length * corruption_rate)
num_spans = max(1, num_tokens_to_corrupt // mean_span_length)
# Sample span lengths from geometric distribution
p = 1 / mean_span_length
span_lengths = np.random.geometric(p, size=num_spans)
# Clip to ensure we don't exceed the corruption budget
total_corrupted = span_lengths.sum()
if total_corrupted > num_tokens_to_corrupt:
# Scale down span lengths proportionally
scale = num_tokens_to_corrupt / total_corrupted
span_lengths = np.maximum(1, (span_lengths * scale).astype(int))
# Randomly place spans without overlap
mask = np.zeros(seq_length, dtype=bool)
available_positions = list(range(seq_length))
for length in span_lengths:
if len(available_positions) < length:
break
# Choose a starting position
valid_starts = [i for i in range(len(available_positions) - length + 1)]
if not valid_starts:
break
start_idx = np.random.choice(valid_starts)
# Mark positions as corrupted
for offset in range(length):
pos = available_positions[start_idx]
mask[pos] = True
available_positions.remove(pos)
return mask# Test the span selection
test_length = 50
mask = select_corruption_mask(
test_length, corruption_rate=0.15, mean_span_length=3
)Sequence length: 50 Tokens corrupted: 3 (6.0%) Corruption mask (1 = corrupted): ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░██░░░░░░░░░░░█
The visualization shows corrupted positions as filled blocks (█) and uncorrupted positions as empty blocks (░). Notice how the corrupted tokens cluster into contiguous groups rather than scattering randomly across the sequence. This is the defining characteristic of span corruption: tokens within each span will be replaced by a single sentinel and reconstructed together.
With approximately 15% of tokens corrupted, we've created distinct "holes" in the sequence that the model must learn to fill. The spans vary in length according to our geometric distribution, some containing just one or two tokens, others stretching across several positions.
Identifying Span Boundaries
The corruption mask tells us which positions are corrupted, but for the next step we need to know where each span begins and ends. This boundary information lets us replace each contiguous span with exactly one sentinel token and extract the corresponding target tokens.
def get_span_boundaries(mask):
"""
Extract start and end indices for each contiguous span in the mask.
Returns list of (start, end) tuples where end is exclusive.
"""
spans = []
in_span = False
start = 0
for i, is_corrupted in enumerate(mask):
if is_corrupted and not in_span:
# Starting a new span
start = i
in_span = True
elif not is_corrupted and in_span:
# Ending current span
spans.append((start, i))
in_span = False
# Handle span at end of sequence
if in_span:
spans.append((start, len(mask)))
return spans# Find spans in our test mask
spans = get_span_boundaries(mask)Found 2 spans: Span 0: positions 36-37 (length 2) Span 1: positions 49-49 (length 1)
The algorithm finds each contiguous run of True values in the mask. Each span gets a unique index (0, 1, 2, ...) that will serve as its identifier when we assign sentinel tokens. Notice how span lengths vary: some spans contain just one token, others two or three, following our geometric distribution.
Building Input and Target Sequences
With span boundaries identified, we can now construct the two sequences that define the training example:
- Corrupted input: The original sequence with each span replaced by its sentinel token (e.g.,
<extra_id_0>,<extra_id_1>) - Target sequence: A concatenation of all corrupted spans, each preceded by its corresponding sentinel
This format creates a clear mapping: the model sees a sentinel in the input and learns to generate the corresponding tokens in the target.
def corrupt_sequence(tokens, mask, sentinel_prefix="<extra_id_"):
"""
Apply span corruption to a token sequence.
Args:
tokens: List of tokens
mask: Boolean array indicating corrupted positions
sentinel_prefix: Prefix for sentinel tokens
Returns:
corrupted_input: Input sequence with sentinels replacing spans
target: Target sequence for reconstruction
"""
spans = get_span_boundaries(mask)
# Build corrupted input
corrupted_input = []
last_end = 0
for span_idx, (start, end) in enumerate(spans):
# Add uncorrupted tokens before this span
corrupted_input.extend(tokens[last_end:start])
# Add sentinel for this span
corrupted_input.append(f"{sentinel_prefix}{span_idx}>")
last_end = end
# Add remaining uncorrupted tokens
corrupted_input.extend(tokens[last_end:])
# Build target sequence
target = []
for span_idx, (start, end) in enumerate(spans):
# Add sentinel
target.append(f"{sentinel_prefix}{span_idx}>")
# Add original span tokens
target.extend(tokens[start:end])
return corrupted_input, targetLet's see this in action with a real sentence:
# Example sentence
sentence = "The quick brown fox jumps over the lazy dog near the riverbank"
tokens = sentence.split()
# Create corruption mask for this specific example
mask = select_corruption_mask(
len(tokens), corruption_rate=0.25, mean_span_length=2
)
# Apply corruption
corrupted_input, target = corrupt_sequence(tokens, mask)Original tokens: ['The', 'quick', 'brown', 'fox', 'jumps', 'over', 'the', 'lazy', 'dog', 'near', 'the', 'riverbank'] Corruption mask: ░░░░░░░░░░██ Corrupted input (11 tokens): ['The', 'quick', 'brown', 'fox', 'jumps', 'over', 'the', 'lazy', 'dog', 'near', '<extra_id_0>'] Target sequence (3 tokens): ['<extra_id_0>', 'the', 'riverbank']
Examine the output carefully. The corrupted input is noticeably shorter than the original: multi-token spans have collapsed into single sentinel tokens. The target sequence contains exactly the tokens that were removed, each group prefixed by its sentinel to maintain the correspondence.
This structure shows how span corruption works in practice. The model must:
- Understand context: Use the uncorrupted tokens surrounding each sentinel to infer what's missing
- Generate coherently: Produce complete phrases, not just isolated words, for each sentinel
- Maintain boundaries: Know when one span ends and the next begins, guided by the sentinel markers
Complete Span Corruption Pipeline
With all the pieces in place, let's combine them into a reusable class that encapsulates the full span corruption logic:
class SpanCorruptor:
"""Applies T5-style span corruption to text sequences."""
def __init__(
self, corruption_rate=0.15, mean_span_length=3, num_sentinels=100
):
self.corruption_rate = corruption_rate
self.mean_span_length = mean_span_length
self.num_sentinels = num_sentinels
self.sentinels = [f"<extra_id_{i}>" for i in range(num_sentinels)]
def corrupt(self, tokens):
"""
Corrupt a token sequence using span corruption.
Args:
tokens: List of string tokens
Returns:
dict with 'input' and 'target' token lists
"""
if len(tokens) < 2:
return {"input": tokens, "target": []}
# Select spans to corrupt
mask = select_corruption_mask(
len(tokens), self.corruption_rate, self.mean_span_length
)
# Get span boundaries
spans = get_span_boundaries(mask)
if len(spans) > self.num_sentinels:
# Too many spans, merge some
spans = spans[: self.num_sentinels]
# Build sequences
corrupted_input = []
target = []
last_end = 0
for span_idx, (start, end) in enumerate(spans):
corrupted_input.extend(tokens[last_end:start])
corrupted_input.append(self.sentinels[span_idx])
target.append(self.sentinels[span_idx])
target.extend(tokens[start:end])
last_end = end
corrupted_input.extend(tokens[last_end:])
return {"input": corrupted_input, "target": target}# Test the complete pipeline
corruptor = SpanCorruptor(corruption_rate=0.15, mean_span_length=3)
text = """Natural language processing enables computers to understand
and generate human language in meaningful ways"""
tokens = text.split()
result = corruptor.corrupt(tokens)Original (14 tokens): Natural language processing enables computers to understand and generate human language in meaningful ways Corrupted input (13 tokens): Natural language processing enables computers to understand and <extra_id_0> language in meaningful ways Target (3 tokens): <extra_id_0> generate human
The results confirm our earlier calculations. The original sequence becomes a shorter corrupted input (the corruption plus sentinels compress the sequence) and a compact target containing only the corrupted material plus sentinel markers.
This completes our span corruption implementation. Starting from the mathematical foundations, we've built each component:
- Span count estimation: spans needed for the target corruption rate
- Length sampling: Geometric distribution with for natural length variation
- Boundary detection: Linear scan to identify contiguous corrupted regions
- Sequence construction: Parallel building of input (with sentinels) and target (spans preceded by sentinels)
The algorithm is efficient, running in linear time relative to sequence length, and produces the exact format expected by T5-style encoder-decoder training.
Worked Example: Tracing a Complete Training Sample
Before moving to the T5 training setup, it helps to trace a single training example from raw text all the way to loss computation. This concrete walk-through ties together every component we have built.
Suppose we have the sentence: "Researchers at Google trained the model on a large text corpus."
Step 1: Tokenization. We split this into 14 word tokens for simplicity (in practice, T5 uses SentencePiece subword tokenization, but the logic is identical):
["Researchers", "at", "Google", "trained", "the", "model", "on", "a", "large", "text", "corpus", "."]
That gives us tokens.
Step 2: Span selection. With and , we expect to corrupt tokens spread across span. Let's say the sampler selects tokens 2 and 3 (indices 2-3, a span of length 2 covering "Google trained"):
Corruption mask: ░░██░░░░░░░░
Step 3: Building the corrupted input. Token positions 0-1 pass through unchanged. Positions 2-3 collapse to <extra_id_0>. Positions 4-11 pass through unchanged:
Corrupted input: ["Researchers", "at", "<extra_id_0>", "the", "model", "on", "a", "large", "text", "corpus", "."]
The input has 11 tokens, down from 12.
Step 4: Building the target sequence. The target begins with the sentinel, followed by the corrupted span's tokens:
Target: ["<extra_id_0>", "Google", "trained"]
Step 5: Encoder processing. The encoder receives all 11 input tokens and produces a hidden state matrix of shape . The hidden state at the <extra_id_0> position encodes everything the model can infer about the missing span from the surrounding context: "Researchers at ??? the model on a large text corpus" strongly implies that the missing words are a proper noun (an organization) followed by a verb. The encoder's bidirectional attention can draw on all 11 positions simultaneously.
Step 6: Decoder generation. The decoder starts with <extra_id_0> as its first input token, attending to all encoder states. It produces a probability distribution over the vocabulary and assigns high probability to "Google" (a proper noun that fits the context). "Google" is fed as input at the next step, and the decoder predicts "trained." Then the decoder sees "trained" and must decide what comes next: since the target ends here, the correct output is the end-of-sequence token (or in T5's format, simply no more tokens for this sentinel).
Step 7: Loss computation. The cross-entropy loss sums over the three target positions:
In practice the first term is trivially minimized since the target literally begins with the sentinel. The meaningful learning happens in the second and third terms, where the model must predict the actual missing words from context.
The main point from this trace is how naturally the task matches real generation. The model learns: "Given a partially visible sentence with a marked gap, produce the words that fill the gap." At inference time on real tasks like summarization or translation, the model does essentially the same thing: given an input (source text or passage), produce an output (summary or translation). Span corruption during pretraining directly practices this input-to-output mapping, which is why T5 transfers so effectively to later generation tasks.
T5-Style Training
T5 uses span corruption within an encoder-decoder framework. The corrupted input feeds into the encoder, while the decoder autoregressively generates the target sequence. This setup naturally handles variable-length span reconstruction.
Encoder-Decoder Architecture
The encoder processes the corrupted input sequence, producing hidden representations that capture the context around each sentinel:
where:
- : the encoder's output hidden states, a matrix of shape (sequence length hidden dimension)
- : the corrupted input tokens, including uncorrupted tokens and sentinel placeholders
- : a sentinel token marking where a span was removed
The encoder uses full bidirectional self-attention: every token can attend to every other token. This means the representation at the <extra_id_0> position has access to all the surrounding context, both the tokens that appear before the sentinel and those that appear after. This bidirectional view is a significant advantage over causal decoder-only models, which can only see leftward context. When predicting what phrase fills a gap, the context to the right of the gap is often just as informative as the context to the left, and the encoder captures both.
The decoder then generates the target sequence autoregressively, conditioned on the encoder's hidden states:
where:
- : the token being predicted at position in the target sequence
- : all previously generated target tokens
- : the probability distribution over the vocabulary for the next token
During training, we use teacher forcing: the decoder receives the true previous tokens rather than its own predictions. Teacher forcing is important for span corruption because the target sequence can be quite long (dozens of tokens across multiple spans), and sampling from the model's own distributions during training would introduce compounding errors that would destabilize learning. By feeding the ground-truth tokens at each decoder step, the model receives clean, unambiguous gradient signal throughout the target sequence.
The training objective is standard cross-entropy loss over the target sequence:
where:
- : the total loss for this training example
- : the length of the target sequence
- : the log probability assigned to the correct token at each position
This loss encourages the model to assign high probability to the actual tokens that were corrupted, learning to reconstruct spans from context. Because the target includes both sentinel tokens and the actual span content, the model learns two things simultaneously: when to begin generating a new span (at sentinel boundaries) and what content to generate within each span.
Why Encoder-Decoder?
Span corruption works well with encoder-decoder models for several interconnected reasons.
Bidirectional encoding is perhaps the most important. The encoder sees all uncorrupted context simultaneously, enabling rich representations of what surrounds each sentinel. Unlike a causal language model that can only look left, the encoder integrates both left and right context at every position, producing representations that carry far more information about the missing span. For named entity recognition, for instance, the context to the right of a person's name can be just as disambiguating as the context to the left.
Autoregressive decoding complements this by enabling the model to generate spans of variable length coherently. The decoder generates spans token by token, learning proper phrase structure and coherence within each span. Because each generated token conditions on all previous tokens in the target, the decoder naturally models intra-span dependencies: the second word of a phrase is influenced by the first, the third by both, and so on.
Natural length handling emerges from this architecture without any special design. Spans of different lengths produce targets of different lengths, which encoder-decoder models handle gracefully through their sequence-to-sequence framing. A decoder-only model using span corruption as a fill-in-the-blank task would need to know in advance how many tokens to generate for each gap, which is not available during inference.
Decoder-only models can also use span corruption, treating the task as infilling. The corrupted input and target concatenate with appropriate separators, and the model learns to continue after each sentinel. Some newer models like FIM (Fill in the Middle) explore exactly this setting. The trade-off is that decoder-only models must still predict a fixed-length span or use special boundary tokens, whereas the encoder-decoder design handles this naturally through its two-part architecture.
Comparison to MLM
Let's compare the training signals from span corruption versus token-level MLM:
def mlm_mask(tokens, mask_rate=0.15, mask_token="[MASK]"):
"""Apply standard MLM masking (token-level, not spans)."""
mask = np.random.random(len(tokens)) < mask_rate
masked_tokens = [mask_token if m else t for t, m in zip(tokens, mask)]
targets = [(i, t) for i, (t, m) in enumerate(zip(tokens, mask)) if m]
return masked_tokens, targets
def compare_corruption_methods(text, seed=42):
"""Compare span corruption vs MLM on the same text."""
tokens = text.split()
# MLM
np.random.seed(seed)
mlm_input, mlm_targets = mlm_mask(tokens)
# Span corruption
np.random.seed(seed)
span_result = SpanCorruptor().corrupt(tokens)
return {
"original": tokens,
"mlm_input": mlm_input,
"mlm_targets": mlm_targets,
"span_input": span_result["input"],
"span_target": span_result["target"],
}sample_text = """The transformer architecture revolutionized natural language
processing by enabling parallel computation and capturing long range dependencies
through self attention mechanisms"""
comparison = compare_corruption_methods(sample_text)=== Original === The transformer architecture revolutionized natural language processing by enabling parallel computation and capturing long range dependencies through self attention mechanisms === MLM (2 masks) === The transformer architecture revolutionized natural language [MASK] by enabling parallel [MASK] and capturing long range dependencies through self attention mechanisms Targets: [(6, 'processing'), (10, 'computation')] === Span Corruption (19 input tokens) === The transformer architecture revolutionized natural language processing by enabling parallel computation and capturing long <extra_id_0> through self attention mechanisms Target: <extra_id_0> range dependencies
With MLM, each mask corresponds to exactly one token. With span corruption, each sentinel may require generating multiple tokens, forcing the model to learn phrase-level coherence.
The difference becomes clearer when we visualize the corruption patterns side by side. Let's create multiple samples and compare how MLM scatters masks across the sequence while span corruption creates contiguous blocks.
def visualize_corruption_patterns(seq_length=60, num_samples=8, seed=42):
"""Generate corruption patterns for visualization."""
np.random.seed(seed)
mlm_patterns = []
span_patterns = []
for i in range(num_samples):
# MLM: independent token masking
mlm_mask = np.random.random(seq_length) < 0.15
mlm_patterns.append(mlm_mask)
# Span corruption
span_mask = select_corruption_mask(
seq_length, corruption_rate=0.15, mean_span_length=3
)
span_patterns.append(span_mask)
return np.array(mlm_patterns), np.array(span_patterns)
mlm_patterns, span_patterns = visualize_corruption_patterns()

The visual contrast is clear. MLM produces a scattered, salt-and-pepper pattern where each dark cell is an isolated prediction task. Span corruption creates horizontal streaks, each representing a multi-token reconstruction challenge. This structural difference explains why span-corrupted models develop stronger phrase-level generation capabilities.
Span Corruption Variants
Researchers have explored several variations of the basic span corruption approach. Each variant modifies one or more of the core design choices: the span length distribution, the corruption rate, or the attention pattern used during training. Understanding these variants illuminates the space of possible pretraining objectives and helps explain design choices in models beyond T5.
Uniform vs. Geometric Span Lengths
While T5 uses geometric distribution, some work explores uniform distributions. With uniform sampling between 1 and , every span length up to is equally likely:
def sample_span_lengths_uniform(num_spans, min_length=1, max_length=5):
"""Sample span lengths from uniform distribution."""
return np.random.randint(min_length, max_length + 1, size=num_spans)
def sample_span_lengths_geometric(num_spans, mean_length=3):
"""Sample span lengths from geometric distribution."""
p = 1 / mean_length
return np.random.geometric(p, size=num_spans)n_samples = 5000
uniform_spans = sample_span_lengths_uniform(n_samples, max_length=5)
geometric_spans = sample_span_lengths_geometric(n_samples, mean_length=3)

Geometric distributions produce more short spans but allow occasional long ones. Uniform distributions guarantee exposure to longer spans but may waste capacity on trivially short reconstructions.
The practical implication is that the choice of distribution should match the downstream task distribution. If the primary downstream tasks involve short phrase completions and entity identification, the geometric distribution with a small mean aligns well. If the tasks involve longer document-level generation or summarization where multi-sentence spans matter, experimenting with higher mean values or distributions with heavier tails may be worthwhile.
We can also explore how different mean span lengths affect the distribution shape. Higher means produce flatter distributions with more emphasis on longer spans.

With , nearly half of all spans are single tokens. This provides mostly token-level signal. With , the distribution flattens clearly. This exposes the model to longer spans more frequently but at the cost of fewer distinct spans per sequence. T5's choice of balances these extremes.
Corruption Rate Variations
The original T5 paper tested corruption rates from 10% to 50%. Higher rates create shorter inputs but longer targets:
def analyze_corruption_rate(
seq_length, corruption_rate, mean_span=3, num_trials=1000
):
"""Analyze the effect of corruption rate on sequence lengths."""
input_lengths = []
target_lengths = []
for _ in range(num_trials):
mask = select_corruption_mask(seq_length, corruption_rate, mean_span)
spans = get_span_boundaries(mask)
# Input length = original - corrupted + num_sentinels
corrupted_count = mask.sum()
input_len = seq_length - corrupted_count + len(spans)
input_lengths.append(input_len)
# Target length = corrupted + num_sentinels
target_len = corrupted_count + len(spans)
target_lengths.append(target_len)
return np.mean(input_lengths), np.mean(target_lengths)seq_length = 512
corruption_rates = [0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5]
results = []
for rate in corruption_rates:
input_len, target_len = analyze_corruption_rate(seq_length, rate)
results.append(
{
"rate": rate,
"input_length": input_len,
"target_length": target_len,
"total": input_len + target_len,
}
)
At 15% corruption, the input shrinks modestly while the target remains manageable. Higher rates shift more content to the target, which can slow training due to longer autoregressive generation.
Prefix LM Variant
Some models combine span corruption with prefix language modeling. The uncorrupted prefix receives bidirectional attention, while corrupted spans use causal attention:
Input: "The quick brown" + <extra_id_0> + "jumps over" + <extra_id_1>
Attention:
- "The quick brown": Bidirectional (full visibility)
- Sentinels and after: Causal (left-to-right only)
This hybrid approach lets the model use bidirectional context for understanding while maintaining generative capability. The intuition is that a strong language model needs both skills: reading and understanding text (served by bidirectional attention) and generating new text (served by causal autoregressive generation). The prefix LM variant approximates this by giving the prefix full attention while treating the generation phase causally.
UL2 (Unifying Language Learning Paradigms), introduced by Google in 2022, extended this idea further by mixing multiple denoising objectives during pretraining, including span corruption, prefix language modeling, and extreme long-span corruption. The UL2 paper showed that no single objective is universally best across all tasks and that a mixture of objectives produces more versatile models. This finding validated the span corruption design while also showing its limitations: a model pretrained only on span corruption will underperform on tasks that require strong causal generation from an open-ended prompt.
Computational Benefits
Span corruption offers surprising computational advantages over token-level masking. These benefits arise not from architectural changes but purely from the way the corruption objective reshapes the sequences that the model must process.
Shorter Sequences
Because multiple tokens collapse into single sentinels, the encoder processes fewer tokens. For a 512-token sequence with 15% corruption and mean span length 3:
- Original tokens corrupted: ~77 tokens
- Sentinels added: ~26 tokens
- Net input length: ~461 tokens (10% reduction)
This reduction compounds across the quadratic attention mechanism. Attention cost scales as , where is the sequence length. A 10% length reduction (from to ) yields roughly 19% savings in attention computation, since .
In practice, at the scale T5 operates, this matters. T5-11B was pretrained on approximately one trillion tokens. A 10% reduction in per-example computation translates directly to fewer GPU-hours needed to process the same amount of text, or equivalently, the ability to process 10% more text within the same compute budget. At billion-dollar pretraining costs, even single-digit percentage efficiency improvements are meaningful.
Shorter Targets
The target sequence contains only corrupted spans plus sentinels. With 15% corruption:
- Target length: ~77 (spans) + ~26 (sentinels) = ~103 tokens
The decoder processes 103 tokens instead of 512, dramatically reducing generation cost during training. Autoregressive generation in the decoder is sequential and cannot be parallelized across positions (each token depends on all previous tokens). Reducing the target from 512 to 103 tokens makes the sequential decoding step roughly 5x faster, which is a much more significant speedup than the encoder efficiency gain.
The combination of a shorter encoder input and a dramatically shorter decoder target means span corruption's computational profile looks quite different from full-sequence causal language modeling. While a causal LM must process the entire 512-token sequence sequentially in the decoder, span corruption processes 461 tokens in parallel in the encoder and 103 tokens sequentially in the decoder. This asymmetric split is computationally favorable because parallel operations are so much cheaper on modern hardware than sequential ones.
Training Efficiency Comparison
Let's quantify the computational savings:
def compute_training_cost(seq_length, corruption_rate, mean_span):
"""
Estimate relative computational cost for different pretraining approaches.
Uses simplified model where cost ~ sequence_length^2 for attention.
"""
# Calculate expected lengths
num_corrupted = int(seq_length * corruption_rate)
num_spans = max(1, num_corrupted // mean_span)
# Span corruption
span_input_len = seq_length - num_corrupted + num_spans
span_target_len = num_corrupted + num_spans
span_cost = span_input_len**2 + span_target_len**2 # Encoder + Decoder
# Standard MLM (encoder only, full sequence)
mlm_cost = seq_length**2
# Causal LM (decoder only, full sequence)
clm_cost = (
seq_length**2
) # Simplified; actual is O(n^2/2) due to causal mask
return {
"span_corruption": span_cost,
"mlm": mlm_cost,
"causal_lm": clm_cost,
"span_input_len": span_input_len,
"span_target_len": span_target_len,
}costs = compute_training_cost(512, 0.15, 3)Sequence length: 512 Corruption rate: 15%, Mean span: 3 Span corruption: Input length: 461 Target length: 101 Relative cost: 84.96% of MLM MLM/Causal LM cost: 100% (baseline)
Let's visualize these computational trade-offs across different corruption rates to understand when span corruption provides the greatest efficiency gains.
# Compute costs across corruption rates
cost_comparison = []
for rate in [0.1, 0.15, 0.2, 0.25, 0.3]:
costs = compute_training_cost(512, rate, 3)
cost_comparison.append(
{
"rate": rate,
"relative_cost": costs["span_corruption"] / costs["mlm"] * 100,
"input_len": costs["span_input_len"],
"target_len": costs["span_target_len"],
}
)
Span corruption achieves meaningful computational savings while still exposing the model to diverse reconstruction challenges. At the standard 15% corruption rate, we save roughly 15% of computation compared to MLM. The encoder sees nearly the full context (minus corrupted tokens), while the decoder focuses only on what needs reconstruction.
In Practice: Using Span Corruption for Fine-Tuning
Understanding span corruption is useful for pretraining from scratch and for fine-tuning T5-based models effectively. The text-to-text framework that span corruption enables means that T5 can be adapted to almost any NLP task simply by framing the task as a text generation problem.
For classification tasks such as sentiment analysis, the input becomes "classify sentiment: [text]" and the target becomes "positive" or "negative." There is no separate classification head; the model generates the label as a text token. For extractive question answering, the input becomes "answer question: [question] context: [passage]" and the target is the answer span extracted from the passage. For abstractive summarization, the input is "summarize: [document]" and the target is the summary. In each case, the model applies exactly the same generation mechanism it learned during pretraining.
This uniformity is both a strength and a constraint. The strength is that a single model architecture and inference pipeline handles all tasks. The constraint is that task framing matters: the input prefix (the "classify sentiment:" or "summarize:" prefix) signals to the model what kind of generation is expected, and choosing the right prefix can substantially affect performance. Practitioners often experiment with several prefix phrasings and select the one that achieves the best validation performance, a process sometimes called prompt engineering even in the fine-tuning context.
Span corruption during pretraining directly trains the skills needed for this fine-tuning pattern. When the pretrained model sees <extra_id_0> in the input, it has learned: "this sentinel marks a location where I need to generate tokens." When fine-tuning reformulates tasks as "given this input, generate this output," the model applies the same generation skill. The sentinel tokens themselves are not present at fine-tuning time, but the underlying capability of context-conditioned generation transfers directly.
In practice, T5 fine-tuning proceeds as follows. You take a pretrained T5 checkpoint and a dataset formatted as input-output text pairs. You run standard seq2seq training with cross-entropy loss, updating all model weights (full fine-tuning) or just a subset (parameter-efficient methods like LoRA). Learning rates are typically much lower than pretraining (1e-4 to 1e-3 range), and training runs for far fewer steps. The model converges quickly because the pretrained representations already encode strong linguistic knowledge; fine-tuning only needs to specialize this knowledge for the specific task distribution.
Limitations and Practical Considerations
Span corruption offers useful capabilities, but it comes with trade-offs that affect model behavior and downstream applications.
The most significant limitation is the mismatch between pretraining and generation tasks. During pretraining, the model learns to infill missing spans given surrounding context. During generation, the model must produce text autoregressively without such scaffolding. This gap means span-corrupted models may struggle with open-ended generation compared to models pretrained with causal language modeling. T5 addresses this partially by framing all tasks as text-to-text, but the underlying tension remains. Fine-tuning helps bridge the gap, yet models pretrained exclusively on span corruption often require more adaptation for generation-heavy applications where no clear "surrounding context" exists to anchor reconstruction.
Another consideration is span boundary artifacts. The model learns that sentinels mark span boundaries, potentially creating implicit assumptions about phrase structure. If spans happen to align with linguistic units (noun phrases, clauses), the model may learn useful structure. If spans cut through units arbitrarily, the model must learn to handle artificial boundaries. The random nature of span selection means both cases occur, which may introduce noise into the learned representations.
The reconstruction ambiguity problem is subtler but worth understanding. When the model is asked to reconstruct "quick brown" in the context "The <extra_id_0> fox jumps," the answer is unambiguous. But many natural text spans are not uniquely determined by context. "The <extra_id_0> walked into the room" could be completed with "man," "woman," "professor," "child," or hundreds of other valid continuations. The training objective penalizes all completions except the exact original text, which means the model is repeatedly penalized for generating plausible but non-original responses. This is similar to the exposure bias problem in sequence generation: the model is trained to match a specific output, but at test time there is no single correct output. Over billions of training examples, this effect averages out because the model encounters so many diverse contexts, but it means the model's generation probabilities should not be interpreted as true posterior probabilities over all valid completions.
Key practical limitations include:
- Sentinel token overhead: Each span requires a unique sentinel, consuming vocabulary space and adding tokens to both input and target sequences.
- Span length sensitivity: Very short spans (length 1) resemble MLM without the multi-prediction benefit. Very long spans remove too much context for reliable reconstruction.
- Reconstruction ambiguity: When spans contain common phrases, multiple valid reconstructions may exist, yet training penalizes all but the original.
- Pretraining-inference mismatch: At inference time on downstream tasks, no sentinel tokens appear in the input. The model must generalize from the pretraining distribution of sentinel-marked gaps to the fine-tuning distribution of task-formatted prompts.
- Dependency on encoder-decoder architecture: The full benefits of span corruption rely on having a bidirectional encoder. Applying span corruption to decoder-only models requires adaptation (such as the fill-in-the-middle variant) and loses the bidirectional encoding advantage.
Despite these limitations, span corruption works well in practice. T5 and its variants achieve strong performance across diverse benchmarks, suggesting that the benefits of efficient training and phrase-level learning outweigh the costs. The span corruption objective has also influenced subsequent work in unexpected ways: document denoising tasks, code infilling models, and even diffusion-based text generation approaches all draw on the insight that reconstruction from partial observations is a powerful self-supervised learning signal.
Key Parameters
When implementing span corruption, the following parameters have the greatest impact on training behavior and model performance:
-
corruption_rate(float, default: 0.15): Fraction of tokens to include in corrupted spans. Higher rates (0.3-0.5) create more challenging reconstruction tasks but produce longer targets that slow training. Lower rates (0.1) may underutilize the model's capacity. The T5 default of 0.15 balances learning signal with computational efficiency. -
mean_span_length(int, default: 3): Controls the average number of tokens per corrupted span via the geometric distribution parameter . Shorter spans (mean 1-2) behave like token-level masking. Longer spans (mean 5+) force phrase-level generation but reduce the number of distinct spans. A mean of 3 provides good coverage of both short and medium-length phrases. -
num_sentinels(int, default: 100): Maximum number of unique sentinel tokens available. This caps the number of spans per sequence. With a 15% corruption rate and mean span length of 3, a 512-token sequence produces ~26 spans, well within the default limit. Increase only for very long sequences or high corruption rates. -
span_length_distribution(str, options: "geometric", "uniform"): Geometric distributions favor shorter spans with occasional long ones, matching T5's approach. Uniform distributions guarantee equal exposure to all span lengths within a range but may waste capacity on trivially short reconstructions.
Summary
Span corruption extends masked language modeling by corrupting contiguous token spans rather than individual tokens. Key takeaways:
- Span selection: Geometric distributions with mean 3 balance short and long spans, exposing the model to both token-level and phrase-level reconstruction.
- Sentinel tokens: Unique sentinels (
<extra_id_0>,<extra_id_1>, ...) replace each span, creating a clear mapping between input positions and target outputs. - Sequence-to-sequence framing: The corrupted input becomes the encoder source; concatenated spans with sentinels become the decoder target.
- Computational efficiency: Shorter input and target sequences reduce attention costs, making training more efficient than full-sequence objectives.
- Trade-offs: The infilling objective may not transfer perfectly to open-ended generation, requiring careful fine-tuning for generation tasks.
Span corruption powers T5 and influenced subsequent models like UL2 and PaLM, which explore mixtures of denoising objectives. Understanding this technique provides insight into how pretraining objectives shape model capabilities and efficiency. The broader lesson is that self-supervised pretraining objectives are not arbitrary: the structure of the pretraining task directly shapes what the model learns to do, and choosing an objective that mirrors the structure of downstream tasks is a key lever for building capable, transferable models.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about span corruption and T5-style pretraining.
Span Corruption Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!