T5 Architecture and Text-to-Text Transfer Learning

Michael BrenndoerferAugust 14, 202551 min read

Part of Language AI Handbook

Covers T5's encoder-decoder architecture, relative position biases, span corruption pretraining, and text-to-text framework for unified NLP tasks.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

T5 Architecture

The Text-to-Text Transfer Transformer (T5) introduced a unifying framework for natural language processing: treat every task as text generation. Translation, summarization, question answering, and classification all become the same problem of mapping input text to output text. This simple approach allowed researchers to study what matters most for transfer learning at scale, and the answers they found reshaped how the field thinks about model design, pretraining objectives, and the relationship between architecture and capability.

Released by Google Research in 2019, T5 emerged from a systematic exploration of pre-training techniques, model architectures, and scaling strategies. The accompanying paper, "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer," tested dozens of design choices to identify which ones mattered. The result was both a powerful model family and a practical guide to building better language models. Rather than proposing a single novel trick and claiming it as the source of improved performance, the T5 researchers ran a controlled ablation study at a scale that few teams could match, making their findings unusually trustworthy.

T5's encoder-decoder architecture handles both understanding and generation in a single model. The encoder processes the full input with bidirectional attention, while the decoder generates output autoregressively. This design excels at tasks requiring deep comprehension of the input before creating structured output. Translation and summarization fit this pattern, as does question answering.

The influence of T5 reaches well beyond the model itself. Relative position biases, which T5 popularized for large-scale Transformers, became a standard component in many subsequent architectures. The span corruption pretraining objective demonstrated that masked language modeling could be made more effective by corrupting spans rather than individual tokens. And the text-to-text framing anticipated the instruction-following paradigm that would later dominate the field: if you can express any task as "here is your instruction, here is your input, now produce the output," then a single model can learn to follow arbitrary natural language directions.

Understanding T5 also gives you a clearer picture of the range of Transformer architectures. By the time T5 was released, the field had split into two main camps: encoder-only models like BERT, optimized for understanding tasks, and decoder-only models like GPT, optimized for generation. T5 offered a third path, one that maintained the encoder-decoder separation from the original Transformer and showed it could compete with specialized designs across a wide range of tasks. The story of why certain architectures suit certain tasks, and how training objectives shape the representations models learn, runs through every section of this chapter.

Historical Context

T5 appeared in late 2019, roughly one year after BERT and at the same time as GPT-2. The NLP community had recently discovered that large-scale pretraining on raw text followed by fine-tuning on labeled data sharply outperformed models trained from scratch on individual tasks. T5's contribution was to ask: given that pretraining works, what design decisions matter most? Its large-scale controlled experiments influenced model design for years afterward, and its text-to-text framing directly foreshadowed the instruction-tuned models that would dominate benchmarks by 2022.

The Text-to-Text Framework

T5's central insight is that natural language gives a universal interface for NLP tasks. Instead of building task-specific architectures with specialized output heads, T5 learns to generate the answer as text. This means the same model, loss function, and training procedure work for any task.

To appreciate why this matters, consider how NLP systems were traditionally built. Classification tasks required a final layer that mapped hidden representations to a fixed set of class probabilities. Question answering systems needed span prediction heads that identified start and end positions in text. Translation models required sequence-to-sequence architectures with dedicated vocabulary handling for each language pair. Each task demanded its own architectural modifications, training objectives, and output processing logic. When practitioners wanted to build a system that handled multiple tasks, they faced a choice: build and maintain several separate models, or construct multi-task architectures with complex shared backbones and task-specific heads. Neither option was particularly elegant.

The text-to-text framework eliminates these distinctions by building on an insight: natural language is already a universal representation system. Humans express classification decisions ("this is negative"), answer questions ("the color is blue"), and produce translations ("Das Haus ist wunderbar") all through the same medium, text. If we train a model to be exceptionally good at generating text, it can express any answer we might need. The format of the output encodes the task semantics, not the architecture. A model that generates "positive" when asked about sentiment is performing classification just as validly as one that outputs a probability vector.

Text-to-Text Transfer

A framework where all NLP tasks are cast as text generation problems. The model receives text input (with a task prefix) and produces text output, eliminating the need for task-specific architectures.

Consider how different tasks map to this framework:

  • Translation: translate English to German: The house is wonderful. → Das Haus ist wunderbar.
  • Summarization: summarize: [long article text] → [concise summary]
  • Classification: sentiment: This movie was terrible. → negative
  • Question answering: question: What color is the sky? context: The sky appears blue during the day. → blue

The task prefix tells the model what to do, and the target text encodes the answer. Classification becomes generating the class label as a word. Regression could output the number as text. Even complex structured outputs like parse trees can be serialized as strings. The key constraint is that the model must learn to associate each prefix with the correct behavior during pretraining and fine-tuning. This is less fragile than it might sound: natural language is rich enough that even short prefixes like "summarize:" or "translate English to German:" carry strong signals about what kind of output is expected.

This uniformity has practical benefits. You can fine-tune a single model on multiple tasks simultaneously. You can add new tasks without changing the architecture. You can also use text generation advances, like beam search and nucleus sampling, across all applications. Perhaps most importantly, unified training allows tasks to share representations. A model that learns to translate English to French and English to German does not learn two unrelated skill sets. It learns representations of English meaning that transfer across both translation directions, and this sharing often improves performance on both tasks.

The text-to-text framing also has a subtle pedagogical advantage: it makes the model's "reasoning" more transparent. When a sentiment classifier outputs a probability distribution over classes, it is hard to interrogate what the model understood. When T5 outputs "negative" or "positive" as text, you can probe why it chose that word, ask follow-up questions, or even request explanations. This interpretability benefit becomes more significant in the era of larger models and chain-of-thought prompting, where the output text itself is a reasoning trace.

In Practice: Designing Task Prefixes

The T5 paper used simple, human-readable prefixes for each task. For the GLUE benchmark tasks, prefixes like "sst2 sentence:" (for sentiment) and "cola sentence:" (for grammatical acceptability) were used. These prefixes were not learned automatically; they were hand-designed to be descriptive of the task. The model's ability to generalize to new tasks at inference time depends on the prefix resembling the task descriptions it saw during training. This is why instruction-tuned successors like Flan-T5 invest heavily in diverse and natural-sounding instruction formats: if you train on hundreds of different phrasings of the same task, the model becomes more reliable to novel instructions that differ from any specific training example.

Notice that this framing requires the vocabulary to include the class labels as tokens. For sentiment classification, "positive" and "negative" must be representable tokens, which they are in T5's SentencePiece vocabulary. For tasks with unusual label spaces (like predicting numerical ratings or specialized codes), you may need to ensure your labels are covered by the tokenizer. This is rarely a problem in practice because T5's 32,128-token vocabulary covers most English words and common numeric representations, but it is worth keeping in mind when adapting the framework to highly specialized domains.

Encoder-Decoder Architecture

T5 uses the original Transformer's encoder-decoder structure, with modifications that improve training stability and performance. This architectural choice shows an insight about how different NLP tasks process information. Some tasks, like classification or sentiment analysis, primarily require understanding an input. Others, like open-ended writing, primarily require generating new content. But tasks that combine understanding with generation require both: deep comprehension of the input followed by structured generation of output. The encoder-decoder architecture gives dedicated machinery for each phase of this process.

The encoder processes the input sequence bidirectionally. This creates rich contextual representations where each token's embedding shows its relationships with every other token in the input. The decoder then generates output tokens one at a time, attending to both the encoder's output and its own previous predictions. This separation allows the encoder to build a complete understanding of the source text before the decoder begins generating. This keeps even the first output token benefits from full context about the input.

Think of the encoder as a reader who reads the entire source document carefully before picking up a pen, and the decoder as the writer who drafts the output word by word while frequently consulting the reader's notes. The notes (the encoder's hidden states) do not change as the writer works; they stand for a fixed, complete summary of the source. At each step of writing, the decoder can ask: "which parts of the source are most relevant to what I am about to say?" The cross-attention mechanism gives the mechanism for this consultation.

Encoder Structure

The encoder consists of stacked Transformer blocks, each containing self-attention and feed-forward layers. Unlike GPT-style decoders, the encoder uses bidirectional attention. Every token can attend to every other token in the input, regardless of position. This bidirectionality is important for tasks like translation and summarization, where understanding a word often requires seeing what comes both before and after it. Consider the sentence "The bank was steep." Determining whether "bank" refers to a financial institution or a riverbank requires seeing "steep," which comes later in the sequence. A unidirectional encoder would have already committed to a representation of "bank" before seeing "steep," and correcting that representation would require it to propagate updates backward through attention, which is less direct than bidirectional attention that simply sees the whole sentence at once.

Each encoder block applies a carefully orchestrated sequence of operations that transform token representations while maintaining training stability:

  1. Layer normalization (applied before attention, not after)
  2. Multi-head self-attention with relative position biases
  3. Residual connection
  4. Layer normalization
  5. Position-wise feed-forward network
  6. Residual connection

T5 uses "pre-norm" placement, where layer normalization comes before each sublayer rather than after. This architectural choice, explored systematically in the T5 paper, improves training stability, especially for deeper models. The intuition is straightforward: normalizing inputs to each sublayer ensures that attention and feed-forward operations receive consistently scaled values, preventing the accumulation of extreme activations that can destabilize training in deep networks. In the original Transformer, layer normalization was applied after the residual connection ("post-norm"). Empirically, pre-norm training is more stable and allows higher learning rates, though post-norm sometimes reaches better final performance when the training does converge. T5's systematic experiments confirmed that pre-norm was the better default for large-scale training.

The feed-forward network inside each encoder block follows a specific structure: a linear expansion to a larger dimension (typically 4 times the hidden size), a nonlinear activation, and a linear projection back to the hidden dimension. In T5, this uses a gated activation function called GeGLU (Gated Linear Unit with GELU), which allows the network to selectively pass or gate information based on a learned gating signal. The result is that the feed-forward layer acts somewhat like a key-value memory: certain input patterns activate certain "memory slots" and retrieve associated information. Research on larger models has found that a significant fraction of factual knowledge is stored in the feed-forward layers rather than the attention layers.

Decoder Structure

The decoder mirrors the encoder's structure but adds cross-attention to incorporate information from the encoded input. This cross-attention mechanism is what allows the decoder to "consult" the encoder's understanding of the input at every step of generation. Each decoder block contains:

  1. Layer normalization
  2. Masked self-attention (causal, so tokens only attend to previous positions)
  3. Residual connection
  4. Layer normalization
  5. Cross-attention to encoder outputs
  6. Residual connection
  7. Layer normalization
  8. Position-wise feed-forward network
  9. Residual connection

The masking in self-attention ensures the decoder can only see tokens it has already generated, maintaining the autoregressive property needed for text generation. Without this mask, the model could "cheat" during training by looking at future tokens, learning to copy rather than predict. Cross-attention allows each decoder position to attend to all encoder positions, integrating the input representation into the generation process. When generating a German translation, each German word can attend to all the English words, determining which parts of the source are most relevant for creating the current output token.

The interaction between masked self-attention and cross-attention is worth examining carefully. The self-attention sublayer lets the decoder build coherent internal context: "given what I have generated so far, what is the current state of my output?" The cross-attention sublayer then anchors that internal context to the source: "given my current generation state, which parts of the input should inform my next word?" These two sources of information, the decoder's own history and the encoder's representation of the source, combine in the feed-forward layers to produce the final output distribution. This two-source structure is what makes encoder-decoder models particularly well-suited for tasks with a clear input-output separation, as opposed to open-ended generation where there is no distinct "source" to attend to.

Information Flow

The complete forward pass proceeds through a two-stage process that first builds understanding and then produces output. First, input tokens are embedded and passed through all encoder layers. Each encoder layer refines the representations, with early layers capturing local syntactic patterns and deeper layers building more abstract semantic representations. The encoder's final hidden states, one vector per input token, become the "memory" that the decoder will reference throughout generation.

During generation, the decoder receives the previously generated tokens (or just a start token initially). These pass through self-attention layers with causal masking, letting the model to consider what it has already said while deciding what to say next. Then cross-attention layers query the encoder memory, determining which parts of the input are relevant for generating the current token. The final decoder hidden state projects to vocabulary logits, and the highest-probability token becomes the next output. This process repeats, with each new token extending the decoder's self-attention context, until the model produces a stop token or reaches a maximum length.

The key insight is that encoder computation and decoder computation are interleaved at training time but separated at inference time. During training, the encoder runs once per example, and the decoder runs in parallel over all target positions (using teacher forcing, where the ground truth target is fed in rather than the model's own predictions). During inference, the encoder runs once, but the decoder must run autoregressively, one step at a time, because each new token depends on all previous outputs. This asymmetry means that inference speed is dominated by decoder steps, and techniques for accelerating autoregressive generation (like key-value caching) are necessary for deploying T5 in production settings.

In[3]:
Code
from transformers import T5ForConditionalGeneration, T5Tokenizer

# Load T5-small for exploration
model_name = "t5-small"
tokenizer = T5Tokenizer.from_pretrained(model_name)
model = T5ForConditionalGeneration.from_pretrained(model_name)

# Examine the architecture structure
print("Encoder blocks:", len(model.encoder.block))
print("Decoder blocks:", len(model.decoder.block))
Out[3]:
Console
Loading weights:   0%|          | 0/131 [00:00<?, ?it/s]
Encoder blocks: 6
Decoder blocks: 6

T5-small uses 6 blocks in both the encoder and decoder. Let's examine a single encoder block to see its components:

In[4]:
Code
# Inspect first encoder block
encoder_block = model.encoder.block[0]
print("Encoder block components:")
for name, module in encoder_block.named_children():
    print(f"  {name}: {module.__class__.__name__}")
Out[4]:
Console
Encoder block components:
  layer: ModuleList

The T5LayerSelfAttention handles self-attention with relative positions, while T5LayerFF implements the feed-forward network. Now let's compare with a decoder block:

In[5]:
Code
# Inspect first decoder block
decoder_block = model.decoder.block[0]
print("Decoder block components:")
for name, module in decoder_block.named_children():
    print(f"  {name}: {module.__class__.__name__}")
Out[5]:
Console
Decoder block components:
  layer: ModuleList

The decoder adds a second attention layer (cross-attention) for attending to encoder outputs. This extra sublayer is what distinguishes encoder-decoder from decoder-only architectures at the implementation level, and it is the source of both the architectural advantage (dedicated source attention) and the additional parameter cost (roughly 50% more attention parameters in each decoder block).

Relative Position Biases

Standard Transformers use absolute position embeddings. Each position in the sequence gets a fixed embedding vector added to the token embedding. T5 takes a different approach with relative position biases, which encode the distance between tokens rather than their absolute locations. This design decision shows a better understanding of what position information the model needs.

Why Relative Positions?

Absolute positions have limitations that become apparent when you consider how humans process language. When reading the phrase "the big red ball," understanding that "big" and "red" both modify "ball" doesn't depend on whether this phrase appears at the beginning or middle of a paragraph. What matters is that these words are adjacent to each other. A model trained on sequences up to 512 tokens has never seen position 513, making generalization to longer sequences difficult with absolute positions. The embedding for position 513 simply doesn't exist, forcing awkward workarounds like position interpolation.

Relative positions encode the offset between the query and key positions in attention, capturing this linguistically real notion of proximity. Whether two words appear at positions 10 and 15 or positions 100 and 105, they're "5 positions apart" in both cases. This translation invariance helps the model generalize across different sequence positions. A relationship learned between adjacent words at the beginning of training examples automatically transfers to adjacent words anywhere in any sequence.

The key insight is that position biases should modify attention scores directly. Rather than adding position information to token embeddings (which then influences attention indirectly through the learned query/key projections), T5 adds learned biases directly to the attention logits. In standard attention, the attention score between positions ii and jj becomes:

score(i,j)=qi⋅kjdk+b(j−i)\text{score}(i, j) = \frac{\mathbf{q}_i \cdot \mathbf{k}_j}{\sqrt{d_k}} + b(j - i)

where:

  • qi\mathbf{q}_i: the query vector at position ii
  • kj\mathbf{k}_j: the key vector at position jj
  • dkd_k: the dimension of the key vectors (used for scaling)
  • b(j−i)b(j - i): the learned bias for relative position j−ij - i

Let's unpack what this formula tells us. The first term, qi⋅kjdk\frac{\mathbf{q}_i \cdot \mathbf{k}_j}{\sqrt{d_k}}, is the standard scaled dot-product attention score that measures content-based similarity between positions. This captures whether the semantic content at position ii wants to attend to the semantic content at position jj. The second term, b(j−i)b(j - i), adds a position-based preference that depends only on how far apart the positions are, not where they sit in the sequence. A token might learn that it generally wants to attend strongly to the immediately preceding token (j−i=−1j - i = -1), regardless of content.

The bias term b(j−i)b(j - i) depends only on the distance between positions, not their absolute values. This allows the same position relationship to receive the same bias regardless of where it appears in the sequence. The model can learn, for example, that in English text, tokens often attend strongly to words 1-3 positions away (local syntactic patterns) while having weaker but still real attention to more distant positions (long-range dependencies).

Notice that the position bias is added directly to the pre-softmax logits, not to the post-softmax weights. This placement matters: adding to logits means the bias operates in the same space as the content-based scores, letting a large positive bias to override weak content similarity and a large negative bias to suppress attention even when content similarity is high. An absolute position embedding added to token representations, by contrast, must be routed through the learned query and key projections before influencing attention scores, giving the model less direct control over position-based attention patterns.

Relative Position Bias

A learned scalar added to attention logits based on the distance between query and key positions. Unlike position embeddings added to token representations, position biases directly modulate attention weights.

T5's Position Bias Implementation

T5 uses a bucketed relative position scheme that balances expressiveness with parameter efficiency. Instead of learning a separate bias for every possible offset (which would require unbounded parameters as sequence length grows), T5 groups offsets into logarithmically spaced buckets. For a query at position ii and a key at position jj, the relative position is computed as:

r=j−ir = j - i

where:

  • rr: the relative position offset (positive when the key comes after the query, negative when before)
  • ii: the position of the query token in the sequence
  • jj: the position of the key token in the sequence

This offset rr is then mapped to a bucket index, and the model learns a bias value for each bucket. The bucketing strategy encodes an important insight: precise position differences matter more for nearby words than for distant ones. Whether a related word is exactly 57 or 62 positions away rarely changes its relevance, but whether it's 1 or 2 positions away often does.

The bucketing works as follows:

  1. Compute the relative position: r=j−ir = j - i where ii is the query position and jj is the key position
  2. For small offsets, use exact values (each offset gets its own bucket)
  3. For larger offsets, use logarithmic bucketing (multiple offsets share a bucket)
  4. Look up the learned bias for that bucket

The logarithmic spacing means nearby positions (which often carry more grammatical signal) get fine-grained distinctions, while distant positions are grouped more coarsely. This keeps the parameter count manageable while still capturing useful position information. With 32 total buckets split between forward and backward directions, the model can stand for a rich set of position relationships without requiring thousands of parameters per attention head.

For larger offsets, the bucket index is computed using logarithmic scaling:

b=bexact+⌊log⁡(r/bexact)log⁡(dmax/bexact)⋅(Bdir−bexact)⌋b = b_{\text{exact}} + \left\lfloor \frac{\log(r / b_{\text{exact}})}{\log(d_{\text{max}} / b_{\text{exact}})} \cdot (B_{\text{dir}} - b_{\text{exact}}) \right\rfloor

where:

  • bb: the final bucket index for this relative position
  • bexactb_{\text{exact}}: the number of buckets reserved for exact (small) offsets
  • rr: the absolute value of the relative position offset
  • dmaxd_{\text{max}}: the maximum distance considered (default 128 in T5)
  • BdirB_{\text{dir}}: the number of buckets per direction (half the total buckets)

This formula deserves careful examination because it reveals the design philosophy behind T5's position encoding. The numerator log⁡(r/bexact)\log(r / b_{\text{exact}}) measures how far beyond the exact-bucket threshold the offset reaches, on a logarithmic scale. Dividing by log⁡(dmax/bexact)\log(d_{\text{max}} / b_{\text{exact}}) normalizes this to a value between 0 and 1 across the range of larger offsets. Multiplying by (Bdir−bexact)(B_{\text{dir}} - b_{\text{exact}}) spreads these normalized values across the available buckets for large offsets. The floor operation ensures we get discrete bucket indices. Adding bexactb_{\text{exact}} shifts the result into the correct range, after the buckets reserved for small exact offsets.

This formula maps offsets beyond bexactb_{\text{exact}} into logarithmically-spaced buckets. This keeps the distinction between positions 1 and 2 is preserved while positions 50 and 55 share the same bucket. The logarithmic spacing means bucket boundaries grow exponentially: perhaps buckets for offsets 1, 2, 3, 4, then 5-7, 8-15, 16-31, 32-63, and so on. This mirrors human perception of distance. We notice fine distinctions between nearby objects but group distant objects more coarsely.

An important practical benefit of this design is length generalization. Because the biases are parameterized by relative distance rather than absolute position, the model can handle sequences longer than those seen during training without requiring any extrapolation of position embeddings. For offsets beyond dmax=128d_{\text{max}} = 128, T5 simply clamps them to the last bucket, assigning the same bias to all very distant positions. This is a mild form of extrapolation, and in practice it works well: the model does not become confused by long documents, it simply treats all tokens beyond 128 positions as "distant."

In[6]:
Code
import numpy as np


def compute_t5_bucket(relative_position, num_buckets=32, max_distance=128):
    """
    Compute T5's relative position bucket.
    T5 uses half the buckets for exact positions, half for log-spaced.
    """
    relative_buckets = 0

    # Handle negative (backward) positions
    # In encoder, positions can be negative (key before query)
    # Use separate buckets for forward and backward
    num_buckets_per_direction = num_buckets // 2

    if relative_position < 0:
        relative_buckets = num_buckets_per_direction
        relative_position = -relative_position

    # Exact buckets for small offsets
    max_exact = num_buckets_per_direction // 2
    if relative_position < max_exact:
        return relative_buckets + relative_position

    # Log buckets for larger offsets
    relative_position_if_large = max_exact + int(
        np.log(relative_position / max_exact)
        / np.log(max_distance / max_exact)
        * (num_buckets_per_direction - max_exact)
    )
    relative_position_if_large = min(
        relative_position_if_large, num_buckets_per_direction - 1
    )

    return relative_buckets + relative_position_if_large


# Show bucket assignments for different offsets
print("Offset -> Bucket mapping:")
offsets = [-10, -5, -1, 0, 1, 2, 5, 10, 20, 50, 100]
for offset in offsets:
    bucket = compute_t5_bucket(offset)
    print(f"  Offset {offset:4d} -> Bucket {bucket:2d}")
Out[7]:
Console
Offset -> Bucket mapping:
  Offset  -10 -> Bucket 24
  Offset   -5 -> Bucket 21
  Offset   -1 -> Bucket 17
  Offset    0 -> Bucket  0
  Offset    1 -> Bucket  1
  Offset    2 -> Bucket  2
  Offset    5 -> Bucket  5
  Offset   10 -> Bucket  8
  Offset   20 -> Bucket 10
  Offset   50 -> Bucket 13
  Offset  100 -> Bucket 15

Notice how small positive offsets (0, 1, 2) each get unique buckets, while larger offsets (20, 50, 100) start collapsing into shared buckets. Negative offsets (key before query) use a separate set of buckets, letting the model to learn different biases for forward vs. backward attention. This asymmetry makes linguistic sense: attending to a word that came before ("I saw the") versus a word that comes after ("the dog ran") often serves different purposes, and the model can learn distinct patterns for each direction.

Out[8]:
Visualization
Staircase plot showing T5 relative position bucket mapping, with fine-grained buckets for small offsets and coarse-grained buckets for larger offsets.
T5's logarithmic bucketing assigns unique buckets to small offsets (fine-grained) while grouping larger offsets together (coarse-grained). The staircase pattern shows bucket boundaries growing exponentially, which shows the intuition that precise distances matter more for nearby tokens.

Visualizing Position Biases

Let's visualize the actual learned position biases from a trained T5 model:

Out[9]:
Visualization
Heatmap showing T5 position bias matrix with stronger attention near the diagonal.
Learned relative position biases for T5-small's first encoder layer, head 0. Positive biases (lighter) encourage attention between those relative positions, while negative biases (darker) suppress it. The diagonal band of positive values shows the model's learned preference for local attention to nearby tokens.

The bias pattern shows the model has learned to encourage attention to nearby positions (near the diagonal) while letting more flexibility for distant positions. This structure emerges purely from training. The model discovers what relative position patterns help solve its pretraining objective. Different attention heads learn different position bias patterns: some heads specialize in local attention (strong diagonal bands), while others develop patterns suited for attending to specific structural positions like the beginning of the sequence or position zero.

Model Sizes

T5 was released in five sizes, letting researchers and practitioners to choose the right trade-off between capability and computational cost. Each size follows the same architecture but varies in its layer depth and width, which together determine the total parameter count.

T5 model family specifications across five sizes.
ModelParametersLayersHidden SizeAttention HeadsFeed-Forward Size
T5-Small60M651282048
T5-Base220M12768123072
T5-Large770M241024164096
T5-3B3B2410243216384
T5-11B11B24102412865536

Several patterns emerge from this scaling progression. Smaller models increase depth (more layers) as they grow, while the largest models hold depth constant and scale width instead. The feed-forward dimension grows proportionally larger at scale. The number of attention heads increases materially for T5-3B and T5-11B. This gives more specialized attention patterns.

The jump from T5-Large to T5-3B is particularly striking: the hidden dimension and number of layers stay the same, but the feed-forward dimension quadruples and the number of attention heads doubles. This asymmetric scaling shows practical findings about where additional capacity is most useful. At large scales, widening the feed-forward network (which is where factual associations tend to be stored) and adding more attention heads (which allows more diverse attention patterns) tends to give more return per parameter than simply adding more layers.

Understanding the model size trade-offs also matters for practical deployment. T5-Small (60M parameters) can run on a laptop CPU and fine-tunes in minutes on modest hardware. T5-Base strikes a balance many practitioners find useful: strong enough for production quality on many tasks, yet fast enough to iterate on. T5-Large and above typically require GPU memory measured in tens of gigabytes and training times measured in hours to days. T5-11B was primarily a research tool to study scaling behavior; for production, the smaller variants are more practical.

Out[10]:
Visualization
Bar chart on a logarithmic scale comparing parameter counts across the T5 model family from T5-Small to T5-11B.
T5 model family parameter counts on a logarithmic scale, showing the large jumps between model sizes. The 180x difference between T5-Small (60M) and T5-11B (11B) lets research across a wide range of computational budgets, from a laptop to a multi-GPU cluster.
In[11]:
Code
# Compare parameter counts across model sizes
from transformers import T5Config

sizes = ["t5-small", "t5-base", "t5-large"]
for size in sizes:
    config = T5Config.from_pretrained(size)
    print(f"\n{size}:")
    print(f"  Layers: {config.num_layers}")
    print(f"  Hidden size: {config.d_model}")
    print(f"  Attention heads: {config.num_heads}")
    print(f"  FF dimension: {config.d_ff}")
    print(f"  Vocab size: {config.vocab_size}")
Out[11]:
Console

t5-small:
  Layers: 6
  Hidden size: 512
  Attention heads: 8
  FF dimension: 2048
  Vocab size: 32128

t5-base:
  Layers: 12
  Hidden size: 768
  Attention heads: 12
  FF dimension: 3072
  Vocab size: 32128

t5-large:
  Layers: 24
  Hidden size: 1024
  Attention heads: 16
  FF dimension: 4096
  Vocab size: 32128

The vocabulary size remains constant at 32,128 tokens across all sizes. This vocabulary was trained using SentencePiece on the C4 dataset, the same corpus used for pretraining. Using a fixed vocabulary across all model sizes simplifies deployment: you need only one tokenizer regardless of which T5 variant you use, and fine-tuned models can be upgraded to larger sizes without re-tokenizing training data.

Pretraining: Span Corruption

T5 uses a "span corruption" objective during pretraining, which the paper found more effective than alternatives like standard language modeling or BERT-style masked language modeling. This objective is a good challenge that forces the model to develop reliable language understanding.

How Span Corruption Works

The objective corrupts the input by replacing contiguous spans of tokens with single sentinel tokens, then asks the model to reconstruct those spans. This approach differs from BERT's masked language modeling, which corrupts individual tokens, and from GPT's causal language modeling, which predicts the next token given previous context. Span corruption strikes a middle ground that encourages the model to understand broader context while still learning to generate coherent multi-token sequences.

Here's the process in detail:

  1. Sample span lengths from a distribution (mean length 3)
  2. Select 15% of tokens total to corrupt
  3. Replace each selected span with a unique sentinel token (<extra_id_0>, <extra_id_1>, etc.)
  4. Create targets that contain the sentinel followed by the original tokens

For example:

  • Original: The quick brown fox jumps over the lazy dog
  • Corrupted input: The <extra_id_0> fox <extra_id_1> the lazy dog
  • Target: <extra_id_0> quick brown <extra_id_1> jumps over

This approach forces the model to understand context deeply. It must determine what type of content belongs in each corrupted span based on surrounding words. Unlike next-token prediction (which only requires predicting one token at a time), span reconstruction requires understanding the complete context. When the model sees The <extra_id_0> fox jumps, it must recognize that the missing span should contain adjectives describing a fox, likely words like "quick brown" or "sly red." This requires understanding both syntax (adjectives precede nouns) and semantics (foxes have certain typical descriptions).

The use of contiguous spans rather than individual tokens adds another dimension of difficulty. The model cannot simply guess each missing token independently; it must generate a coherent sequence that fits grammatically and semantically as a unit. This trains the model for the kind of fluent generation required in downstream tasks like summarization and translation. Single-token masking (as in BERT) lets the model treat each masked position independently. Span masking forces the model to plan multi-token outputs, which is a closer match to the generation tasks T5 will be fine-tuned on.

Why Span Corruption Outperforms Alternatives

The T5 paper systematically compared span corruption against several other pretraining objectives:

  • Standard language modeling: predicting the next token given previous context, as in GPT
  • BERT-style masked language modeling: independently masking 15% of tokens and predicting them
  • Deshuffling: restoring a shuffled sentence to its original order
  • Prefix language modeling: a compromise where a prefix is fed to the encoder and the suffix must be generated

Span corruption consistently outperformed these alternatives on downstream benchmarks. The key advantage over standard language modeling is that span corruption uses a bidirectional encoder, letting the model to see the full input context on both sides of each masked span. The key advantage over BERT-style masking is that span corruption requires generating multi-token outputs, which directly trains the generation capability that downstream tasks require. The key advantage over deshuffling is that span corruption preserves grammatical structure in the visible tokens, keeping the pretraining input closer to natural language.

The sentinel token approach also has an efficiency advantage. Because the target sequence contains only the masked spans (plus sentinels), it is much shorter than the full input. This reduces the computational cost of the decoder during pretraining and means the model spends most of its capacity on the interesting parts (the corrupted regions) rather than trivially copying uncorrupted text.

Sentinel Tokens

Special tokens like <extra_id_0>, <extra_id_1>, etc., used in T5's span corruption objective. Each sentinel replaces one span in the input and anchors the corresponding target reconstruction. T5's vocabulary includes 100 sentinel tokens, supporting up to 100 distinct spans per training example.

In[12]:
Code
# Demonstrate span corruption format
text = "Natural language processing lets computers to understand text."

# Simulated corruption (T5 pretraining would do this automatically)
corrupted = (
    "Natural language <extra_id_0> lets <extra_id_1> to understand text."
)
target = "<extra_id_0> processing <extra_id_1> computers"

print("Original:", text)
print("Corrupted input:", corrupted)
print("Target:", target)
Out[12]:
Console
Original: Natural language processing lets computers to understand text.
Corrupted input: Natural language <extra_id_0> lets <extra_id_1> to understand text.
Target: <extra_id_0> processing <extra_id_1> computers
Out[13]:
Visualization
Diagram comparing BERT token masking, GPT next-token prediction, and T5 span corruption pretraining objectives side by side.
Comparison of three pretraining objectives used in language model research. BERT masks individual tokens independently, GPT predicts each token from its left context only, and T5 replaces contiguous spans with sentinel tokens and reconstructs them. Span corruption (T5) combines bidirectional context with multi-token generation, which makes it better suited for downstream tasks that require both comprehension and generation.

The sentinel tokens serve as placeholders in the input and anchors in the output, letting the model to learn which span corresponds to which sentinel. This correspondence is important. By seeing <extra_id_0> in both input and output, the model learns that whatever follows <extra_id_0> in the target is what should fill the <extra_id_0> position in the input. The sentinel approach also makes the target sequence much shorter than the original input, improving training efficiency since the model only needs to generate the corrupted portions rather than reconstructing the entire input.

C4 Dataset

T5 was trained on the Colossal Clean Crawled Corpus (C4), a 750GB dataset derived from Common Crawl. The researchers applied extensive filtering to improve quality:

  • Remove pages with fewer than 5 sentences
  • Discard pages containing words from a blocklist
  • Remove duplicate lines across the corpus
  • Keep only English text (detected by language ID)
  • Remove pages with too many repetitive patterns

This cleaning produced a dataset much larger than typical pretraining corpora at the time, letting the scale experiments that T5 aimed to explore. The C4 dataset was itself a contribution of the T5 paper, and it has since been widely used as a benchmark pretraining corpus. Subsequent work on data quality, including analyses of the effect of de-duplication, quality filtering, and domain distribution, built directly on C4 as a baseline.

The scale of C4 matters for raw data volume and diversity. Common Crawl contains web pages from an enormous range of topics, writing styles, and domains. A model trained on C4 sees scientific text, news articles, forum discussions, product reviews, legal documents, and much more. This diversity is part of why T5's representations transfer well across such different downstream tasks. The breadth of the pretraining distribution means the model develops representations that generalize, rather than representations optimized for a specific domain.

One important limitation of C4 is its English-centric nature. By filtering for English text, T5's pretraining corpus excludes the majority of the world's languages. This limitation motivated the development of mT5, which we discuss below. The C4 filtering pipeline also made deliberate choices about what constitutes "quality" web text, and these choices embed assumptions about language and content that may not generalize to all use cases.

Worked Example: Fine-Tuning T5 for Question Answering

To see the text-to-text framework in action end-to-end, let's trace through what happens when T5 is fine-tuned and then used for a specific task. We will use extractive question answering as our example: given a passage and a question, generate the answer as text.

During fine-tuning, each training example is formatted as follows. The input becomes something like: question: What is the capital of France? context: France is a country in Western Europe. Its capital city is Paris, which is also the country's largest city. The target is: Paris. The model receives this formatted string as its encoder input, and the decoder is trained to generate the target string with teacher forcing.

The model already understands English from pretraining. Fine-tuning teaches it to associate the question: and context: prefix pattern with the specific behavior of extracting or generating a concise answer. After a relatively small number of fine-tuning steps on labeled question-answer pairs, the model becomes proficient at this task. This is the transfer learning payoff: pretraining builds general language understanding, and fine-tuning specializes it.

At inference time, the encoder processes the question-context string and builds a bidirectional representation. The decoder then generates an answer token by token, using cross-attention to reference the encoder's representation of the passage at each step. The first generated token might be a capital letter ("P"), the next the rest of the word ("aris"), and the decoder then produces an end-of-sequence token. The final output, "Paris," is returned as the answer.

The key insight from this example is that T5 does not need to be told explicitly where in the passage the answer is located (as span-extraction models like BERT must do). Instead, it generates the answer text directly, which means it can handle questions whose answers require paraphrasing or synthesis rather than verbatim extraction. This flexibility comes at a cost: the model's output must be post-processed to verify it matches an expected answer, whereas span extraction guarantees the output is a substring of the input. In practice, researchers use both approaches depending on the task requirements.

Working with T5

Let's use T5 for various tasks to see the text-to-text framework in action. We'll use T5-small for these examples to keep computational requirements modest.

Translation

In[14]:
Code
import torch

# Translation example
input_text = "translate English to German: The weather is beautiful today."

# Tokenize
inputs = tokenizer(input_text, return_tensors="pt", padding=True)

# Generate
with torch.no_grad():
    outputs = model.generate(
        inputs.input_ids, max_length=50, num_beams=4, early_stopping=True
    )

# Decode
translation = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Input: {input_text}")
print(f"Output: {translation}")
Out[14]:
Console
Input: translate English to German: The weather is beautiful today.
Output: Das Wetter ist heute schön.

Summarization

In[15]:
Code
# Summarization example
article = """
The Amazon rainforest produces about 20% of the world's oxygen.
It spans across nine countries in South America and contains
10% of all species on Earth. Deforestation threatens this vital
ecosystem, with an area the size of a football field being cleared
every minute. Conservation efforts are necessary to preserve
biodiversity and combat climate change.
"""

input_text = f"summarize: {article}"
inputs = tokenizer(
    input_text, return_tensors="pt", padding=True, truncation=True
)

with torch.no_grad():
    outputs = model.generate(
        inputs.input_ids, max_length=50, num_beams=4, early_stopping=True
    )

summary = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Summary: {summary}")
Out[15]:
Console
Summary: the amazon rainforest produces 20% of the world's oxygen. it spans across nine countries in south America and contains 10% of all species on earth.

Examining Internal Representations

Let's look at how T5 processes input through its encoder:

In[16]:
Code
# Get encoder hidden states
text = "The model learns to understand language through pretraining."
inputs = tokenizer(text, return_tensors="pt")

with torch.no_grad():
    encoder_outputs = model.encoder(
        input_ids=inputs.input_ids, output_hidden_states=True
    )

# Shape of hidden states from each layer
print("Encoder hidden state shapes:")
for i, hidden in enumerate(encoder_outputs.hidden_states):
    print(f"  Layer {i}: {hidden.shape}")
Out[16]:
Console
Encoder hidden state shapes:
  Layer 0: torch.Size([1, 12, 512])
  Layer 1: torch.Size([1, 12, 512])
  Layer 2: torch.Size([1, 12, 512])
  Layer 3: torch.Size([1, 12, 512])
  Layer 4: torch.Size([1, 12, 512])
  Layer 5: torch.Size([1, 12, 512])
  Layer 6: torch.Size([1, 12, 512])

Each layer produces a tensor of shape (batch_size, sequence_length, hidden_dim). The sequence has 11 tokens, and each token gets a 512-dimensional representation. Layer 0 is the embedding layer, and layers 1-6 are the transformer blocks.

The progression across layers shows how the model builds meaning. Early layers capture surface-level features: token identity, simple co-occurrence patterns, and basic syntax. Middle layers develop richer syntactic structure, resolving phrase boundaries and grammatical relationships. Late layers encode semantic content most strongly, bringing together context from across the sequence into representations that encode meaning rather than form. This layer-by-layer refinement is not unique to T5 but is a general property of deep transformer stacks; it makes representations from different layers useful for different downstream tasks.

Attention Pattern Visualization

Let's visualize how attention patterns differ between encoder and decoder:

Out[17]:
Visualization
Heatmap of T5 encoder attention weights showing full bidirectional pattern.
Encoder self-attention weights for the first layer and head, showing the bidirectional pattern where every token can attend to every other token. Bright cells indicate high attention weight between those positions.
Heatmap of T5 decoder attention weights showing lower triangular causal pattern.
Decoder self-attention weights for the first layer and head, showing the causal lower-triangular pattern enforced by masking. Each token can only attend to itself and tokens that came before it, preventing information leakage from future positions during generation.

The encoder attention shows bidirectional patterns where tokens can attend freely to any position. The decoder attention shows the characteristic lower-triangular pattern from causal masking. Each token can only attend to itself and previous tokens, preventing information leakage from future positions during generation.

Out[18]:
Visualization
Heatmap of T5 cross-attention weights showing decoder output tokens attending to encoder input tokens, with strong focus on semantically aligned words.
Cross-attention from the decoder to the encoder shows how each output token attends to the input. When generating 'Hallo Welt', the decoder focuses on the relevant English words 'Hello world' while largely ignoring the task prefix 'translate English to German:'. This selective focus is what allows the model to perform accurate translation without being distracted by the instruction tokens.

T5 Variants and Descendants

The T5 architecture inspired several important follow-up models that extended its capabilities in different directions. Together, these variants show both the strengths and the limitations of the original design, and they illustrate how the community builds on foundational work to push state-of-the-art performance.

Flan-T5 applied instruction tuning to T5, training on a diverse mixture of tasks phrased as natural language instructions. This sharply improved zero-shot and few-shot performance, making the model more useful for novel tasks without task-specific fine-tuning. The Flan-T5 paper trained on over 1,800 tasks formatted as instructions, covering a much wider range of task types and phrasings than the original T5 fine-tuning. The result was a model that could follow natural language instructions it had never seen during pretraining, generalizing the text-to-text framework to its logical conclusion. Flan-T5 models are often the practical choice when you want strong zero-shot performance without the computational cost of GPT-4-scale models.

mT5 (multilingual T5) extended pretraining to 101 languages using the mC4 dataset, a multilingual counterpart to C4. This enabled cross-lingual transfer, where a model fine-tuned on English data can perform the same task in other languages. mT5 demonstrated that the T5 architecture scales well to multilingual settings, though performance on lower-resource languages remains substantially below English performance. The architecture is identical to T5; only the pretraining data and vocabulary change. mT5 uses a larger vocabulary (250,000 tokens) to accommodate the diverse character sets and morphological patterns of 101 languages.

LongT5 addressed T5's context length limitations by incorporating efficient attention mechanisms. The standard T5 architecture has quadratic complexity in sequence length due to full self-attention, which makes processing long documents slow and memory-intensive. Using transient global attention patterns, LongT5 handles documents up to 16,384 tokens while maintaining the encoder-decoder structure. This makes it practical for tasks like summarizing long scientific papers or answering questions over book-length documents. The global attention tokens serve as summary representations that allow distant tokens to communicate without requiring full pairwise attention.

UL2 (Unified Language Learner) combined multiple pretraining objectives, including span corruption, prefix language modeling, and causal language modeling. This mixture of denoisers improved performance across diverse downstream tasks by exposing the model to different ways of consuming and generating text during pretraining. UL2 also introduced a conditioning mechanism that tells the model which pretraining mode to use at fine-tuning time, letting practitioners to select the mode best suited for their task. The success of UL2 suggests that T5's single-objective pretraining, while effective, leaves room for improvement by diversifying the pretraining signal.

The broader family of sequence-to-sequence models beyond strict T5 descendants includes BART (which uses a different corruption strategy and denoising objective), mBART (the multilingual extension of BART), and Pegasus (which uses a gap-sentence generation objective specifically designed for summarization). Each of these explores different points in the design space that T5 mapped out, and comparing them illustrates how much the field learned from T5's systematic ablations.

Limitations and Impact

T5's encoder-decoder architecture offers advantages for certain task types but introduces trade-offs compared to decoder-only alternatives. The bidirectional encoder excels when the full input must be processed before generating output. Summarization and translation benefit from understanding the complete context first; question answering does too. However, this architecture requires separate encoder and decoder computations, increasing memory requirements compared to decoder-only models of similar parameter counts. For the same total parameters, a decoder-only model dedicates all capacity to a single transformer stack, while an encoder-decoder splits parameters between two stacks. In practice, for a fixed parameter budget, decoder-only models often outperform encoder-decoder models on generation tasks because they can allocate all capacity to the generative direction.

The text-to-text framework, while elegant, has practical limitations. Classification tasks produce output tokens that must be mapped back to discrete labels, adding a parsing step that can fail if the model generates unexpected text. For high-throughput classification in production, task-specific heads on BERT-style models often prove more efficient. Additionally, regression tasks require outputting numbers as text strings, which is less numerically precise than dedicated regression heads. The model might output "3.7" when the correct answer is "3.71," a rounding error that would never occur with a regression head. Similarly, structured prediction tasks (like named entity recognition with overlapping spans) require serializing complex structures as text, which introduces format learning overhead and potential parsing errors.

T5 also shows the limitations of its training corpus. C4's English-centric, web-derived content means T5 performs best on the kinds of text that appear frequently on the web: news articles, encyclopedia-style text, and conversational English. Tasks requiring specialized domain knowledge (legal, medical, scientific) benefit from domain-adaptive pretraining or fine-tuning on domain-specific data. The filtering pipeline used to create C4, while effective at removing low-quality content, also removed content in non-English languages and may have inadvertently filtered out content from underrepresented communities whose writing style differs from mainstream web English.

Context length is another structural limitation. T5's relative position biases handle sequences of up to 512 tokens naturally, with graceful degradation beyond that. Many real-world tasks, summarizing a research paper, answering questions about a book chapter, processing long legal documents, exceed this limit. LongT5 addresses this directly, but the standard T5 architecture requires chunking or truncation strategies that can disrupt cross-chunk context.

T5 affected the field in several ways. The systematic ablation study in the original paper influenced many subsequent design decisions. Researchers could consult T5's experiments rather than re-running their own, with justified confidence that the findings generalized beyond the specific models tested. The text-to-text framework demonstrated that unified architectures could match or exceed task-specific approaches, paving the way for general-purpose instruction-following models. T5's pretraining recipe, combining span corruption with large-scale data, informed the development of models like PaLM and Flan-PaLM. Its relative position biases appeared in subsequent models from multiple research groups, often cited as superior to absolute positions for general-purpose language understanding.

T5 also established that encoder-decoder architectures remained competitive even as decoder-only models (the GPT family) gained prominence. This architectural diversity has proven useful. Encoder-decoder models continue to excel at translation and summarization, while decoder-only models dominate open-ended generation. Understanding both paradigms remains needed for practitioners choosing the right architecture for their specific application.

Summary

T5 unified NLP around a simple principle: treat every task as text-to-text generation. This framework eliminated the need for task-specific architectures, letting a single model handle these tasks through the same interface. The text-to-text design also anticipated instruction-following models: once you commit to text as the universal interface, adding new tasks requires nothing more than formatting examples with the right prefix and target.

The architecture builds on the original Transformer's encoder-decoder design with key modifications. Pre-norm layer placement improves training stability. Relative position biases replace absolute position embeddings, encoding distances between tokens rather than their absolute locations. The bucketized position scheme keeps parameters bounded while capturing both fine-grained local and coarser global position information. These design choices were not arbitrary; each was validated by the systematic ablation study in the T5 paper, making T5 one of the best-evidenced architectural designs in the transformer era.

T5's five model sizes span from 60 million to 11 billion parameters, with systematic scaling that increases depth for smaller models and width for larger ones. The span corruption pretraining objective, replacing contiguous token spans with sentinels, proved more effective than alternatives like standard language modeling or masked language modeling. The C4 dataset, a 750GB filtered web corpus, provided the scale needed to study how pretraining data quantity and quality affect downstream performance.

The encoder-decoder structure particularly suits tasks requiring deep comprehension before generation. The encoder processes input bidirectionally, and the decoder generates output autoregressively while attending to encoded representations via cross-attention. This two-stage approach excels at translation and summarization, where understanding the full source is needed before creating the target. The cross-attention heatmaps we examined visually confirm this: the decoder learns to focus on semantically relevant parts of the input at each generation step.

T5's influence extended beyond its direct applications. The systematic ablation study guided subsequent architectural decisions. The text-to-text framework inspired instruction-tuning approaches that became central to modern language models. Variants like Flan-T5, mT5, and LongT5 extended its capabilities to instruction following, multilingual processing, and long-context understanding. By showing that encoder-decoder models could compete with and sometimes exceed specialized architectures, T5 ensured this architectural family remained part of the practitioner's toolkit. As you work with modern language models, the design choices T5 explored and validated will appear repeatedly: in how models handle position information, how pretraining objectives are formulated, and how unified frameworks replace task-specific engineering.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about T5's architecture and design principles.

T5 Architecture

Question 1 of 80 of 8 completed
What is the core insight behind T5's text-to-text framework?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025t5architecture, author = {Michael Brenndoerfer}, title = {T5 Architecture and Text-to-Text Transfer Learning}, year = {2025}, url = {https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). T5 Architecture and Text-to-Text Transfer Learning. Retrieved from https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer
MLAAcademic
Michael Brenndoerfer. "T5 Architecture and Text-to-Text Transfer Learning." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer>.
CHICAGOAcademic
Michael Brenndoerfer. "T5 Architecture and Text-to-Text Transfer Learning." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer.
HARVARDAcademic
Michael Brenndoerfer (2025) 'T5 Architecture and Text-to-Text Transfer Learning'. Available at: https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). T5 Architecture and Text-to-Text Transfer Learning. https://mbrenndoerfer.com/writing/t5-architecture-text-to-text-transformer

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.