T5 Task Formatting: Text-to-Text NLP Unification

Michael BrenndoerferOctober 15, 202557 min read

Part of Language AI Handbook

Explains how T5 reformulates all NLP tasks as text-to-text problems. Topics include task prefixes, classification, NER.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

T5 Task Formatting: Text-to-Text NLP

In the previous chapters, we explored T5's encoder-decoder architecture and its span corruption pre-training objective. We saw how T5 learns rich contextual representations by encoding corrupted text and then decoding the missing spans. But what makes T5 truly distinctive is not just its architecture or its pre-training recipe. The key insight is that every NLP task can be reformulated as a text-to-text problem. Classification, translation, summarization, question answering, and even structured prediction tasks like named entity recognition can all be expressed as "given this input text, produce this output text." This unification is not a minor implementation convenience. It is a philosophical reframing of what natural language processing is, collapsing a diverse zoo of specialized methods into a single, coherent framework governed by one principle: language tasks are transformations from one text string to another.

Think of the text-to-text paradigm as a universal adapter. Before it existed, every NLP task required its own connector, its own special plug that linked input data to output predictions. A sentiment classifier needed a classification head. A sequence tagger needed a CRF layer. A translator needed an encoder-decoder with attention over source tokens. Each plug had a different shape, so you could not simply swap tasks without swapping the connector. The text-to-text paradigm replaces all these specialized connectors with a single universal interface: "given a string, generate a string." Every task plugs into the same socket, and the model's job is simply to learn which kind of transformation each task requires.

This chapter explores how T5 reaches this unification through clever task formatting. We will examine how task prefixes condition the model's behavior, how traditional NLP tasks are reformulated as text generation, and what this means for building versatile language models that can handle diverse workloads from a single checkpoint. Along the way, we will develop working code to format inputs, generate outputs, and parse the results back into structured data that applications can consume. We will also look at the engineering challenges this paradigm introduces and how to address them in practice.

The implications extend well beyond T5 itself. The text-to-text framing anticipates the instruction-following paradigm that would come to define the next generation of large language models. When you see a model like GPT-4 following complex natural language instructions, you are seeing the logical continuation of the insight T5 demonstrated: language is both what models process and how they are directed. This chapter uses T5 to explain the conceptual foundation on which modern generative AI is built.

Historical Context

The text-to-text paradigm was introduced in Google's 2019 paper "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer" by Colin Raffel and colleagues. The paper was specific for introducing T5 and for its systematic empirical approach: the authors conducted hundreds of experiments comparing model architectures, pre-training objectives, training data scales, and task formatting strategies. The paper's contribution was therefore both a specific model and a methodology for rigorously evaluating design choices in transfer learning. T5 was pre-trained on the Colossal Clean Crawled Corpus (C4), a cleaned version of Common Crawl roughly 750 GB in size, and trained at scales from 60 million to 11 billion parameters. The text-to-text framing directly influenced subsequent work on instruction tuning, FLAN, T0, and ultimately the instruction-following paradigm that defines modern large language model deployment.

The Text-to-Text Paradigm

Traditional NLP systems treat different tasks as fundamentally different problems, each requiring its own architecture, training procedure, and inference logic. A classifier predicts discrete labels using a softmax output layer. A sequence tagger assigns one label per input token, often using a CRF to capture label dependencies. A machine translation system uses an encoder-decoder with cross-attention to align source and target tokens. A summarizer generates text autoregressively conditioned on a long document. Each of these task types has its own form of output, its own loss function, and often its own architecture. The practical consequence was that NLP researchers and engineers maintained separate systems for separate tasks, with each system requiring its own expertise to build and maintain.

T5 rejects this fragmentation entirely by recognizing a simple fact: text is a universal output format. Any output a model might need to produce, whether a single discrete label, a sequence of tags, a numerical score, a translated sentence, or a structured annotation, can be serialized as a string. The output "positive" is text. The output "B-PER I-PER O B-LOC O" is text. The translated sentence "Das Wetter ist schön" is text. If every output can be expressed as text, and every input is already text, then every NLP task becomes a sequence-to-sequence problem: given an input string, generate an output string. A single model with a single architecture, a single loss function, and a single training procedure can handle them all.

To appreciate why this matters, consider the practical situation before unified frameworks like T5. A practitioner building a multi-task NLP system would need a BERT-based classifier with a task-specific classification head for sentiment analysis, a separate BiLSTM-CRF for named entity recognition, a transformer encoder-decoder for machine translation, and yet another model for abstractive summarization. Each model requires its own training pipeline, its own hyperparameter tuning cycle, and its own deployment infrastructure. Maintaining four separate models is four times the engineering work, four times the failure points, and four times the serving cost. The text-to-text paradigm collapses this complexity into a single model that learns to perform all tasks through the same mechanism: autoregressive text generation.

Beyond the practical benefits, this unification reveals something philosophically interesting about the nature of language tasks. Generation is a general form of prediction. When a model generates the word "positive" in response to a movie review, it is making a prediction equivalent to what a classifier makes by selecting class 1. When it generates "B-PER" in response to the name "Marie Curie," it is making a labeling decision equivalent to what a sequence tagger makes token by token. The text-to-text framing is a software engineering convenience with a broader consequence. It suggests that the boundary between "understanding" and "generation" is more fluid than traditional architectures imply. A model that can generate the correct answer to a question shows understanding of the question and the relevant context. Generation becomes the unified test of linguistic competence.

Out[3]:
Visualization
Diagram showing traditional NLP with separate specialized architectures for classification, tagging, translation, and summarization tasks.
Traditional NLP requires a separate specialized architecture for each task type, with task-specific output layers, loss functions, and inference procedures.
Diagram showing T5's unified text-to-text framework where a single encoder-decoder model handles all NLP tasks by converting inputs and outputs to text strings.
T5's text-to-text approach routes all tasks through a single encoder-decoder model, converting each task's inputs and outputs into text strings for uniform processing.
Text-to-Text Transfer

The text-to-text approach treats every NLP task as translating from one text string to another. The model learns a unified mapping from input sequences to output sequences, regardless of whether the underlying task is classification, extraction, or generation. This shared format supports shared pre-training, multitask learning, and zero-shot generalization across tasks.

This unification brings several concrete benefits that practitioners can take advantage of immediately:

  • Single architecture: The same encoder-decoder model handles all tasks without modification to its structure
  • Shared pre-training: Knowledge learned during pre-training on large text corpora transfers to any downstream task, regardless of what that task looks like
  • Multitask learning: Multiple tasks can be mixed in a single training batch, letting positive transfer between related skills
  • Zero-shot generalization: The model can attempt new tasks if given appropriate formatting, because it has learned to follow natural language instructions embedded in the input

Beyond these practical benefits, this approach suggests that the boundary between "understanding" and "generation" may be more fluid than traditional NLP architectures imply. A model that can generate the correct answer to a question shows understanding of both the question and the relevant context. A model that generates appropriate entity labels has learned to recognize those entities. Generation becomes the unified test of linguistic competence.

Task Prefixes

How does T5 know which task to perform when given an input? The answer is task prefixes, short text strings prepended to the input that signal the desired operation. This mechanism is remarkably simple yet surprisingly powerful, in effect teaching the model to follow instructions embedded in natural language. The prefix acts as a routing instruction, telling the model how to process the input and what kind of output to generate.

Consider these examples:

  • "translate English to German: The house is wonderful."
  • "summarize: The stock market fell sharply today after..."
  • "question: What is the capital of France? context: Paris is the capital..."

In each case, the colon acts as a natural delimiter between the instruction and the content. The model sees the prefix first, uses it to establish the task context, and then processes the remaining text accordingly. When the model sees "translate English to German:", it knows to treat the following text as English source material and produce German output. When it sees "summarize:", it understands that it should produce a condensed version of the following content. This is conceptually similar to how special tokens work in architectures we discussed in earlier chapters, but instead of learned embeddings with special meanings, T5 uses natural language instructions that the model learns to interpret through training.

This approach exploits the model's core competency in understanding language to specify its behavior. Rather than introducing specialized tokens or architectural modifications for each task, T5 uses the same linguistic representations it applies elsewhere. The model learns that certain word patterns at the start of the input correlate with certain expected output patterns, much as it learns any other linguistic regularity. Think of it as teaching the model a meta-skill: the ability to read a natural language description of a task and then perform that task. This meta-skill is exactly what instruction tuning in later models formalizes and scales up.

The key insight is that natural language prefixes work because they are convenient and real. T5 was pre-trained on large amounts of text that includes imperative sentences and task descriptions. The model has already encountered phrases like "translate this sentence" and "summarize the following" in its pre-training data. These phrases carry semantic content the model understands. When a prefix like "translate English to German:" appears, the model activates linguistic representations associated with translation, source and target languages, and cross-lingual equivalence. This semantic activation is what makes natural language prefixes more effective than arbitrary codes or symbols.

In[4]:
Code
import warnings

from transformers import T5ForConditionalGeneration, T5Tokenizer

warnings.filterwarnings("ignore")

# Load T5 model and tokenizer
model_name = "t5-small"
tokenizer = T5Tokenizer.from_pretrained(model_name, legacy=False)
model = T5ForConditionalGeneration.from_pretrained(model_name)
In[5]:
Code
# Different tasks with different prefixes
tasks = [
    "translate English to German: The weather is nice today.",
    "translate English to French: Hello, how are you?",
    "summarize: Machine learning is a subset of artificial intelligence that lets computers to learn from data without being explicitly programmed. It has applications in image recognition, natural language processing, and many other fields.",
]
In[6]:
Code
# Generate outputs for each task
results = []
for task_input in tasks:
    inputs = tokenizer(
        task_input, return_tensors="pt", max_length=512, truncation=True
    )
    outputs = model.generate(**inputs, max_new_tokens=64)
    decoded = tokenizer.decode(outputs[0], skip_special_tokens=True)
    results.append((task_input[:50] + "...", decoded))
Out[7]:
Console
Input: translate English to German: The weather is nice t...
Output: Das Wetter ist heute schön.
------------------------------------------------------------
Input: translate English to French: Hello, how are you?...
Output: Bonjour, comment êtes-vous?
------------------------------------------------------------
Input: summarize: Machine learning is a subset of artific...
Output: machine learning is a subset of artificial intelligence that lets computers learn from data without being explicitly programmed. it has applications in image recognition, natural language processing, and many other fields.
------------------------------------------------------------

The model produces task-appropriate outputs based solely on the prefix: German translation for the first, French for the second, and a summary for the third. The prefix creates an implicit conditioning that shapes the entire generation process. Note that T5-small is a relatively small model, so translation quality may be limited compared to larger variants. The key insight is that the same architecture handles fundamentally different tasks through prefix-based routing.

Prefix Design Principles

The choice of prefix matters more than one might initially expect. T5's authors experimented with different prefix styles and found that clarity and consistency are more important than brevity. The prefix is a task specification, not a mere tag, that the model must parse and act upon. Ambiguous or inconsistent prefixes lead to ambiguous or inconsistent outputs.

Effective prefixes share several characteristics. First, they use explicit task naming: words like "translate", "summarize", and "classify" clearly indicate the operation. Second, they include relevant parameters when needed: "English to German" specifies source and target languages, and these language names carry meaning the model understands from its pre-training corpus. Third, they maintain consistent formatting by using the same pattern "task: input" across all instances. Fourth, they read as natural language instructions that a human would understand.

Natural language prefixes work well because T5 has strong priors about how language works. A prefix like "translate English to German:" uses the model's understanding of what translation means, what English and German are, and how to interpret the colon as a delimiter. This is why natural language prefixes often outperform arbitrary codes or symbols. They activate knowledge the model already has. A prefix like "X12: The house is wonderful." carries no semantic content for the model to use. It would need to learn the mapping from scratch, requiring more task-specific training data. In contrast, a human-readable prefix bootstraps the model's existing language understanding.

Here is how different prefix choices affect the same underlying task:

In[8]:
Code
# Same task, different prefix formulations
prefix_variations = [
    "cola sentence: The cat sat on the mat.",  # T5's actual format for CoLA
    "Is this sentence grammatically correct? The cat sat on the mat.",
    "grammar check: The cat sat on the mat.",
]

The comments highlight an important point. T5 was trained with specific prefixes like "cola sentence:", so using different formulations at inference time may produce unexpected results unless the model has been fine-tuned on the new format.

The specific prefixes used during training become part of the model's learned behavior. T5's original training used prefixes like "cola sentence:" for the CoLA grammaticality task, "sst2 sentence:" for sentiment analysis, and "translate English to German:" for translation. Using different prefixes at inference time may produce unexpected results unless the model has been fine-tuned on the new format. This highlights an important principle: the text-to-text paradigm is flexible but not magical. The model can only reliably perform tasks it has been trained to recognize through their prefixes. For novel tasks or custom prefix formats, fine-tuning on appropriately formatted data is necessary.

There is a deeper point here about the relationship between prefix design and model capability. When you create a new prefix format for a fine-tuning task, you are making a choice about what the model needs to learn. A very explicit prefix like "classify the sentiment of this product review as positive, negative, or neutral:" contains more information than "classify:", but it also means more of the capacity is used on parsing the instruction rather than reasoning about the content. The optimal balance depends on how much data you have for fine-tuning and how diverse the tasks are. For production systems with abundant fine-tuning data, concise and consistent prefixes are usually better.

Classification as Generation

Converting classification to generation might seem wasteful. Why generate text when you only need a label? This question highlights the efficiency versus flexibility trade-off that defines the text-to-text paradigm. But the benefits of unification generally outweigh the slight overhead, and the approach proves surprisingly powerful in practice.

To understand why, consider what classification entails. A classifier takes an input and selects one of several possible outputs from a fixed set. In traditional systems, this selection happens through a softmax over a fixed set of class vectors. In T5, the selection happens through generation. The "selection" is the model's choice of which label text to produce. The mechanism differs, but the underlying computation is conceptually similar: the model must encode the input, reason about its semantic content, and produce an appropriate response. What changes is not the reasoning required, but the form of the output.

For binary classification, the model simply generates a label token. The label text itself can encode semantic meaning. Using descriptive labels like "positive" and "negative" rather than arbitrary integers like "0" and "1" means the model can use its understanding of what these words mean. A model that understands English knows that "positive" is associated with favorable evaluations and "negative" with unfavorable ones. This semantic grounding potentially allows the model to classify correctly even for inputs unlike anything it saw during fine-tuning.

In[9]:
Code
# Sentiment classification as text generation
sentiment_examples = [
    ("sst2 sentence: This movie was absolutely fantastic!", "positive"),
    ("sst2 sentence: What a waste of time and money.", "negative"),
    ("sst2 sentence: The acting was superb but the plot was confusing.", "?"),
]
Out[10]:
Console
Input:    sst2 sentence: This movie was absolutely fantastic!
Expected: positive

Input:    sst2 sentence: What a waste of time and money.
Expected: negative

Input:    sst2 sentence: The acting was superb but the plot was confusing.
Expected: ?

These examples show how sentiment classification maps to text generation. The model learns to produce "positive" or "negative" based on the input text. The third example with "?" as expected output illustrates an ambiguous case where mixed sentiment makes the correct label unclear. Real training data would need a consistent strategy for handling such cases, whether by requiring the model to pick one label, generate a fine-grained response, or output a special token showing ambiguity.

During training, the model learns to generate the appropriate label text. For SST-2 (Stanford Sentiment Treebank), T5 was trained to output "positive" or "negative". For CoLA (grammaticality), it outputs "acceptable" or "unacceptable". The choice of label text is arbitrary in principle: the model could learn to output "1" for positive and "0" for negative. But descriptive labels let the model use its semantic understanding and tend to work better in practice, especially with limited fine-tuning data, because they are semantically coherent rather than arbitrary symbols.

Multi-class and Multi-label Classification

The text-to-text format naturally extends to more complex classification scenarios without requiring any architectural changes. This is where the paradigm's flexibility shines most clearly. A traditional multi-class classifier needs its output layer sized to the number of classes. Adding a new class means modifying the architecture and retraining. In T5, adding a new class just means introducing a new label string in the training data. The model learns to generate the new label through the same training procedure it already uses. No architectural changes are needed.

In[11]:
Code
# Multi-class: News topic classification
multiclass_examples = [
    (
        "classify topic: Apple announces new iPhone with revolutionary camera.",
        "technology",
    ),
    ("classify topic: Lakers defeat Warriors in overtime thriller.", "sports"),
    (
        "classify topic: Federal Reserve raises interest rates by 0.25%.",
        "business",
    ),
]

# Multi-label: Multiple applicable tags
multilabel_examples = [
    (
        "tags: A new AI chip powers the latest smartphone.",
        "technology, electronics",
    ),
    (
        "tags: Tech stocks surge after earnings reports.",
        "technology, business, finance",
    ),
]

For multi-label classification, the model generates multiple labels separated by a delimiter. This approach is more flexible than traditional multi-hot encoding because the model can generate any number of labels and even combinations of labels it has not explicitly seen during training. The model learns which labels to generate, the delimiter pattern, and the approximate number of labels appropriate for different inputs. A document about an AI startup going public might generate "technology, business, finance," combining labels it has seen in combination before.

The multi-label case also illustrates a subtle but important point about output ordering. Traditional multi-hot vectors have no ordering: the presence or absence of each class is independent. In the generative approach, labels must be listed in some order. The model will learn to follow the ordering pattern in its training data, which could be alphabetical, by salience, or by the order labels appear in the source text. You should choose an ordering convention and apply it consistently in your training data, because inconsistent ordering can degrade performance by requiring the model to learn multiple valid output formats for the same conceptual answer.

Extracting Probabilities

One apparent limitation of generation-based classification is losing access to prediction probabilities. Traditional classifiers output a probability distribution over labels. This supports calibration and thresholding while also providing an estimate of uncertainty. If a traditional classifier is 51% confident about "positive" and 49% about "negative", you know the prediction is uncertain and might want to abstain or route to a human reviewer. With greedy generation, you just get "positive" with no indication of the model's uncertainty.

T5 solves this by giving access to the decoder's output logits. Since generation proceeds token by token, and each token is selected based on a probability distribution over the vocabulary, we can recover probability information by examining these distributions directly. The key insight is that the generated text is only one sample path; the probability distribution behind it contains rich information about the model's confidence and the relative likelihood of alternative outputs.

In[12]:
Code
import torch


def get_label_probability(model, tokenizer, input_text, labels):
    """Get probability of each label for a classification task."""
    # Tokenize input
    inputs = tokenizer(input_text, return_tensors="pt")

    # Get logits for each label
    label_probs = {}
    for label in labels:
        label_ids = tokenizer(label, return_tensors="pt").input_ids

        # Forward pass to get logits
        with torch.no_grad():
            outputs = model(**inputs, decoder_input_ids=label_ids[:, :-1])
            logits = outputs.logits

        # Calculate probability of generating this label
        # (simplified - real implementation handles multiple tokens)
        probs = torch.softmax(logits[0, -1, :], dim=-1)
        # Handle case where label tokenizes to single token (index 0 after BOS)
        target_idx = label_ids[0, 0].item()
        label_probs[label] = probs[target_idx].item()

    return label_probs

This approach scores how likely the model is to generate each candidate label given the input. By normalizing these scores across the set of valid labels, we recover a probability distribution equivalent to what a traditional classifier would output. The approach requires knowing the set of valid labels in advance, which is always the case for classification tasks. For tasks with open-ended output spaces, like generation or translation, this probability recovery technique does not apply directly, but for classification it works well.

Out[13]:
Visualization
Bar chart showing a traditional classifier outputting a softmax probability distribution over fixed class labels such as positive, negative, and neutral.
Traditional classifiers output a softmax probability distribution over a fixed set of class labels, giving direct uncertainty estimates for all possible classes simultaneously.
Bar chart showing T5's generative approach where token probabilities from the vocabulary distribution are used to produce label text.
T5's generative approach assigns probabilities from the full vocabulary distribution, letting label probabilities to be recovered by examining the likelihood of generating each candidate label string.

Named Entity Recognition as Generation

Named entity recognition, as we covered in earlier chapters on sequence labeling, traditionally uses BIO tagging to assign labels to each token in sequence. This approach is elegant for sequence labeling architectures that produce one output per input position. The model reads the input left to right and, at each token, decides its label based on context and the previous label. However, this approach does not naturally fit the text-to-text paradigm where input and output lengths can differ arbitrarily. Converting NER to text-to-text requires reformulating the task: instead of predicting per-token labels, the model generates the entities directly.

This reformulation changes what the model needs to learn. Instead of learning to classify each token independently with CRF dependencies capturing label transitions, the model learns to read the input text, identify entity spans, and express those findings in a structured output format. This is arguably closer to how humans perform entity recognition. We do not mentally label each word as B-PER or I-PER in sequence. We read the sentence as a whole and notice that "Marie Curie" is a person's name, that "Paris" is a location, and that "Nobel Prize" is an award. The generative approach allows the model to discover entities through global comprehension rather than local token classification.

The generative approach also handles a case that BIO tagging handles awkwardly: entities that span discontinuous text, or situations where one entity type overlaps with another in nested structures. In BIO tagging, the standard approach requires either flattening these relationships or using multiple passes. In the generative approach, nested entities can be listed explicitly in the output without any special treatment. The model simply generates the nested structure as part of its output string.

There are two main approaches to formatting NER as text generation, and each makes different tradeoffs between parsability, information preservation, and generation complexity.

Approach 1: Entity Listing

The model generates a structured list of entities and their types. This format separates entity recognition from positional information, focusing purely on what entities exist and what types they belong to. Think of it as asking the model: "What entities are mentioned in this text, and what are they?"

In[14]:
Code
# NER as entity extraction
ner_examples = [
    {
        "input": "ner: Barack Obama was born in Honolulu, Hawaii.",
        "output": "person: Barack Obama, location: Honolulu, location: Hawaii",
    },
    {
        "input": "ner: Apple Inc. announced the iPhone 15 in Cupertino.",
        "output": "organization: Apple Inc., product: iPhone 15, location: Cupertino",
    },
    {
        "input": "ner: The patient was prescribed 500mg of ibuprofen.",
        "output": "dosage: 500mg, medication: ibuprofen",
    },
]
Out[15]:
Console
Input:  ner: Barack Obama was born in Honolulu, Hawaii.
Output: person: Barack Obama, location: Honolulu, location: Hawaii

Input:  ner: Apple Inc. announced the iPhone 15 in Cupertino.
Output: organization: Apple Inc., product: iPhone 15, location: Cupertino

Input:  ner: The patient was prescribed 500mg of ibuprofen.
Output: dosage: 500mg, medication: ibuprofen

This format is intuitive and handles overlapping entities naturally, since each entity is listed separately regardless of its position in the source text. The model learns a consistent pattern: type followed by a colon, then entity text, with entities separated by commas. This regularity helps the model generalize to new entity types and new texts, because the structural pattern is consistent across all examples. Adding a new entity type like "medication" or "product" requires only that the training data include examples with that label. No architectural changes are needed, and the model naturally learns to use the new label in the same structural format as all the others.

Approach 2: Inline Markup

Alternatively, entities can be marked inline with special delimiters. This approach preserves the original text structure while adding entity annotations as in-place markup. The output reads like the input text but with entity spans wrapped in type tags.

In[16]:
Code
# NER with inline markup
inline_ner_examples = [
    {
        "input": "extract entities: Barack Obama was born in Honolulu, Hawaii.",
        "output": "[PER Barack Obama] was born in [LOC Honolulu], [LOC Hawaii].",
    },
    {
        "input": "extract entities: Microsoft acquired GitHub for $7.5 billion.",
        "output": "[ORG Microsoft] acquired [ORG GitHub] for [MONEY $7.5 billion].",
    },
]

This approach preserves entity positions in the original text and makes nested entities explicit, but requires careful parsing of the output. The inline format has the advantage of maintaining context around entities, which can help with certain downstream applications where the position or surrounding context matters. For example, relation extraction systems that need to know how two entities are related in the sentence can benefit from the inline representation, which makes the structural relationship visible. However, this approach requires the model to regenerate the entire input with markup added, which is more computationally expensive and introduces more opportunities for the model to make mistakes in recreating the original text.

Out[17]:
Visualization
Side-by-side comparison of two NER output formats: entity listing on the left shows extracted entities as type-value pairs, and inline markup on the right shows the original text annotated with entity tags around spans.
Two approaches to NER as text generation. Entity listing extracts entities as type-value pairs and is easy to parse but loses positional information. Inline markup annotates entities within the original text structure, preserving position at the cost of more complex output generation.

Handling Edge Cases

The generative approach to NER introduces challenges that traditional sequence labeling does not face. Because the model generates entities as discrete items rather than labeling positions, it must handle various edge cases that require explicit decisions in how you format your training data.

In[18]:
Code
# Edge cases in generative NER
edge_cases = [
    # Entity spans multiple mentions of same text
    {
        "input": "ner: New York hosted the New York Marathon.",
        "output": "location: New York, event: New York Marathon",
    },
    # Nested entities
    {
        "input": "ner: University of California, Berkeley researchers published...",
        "output": "organization: University of California, Berkeley, location: California, location: Berkeley",
    },
    # No entities
    {
        "input": "ner: It was a beautiful day.",
        "output": "",  # Or "none" depending on format
    },
]

Training data must cover these edge cases consistently for the model to handle them reliably. The format chosen during training becomes necessary. Mixing formats leads to inconsistent outputs. If some training examples list nested entities and others do not, the model will not learn a consistent policy for handling nesting. Practitioners must make deliberate choices about how to handle these edge cases and apply those choices uniformly across training data.

The case of repeated entity surface forms deserves special attention. In the sentence "New York hosted the New York Marathon," both "New York" (as a location) and "New York Marathon" (as an event) contain the substring "New York." A BIO tagger handles this straightforwardly by labeling each token position. The generative approach must decide how to stand for the relationship between the overlapping spans in its output. There is no universally correct answer. The right approach depends on your downstream application and what information you need to extract. The necessary principle is consistency: whatever convention you choose, apply it uniformly in your training data.

Question Answering as Generation

Question answering fits the text-to-text paradigm naturally, arguably more naturally than any other classical NLP task. The core operation involves reading a question, consulting some source of knowledge (either a provided context or the model's internal parametric memory), and creating an answer as text. This is inherently about transforming one text into another. Given a question and optional context, generate the answer.

Question answering has historically been divided into several subtypes that previously required different architectures. Span extraction QA, as in the SQuAD benchmark, required models that could identify the start and end positions of the answer span in the provided context. This produced architectures with two classification heads over the encoder output: one for the start position, one for the end. Multi-hop QA required architectures that could reason across multiple documents or passages. Knowledge base QA required linking natural language questions to structured database queries. The text-to-text paradigm subsumes all of these under a single formulation.

In[19]:
Code
# Extractive QA: answer is a span from the context
extractive_qa = {
    "input": "question: What is the capital of France? context: Paris is the capital and largest city of France. It is located on the Seine River.",
    "output": "Paris",
}

# Multi-hop QA: requires reasoning across facts
multihop_qa = {
    "input": "question: Who founded the company that makes the iPhone? context: Apple Inc. manufactures the iPhone. Apple was founded by Steve Jobs, Steve Wozniak, and Ronald Wayne.",
    "output": "Steve Jobs, Steve Wozniak, and Ronald Wayne",
}
Out[20]:
Console
Extractive QA Input: question: What is the capital of France? context: Paris is t...
Extractive QA Output: Paris

Multi-hop QA Input: question: Who founded the company that makes the iPhone? con...
Multi-hop QA Output: Steve Jobs, Steve Wozniak, and Ronald Wayne

The extractive example shows a simple fact lookup where the answer appears directly in the context. The multi-hop example requires combining information from multiple sentences, identifying that Apple makes the iPhone, then finding who founded Apple. This chaining of facts is a form of reasoning that the model must perform through its sequence-to-sequence generation process. The T5 encoder must capture both facts from the context and their relationship, and the decoder must synthesize them into a coherent answer.

In[21]:
Code
# Let's see T5 handle a question
qa_input = "question: What is the capital of Germany? context: Berlin is the capital and largest city of Germany. Munich is the third largest city."
inputs = tokenizer(
    qa_input, return_tensors="pt", max_length=512, truncation=True
)
outputs = model.generate(**inputs, max_new_tokens=32)
answer = tokenizer.decode(outputs[0], skip_special_tokens=True)
Out[22]:
Console
Question: What is the capital of Germany?
Answer:   Berlin

T5 correctly extracts "Berlin" from the provided context. This shows the open-book QA pattern where the model locates and returns the relevant information from the given passage rather than relying on memorized knowledge. The answer is grounded in the context, which reduces the risk of factual hallucination that plagues purely parametric generation.

Open-book vs Closed-book QA

T5's text-to-text format supports a distinction that reveals something important about what large language models learn during pre-training. The format supports both open-book QA, where context is provided, and closed-book QA, where the model must rely entirely on knowledge encoded in its parameters. This distinction matters both for understanding model capabilities and for designing production systems.

In[23]:
Code
# Open-book QA: context provided
open_book = "question: When was Python created? context: Python was created by Guido van Rossum and first released in 1991."

# Closed-book QA: no context, rely on parametric knowledge
closed_book = "question: When was Python created?"

# Both use the same format, but closed-book relies on knowledge stored in model weights

Open-book QA gives the answer source explicitly. The model's job is comprehension and extraction: finding the relevant information in the provided text. This is a relatively safe form of QA because the answer is grounded in a verifiable source. If the context says "1991," the model should say "1991," and you can check its work.

Closed-book QA tests whether the model memorized relevant facts during pre-training. T5 demonstrated surprisingly strong closed-book QA performance, suggesting that large-scale pre-training on diverse text encodes substantial world knowledge. This finding was influential in showing that large language models do not just learn linguistic patterns. They absorb and can retrieve factual information. A model trained on enough text will have encountered facts about Python's creation date, Nobel Prize winners, historical events, and scientific discoveries, and can reproduce that information on demand.

The closed-book capability is impressive but comes with significant caveats. Parametric knowledge is static: the model cannot update what it learned at training time. If Python released a new major version after T5's training cutoff, closed-book T5 would not know about it. Knowledge is also imprecise: the model might "remember" a fact but confabulate details around it, creating plausible-sounding but incorrect answers. These limitations of parametric memory motivated the development of retrieval-augmented generation approaches, where models are equipped with access to external knowledge sources and use open-book QA formulations to answer questions grounded in retrieved passages. We will explore these approaches in later chapters.

Abstractive vs Extractive Answers

Unlike span-based extractive QA models that can only copy text from the context, T5 can generate abstractive answers by combining information from the source and drawing conclusions that do not appear verbatim in it. This changes what question answering can accomplish. Traditional extractive models are limited to answers that appear verbatim in the source text. T5 can produce answers that the source text implies but does not explicitly state.

In[24]:
Code
# Extractive: answer is verbatim from context
extractive_example = {
    "context": "The Eiffel Tower is 330 meters tall.",
    "question": "How tall is the Eiffel Tower?",
    "extractive_answer": "330 meters tall",
    "abstractive_answer": "It stands at 330 meters.",
}

# T5 might paraphrase or synthesize information
synthesis_example = {
    "context": "Marie Curie won the Nobel Prize in Physics in 1903. She also won the Nobel Prize in Chemistry in 1911.",
    "question": "How many Nobel Prizes did Marie Curie win?",
    "answer": "two",  # Not a direct span!
}

The synthesis example shows this powerfully. The word "two" does not appear anywhere in the context, yet it is the correct answer. The model must count the Nobel Prizes mentioned and express that count as a word. This kind of reasoning and synthesis goes beyond pattern matching. It requires comprehension and inference.

This abstractive capability is powerful but requires careful training data design. If training answers are always extracted verbatim from context, the model will learn to copy rather than synthesize. If training answers are always abstractive paraphrases, the model may introduce unnecessary rephrasing that deviates from the source material. The nature of the training data shapes the model's generation behavior in subtle but important ways. For most production QA systems, a mix of extractive and abstractive training examples produces models that can handle both types of questions gracefully.

Additional Task Formatting Examples

The text-to-text paradigm extends far beyond classification and NER; it also handles QA. Its power is its universality. Essentially any NLP task that can be described as "given this text, produce that text" fits naturally into the framework. Let us examine formatting for additional task types, both to show the breadth of the paradigm and to give you concrete patterns for your own tasks.

Semantic Similarity

Sentence similarity tasks ask whether two sentences mean the same thing. This is a basic capability underlying many applications, from duplicate question detection to paraphrase identification to information retrieval. The text-to-text formulation is straightforward: given two sentences as input, output a label or score.

In[25]:
Code
similarity_examples = [
    {
        "input": "mrpc sentence1: The company posted revenue of $1.2 billion. sentence2: Revenue reached $1.2 billion.",
        "output": "equivalent",
    },
    {
        "input": "mrpc sentence1: The stock fell 5%. sentence2: The stock rose 5%.",
        "output": "not_equivalent",
    },
]

For graded similarity on benchmarks like STS-B, the output is a numerical score. This shows how the text-to-text format handles continuous outputs by simply generating the number as text. The model learns to map pairs of sentences to scores in the range 0 to 5, where 0 indicates completely unrelated sentences and 5 indicates semantically equivalent sentences.

In[26]:
Code
graded_similarity = {
    "input": "stsb sentence1: A man is playing guitar. sentence2: A person is playing a musical instrument.",
    "output": "4.2",  # Score from 0-5
}

Generating numerical scores as text is a subtle but important capability. The model must distinguish related from unrelated pairs and calibrate its outputs to match the scale of the training labels. This requires understanding that "4.2" is a higher similarity than "2.1," which requires some arithmetic reasoning within a language generation framework. Research has shown that T5 handles this reasonably well, partly because numbers appear extensively in pre-training text and the model learns their relative magnitudes.

Natural Language Inference

NLI determines whether a hypothesis follows logically from a premise. This task tests logical reasoning and semantic understanding, asking the model to evaluate three possible relationships: the premise entails the hypothesis, the hypothesis contradicts the premise, or neither relationship holds (neutral).

In[27]:
Code
nli_examples = [
    {
        "input": "mnli hypothesis: The animal is sleeping. premise: The cat is curled up on the couch with its eyes closed.",
        "output": "entailment",
    },
    {
        "input": "mnli hypothesis: It is raining. premise: People are carrying umbrellas.",
        "output": "neutral",  # Possible but not certain
    },
    {
        "input": "mnli hypothesis: The restaurant is empty. premise: The restaurant is crowded with diners.",
        "output": "contradiction",
    },
]

NLI is particularly interesting as a text-to-text task because the required reasoning is not about language patterns but about logical relationships between statements. The entailment example requires recognizing that "curled up with eyes closed" implies sleeping. The neutral example requires recognizing that umbrella use is consistent with rain but does not prove it. The contradiction example requires recognizing that "empty" and "crowded" are mutually exclusive. All three require world knowledge and inference, not just pattern matching.

Text Correction

Grammar and spelling correction fit the text-to-text format naturally. The task is to transform erroneous text into corrected text. This is a sequence-to-sequence problem in the most direct sense: the input is a sentence with errors, and the output is the corrected version. This makes it one of the most natural fits for the text-to-text paradigm.

In[28]:
Code
correction_examples = [
    {
        "input": "correct: Their going to the store tommorrow.",
        "output": "They're going to the store tomorrow.",
    },
    {
        "input": "correct: The quick brown fox jump over the lazy dog.",
        "output": "The quick brown fox jumps over the lazy dog.",
    },
]

Text correction shows how the model must simultaneously understand what the intended meaning is (inferring from context what "their" should be) and know the correct form (the contraction "they're"). This requires both semantic understanding and grammatical knowledge. The text-to-text formulation handles this gracefully: the model learns to produce corrected text through the same generation mechanism it uses for everything else.

Summarization with Length Control

T5 can be trained to follow length specifications embedded in the prefix. By incorporating length requirements into the task instruction, the model learns to produce summaries of varying granularity on demand.

In[29]:
Code
# Summarization with different length targets
summarization_variants = [
    {
        "input": "summarize in one sentence: [long article text]",
        "output": "Brief one-line summary.",
    },
    {
        "input": "summarize in 3 sentences: [long article text]",
        "output": "First key point. Second important detail. Final takeaway.",
    },
    {
        "input": "summarize for twitter: [long article text]",
        "output": "Ultra-brief summary under 280 chars.",
    },
]

Length-controlled summarization illustrates the generality of the prefix mechanism. The prefix can encode parameters, not merely name the task. It can specify how the task should be performed. The model learns to interpret natural language length specifications and adjust its output accordingly. This is a primitive form of instruction following: the model reads a natural language constraint and tries to satisfy it during generation.

Structured Data Generation

Even structured outputs can be serialized as text. This shows the full power of the text-to-text paradigm. Any output that can be expressed as a string, including JSON objects, SQL queries, and code, becomes a valid generation target.

In[30]:
Code
structured_outputs = [
    # JSON-like output
    {
        "input": "extract info: John Smith is a 35-year-old software engineer at Google.",
        "output": '{"name": "John Smith", "age": 35, "occupation": "software engineer", "employer": "Google"}',
    },
    # SQL generation
    {
        "input": "translate to SQL: Show me all customers from New York",
        "output": "SELECT * FROM customers WHERE city = 'New York'",
    },
    # Code generation
    {
        "input": "python function: calculate factorial of n",
        "output": "def factorial(n): return 1 if n <= 1 else n * factorial(n-1)",
    },
]

These structured output examples hint at capabilities that would become central to later models. The ability to generate JSON or SQL from natural language descriptions, as well as executable code, anticipates the code-generating and structured-reasoning capabilities of models like Codex, GPT-4, and Claude. T5 demonstrated these capabilities at modest scale, and the insight that "any structured output is just text" proved enormously influential in how subsequent systems were designed and trained.

Worked Example: Tracing a Full Task Formatting Pipeline

To solidify how all these pieces fit together, let us trace through a complete end-to-end example. We will take a natural language inference problem from raw input through formatting and generation, then parse its output.

The task: Given the premise "The professor wrote equations on the whiteboard," determine whether the hypothesis "Someone taught a class" is entailed, neutral, or a contradiction.

Step 1: Format the input. We concatenate the task prefix with the hypothesis and premise in the T5 NLI format:

mnli hypothesis: Someone taught a class. premise: The professor wrote equations on the whiteboard.

Step 2: Tokenize. The tokenizer converts this string to a sequence of SentencePiece token IDs. The prefix words "mnli," "hypothesis," and "premise" are common enough to be single tokens or short sequences. The full input might tokenize to roughly 25-30 tokens.

Step 3: Encode. The T5 encoder processes all input tokens with full bidirectional attention. The encoder learns representations that capture the individual words and the logical relationship between the hypothesis and premise. The encoder has access to both the task prefix and the content simultaneously, so the representation of "professor wrote equations" is influenced by the knowledge that we are performing NLI.

Step 4: Decode. The decoder generates the output token by token. At the first decoding step, it attends to the entire encoder output through cross-attention and generates a probability distribution over the vocabulary. The top candidates at this step should be "entailment," "neutral," and "contradiction," because the model has learned to associate the "mnli" prefix with these output labels.

Step 5: Parse. The generated text "entailment" is returned. Our output parser checks it against the set of valid NLI labels and returns the structured result.

Reasoning check. The correct answer is "entailment" because professors writing equations on whiteboards is a form of teaching, and writing equations is a common activity in educational settings. A model that has learned these associations will assign high probability to "entailment." This reasoning relies on world knowledge absorbed during pre-training, not just pattern matching in the immediate context.

This trace reveals something important: the text-to-text format is a convenient interface, and it also shapes the entire computational process from encoding to decoding. The task prefix is not a post-hoc label. It conditions every step of the computation.

Building a Task Formatter

The following utility class handles task formatting consistently, encapsulating the formatting conventions we have discussed into a clean interface for preparing inputs across different task types. In production systems, centralized formatting logic like this is needed for consistency: all parts of the system that construct T5 inputs should go through the same formatter to ensure they use the same prefix conventions.

In[31]:
Code
class T5TaskFormatter:
    """Format various NLP tasks for T5 text-to-text processing."""

    @staticmethod
    def sentiment(text: str) -> str:
        """Format text for sentiment classification."""
        return f"sst2 sentence: {text}"

    @staticmethod
    def translation(text: str, source_lang: str, target_lang: str) -> str:
        """Format text for translation."""
        return f"translate {source_lang} to {target_lang}: {text}"

    @staticmethod
    def summarization(text: str, max_words: int = None) -> str:
        """Format text for summarization with optional length control."""
        if max_words:
            return f"summarize to {max_words} words: {text}"
        return f"summarize: {text}"

    @staticmethod
    def qa(question: str, context: str = None) -> str:
        """Format for question answering (open or closed book)."""
        if context:
            return f"question: {question} context: {context}"
        return f"question: {question}"

    @staticmethod
    def ner(text: str) -> str:
        """Format text for named entity recognition."""
        return f"extract entities: {text}"

    @staticmethod
    def nli(premise: str, hypothesis: str) -> str:
        """Format for natural language inference."""
        return f"mnli premise: {premise} hypothesis: {hypothesis}"

    @staticmethod
    def grammar_correction(text: str) -> str:
        """Format for grammar and spelling correction."""
        return f"correct: {text}"

    @staticmethod
    def similarity(sentence1: str, sentence2: str) -> str:
        """Format for semantic similarity."""
        return f"stsb sentence1: {sentence1} sentence2: {sentence2}"
In[32]:
Code
# Using the formatter
formatter = T5TaskFormatter()

formatted_examples = [
    formatter.sentiment("This book was incredibly engaging!"),
    formatter.translation("Good morning", "English", "Spanish"),
    formatter.qa(
        "What color is the sky?", "The sky appears blue during the day."
    ),
    formatter.nli("All birds can fly.", "Penguins are birds."),
]
Out[33]:
Console
Formatted task examples:
  sst2 sentence: This book was incredibly engaging!
  translate English to Spanish: Good morning
  question: What color is the sky? context: The sky appears blue during the day.
  mnli premise: All birds can fly. hypothesis: Penguins are birds.

Each formatted string follows a consistent pattern: a task-specific prefix followed by the input content. This consistency is key to T5's ability to route inputs to the appropriate behavior. The formatter class encapsulates these patterns, which makes it easy to ensure correct formatting across an application. By centralizing the formatting logic, we also make it easier to update formats if needed during iterative development. Changing the formatter updates all usages automatically, preventing the subtle inconsistencies that arise when formatting logic is scattered across a codebase.

Parsing Generated Outputs

Generating text is only half the problem. For tasks with structured outputs, you need to parse the model's generations back into usable data. The model produces strings, but your application likely needs Python objects, database entries, or API responses. This parsing step bridges the gap between the model's text-based interface and the structured world of software systems.

Reliable output parsing is one of the most practically important aspects of deploying text-to-text models. Model outputs will occasionally deviate from expected formats, especially for novel inputs that differ from training data. A well-designed parser handles variations gracefully, normalizes outputs where possible, and fails in controlled, informative ways when the output is truly unparseable. The goal is not perfect parsing, but graceful degradation.

In[34]:
Code
import re
from typing import Any, Dict, List, Tuple


class T5OutputParser:
    """Parse T5 generated outputs back into structured data."""

    @staticmethod
    def parse_classification(output: str, valid_labels: List[str]) -> str:
        """Parse classification output, handling slight variations."""
        output = output.strip().lower()
        for label in valid_labels:
            if label.lower() in output:
                return label
        return output  # Return raw if no match

    @staticmethod
    def parse_entities(output: str) -> List[Tuple[str, str]]:
        """Parse entity list format: 'type1: entity1, type2: entity2'"""
        if not output.strip():
            return []

        entities = []
        for part in output.split(", "):
            if ": " in part:
                entity_type, entity_text = part.split(": ", 1)
                entities.append((entity_type.strip(), entity_text.strip()))

        return entities

    @staticmethod
    def parse_inline_entities(output: str) -> List[Dict[str, Any]]:
        """Parse inline markup format: '[TYPE entity text]'"""
        pattern = r"\[(\w+)\s+([^\]]+)\]"
        matches = re.findall(pattern, output)
        return [{"type": m[0], "text": m[1]} for m in matches]

    @staticmethod
    def parse_similarity_score(output: str) -> float:
        """Parse numeric similarity score."""
        try:
            # Handle outputs like "4.2" or "score: 4.2"
            numbers = re.findall(r"\d+\.?\d*", output)
            if numbers:
                return float(numbers[0])
        except ValueError:
            pass
        return 0.0
In[35]:
Code
# Test the parsers
parser = T5OutputParser()

# Classification parsing
sentiment_output = "positive"
parsed_sentiment = parser.parse_classification(
    sentiment_output, ["positive", "negative"]
)

# Entity parsing
ner_output = (
    "person: Marie Curie, location: Paris, organization: Sorbonne University"
)
parsed_entities = parser.parse_entities(ner_output)

# Similarity parsing
similarity_output = "3.8"
parsed_score = parser.parse_similarity_score(similarity_output)
Out[36]:
Console
Sentiment: 'positive' -> positive

Entities: 'person: Marie Curie, location: Paris, organization: Sorbonne University'
  - person: Marie Curie
  - location: Paris
  - organization: Sorbonne University

Similarity score: '3.8' -> 3.8

The parsers successfully convert raw text outputs back into structured Python data. Classification outputs map to discrete labels, entity strings become lists of (type, text) tuples, and numeric scores parse to floats. Notice that parse_classification uses substring matching rather than exact matching: it checks whether any valid label appears in the output rather than requiring an exact string match. This handles common variations like leading or trailing whitespace, extra qualifying words ("the sentiment is positive"), or minor phrasing differences. The tradeoff is that it might incorrectly match if one label is a substring of another. Designing parsing logic requires thinking carefully about what kinds of variation your model is likely to produce for your specific training data.

Multitask Training

A key consequence of unified task formatting is the ability to train on multiple tasks simultaneously. This forms a different training paradigm: learning one task can improve performance on others through positive transfer. When all tasks use the same input-output format, they can be mixed freely in the same training batch without any changes to the training loop.

Think of multitask training as cross-training for a language model. A runner who only trains on flat ground may be fast on flat courses but struggle on hills. A runner who trains on both hills and flat surfaces develops leg strength and cardiovascular fitness that transfers between environments. Similarly, a model that trains only on sentiment classification learns to identify opinion words and sentence-level sentiment polarity. A model that also trains on translation develops a deeper understanding of semantic equivalence across languages, which benefits its sentiment understanding. A model that trains on QA develops reasoning skills that benefit its comprehension across all tasks.

In[37]:
Code
# Multitask training batch
multitask_batch = [
    # Translation
    {
        "input": "translate English to German: Hello world",
        "output": "Hallo Welt",
    },
    # Classification
    {"input": "sst2 sentence: Great movie!", "output": "positive"},
    # Summarization
    {
        "input": "summarize: Long article about climate change...",
        "output": "Climate change accelerates...",
    },
    # QA
    {
        "input": "question: What year? context: Founded in 1998.",
        "output": "1998",
    },
]

# All examples use the same model, loss function, and update procedure
Out[38]:
Visualization
Pie chart showing the proportional distribution of tasks in T5's multitask training data, including translation, summarization, question answering, classification, and other tasks.
Approximate task proportions in T5's multitask training data, showing that translation and summarization receive larger data shares while classification and NLI contribute smaller but important fractions.
Horizontal bar chart illustrating the composition of a sample training batch, showing the counts of translation, classification, summarization, and question-answering examples.
Composition of a sample training batch illustrating how different tasks are mixed together, with translation and classification examples appearing more frequently than summarization and QA in this batch.

T5's original training mixed examples from many tasks, with task proportions tuned to balance learning across all tasks. This multitask setup lets the model to develop reliable representations that transfer across tasks. Skills learned from translation help summarization because both require understanding semantic equivalence. Reasoning developed through QA improves NLI performance because both require logical inference. The shared encoder learns representations that capture linguistic information useful across all tasks, while the decoder learns flexible generation strategies that adapt to different output formats based on the prefix.

The intuition behind why multitask learning helps is that different tasks emphasize different aspects of language understanding. Translation forces the model to deeply understand semantics in order to preserve meaning across languages. Summarization teaches the model to identify salient information and compress it without losing key facts. QA develops fact retrieval and reasoning. Classification hones sensitivity to document-level properties such as sentiment and topic, including stylistic properties. When these tasks share parameters, the model must develop representations that serve all these needs simultaneously. Such representations tend to be richer and more generalizable than those learned for any single task.

Task proportions in multitask training are not arbitrary. T5's authors experimented with different mixing strategies and found that equal mixing does not necessarily yield the best results. Tasks with large datasets, like translation, can dominate if examples are mixed uniformly, potentially causing the model to over-optimize for those tasks at the expense of smaller tasks. Strategies like temperature-based mixing, where task sampling probability is adjusted by a temperature parameter, allow practitioners to balance learning across tasks of different sizes. This is an important hyperparameter in T5-style multitask training that deserves careful tuning for any custom application.

Limitations and Practical Considerations

While the text-to-text paradigm is elegant and powerful, it introduces several challenges that practitioners must handle carefully. Understanding these limitations is needed for building reliable systems and for deciding when the text-to-text approach is appropriate versus when a more traditional task-specific architecture might be preferable.

The most immediate limitation is generation overhead. For simple classification tasks, generating text tokens autoregressively is materially more expensive than predicting a single softmax over labels. A traditional classifier might output a 3-dimensional probability vector in a single forward pass. T5 must autoregressively generate the label text, requiring multiple decoder steps even for single-token labels. For high-volume classification applications, this overhead compounds quickly. A traditional classifier might process tens of thousands of examples per second on a modern GPU. A generative approach might be an order of magnitude slower. If your application requires classifying millions of documents daily, this cost difference matters considerably, and you may need to either use the probability-extraction technique described earlier or distill the T5 model into a task-specific classifier for production.

Output format consistency is a second major challenge. The model might generate valid but unexpected outputs depending on what patterns it learned from training data. When asked for sentiment, it might produce "positive," "very positive," "I think it's positive," or even a rephrasing of the input. Reliable parsing and validation are needed for production systems. Well-designed parsers handle common variations gracefully, but truly malformed outputs require fallback strategies: retry with different decoding parameters, flag for human review, or fall back to a simpler model. The inconsistency of generative outputs is one reason why task-specific architectures remain popular for high-precision applications even as generative models have improved sharply.

Error accumulation in structured outputs is a more subtle limitation. When generating complex structured outputs like entity lists, JSON objects, or multi-hop reasoning chains, early errors can cascade. If the model makes a formatting mistake partway through a long entity list, the remainder may be entirely unparseable. This differs fundamentally from traditional structured prediction, where each output position is typically independent or depends on a fixed set of prior positions through a defined structure. In autoregressive generation, each token depends on all previous tokens, so a single error can corrupt the remaining output. Constrained decoding techniques, which force the model to generate only valid tokens at each step, can help with this but add implementation complexity.

Vocabulary constraints impose another subtle limitation. T5's SentencePiece vocabulary affects what outputs are efficient to generate. Uncommon labels or domain-specific terms may require multiple tokens, each of which is an opportunity for error. If your NER task uses entity types like "regulatory_agency" or "pharmaceutical_compound," these multi-token labels require the model to correctly generate every subword in sequence. Each subword generation step introduces a probability of error, and these probabilities multiply across the sequence. For high-precision applications with many such labels, this can noticeably degrade performance compared to architectures that treat labels as atomic units.

Finally, the text-to-text paradigm conflates the model's language understanding with its task-following behavior. When a model fails to produce the correct output, it is often unclear whether the failure is due to insufficient language understanding or insufficient exposure to the task format during training. This ambiguity makes debugging and improvement harder than in task-specific architectures where these concerns are more cleanly separated. A traditional classifier that fails can be retrained with more data, a better architecture, or different features. A generative model that fails might need more task-specific fine-tuning, a different prefix format, better decoding parameters, or a fundamentally different training approach.

Despite these challenges, the benefits of unification generally outweigh the costs for most applications. The ability to fine-tune a single model on diverse tasks, share representations across domains, and handle new tasks with minimal architectural changes has transformed NLP system design. The text-to-text approach laid the groundwork for the instruction-following capabilities we will explore in later parts of this book, and understanding its mechanics deeply is needed for understanding why modern large language models behave the way they do.

Key Parameters

When deploying T5 for text-to-text tasks, several parameters materially affect both performance and efficiency:

  • max_length: Maximum input sequence length for tokenization. Longer inputs are truncated. T5 variants support different maximum lengths (512 for t5-small and t5-base, 1024 for larger variants). Truncating important context can severely degrade performance on tasks like QA where the answer may appear at the end of a long passage.
  • max_new_tokens: Maximum number of tokens to generate in the output. Controls output length and inference time. Setting this too small will cause the model to cut off its output mid-generation, while setting it too large wastes computation.
  • num_beams: Number of beams for beam search during generation. Higher values explore more candidate sequences but increase computation proportionally. For simple tasks like single-label classification, beam search with num_beams=1 (greedy decoding) is often sufficient. For longer, more complex outputs, beam search with 4-8 beams can noticeably improve quality.
  • skip_special_tokens: Whether to remove special tokens like </s> when decoding generated output to text. This should almost always be True for the final human-readable output, but should be False if you need to inspect the raw token sequence.

Summary

T5's text-to-text paradigm turns NLP by unifying all tasks into a single sequence-to-sequence generation framework. Every task becomes a text transformation: given an input string, produce an output string. Task prefixes signal which operation to perform, while consistent formatting lets multitask training and positive transfer between tasks. Classification becomes generating label text, NER becomes listing or marking entities, QA becomes generating answers conditioned on a question and optional context, and structured outputs like JSON or SQL become text generation targets.

The key insights from this chapter:

  • Task prefixes act as routing instructions, conditioning the model on which transformation to apply. Natural language prefixes work better than arbitrary codes because they activate the model's existing linguistic knowledge.
  • Classification as generation works by having the model output label text, with prediction probabilities recoverable from decoder logits by scoring the likelihood of each candidate label.
  • NER as generation can use either entity listing or inline markup formats, each with different tradeoffs between parsability, positional information, and generation complexity.
  • QA naturally fits text-to-text, letting both extractive and abstractive answers, as well as both open-book QA with provided context and closed-book QA that draws on parametric memory from pre-training.
  • Multitask training becomes trivial since all tasks share the same input/output format, and mixing diverse tasks lets positive transfer that benefits all tasks simultaneously.
  • Output parsing is needed in production systems because generated text may vary from expected formats, and good parsers handle common variations gracefully while failing informatively on truly malformed outputs.
  • Limitations include generation overhead for simple tasks, output format inconsistency, error accumulation in structured outputs, and vocabulary constraints for uncommon labels.

The next chapter introduces BART, another encoder-decoder model that takes a different approach to pre-training by corrupting text in diverse ways and learning to reconstruct the original, while maintaining similar flexibility for downstream text generation and understanding tasks.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about T5's text-to-text task formatting.

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025t5task, author = {Michael Brenndoerfer}, title = {T5 Task Formatting: Text-to-Text NLP Unification}, year = {2025}, url = {https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). T5 Task Formatting: Text-to-Text NLP Unification. Retrieved from https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp
MLAAcademic
Michael Brenndoerfer. "T5 Task Formatting: Text-to-Text NLP Unification." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp>.
CHICAGOAcademic
Michael Brenndoerfer. "T5 Task Formatting: Text-to-Text NLP Unification." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp.
HARVARDAcademic
Michael Brenndoerfer (2025) 'T5 Task Formatting: Text-to-Text NLP Unification'. Available at: https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). T5 Task Formatting: Text-to-Text NLP Unification. https://mbrenndoerfer.com/writing/t5-task-formatting-text-to-text-nlp

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.