Part of Language AI Handbook
Explains how mT5 extends T5 to 101 languages using temperature-based sampling, the mC4 corpus, and 250K vocabulary for effective cross-lingual transfer.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
mT5
T5's text-to-text framework demonstrated that a single model could handle diverse NLP tasks through unified formatting. However, T5 was trained exclusively on English text from the C4 corpus. This English-only focus meant the model could not process other languages without substantial additional training. mT5 (multilingual T5) addresses this limitation by extending the T5 paradigm to 101 languages while preserving the elegant text-to-text approach we covered in the T5 chapters.
The challenge of building a truly multilingual model goes beyond simply adding more training data. Languages differ dramatically in their available text resources. English dominates the web, while many languages have orders of magnitude less content. Training naively on this imbalanced data would produce a model that excels at English but performs poorly for low-resource languages. mT5 addresses this imbalance through careful corpus curation, temperature-based language sampling, and an expanded multilingual vocabulary.
To appreciate why mT5 matters, consider the languages it is designed to support. Roughly 7,000 languages are spoken in the world today, yet the overwhelming majority of NLP research has focused on fewer than a dozen of them. Even among languages with significant online presence, quality NLP tools are often limited to a handful. A speaker of Swahili, Telugu, or Amharic would find that most conversational AI systems, search engines with query understanding, and text classifiers simply do not work well in their language. The aspiration behind mT5 is that a single well-trained model could serve as the backbone for all of these use cases, letting researchers and engineers build on a common multilingual foundation rather than training separate models for every language they need to support.
The key insight motivating the entire mT5 design is that language understanding at a deep level should be largely language-agnostic. When a model learns that certain patterns of context predict certain kinds of spans, it is learning patterns in how information is structured in text, rather than properties of the English words used to express that information. This universality hypothesis drives the decision to train a single model across all 101 languages simultaneously rather than training separate models and hoping they stay compatible. The T5 chapters established the text-to-text framework and the span corruption pre-training objective. mT5 takes both of these and asks: how much can we generalize them across the world's languages while preserving the properties that made T5 effective?
Building on what we covered about T5's pre-training pipeline, the three central challenges mT5 had to solve were: where to get text in 101 languages, how to ensure all languages receive adequate training signal despite wildly different data quantities, and how to build a tokenizer that handles 101 different writing systems without exploding the model's vocabulary. Each of these challenges has a principled solution, and understanding those solutions is the subject of this chapter.
The mC4 Corpus
Building a multilingual model requires multilingual data. The mT5 team created mC4 (multilingual Colossal Clean Crawled Corpus) by applying language-specific filtering to Common Crawl web data. This process extracted text in 101 languages, resulting in a corpus orders of magnitude larger than previous multilingual datasets.
Think of Common Crawl as a photograph of the web taken repeatedly over many years, containing billions of documents in raw, unfiltered form. The mC4 pipeline takes that raw snapshot and subjects each document to a series of language detection and quality checks before keeping it. The goal is to produce a corpus where each language's subcorpus contains useful prose rather than spam, auto-generated content, or boilerplate navigation menus. The quality of the resulting corpus directly determines the quality of the representations the model can learn, which is why the construction steps deserve careful attention.
The corpus construction followed similar quality filtering steps to the English C4 we discussed in the T5 pre-training chapter, but applied language detection to separate content:
- Language identification: Each page was classified using CLD3 (Compact Language Detector), keeping only pages where the primary language exceeded a confidence threshold
- Line-level deduplication: Removed duplicate lines within each language's subcorpus to reduce boilerplate and repeated content
- Quality filtering: Applied heuristics to remove pages with too few words, excessive punctuation, or other quality issues
The resulting corpus shows extreme size variation across languages:

English contains roughly 2.7 trillion tokens, while some African languages like Yoruba have only 60 million tokens, a ratio of over 45,000:1. This imbalance creates a basic problem: training proportionally to data size would essentially ignore low-resource languages. Equal sampling would massively oversample (and overfit to) low-resource data while underutilizing high-resource content.
Notice that the size disparity has consequences beyond the number of examples. The English portion of mC4 is so large that the model would encounter many completely distinct documents during training, exposing it to enormous lexical and topical diversity. The Yoruba portion, by contrast, is small enough that the model would cycle through the same documents many times per training epoch. Cycling creates a memorization risk: the model begins to overfit specific sentences rather than learning generalizable patterns. The deeper reason why naive proportional sampling fails for low-resource languages is that repeated exposure to the same documents gives them a poor training signal.
Temperature-Based Language Sampling
The large disparity in corpus sizes across languages presents a basic challenge for multilingual model training. If we simply train on data in proportion to how much exists, English would dominate the training process. The model would see English text roughly 45,000 times more often than Yoruba text. Such a model might achieve excellent English performance, but would perform poorly for speakers of low-resource languages. On the other hand, if we sample equally from all languages, we would cycle through the entire Yoruba corpus thousands of times while barely scratching the surface of available English data. This leads to severe overfitting on low-resource languages and underutilization of high-resource ones.
To address the resource imbalance, mT5 uses temperature-based sampling that interpolates between proportional and uniform sampling. This approach allows practitioners to control the tradeoff between respecting natural data proportions and so adequate representation for all languages. Let represent the probability of sampling from language during training. If we sample proportionally to corpus size, we have:
where:
- : the probability of sampling from language during training
- : the number of tokens in language 's subcorpus
- : the total number of tokens across all languages, serving as a normalizing constant
This straightforward proportional approach would give English roughly 27% of all samples while many languages would appear in less than 0.01% of batches, essentially invisible during training.
The key insight behind temperature sampling is that we can systematically compress the differences between corpus sizes by applying a mathematical transformation. Temperature sampling modifies these probabilities by raising them to a power , where is the temperature. The intuition is that the exponent acts as a "flattening" operator on the distribution. When we raise numbers of vastly different magnitudes to a small power, their differences shrink dramatically. Consider what happens when you raise both 1,000,000 and 1 to the power 0.01: you get approximately 1.15 and 1.0 respectively. The millionfold difference has compressed to just 15%. This compression is precisely what temperature sampling exploits:
where:
- : the temperature-adjusted sampling probability for language
- : the temperature parameter controlling the balance between proportional and uniform sampling
- : the corpus size raised to power , which compresses differences between languages as increases
- : the sum of adjusted corpus sizes across all languages, so probabilities sum to 1
The key insight is that as temperature increases, the exponent approaches zero, making approach 1 for all languages regardless of their original corpus size. This progressively flattens the distribution toward uniform sampling. Understanding this behavior requires thinking carefully about what happens to the exponentiation as the temperature parameter changes.
To see why this works mathematically, consider two extreme cases:
- When : We have , so probabilities are exactly proportional to corpus size
- When : We have for all languages, giving uniform probabilities of where is the number of languages
The benefit of this approach becomes clear when we consider intermediate temperatures. Rather than requiring a separate mechanism to interpolate between proportional and uniform sampling, the temperature parameter provides a continuous dial that smoothly transitions between these extremes. For example, with English at 2749B tokens and Yoruba at 0.06B tokens, at the ratio is 45,817:1, but at the ratio becomes , a significant compression. This means that even with temperature sampling, high-resource languages still receive slightly more training signal. This reflects their richer and more diverse content, but the gap narrows enough that low-resource languages can learn meaningful representations.
Temperature controls interpolation between sampling strategies. At , sampling is proportional to corpus size. As , sampling approaches uniform across languages. The mT5 authors found to work well, significantly boosting low-resource languages while still favoring high-resource languages.
Temperature-based sampling for multilingual training was not invented by the mT5 team. Earlier multilingual machine translation systems, including Google's massively multilingual neural MT work from 2019, faced the same resource imbalance problem and adopted similar sampling strategies. The mT5 contribution was showing that this technique transfers cleanly to encoder-decoder pre-training and produces cross-lingual representations powerful enough for zero-shot transfer across diverse downstream tasks, not just translation.
Let's see how temperature affects sampling probabilities:
import numpy as np
# Corpus sizes for selected languages (in billions of tokens)
languages = [
"English",
"Russian",
"Spanish",
"Japanese",
"Swahili",
"Telugu",
"Yoruba",
]
corpus_sizes = np.array([2749, 743, 416, 261, 1.8, 1.2, 0.06])
def temperature_sampling(sizes, temperature):
"""Compute sampling probabilities with temperature."""
# Raise to power 1/T
adjusted = sizes ** (1 / temperature)
# Normalize to probabilities
return adjusted / adjusted.sum()
# Compare different temperatures
temperatures = [1, 10, 100, float("inf")]
# Calculate probabilities for each temperature
results = {}
for t in temperatures:
if t == float("inf"):
results[t] = np.ones(len(languages)) / len(languages)
else:
results[t] = temperature_sampling(corpus_sizes, t)Sampling probabilities at different temperatures: Language T=1 T=10 T=100 T=∞ ------------------------------------------------------------ English 65.8907 20.9243 14.9296 14.2857 Russian 17.8089 18.3583 14.7355 14.2857 Spanish 9.9711 17.3238 14.6503 14.2857 Japanese 6.2559 16.5347 14.5822 14.2857 Swahili 0.0431 10.0522 13.8742 14.2857 Telugu 0.0288 9.6528 13.8181 14.2857 Yoruba 0.0014 7.1540 13.4102 14.2857 ------------------------------------------------------------ Total 100.0 100.0 100.0 100.0
The table reveals the substantial effect of temperature on sampling distribution. At (proportional sampling), English would comprise over 65% of training data while Yoruba gets essentially zero. At (used by mT5), English is reduced to about 4.5% while Yoruba increases to 2.5%. This represents a large boost for low-resource languages. Yoruba's sampling probability increases by a factor of over 1000. The practical consequence is important: a model trained with proportional sampling would see Yoruba text so rarely that it could never learn the language's patterns, while temperature sampling ensures Yoruba appears frequently enough to develop useful representations of the language.


The temperature-based approach allows low-resource languages to receive meaningful training signal without completely ignoring the abundant high-resource data. However, this comes with a tradeoff: high-resource language performance slightly decreases compared to what could be achieved with proportional sampling. The mT5 authors found provided a good balance between these competing objectives. This choice reflects careful empirical tuning. Lower temperatures would still underrepresent low-resource languages, while higher temperatures would waste the rich diversity of high-resource data by undersampling it.
The code demonstrates that at , English's sampling probability drops from over 65% to around 14%, while Yoruba increases from nearly 0% to about 2%. This represents a boost factor of over 1000x for the low-resource language.
In practice, temperature sampling is applied both during model pre-training and during vocabulary training. Applying it only during model training but not vocabulary training would create a mismatch: the tokenizer would learn subword units primarily from English text. This produces a vocabulary ill-suited to the languages that the model then trains on more heavily. The mT5 team applied the same schedule to both stages. This keeps the vocabulary reflects the same language balance that the model itself encounters during pre-training. This consistency is easy to overlook but important for low-resource language quality.
Multilingual Tokenization
Training a single model on 101 languages requires a vocabulary that can effectively tokenize all of them. This presents a challenging problem combining computational linguistics and machine learning efficiency. As we covered in the SentencePiece chapter, subword tokenization algorithms learn vocabularies from data by identifying frequently occurring character sequences. The challenge for multilingual models is that vocabulary slots are finite. A larger vocabulary means more parameters in the embedding layer and slower softmax computation during training and inference. Yet a vocabulary that is too small cannot adequately represent the diverse morphological patterns and writing systems found across 101 languages.
mT5 uses a SentencePiece unigram model with a vocabulary of 250,000 subword tokens, compared to T5's 32,000 tokens for English only. This 8x increase accommodates the diverse character sets and morphological patterns across 101 languages. The expansion is necessary because different language families have fundamentally different word formation rules: agglutinative languages like Turkish build complex words by chaining morphemes, while isolating languages like Chinese use single characters to represent concepts. A vocabulary optimized for English would fragment Turkish words into unrecognizable pieces while failing to provide useful decompositions for Chinese characters.
The vocabulary training process samples from mC4 using the same temperature-based sampling as model training. This design choice is important for so low-resource languages contribute meaningful vocabulary entries rather than being drowned out by English. Without temperature sampling during vocabulary construction, the tokenizer would learn subword patterns primarily from English text, leading to poor tokenization quality for low-resource languages.
# Conceptual vocabulary training with temperature sampling
# (Actual mT5 used proprietary SentencePiece training)
def sample_training_text(corpus_sizes, temperature, total_chars=10_000_000):
"""
Sample text for vocabulary training with temperature.
Returns approximate character counts per language.
"""
probs = corpus_sizes ** (1 / temperature)
probs = probs / probs.sum()
# Allocate characters proportionally to tempered probabilities
char_counts = (probs * total_chars).astype(int)
return char_counts
# Compare vocabulary training samples
languages_vocab = ["English", "Russian", "Japanese", "Swahili", "Yoruba"]
sizes_vocab = np.array([2749, 743, 261, 1.8, 0.06])
# Calculate allocations for both sampling strategies
proportional = sample_training_text(sizes_vocab, temperature=1)
tempered = sample_training_text(sizes_vocab, temperature=100)Character allocation for vocabulary training (10M total): Language Proportional T=100 Boost ------------------------------------------------------- English 7,321,178 2,087,125 0.3x Russian 1,978,768 2,059,997 1.0x Japanese 695,099 2,038,559 2.9x Swahili 4,793 1,939,588 404.7x Yoruba 159 1,874,728 11790.7x
Temperature sampling ensures that even low-resource languages contribute substantial training data for vocabulary learning. Without this adjustment, languages like Yoruba might have fewer than 200 characters in the training sample. This is far too few for meaningful subword discovery. The SentencePiece algorithm needs sufficient examples of each language to identify common character patterns and build effective subword units.
Script Coverage
The 250K vocabulary must cover many writing systems. These include Latin and Cyrillic, Arabic and Hebrew, plus Chinese and Japanese. Korean and numerous other writing systems also need representation. Each writing system brings its own characteristics: alphabetic scripts like Latin and Cyrillic build words from individual letters, syllabic scripts like Japanese hiragana represent syllables, and logographic scripts like Chinese use characters that represent morphemes or words. The vocabulary breakdown reflects this diversity, with capacity allocated across script families to ensure adequate coverage:

The vocabulary expansion from 32K to 250K tokens has measurable implications for model efficiency. Each token requires an embedding vector that maps the discrete token to a continuous representation the model can process. The embedding matrix size therefore grows proportionally with vocabulary size, creating a direct tradeoff between linguistic coverage and parameter efficiency:
# Compare embedding matrix sizes
embedding_dim = 1024 # mT5-Large uses 1024-dimensional embeddings
t5_vocab = 32_000
mt5_vocab = 250_000
t5_params = t5_vocab * embedding_dim
mt5_params = mt5_vocab * embedding_dimT5 embedding parameters: 32,768,000 (32.8M) mT5 embedding parameters: 256,000,000 (256.0M) Increase factor: 7.81x
The embedding layer expansion from 32.8M to 256M parameters represents a significant overhead, particularly for smaller model variants. For mT5-Small with approximately 300M total parameters, the embeddings alone account for a substantial fraction of model capacity. This means that a non-trivial portion of the model's learning capacity is dedicated purely to representing the expanded vocabulary, leaving less capacity for learning language understanding and generation. The designers of mT5 judged this tradeoff worthwhile because adequate vocabulary coverage is foundational. A model cannot learn patterns in text it cannot properly tokenize.
Tokenization Efficiency
Multilingual tokenizers face a fertility tradeoff. Tokens optimized for one language may fragment words in another, leading to longer sequences and slower processing. This phenomenon occurs because subword patterns that are common in one language may be rare or nonexistent in another. For instance, the English suffix "-tion" appears frequently and would likely become a single token. However, this character sequence rarely occurs in Japanese. When mT5's tokenizer encounters Japanese text, it must use different subword patterns entirely. Let's examine how mT5's tokenizer handles different languages:
from transformers import T5Tokenizer
# Load mT5 tokenizer (use slow tokenizer to avoid conversion issues)
tokenizer = T5Tokenizer.from_pretrained("google/mt5-small")
# Test sentences (translations of "The quick brown fox jumps over the lazy dog")
test_sentences = {
"English": "The quick brown fox jumps over the lazy dog.",
"Spanish": "El rápido zorro marrón salta sobre el perro perezoso.",
"German": "Der schnelle braune Fuchs springt über den faulen Hund.",
"Russian": "Быстрая коричневая лиса перепрыгивает через ленивую собаку.",
"Japanese": "素早い茶色の狐が怠惰な犬を飛び越える。",
"Chinese": "敏捷的棕色狐狸跳过懒狗。",
"Arabic": "الثعلب البني السريع يقفز فوق الكلب الكسول.",
"Swahili": "Mbweha mwepesi wa kahawia anaruka juu ya mbwa mvivu.",
}mT5 Tokenization Efficiency by Language: Language Chars Tokens Chars/Tok ------------------------------------------ English 44 14 3.14 Spanish 53 19 2.79 German 55 19 2.89 Russian 59 20 2.95 Japanese 19 17 1.12 Chinese 12 14 0.86 Arabic 42 17 2.47 Swahili 52 20 2.60 Average - - 2.35

The tokenization efficiency results show that Latin-script languages achieve roughly 4-6 characters per token. This holds for English and Spanish as well as German, while languages like Japanese and Chinese with unique scripts show different patterns. Arabic script languages fall somewhere in between. These efficiency differences directly impact sequence lengths. Less efficient tokenization means longer sequences for the same content, which affects both computational cost and the model's ability to capture long-range dependencies within its context window. A sentence that tokenizes to 10 tokens in English might require 20 tokens in another language, effectively halving the amount of context the model can consider for that language within a fixed context window.
In practice, this fertility gap has downstream consequences for fine-tuning and inference that are easy to miss. When you set a maximum input length of 512 tokens for an mT5 fine-tuning job, a document that fits comfortably in English might get truncated when expressed in a language with higher fertility. This means that languages already disadvantaged by lower pre-training data volumes can also end up with shorter effective context during fine-tuning, compounding the disadvantage. When deploying mT5 for a multilingual application, it is worth measuring the token-to-word ratio for each target language and adjusting the maximum sequence length accordingly, or at least being aware that different languages effectively see different context windows at the same token limit.
Cross-Lingual Transfer
One of mT5's most powerful capabilities is cross-lingual transfer: the ability to fine-tune on data in one language and achieve reasonable performance in others. This property emerges from the shared multilingual representations learned during pre-training. When the model learns to predict masked spans across 101 languages simultaneously, it develops internal representations that capture language-universal patterns in how text structures information and expresses meaning.
The key insight here is that pre-training creates a shared semantic space across languages. Fine-tuning on English data teaches the model to perform a task, such as extracting answers from passages. Because the underlying representations for English and, say, Vietnamese partially overlap in this shared space, the model can apply what it learned about the task structure to Vietnamese text without ever having seen a Vietnamese training example for that task. This is not magic; it is a consequence of the model having learned, during pre-training, to encode semantically similar content in semantically similar positions in representation space, regardless of language.
Cross-lingual transfer comes in several variants depending on how much target-language data is available during fine-tuning:
- Zero-shot transfer: Fine-tune only on source language (usually English), evaluate directly on target languages
- Few-shot transfer: Fine-tune on source language, then provide a small number of target-language examples to adapt
- Translate-train: Automatically translate training data into target languages and fine-tune on the translated data
- Multilingual fine-tuning: Fine-tune on labeled data in multiple languages simultaneously when such data exists
Each strategy has different resource requirements and performance characteristics. Zero-shot transfer is the most practically useful because it requires no target-language annotations, but it also produces the largest performance gap compared to models trained on target-language data. Few-shot transfer dramatically narrows this gap with only a handful of examples.
Zero-Shot Cross-Lingual Transfer
In zero-shot transfer, a model is fine-tuned on task data in one language (typically English, where labeled data is abundant) and evaluated on the same task in other languages without seeing any target-language training examples. This capability is valuable because practitioners can use English datasets and achieve reasonable performance across dozens of languages without requiring labeled data in every language:

The visualization shows that mT5-Large achieves strong cross-lingual transfer, with Spanish and German (both related to English) reaching F1 scores above 68, while more distant languages like Hindi and Arabic show scores around 59-62. mT5 outperforms XLM-R across all languages, with improvements ranging from 5-7 F1 points. The performance gradient from English through related languages to distant languages reveals how linguistic similarity affects transfer success.
Transfer performance correlates with several factors:
- Linguistic similarity: Languages related to English (like German and Spanish) typically show better transfer than distant languages
- Script overlap: Languages sharing Latin script often transfer better due to shared subword tokens
- Pre-training data quantity: Languages with more mC4 data develop richer representations that support better transfer
Mechanisms of Cross-Lingual Transfer
Cross-lingual transfer works because mT5 learns language-agnostic representations during pre-training. The model doesn't learn 101 separate languages in isolation. Instead, it learns a unified representation space where similar concepts across languages map to similar regions, regardless of how those concepts are expressed on the surface. Several factors contribute to this alignment:
Shared vocabulary: When languages share subword tokens, especially cognates and loanwords, knowledge about these tokens transfers directly. For example, "computer" appears in similar forms in many languages. This allows the model to use what it learns about technology concepts in English text when processing Spanish or German text about the same topics. This lexical overlap creates anchor points that align representations across languages. Think of these shared tokens as bridges: when the model encounters the subword "comput" in an English text about programming, it encodes contextual meaning around that token. When it later encounters the same subword in a Spanish sentence, the representation it builds is not starting from zero. It inherits whatever the model has already learned about that token's typical contexts across all languages that share it.
Parallel structure learning: The span corruption objective forces the model to learn syntactic and semantic patterns. Many of these patterns, such as subject-verb-object ordering, generalize across languages. When the model learns that a certain span position typically contains an action word in English, this knowledge can transfer to languages with similar sentence structure. Even when word orders differ, the model learns abstract notions of "what information completes this context" that transcend specific grammatical rules. The span corruption objective is particularly well suited to encouraging this kind of structural abstraction because it requires the model to reason about what a span should contain based on surrounding context, which is at heart a semantic question rather than a lexical one.
Semantic alignment: By processing text in multiple languages about similar topics, the model learns that certain concepts are expressed similarly across languages, even when the surface forms differ. News articles about international events, Wikipedia pages about scientific concepts, and web content about popular topics appear in many languages. This provides implicit supervision for semantic alignment. This is a form of distant supervision: the model never explicitly sees pairs of translated sentences, yet the statistical regularities in multilingual web text teach it that the Spanish representation of "democratic elections" should live near the English one in representation space.
The strength of cross-lingual transfer is not uniform across all language pairs. Language pairs that are typologically similar, share script, and co-occur frequently in the same documents exhibit much stronger transfer than typologically distant pairs with unique scripts and little co-occurrence. This creates a spectrum of transfer quality across the 101 languages in mT5, from near-native performance for closely related European languages to more modest transfer for low-resource languages with distinct scripts. Understanding this spectrum helps you set realistic expectations when deploying mT5 for a given language combination.
# Examine shared tokens across languages
def find_shared_subwords(tokenizer, words_by_language):
"""Find subword tokens shared across language-specific words."""
shared_tokens = {}
for concept, translations in words_by_language.items():
all_subwords = set()
for lang, word in translations.items():
tokens = tokenizer.tokenize(word)
all_subwords.update(tokens)
shared_tokens[concept] = all_subwords
return shared_tokens
# Example: words for "computer" in different languages
computer_words = {
"computer": {
"English": "computer",
"Spanish": "computadora",
"German": "Computer",
"French": "ordinateur",
"Italian": "computer",
"Portuguese": "computador",
}
}Subword tokens for 'computer' across languages: English : computer → ['▁computer'] Spanish : computadora → ['▁', 'computador', 'a'] German : Computer → ['▁Computer'] French : ordinateur → ['▁', 'ordinateur'] Italian : computer → ['▁computer'] Portuguese : computador → ['▁', 'computador'] Shared tokens: ['▁computer', '▁', 'computador']
The shared tokens analysis reveals how mT5's vocabulary captures common subword patterns across related languages. Words derived from the same root often share subword components, letting direct knowledge transfer. The word 'computer' illustrates this overlap in English and German, and Italian retains related components. Languages with unique scripts like Japanese require entirely distinct tokens, which is why cross-lingual transfer to such languages relies more heavily on semantic alignment learned during pre-training rather than surface-level lexical overlap. The model must learn that the Japanese concept corresponding to "computer" should map to the same region of representation space as the English word, even though they share no characters.
The implication is that for languages with Latin-script overlap, cross-lingual transfer gets a significant boost from shared vocabulary. For script-isolated languages, the model must do more work to learn semantic alignment through context alone. Japanese and Chinese fall into this group, as do Arabic and languages using Devanagari or other distinctive alphabets. This is one reason why multilingual benchmarks consistently show stronger transfer performance for European languages than for languages written in non-Latin scripts, even controlling for corpus size differences.
mT5 vs T5 Performance
Comparing mT5 and T5 reveals the tradeoffs involved in multilingual training. On English-only benchmarks, mT5 slightly underperforms T5. This reflects the "curse of multilinguality": the model must divide its capacity across many languages.

The benchmark comparison reveals consistent but modest performance gaps: mT5-Large scores approximately 3 points lower on GLUE (86.4 vs 89.7), SuperGLUE (80.9 vs 84.6), and both SQuAD variants. The ROUGE-L gap on CNN/DM is about 2.8 points. This English performance gap is relatively small, typically 3-4 points, while mT5 gains the ability to process 100 additional languages. For applications requiring multilingual support, this tradeoff is highly favorable.
Notice that the performance gap narrows with model scale. At the XXL level, the sheer number of parameters provides enough capacity that the curse of multilinguality becomes less severe. This is an important empirical observation: the English penalty you pay for multilingual training is not a fixed tax, it is a function of model size relative to the number of languages. With enough parameters, a single multilingual model can come very close to monolingual performance while still supporting dozens of additional languages. This finding has important practical implications for system design, particularly as very large models become more accessible for fine-tuning through parameter-efficient methods like LoRA and prefix tuning, which we will explore in later chapters.
Performance Across Model Sizes
mT5 was released in multiple sizes, following T5's scaling approach:
# mT5 model variant specifications
# Format: (name, parameters, layers, d_model, d_ff)
variants = [
("Small", 300_000_000, 8, 512, 1024),
("Base", 580_000_000, 12, 768, 2048),
("Large", 1_200_000_000, 24, 1024, 2816),
("XL", 3_700_000_000, 24, 2048, 5120),
("XXL", 13_000_000_000, 24, 4096, 10240),
]
# Calculate total embedding parameters for largest variant
xxl_vocab_params = 250_000 * variants[-1][3] # vocab_size * d_modelmT5 Model Variants: Variant Parameters Layers d_model d_ff -------------------------------------------------------- Small 300,000,000 8 512 1024 Base 580,000,000 12 768 2048 Large 1,200,000,000 24 1024 2816 XL 3,700,000,000 24 2048 5120 XXL 13,000,000,000 24 4096 10240

The table shows how mT5 scales from 300M to 13B parameters. Each size increase brings proportionally larger hidden dimensions (d_model) and feed-forward dimensions (d_ff), with the largest models using 24 layers consistently. Larger variants show better cross-lingual transfer, suggesting that additional capacity helps the model maintain stronger representations across more languages. The relationship between model size and multilingual performance is an area we'll explore further in the scaling laws chapters.
The parameter counts show that the embedding layer grows faster than the total parameter count as d_model increases. For mT5-Small, the 250K vocabulary with 512-dimensional embeddings produces 128M embedding parameters, roughly 43% of the total 300M. For mT5-XXL with 4096-dimensional embeddings, the vocabulary accounts for 1.024B parameters out of 13B total, a smaller fraction but still a substantial absolute count. This explains why efficient embedding representations, such as tied input and output embeddings (which mT5 uses, following T5), matter especially in multilingual settings where the vocabulary is large.
Working with mT5
Let's implement a practical example using mT5 for multilingual text generation. We'll use the Hugging Face Transformers library to demonstrate fine-tuning on a simple translation-like task.
Working with mT5 in practice follows the same pattern as T5: format your task as a text-to-text problem, tokenize the input and target sequences, and train the model with a cross-entropy loss over the output tokens. The main difference is that your data can now span multiple languages, and you need to be thoughtful about how your task prefix interacts with the model's multilingual representations. For classification or extraction tasks, using the task prefix in the source language of the input text typically works well. For translation tasks, the convention is to specify the source and target languages explicitly in the prefix, similar to how Google's multilingual MT systems format their prompts.
One practical consideration when loading mT5 is that the model is larger than its T5 counterpart of the same nominal size, primarily because of the expanded vocabulary. mT5-Small has 300M parameters compared to T5-Small's approximately 60M, largely due to the 250K token vocabulary. For production deployments where memory is constrained, this vocabulary overhead is worth factoring into your infrastructure planning.
import torch
from transformers import AutoModelForSeq2SeqLM, T5Tokenizer
# Load mT5-small for demonstration
model_name = "google/mt5-small"
tokenizer = T5Tokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
# Move to GPU if available
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device)
model_params = sum(p.numel() for p in model.parameters())Model: google/mt5-small Device: cpu Parameters: 300,176,768 Vocabulary size: 250,100 Embedding parameters: 128,051,200
The mT5-Small model loads successfully with its full 250K vocabulary. The embedding layer alone accounts for a significant portion of the total parameters. This reflects the vocabulary expansion needed for multilingual support. Despite being the smallest variant, it still provides strong multilingual capabilities for experimentation and deployment in resource-constrained settings.
Multilingual Span Corruption Example
mT5 uses the same span corruption objective as T5. Let's examine how it works on different languages:
def demonstrate_span_corruption(text, language, tokenizer):
"""
Simulate span corruption for demonstration.
In practice, this happens during pre-training data preparation.
"""
tokens = tokenizer.tokenize(text)
# Select spans to corrupt (simplified: corrupt every 4th-7th token)
corrupted = []
targets = []
sentinel_id = 0
i = 0
while i < len(tokens):
if i > 0 and i % 5 == 0 and i + 2 < len(tokens):
# Corrupt a span of 2-3 tokens
span_len = min(3, len(tokens) - i)
corrupted.append(f"<extra_id_{sentinel_id}>")
targets.extend(
[f"<extra_id_{sentinel_id}>"] + tokens[i : i + span_len]
)
sentinel_id += 1
i += span_len
else:
corrupted.append(tokens[i])
i += 1
if sentinel_id > 0:
targets.append(f"<extra_id_{sentinel_id}>")
return " ".join(corrupted), " ".join(targets)
# Test on multiple languages
test_texts = {
"English": "Natural language processing enables computers to understand human language.",
"Spanish": "El procesamiento del lenguaje natural permite que las computadoras entiendan el idioma humano.",
"German": "Die Verarbeitung natürlicher Sprache ermöglicht es Computern, menschliche Sprache zu verstehen.",
"Japanese": "自然言語処理により、コンピュータは人間の言語を理解できるようになります。",
}Span Corruption Examples: === English === Original: Natural language processing enables computers to understand human language. Corrupted: ▁Natural ▁language ▁processing ▁en ables <extra_id_0> ▁human ▁language . Target: <extra_id_0> ▁computers ▁to ▁understand <extra_id_1> === Spanish === Original: El procesamiento del lenguaje natural permite que las computadoras entiendan el idioma humano. Corrupted: ▁El ▁proces amiento ▁del ▁ <extra_id_0> ▁permit e <extra_id_1> computador as <extra_id_2> ▁el ▁ <extra_id_3> Target: <extra_id_0> lengua je ▁natural <extra_id_1> ▁que ▁las ▁ <extra_id_2> ▁en tienda n <extra_id_3> idioma ▁humano . <extra_id_4> === German === Original: Die Verarbeitung natürlicher Sprache ermöglicht es Computern, menschliche Sprache zu verstehen. Corrupted: ▁Die ▁ Verarbeitung ▁ n <extra_id_0> Sprache ▁er <extra_id_1> ▁Computer n <extra_id_2> ▁ Sprache <extra_id_3> . Target: <extra_id_0> atürlich er ▁ <extra_id_1> möglich t ▁es <extra_id_2> , ▁mens chliche <extra_id_3> ▁zu ▁ver stehen <extra_id_4> === Japanese === Original: 自然言語処理により、コンピュータは人間の言語を理解できるようになります。 Corrupted: ▁ 自然 言語 処理 により <extra_id_0> 人間の 言語 <extra_id_1> ようになります 。 Target: <extra_id_0> 、 コンピュータ は <extra_id_1> を 理解 できる <extra_id_2>
The span corruption examples demonstrate how mT5's pre-training objective works uniformly across languages. Regardless of script or language family, the model learns to predict masked spans given surrounding context. This consistent training signal across all 101 languages encourages the model to develop language-agnostic representations that capture universal patterns in how information is structured in text.
Fine-Tuning for Multilingual Tasks
Fine-tuning mT5 follows the same text-to-text format as T5. Here's an example setup for multilingual question answering:
from torch.utils.data import Dataset
class MultilingualQADataset(Dataset):
"""
Dataset for multilingual question answering.
Formats data as text-to-text for mT5.
"""
def __init__(
self, examples, tokenizer, max_input_length=512, max_target_length=128
):
self.examples = examples
self.tokenizer = tokenizer
self.max_input_length = max_input_length
self.max_target_length = max_target_length
def __len__(self):
return len(self.examples)
def __getitem__(self, idx):
example = self.examples[idx]
# Format input as: "question: {question} context: {context}"
input_text = (
f"question: {example['question']} context: {example['context']}"
)
target_text = example["answer"]
# Tokenize
inputs = self.tokenizer(
input_text,
max_length=self.max_input_length,
truncation=True,
padding="max_length",
return_tensors="pt",
)
targets = self.tokenizer(
target_text,
max_length=self.max_target_length,
truncation=True,
padding="max_length",
return_tensors="pt",
)
return {
"input_ids": inputs["input_ids"].squeeze(),
"attention_mask": inputs["attention_mask"].squeeze(),
"labels": targets["input_ids"].squeeze(),
}
# Example multilingual QA data
multilingual_qa_examples = [
{
"question": "What is machine learning?",
"context": "Machine learning is a subset of artificial intelligence that enables systems to learn from data.",
"answer": "a subset of artificial intelligence",
},
{
"question": "¿Qué es el aprendizaje automático?",
"context": "El aprendizaje automático es un subconjunto de la inteligencia artificial que permite a los sistemas aprender de los datos.",
"answer": "un subconjunto de la inteligencia artificial",
},
{
"question": "Was ist maschinelles Lernen?",
"context": "Maschinelles Lernen ist ein Teilgebiet der künstlichen Intelligenz, das Systemen ermöglicht, aus Daten zu lernen.",
"answer": "ein Teilgebiet der künstlichen Intelligenz",
},
]Multilingual QA Dataset Examples: Example 1: Question: What is machine learning? Input tokens: 512 Target: a subset of artificial intelligence Example 2: Question: ¿Qué es el aprendizaje automático? Input tokens: 512 Target: un subconjunto de la inteligencia artificial Example 3: Question: Was ist maschinelles Lernen? Input tokens: 512 Target: ein Teilgebiet der künstlichen Intelligenz
The examples demonstrate mT5's text-to-text format for question answering: questions and contexts are combined into a single input string, and the model learns to generate the answer span. Using the same format for English and Spanish, and likewise for German, allows the model to learn the task structure and use cross-lingual representations.
Generation Across Languages
Let's examine how mT5 generates text in different languages:
def generate_text(model, tokenizer, prompt, max_length=50):
"""Generate text from a prompt using mT5."""
inputs = tokenizer(prompt, return_tensors="pt").to(device)
outputs = model.generate(
**inputs,
max_length=max_length,
num_beams=4,
early_stopping=True,
no_repeat_ngram_size=2,
)
return tokenizer.decode(outputs[0], skip_special_tokens=True)
# Test prompts in different languages
prompts = [
"translate English to German: Hello, how are you?",
"translate English to Spanish: The weather is nice today.",
"summarize: Machine learning is a branch of artificial intelligence focused on building systems that learn from data.",
]mT5 Generation Examples: (Note: mT5-small is not fine-tuned, results are from pre-training only) Input: translate English to German: Hello, how are you?
model.safetensors: reconstructing file: 0%| | 0.00B / 1.20GB
model.safetensors: downloading bytes: | 0.00B
Output: <extra_id_0> Input: translate English to Spanish: The weather is nice today. Output: <extra_id_0> Input: summarize: Machine learning is a branch of artificial intelligence focused on building systems that learn from data. Output: <extra_id_0>.
Note that the base mT5 model requires fine-tuning on specific tasks to produce high-quality outputs. The pre-trained model has learned multilingual representations through span corruption but hasn't been trained to follow specific task instructions. This is the same pattern we observed with T5: pre-training builds powerful general representations, and fine-tuning teaches the model to apply those representations to a specific task format. For production multilingual applications, you would fine-tune mT5 on task-specific multilingual data, potentially combining English-only labeled data with any available target-language examples to maximize transfer.
Worked Example: Temperature Sampling in Action
Let's trace through a concrete numerical example to solidify the temperature sampling concept. Suppose we have a corpus with just three languages: English (1,000 billion tokens), Turkish (10 billion tokens), and Swahili (1 billion tokens). These numbers are round but representative of the order-of-magnitude differences in the real mC4 corpus.
With proportional sampling (), the sampling probabilities are:
Under this scheme, for every 1000 training steps, Swahili appears roughly once. A typical pre-training run involves hundreds of billions of steps, so Swahili does appear many times in absolute terms, but relative to English it is almost invisible. The model's Swahili representations will be weak.
Now apply temperature sampling with . The adjusted corpus sizes are:
The new sum is approximately , giving:
Swahili's sampling probability jumped from 0.1% to 32.3%, a boost factor of over 300. Turkish went from 1.0% to 33.1%, a 33x boost. English dropped from 98.9% to 34.6%. The distribution is now nearly uniform across the three languages, which is exactly what temperature sampling is designed to achieve at high temperatures.
The practical takeaway is striking: temperature sampling turns a situation where Swahili would be nearly invisible into one where it receives roughly equal training signal to English. The model trained with temperature sampling will develop meaningful Swahili representations, while the proportionally sampled model would produce representations so weak as to be nearly useless for Swahili tasks.
Limitations and Impact
mT5 represented a significant advance in multilingual NLP, but several limitations affect its practical deployment and performance. Understanding these limitations clearly helps you make good decisions about when to reach for mT5 versus a monolingual model or a more specialized multilingual architecture.
The most basic challenge is what researchers call the curse of multilinguality. As more languages are added to a fixed-capacity model, each language must share the same parameter space with all the others. The result is a capacity tradeoff: mT5-Large has roughly the same number of non-embedding parameters as T5-Large, but those parameters must represent the grammatical patterns and vocabulary of 101 languages while also encoding world knowledge. For high-resource languages like English, this dilution produces a modest but measurable drop in performance, typically 3-5 points on standard benchmarks. For applications requiring the best possible English performance, a dedicated English model like T5 or a larger mT5 variant trained with higher English weighting might be preferable. The practical message is that multilingual capability comes at a cost, and that cost grows as the number of languages increases.
Resource imbalance persists even after temperature sampling. Temperature at dramatically compresses the difference between English and Yoruba in sampling probability, but it cannot substitute for the underlying fact that the model sees far less total Yoruba text than English text in absolute terms. Languages with only tens of millions of tokens in mC4 develop weaker, more brittle representations. Cross-lingual transfer from English to Yoruba is accordingly less reliable than transfer to Spanish, because the Yoruba representations are less well-grounded. This perpetuates existing gaps in NLP system quality across languages, even as it narrows them. A model like mT5 is therefore not a complete solution to language inequality in AI; it is a meaningful step forward that leaves substantial gaps, particularly for the world's lowest-resource languages.
The 250K vocabulary, while much larger than T5's 32K, still cannot achieve optimal tokenization for all 101 languages simultaneously. The SentencePiece training process optimizes for the sampled training distribution, which means languages with higher sampling probability at vocabulary training time tend to get better tokenization quality. Languages where the tokenizer fragments words excessively face longer sequence lengths, higher computational costs, and effectively shorter context windows for a given maximum token budget. This tokenization efficiency gap is a concrete practical limitation when deploying mT5 in production systems where latency and cost matter.
Evaluation is also hampered by a scarcity of high-quality benchmarks for most of the languages mT5 supports. XNLI and MLQA are among the multilingual benchmarks available, alongside XTREME, but they cover at most a few dozen languages and the quality of annotations varies considerably. For the approximately 60 languages in mT5's coverage that have little or no benchmark presence, it is difficult to know how well the model performs or whether improvements to training translate into real-world benefits. This benchmark scarcity creates a measurement problem and reflects a deeper asymmetry in research investment that affects the entire multilingual NLP ecosystem.
Despite these limitations, mT5 had a major impact on the field. The model demonstrated conclusively that the text-to-text framework could scale to massive multilingual settings without basic architectural changes, validating the universality of T5's design. The mC4 corpus became a standard resource for multilingual training, used by numerous subsequent models. The zero-shot cross-lingual transfer results established new baselines that pushed the community toward more ambitious multilingual evaluation standards. For practitioners, mT5 provided an accessible off-the-shelf foundation that could be fine-tuned for multilingual classification, question answering, and generation tasks without requiring the substantial resources needed to train a multilingual model from scratch. Many production multilingual applications today, spanning translation support tools, multilingual search, and cross-lingual document classification, trace their foundations to mT5 or models directly influenced by its design.
The scaling laws we'll explore in upcoming chapters suggest that many of mT5's limitations can be addressed through increased model scale, improved training data curation, and more sophisticated sampling strategies. Larger models can support more languages with less capacity interference, better-curated corpora can improve low-resource language quality, and adaptive sampling schemes can respond dynamically to training progress rather than using a fixed temperature throughout.
Summary
mT5 extends the T5 text-to-text paradigm to 101 languages through several interconnected innovations that address the unique challenges of multilingual pre-training:
- mC4 corpus: A massive multilingual dataset extracted from Common Crawl, with language-specific filtering and CLD3 language detection applied to create quality-filtered subcorpora for each supported language
- Temperature-based sampling: Uses with to balance between proportional and uniform language sampling, boosting low-resource languages by orders of magnitude relative to their raw corpus sizes
- Expanded vocabulary: 250K SentencePiece tokens (vs. 32K for T5) to cover diverse scripts and morphological patterns, trained with the same temperature sampling to ensure adequate vocabulary coverage for low-resource languages
- Cross-lingual transfer: Learns language-agnostic representations through span corruption across all 101 languages simultaneously, letting fine-tuning on English data and reasonable evaluation performance on other languages
- Performance tradeoffs: Slightly lower English performance compared to T5 due to the curse of multilinguality, but gains 100 additional languages with strong multilingual and cross-lingual transfer capabilities
The central design philosophy of mT5 is that the challenges of multilingual learning are engineering problems that can be addressed with principled solutions, rather than basic barriers requiring architectural changes. Temperature sampling handles resource imbalance. Vocabulary expansion handles script diversity. The span corruption objective, unchanged from T5, handles representation learning across all languages simultaneously. This philosophy has proven highly influential: the temperature sampling formula and multilingual tokenization strategies pioneered by mT5 appear in nearly every subsequent large-scale multilingual model, from BLOOM to mBART to modern multilingual instruction-tuned systems. The core insight that a single encoder-decoder model can serve as a multilingual foundation for dozens of downstream tasks remains one of the most practically significant contributions of the mT5 paper.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about mT5 and multilingual language modeling.
mT5 Multilingual Modeling Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!