Part of History of Language AI
Covers XLM (Cross-lingual Language Model) introduced by Facebook AI Research in 2019. Explains how cross-lingual pretraining.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
2019: XLM
In 2019, Facebook AI Research introduced XLM (Cross-lingual Language Model), a breakthrough in multilingual natural language processing that demonstrated how cross-lingual pretraining with translation language modeling could enable strong zero-shot and few-shot transfer across languages. XLM's ability to learn cross-lingual representations that captured semantic similarities across different languages opened up new possibilities for multilingual AI applications and influenced the development of many subsequent multilingual language models. The model's success demonstrated that neural language models could be trained to understand and generate text in multiple languages simultaneously, establishing new standards for multilingual NLP and influencing the development of many subsequent systems that could handle multiple languages with a single model.
XLM appeared as the field was moving from monolingual language models to multilingual systems. Models like BERT and GPT had performed well when trained on large monolingual corpora, but extending those results to multiple languages required new methods. Cross-lingual pretraining provided a bridge by transferring knowledge between languages and improving performance on low-resource languages.
The Problem
The traditional approach to multilingual natural language processing had relied on training separate models for each language or using translation-based approaches that required intermediate translation steps. This approach was resource-intensive and often resulted in inconsistent performance across languages, especially for low-resource languages with limited training data. Training separate models for each language meant that knowledge learned in one language couldn't be transferred to another, requiring substantial computational resources and data for each language pair.
The statistical translation-based approaches had additional limitations. They required intermediate translation steps, often involving English as a pivot language, which introduced errors and inefficiencies. A question in Spanish might need to be translated to English first, processed, then translated back to Spanish, with each translation step potentially introducing errors. These approaches struggled to capture cross-lingual semantic relationships and often required significant adaptation for each new language.
Monolingual language models like BERT had achieved remarkable success, but they were language-specific. A BERT model trained on English couldn't understand French or Spanish without retraining, and the learned representations didn't capture relationships between words in different languages. This limitation meant that models needed to be retrained from scratch for each language, wasting computational resources and preventing transfer of knowledge between languages.
For low-resource languages with limited training data, the problem was even more severe. High-resource languages like English had abundant training data, enabling sophisticated language models. Low-resource languages might have only a fraction of this data, making it difficult or impossible to train effective language models. The inability to transfer knowledge from high-resource to low-resource languages meant that the gap between language capabilities would continue to grow.
The field needed a solution that could enable language models to work across multiple languages simultaneously, sharing knowledge between languages while maintaining the performance benefits of large-scale pretraining. This required new training objectives that explicitly encouraged the model to learn cross-lingual representations and new architectures that could handle multilingual data effectively.
The Solution
XLM addressed these limitations by using a single model architecture trained on multilingual data that included text from many different languages. The model used a shared vocabulary and embedding space across all languages, allowing it to learn representations that captured semantic similarities between words and phrases in different languages. The key innovation was the use of translation language modeling, which trained the model to predict words in one language given context in another language, encouraging the model to learn cross-lingual representations.
Shared Multilingual Architecture
The model's architecture was based on the transformer, with shared parameters across all languages. Unlike monolingual models that had separate parameters for each language, XLM used the same transformer layers for all languages, forcing the model to learn representations that worked across linguistic boundaries. This shared architecture meant that the model had to find commonalities between languages, learning abstract representations that captured universal linguistic patterns.
The shared vocabulary included subword tokens that were common across languages, as well as language-specific tokens for words that were unique to particular languages. Byte Pair Encoding (BPE) was applied to the concatenated corpora from all languages, creating a unified vocabulary where frequent subword units were shared across languages. This shared vocabulary enabled the model to recognize that "cat" in English and "chat" in French referred to similar concepts, even though they were spelled differently.
The shared embedding space enabled XLM's cross-lingual transfer. By mapping words from different languages into the same vector space, the model could learn that semantically similar words across languages would be close together in the embedding space. For example, "dog" in English, "chien" in French, and "perro" in Spanish would all map to similar regions of the embedding space. This enabled the model to transfer knowledge learned in one language to another, as the shared representations captured universal semantic relationships.
Translation Language Modeling
The key innovation in XLM was translation language modeling (TLM), a training objective that explicitly encouraged cross-lingual learning. In addition to the standard masked language modeling used in BERT, XLM used parallel text data where the same content was available in multiple languages. The model was trained to predict words in one language given context from both languages, forcing it to learn cross-lingual correspondences.
For example, given a parallel sentence pair "The cat sat on the mat" (English) and "Le chat s'est assis sur le tapis" (French), the model might be asked to predict "chat" in the French sentence given context from both languages. This training objective explicitly taught the model that "cat" and "chat" were related, encouraging the learning of cross-lingual representations that captured semantic similarities.
The training process for XLM involved several key components. First, the model was trained on large amounts of monolingual text from many different languages, learning to predict the next word in each language using causal language modeling. Second, the model was trained on parallel text data using translation language modeling, learning to predict words in one language given context in another. This cross-lingual training encouraged the model to learn representations that captured semantic similarities across languages.
Cross-Lingual Transfer Mechanisms
XLM's architecture enabled several mechanisms for cross-lingual transfer. The shared parameters meant that improvements learned from high-resource languages could benefit low-resource languages. When the model learned to recognize grammatical patterns from English, these patterns could transfer to other languages with similar structures. The shared embedding space allowed the model to map similar concepts across languages, enabling knowledge transfer at the semantic level.
The model also used language embeddings to indicate which language each token belonged to, allowing it to learn language-specific patterns while maintaining cross-lingual representations. These language embeddings enabled the model to adapt its behavior based on the language, while the shared transformer layers ensured that knowledge could be transferred across languages.
Applications and Impact
XLM's success demonstrated several key advantages of cross-lingual pretraining for multilingual NLP. First, the model's ability to learn cross-lingual representations enabled zero-shot transfer, where the model could perform tasks in languages it had never seen during training. For example, a model trained on English and Spanish data could perform question answering in Italian without Italian task-specific training by using representations learned from related languages.
Second, the model performed better on low-resource languages than previous approaches because shared representations let it draw on high-resource language data. A language with only thousands of training examples could benefit from the millions of examples available for high-resource languages. This capability was particularly useful for languages spoken by smaller populations or represented by little digital text.
Third, the model's ability to handle multiple languages with a single architecture made it much more efficient and practical than training separate models for each language. Instead of maintaining dozens of language-specific models, a single XLM model could serve all languages, reducing computational requirements and simplifying deployment.
Cross-Lingual Tasks
The model's cross-lingual capabilities were particularly impressive for tasks that required understanding semantic relationships across languages. XLM could perform cross-lingual information retrieval, where queries in one language could retrieve relevant documents in another language. A search query in English could find relevant documents in French or Spanish, even if those documents didn't contain any of the English query words, by matching based on semantic similarity in the shared embedding space.
XLM could also perform cross-lingual question answering, where questions in one language could be answered using information in another language. A question in French could be answered using English Wikipedia articles, enabling users to access information across language barriers. This supported multilingual applications without a separate retrieval system for every language pair.
Influence on Subsequent Models
XLM's success influenced the development of many subsequent multilingual language models and established new standards for cross-lingual NLP. The model's architecture and training approach became a template for other multilingual language model projects, including mBERT (multilingual BERT), XLM-R (XLM-RoBERTa), and many others. XLM's performance benchmarks became standard evaluation metrics for new multilingual systems, establishing clear targets for cross-lingual performance.
The work also influenced the development of other cross-lingual AI systems that could handle multiple languages with a single model. The principles of shared architectures, unified vocabularies, and cross-lingual training objectives became standard approaches for building multilingual systems. Modern multilingual models like mT5, mBERT, and multilingual versions of GPT all build on the foundations established by XLM.
Open-Source Impact
The model's open-source release made it accessible to researchers and developers worldwide, enabling rapid adoption and further development. The availability of the model weights and training code allowed others to build upon the work and develop specialized versions for specific language pairs or tasks. This open approach accelerated research and development in multilingual NLP and related fields, enabling researchers without access to large computational resources to experiment with multilingual language models.
XLM also demonstrated the value of diverse, high-quality multilingual training data. Data quality and language coverage contributed directly to consistent cross-lingual performance. This result influenced later multilingual models and their data collection and curation practices.
Limitations
Despite its significant contributions, XLM faced several limitations that would shape subsequent research directions. The model's performance varied significantly across language pairs, with stronger performance for languages that were typologically similar or had abundant training data. Languages that were very different from those in the training data, or languages with limited representation in the training corpus, showed weaker cross-lingual transfer.
The model's reliance on parallel text data for translation language modeling was also a limitation. While parallel data enabled strong cross-lingual learning, such data was not available for all language pairs, and creating parallel corpora was expensive and time-consuming. Language pairs without parallel data couldn't benefit from the translation language modeling objective, limiting the model's applicability to language pairs with existing translation resources.
The shared vocabulary approach, while effective for related languages, sometimes struggled with languages that used different writing systems or had very different morphological structures. Languages with non-Latin scripts or complex morphology might not benefit as much from the shared subword vocabulary, limiting the effectiveness of cross-lingual transfer.
The model's performance on low-resource languages, while improved compared to monolingual approaches, still lagged behind high-resource languages. The cross-lingual transfer helped but didn't completely eliminate the gap between languages with abundant training data and those with limited resources. This limitation highlighted the continuing importance of having adequate training data for each language.
The computational requirements for training multilingual models were also substantial. Training on multiple languages required processing much more data than monolingual models, increasing training time and computational costs. The need for parallel text data added complexity to the data preparation process, requiring alignment and preprocessing of multilingual corpora.
Legacy
XLM established cross-lingual pretraining as a fundamental approach for building multilingual language models. This showed that neural language models could learn to understand and generate text in multiple languages simultaneously. The model's innovations, including cross-lingual pretraining, translation language modeling, and shared multilingual representations, established new standards for multilingual NLP that continue to influence the field today.
The impact of XLM extended beyond multilingual NLP to influence how researchers approach language model training more broadly. The model's ability to handle multiple languages and tasks with a single model influenced the development of other multimodal AI systems. The idea of using a single architecture for multiple related tasks became a standard approach in modern AI systems, enabling more efficient training and deployment. This principle influenced the development of many subsequent systems that could handle multiple modalities and tasks.
XLM's varied test sets showed the need to evaluate multilingual models systematically across languages and tasks. This influenced later evaluation frameworks and benchmarks for cross-lingual systems.
Modern Multilingual Models
Modern multilingual models build directly on XLM's foundations while addressing its limitations. Models like XLM-RoBERTa improved on XLM by using larger training corpora and more reliable pretraining objectives. mT5 extended the text-to-text framework to multiple languages. This allowed unified modeling of diverse NLP tasks in languages. These developments have made multilingual language models more capable and accessible, but they all build on the cross-lingual pretraining paradigm that XLM established.
The principles of shared architectures and cross-lingual transfer have become fundamental to how modern language models are built. Today's large language models are typically multilingual, trained on data from many languages simultaneously, and capable of cross-lingual understanding and generation. This multilingual capability is considered a standard feature rather than an optional add-on, thanks in large part to the path that XLM helped to establish.
Long-Term Impact
XLM made cross-lingual pretraining a practical method for building multilingual systems. Subsequent projects adopted its shared representations and translation language modeling objective, patterns that continue to inform multilingual NLP research.
Modern language models routinely include multilingual capabilities that build on XLM's approach. XLM showed that one neural model could learn representations linking semantic relationships across languages. This made it easier to build systems that serve users in one language with information written in another.
XLM was a notable point in the history of multilingual natural language processing and artificial intelligence. This showed that neural language models could be trained to understand and generate text in multiple languages simultaneously. The model's innovations established new standards for multilingual NLP and influenced the development of many subsequent multilingual language models. The work demonstrated the potential for cross-lingual AI systems that could handle multiple languages with a single model, making possible international applications and cross-lingual research that continue to shape the field today.
Quiz
Test your understanding of XLM's cross-lingual pretraining and translation language modeling objective.
XLM Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of History of Language AI. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore History of Language AIStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!