Language Identification: Models, Multilingual Handling

Michael BrenndoerferJanuary 10, 202652 min read

Part of Language AI Handbook

Explains how language identification works in NLP pipelines. Topics include n-gram models, FastText LID, code-switching, confidence thresholds.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Language Identification

Before a language model can learn from a document, it needs to know what language that document is in. This may seem like a trivial prerequisite, but language identification sits at the foundation of nearly every modern NLP pipeline. When you are assembling a training corpus from hundreds of billions of web pages, the question of which documents belong to which language is not obvious: pages mix languages, metadata lies, encodings break, and domain-specific jargon obscures natural language boundaries.

Language identification (LangID) is the task of automatically determining the language of a piece of text. The output is typically an ISO 639 language code such as en for English, de for German, or zh for Chinese. At first glance this looks straightforward: English text looks different from Arabic text, which looks different from Japanese text. But in practice LangID must handle several challenges: multilingual documents, code-switching, very short inputs, rare languages, and adversarial noise from web scraping.

The data curation pipeline for a large language model processes enormous volumes of raw text. Filtering by language is one of the first and most impactful decisions a practitioner makes. Keep the wrong languages and you dilute signal for the languages you care about. Discard too aggressively and you starve low-resource languages of their representation. LangID is the gatekeeper, and its errors compound at scale. Misclassifying one percent of two trillion tokens means twenty billion tokens in the wrong bucket, with measurable downstream consequences for model quality.

It is worth grounding this task in history. Early computational linguists recognized that languages leave statistical traces at the character level as far back as the 1980s, and the first principled LangID systems appeared in the 1990s. The fundamental insight has not changed: every language has a characteristic distribution of character sequences. What has changed is the scale at which LangID is applied, the number of languages covered, and the robustness required for noisy web text.

Why Language Identification is Hard

Language identification looks easy when you have long, well-formed text in a common language. A paragraph of French is unmistakably French. But real-world text rarely cooperates, and production pipelines encounter failure modes that simple demonstrations do not reveal.

Short Text

Many LangID systems degrade significantly on short inputs. A tweet, a product title, a code comment, or a metadata field may be only three to fifteen words long. Statistical models need character n-gram evidence to make confident predictions, and that evidence is sparse for short text. Even a human might struggle: "Hello world!" is technically English, but the greeting "Hallo Welt" in German is nearly identical in information content, and neither phrase contains any language-specific structure beyond convention.

The short-text problem is especially acute for social media. Twitter and similar platforms impose strict length limits, which means that the text units entering a pipeline may be far shorter than the training examples LangID models were built on. A model trained on Wikipedia paragraphs will have seen n-gram distributions calibrated to long, coherent text. A five-word tweet can activate many of the same n-grams regardless of language.

The practical workaround is to impose a minimum document length requirement before language filtering, or to accept that short documents will have lower confidence and filter them separately. We will return to this in the visualizations section.

Code-Switching and Mixed Language

Code-switching occurs when a speaker alternates between two or more languages within a single utterance or across consecutive utterances. This is extremely common in bilingual and multilingual communities, and it is not a defect: it reflects observed patterns of human communication.

Some concrete examples illustrate the challenge:

  • A Spanish-English bilingual might write: "No puedo creer que this actually worked!"
  • A Hindi-English speaker might write: "Kal meeting hai, can you please reschedule?"
  • Social media posts from multilingual users routinely blend languages within a single sentence.

For a document-level LangID model that returns a single label, code-switched text is fundamentally ambiguous. The most common workaround is to identify the "dominant" language, but this loses information about multilingual content. A data pipeline that routes a Spanish-English code-switched tweet to the Spanish bucket will include text that contains substantial English, which may or may not be desirable depending on the target use case.

There is also intra-sentential code-switching, where language boundaries fall within a single sentence, and inter-sentential code-switching, where whole sentences alternate. Document-level models handle the latter better, since at least some sentences are monolingual, while intra-sentential mixing is difficult even for sentence-level systems.

Script Ambiguity and Shared Vocabulary

Many languages share a writing system. Latin script is used for hundreds of languages, from English to Swahili to Vietnamese. Short inputs in closely related languages (Portuguese vs. Spanish, Norwegian vs. Danish, Hindi vs. Nepali in Devanagari script) share large portions of their vocabulary, making them difficult to distinguish. The word "hotel" or "taxi" appears in dozens of languages with identical or near-identical form.

The problem is sharpest for language clusters with high mutual intelligibility. Malay and Indonesian are both written in Latin script and are so closely related that they are sometimes treated as registers of the same language. Separating them requires identifying specific function words, orthographic conventions, and vocabulary items that exist in one but not the other. Similarly, Bosnian, Croatian, and Serbian are so similar in written form that distinguishing them may be infeasible for short texts.

On the opposite end, some languages that appear visually similar to non-speakers are quite distinct. Japanese uses multiple scripts simultaneously (hiragana, katakana, and kanji), while Chinese uses only characters. A LangID system can exploit these script mixing patterns as strong signals.

Rare and Low-Resource Languages

Most LangID systems support a few hundred languages, but the web contains text in thousands. A language with few training examples will either be misclassified as a related language or simply absent from the model's vocabulary, resulting in its text being assigned to whatever language the model finds most plausible. This is not a minor issue: the languages most underrepresented in LangID systems are precisely those most vulnerable to being discarded from training corpora, which in turn means language models perform poorly on them, perpetuating a cycle of underrepresentation.

The cycle matters because language model capabilities are often evaluated on benchmark tasks that exist only in English and a handful of other high-resource languages. Low-resource languages that never make it into training data also never develop benchmark tasks, so the deficiency remains invisible in standard evaluations.

ISO 639 Language Codes

ISO 639 is an international standard for language codes. ISO 639-1 provides two-letter codes for 184 languages (e.g., en, fr, de). ISO 639-3 extends this to over 7,000 languages using three-letter codes (e.g., eng, fra, deu). LangID systems often use ISO 639-1 for common languages and ISO 639-3 for less common ones. When comparing LangID tool outputs, confirm which standard is being used, since the same language may appear as zh in one system and zho in another.

Web Noise and Encoding Issues

Web-crawled text introduces LangID challenges beyond linguistics. A page might contain HTML entities that obscure language-specific characters (é instead of é). Encoding errors may corrupt diacritics, turning German ü into a byte sequence that no LangID model recognizes. Scraped text may include navigation menus, footer text, or boilerplate that belongs to a different language than the article content. A French news site may have English-language cookie banners, English SEO metadata, and French article text all in a single crawled document.

Production pipelines often apply text extraction and cleaning before LangID to remove HTML artifacts, boilerplate, and encoding noise. The quality of this preprocessing directly affects LangID accuracy.

Classical Approaches: N-gram Language Models

Before neural methods, the dominant approach to LangID was character-level n-gram language models. The intuition is clean: every language has a characteristic distribution of character sequences. The trigram sch is much more common in German than in English. The sequence zione is quintessentially Italian. The bigram th is a hallmark of English. The character sequence ão signals Portuguese. These distributions are stable, compact, and highly discriminative even on modest amounts of training data.

The Mathematical Foundation

For each candidate language LL in a set of languages L\mathcal{L}, we train a character-level n-gram language model on a corpus of text in that language. Given a query document dd, we compute the log-probability of the document under each language model and classify by maximum likelihood:

L^=arg⁡max⁡L∈Llog⁡PL(d)\hat{L} = \arg\max_{L \in \mathcal{L}} \log P_L(d)

where:

  • L^\hat{L}: the predicted language label
  • L\mathcal{L}: the set of all candidate languages
  • PL(d)P_L(d): the probability assigned to document dd by the language model for language LL
  • log⁡PL(d)\log P_L(d): the log-probability, computed as the sum of log-probabilities of each n-gram in dd

The log-probability of a document decomposes over its character n-grams using the chain rule:

log⁡PL(d)=∑i=1∣d∣log⁡PL(ci∣ci−n+1,…,ci−1)\log P_L(d) = \sum_{i=1}^{|d|} \log P_L(c_i \mid c_{i-n+1}, \ldots, c_{i-1})

where:

  • cic_i: the ii-th character in the document
  • PL(ci∣ci−n+1,…,ci−1)P_L(c_i \mid c_{i-n+1}, \ldots, c_{i-1}): the conditional probability of character cic_i given the preceding n−1n-1 characters, according to language model LL

Each language model is a lookup table of n-gram counts, normalized by the total count of all trigrams sharing the same two-character prefix. Smoothing is essential: unseen n-grams would receive zero probability under a maximum likelihood model, making the log-probability −∞-\infty for any document containing an unseen trigram. Laplace (add-one) smoothing assigns a small non-zero probability to every n-gram, even those not observed in training.

The choice of nn matters. Unigrams (single characters) carry almost no language-specific information beyond script identity. Bigrams start to capture language-specific patterns but are still quite general. Trigrams represent the sweet spot for most languages: they are specific enough to capture morphological signatures but common enough to appear frequently in training data. Quadgrams and longer n-grams are more discriminative but require larger training corpora to estimate reliably.

Normalization by document length is necessary when comparing documents of different lengths. A long document will accumulate a lower total log-probability simply because it is longer, not because it is less likely under any particular language model. Dividing by the number of n-grams converts absolute log-probability to average per-n-gram log-probability, making scores comparable across document lengths.

Caveat: Rank-Based Methods

An influential alternative, introduced by Cavnar and Trenkle in 1994, avoids probability estimation entirely. Instead, it creates a ranked list of the 300 most frequent n-grams for each language. To classify a document, the algorithm computes the ranked n-gram list for the document and compares it to each language profile using a rank-distance metric called the "out-of-place" measure:

Δ(d,L)=∑g∈topk(d)∣rankd(g)−rankL(g)∣\Delta(d, L) = \sum_{g \in \text{top}_k(d)} \left| \text{rank}_d(g) - \text{rank}_L(g) \right|

where:

  • gg: a character n-gram from the document's top-kk most frequent n-grams
  • rankd(g)\text{rank}_d(g): the rank of n-gram gg in the document's frequency list
  • rankL(g)\text{rank}_L(g): the rank of n-gram gg in language LL's profile (set to a large penalty value if gg does not appear in the profile)

The language with the smallest out-of-place distance wins the classification. This approach is robust to short text because it depends on relative frequencies rather than absolute counts, and it does not require careful probability smoothing. The approach remains competitive for clean, well-formed text in common languages.

Why Character N-grams Beat Word-Level Models for LangID

A word-level language model for LangID would assign probabilities based on which words appear in the document. The problem is that word-level models require a fixed vocabulary, and any word outside the vocabulary contributes nothing to classification. For a new document in a rare language, or for a document containing technical vocabulary, many words may be unknown.

Character n-grams sidestep this entirely. Every character sequence of length nn is a valid n-gram. A medical article in German will contain medical vocabulary that does not appear in a general German training corpus, but it will still contain German morphological patterns (the suffixes -ung, -keit, -lich, characteristic prepositions), function words spelled out character by character, and orthographic conventions like the use of ß. These signals persist regardless of domain shift.

Character n-gram models also handle agglutinative languages well. Finnish, Turkish, Hungarian, and many other languages form very long words by combining morphemes, meaning that word-level models face severe vocabulary explosion while character n-gram models see stable distributions.

Modern Approaches: FastText-Based Models

The current standard for production LangID is the fastText library developed by Meta AI Research, with the LID.176 model covering 176 languages. Alongside this, tools like langdetect, lingua, and Google's Compact Language Detector (CLD2/CLD3) are widely used. Each makes different design decisions around speed, accuracy, language coverage, and handling of short text and multilingual content.

Why FastText Works Well for LangID

FastText represents each document as the average of its word and character n-gram embeddings. This is different from the probabilistic n-gram language model approach: FastText is a discriminative classifier, not a generative language model. It learns to separate languages from each other rather than modeling the probability distribution of each language independently.

For LangID specifically, the character n-gram component of FastText is particularly important. It captures morphological and orthographic patterns that are language-specific, even for out-of-vocabulary words. A word like "nieprzekraczalnych" (Polish for "impassable" in genitive plural) contains n-gram patterns that are unmistakably Polish, even if that specific word never appeared in training.

The discriminative training objective sharpens language boundaries. Consider two closely related languages like Portuguese and Spanish. A generative n-gram model trained on Portuguese will assign reasonable probability to Spanish text because the two languages overlap heavily. A discriminative classifier trained on the task of distinguishing them will learn the specific features that differentiate them, giving it a sharper decision boundary.

FastText's LID.176 model is trained on text from Wikipedia, Tatoeba, and SETimes in 176 languages. The model takes as input a short text string and produces a probability distribution over all 176 languages using softmax activation. The predicted label is the language with highest probability, and the probability value itself is a confidence score. Training is fast, inference is extremely fast (millions of documents per second on a single CPU core), and the binary model file is only about 130 MB.

CLD2 and CLD3

Google's Compact Language Detector (CLD2, later CLD3) is another widely used system, specifically designed for web content. CLD2 is a Naïve Bayes classifier over character quadgrams with additional heuristics for scripts, character encoding, and HTML structure. CLD3 replaces the Naïve Bayes classifier with a small neural network that processes character n-grams.

CLD2 and CLD3 differ from FastText in one critical respect: they explicitly handle multilingual documents. Rather than returning a single language label, they can return a ranked list of detected languages with their estimated byte fractions, which is valuable when a page contains substantial content in multiple languages. For a data curation pipeline that wants to extract only the target-language content from mixed pages, this per-language byte fraction is directly actionable.

CLD2 also has well-calibrated confidence scores, a minimum confidence feature that suppresses unreliable predictions, and an explicit fallback to "unknown" when the input is too short or ambiguous. These properties make it suitable for pipelines that must handle the full diversity of web content.

lingua

lingua is a Python library that takes a different design philosophy: precision over coverage and speed. Where FastText covers 176 languages with high throughput but moderate accuracy on very short text, lingua focuses on fewer languages (75 as of recent versions) and uses multiple statistical models, rule-based heuristics, and language-specific knowledge to achieve higher precision on short inputs.

lingua applies a cascade of detection methods. For very short text (one to three words), it uses a list of unambiguous language-specific words. For longer text, it uses trigram and quadgram statistics. It also applies a "confidence threshold" that prevents it from making a guess below a minimum confidence level, returning an "unknown" prediction instead of a low-confidence label. This is useful in pipelines where a conservative non-prediction is better than an incorrect prediction.

The tradeoff is speed and coverage. lingua is slower than FastText by a significant margin, and it does not support most of the languages FastText covers. For a pipeline focused on a small set of languages where accuracy on short text matters more than throughput, lingua is a strong choice. For large-scale web filtering across many languages, FastText is the more practical option.

openlid and Language Coverage

A notable recent development is OpenLID (Burchell et al., 2023), a LangID system trained on data from the GlotLID project covering over 1,600 languages. As research attention has turned to low-resource and endangered languages, the community has recognized that 176-language coverage, while impressive, excludes the vast majority of the world's languages. OpenLID addresses this gap, though its accuracy on very low-resource languages is necessarily limited by the small amount of available training data.

Implementation: Building a Language Detection Pipeline

Let's build a language detection pipeline that demonstrates the key techniques from both classical and modern approaches.

Setup

First, install and import the required libraries. We will use fasttext for the production-grade model and implement a classical character n-gram approach from scratch to build intuition before relying on a black-box tool.

In[3]:
Code
import subprocess
import sys

# Install required packages
subprocess.run(
    [
        sys.executable,
        "-m",
        "pip",
        "install",
        "fasttext-wheel",
        "langdetect",
        "lingua-language-detector",
        "--quiet",
    ],
    capture_output=True,
)

Building a Character N-gram Language Model from Scratch

Before using a pre-built library, let's implement a character trigram language model for language identification. This gives intuition for what production systems do under the hood.

We will use short representative sentences to build profiles for five languages: English, French, German, Spanish, and Italian. The training sentences are intentionally in the NLP domain so they share vocabulary, making the classification task harder and more representative of real-world scenarios.

In[5]:
Code
# Small training corpora for five languages
TRAINING_DATA = {
    "en": [
        "the quick brown fox jumps over the lazy dog",
        "language identification is a fundamental task in natural language processing",
        "machine learning models can process text in many different languages",
        "the algorithm analyzes character sequences to determine the source language",
        "natural language processing has advanced significantly in recent years",
        "deep learning has transformed how we approach language understanding",
        "the model learns statistical patterns from large text corpora",
        "word embeddings capture semantic relationships between vocabulary items",
    ],
    "fr": [
        "le rapide renard brun saute par dessus le chien paresseux",
        "l identification de langue est une tache fondamentale en traitement automatique",
        "les modeles d apprentissage automatique peuvent traiter du texte en plusieurs langues",
        "l algorithme analyse les sequences de caracteres pour determiner la langue source",
        "le traitement automatique des langues a considerablement progresse ces dernieres annees",
        "l apprentissage profond a transforme notre approche de la comprehension du langage",
        "le modele apprend des motifs statistiques a partir de grands corpus de textes",
        "les plongements de mots capturent les relations semantiques entre les mots du vocabulaire",
    ],
    "de": [
        "der schnelle braune fuchs springt ueber den faulen hund",
        "sprachidentifikation ist eine grundlegende aufgabe in der verarbeitung natuerlicher sprache",
        "maschinelle lernmodelle koennen text in vielen verschiedenen sprachen verarbeiten",
        "der algorithmus analysiert zeichenfolgen um die quellsprache zu bestimmen",
        "die verarbeitung natuerlicher sprache hat sich in den letzten jahren erheblich weiterentwickelt",
        "tiefes lernen hat unseren ansatz zum sprachverstaendnis veraendert",
        "das modell lernt statistische muster aus grossen textkorpora",
        "worteinbettungen erfassen semantische beziehungen zwischen vokabulareintraegen",
    ],
    "es": [
        "el rapido zorro marron salta sobre el perro perezoso",
        "la identificacion de idiomas es una tarea fundamental en el procesamiento del lenguaje natural",
        "los modelos de aprendizaje automatico pueden procesar texto en muchos idiomas diferentes",
        "el algoritmo analiza secuencias de caracteres para determinar el idioma fuente",
        "el procesamiento del lenguaje natural ha avanzado significativamente en los ultimos anos",
        "el aprendizaje profundo ha transformado como abordamos la comprension del lenguaje",
        "el modelo aprende patrones estadisticos de grandes corpus de textos",
        "las incrustaciones de palabras capturan relaciones semanticas entre elementos del vocabulario",
    ],
    "it": [
        "la veloce volpe marrone salta sopra il cane pigro",
        "l identificazione della lingua e un compito fondamentale nel trattamento del linguaggio naturale",
        "i modelli di apprendimento automatico possono elaborare testo in molte lingue diverse",
        "l algoritmo analizza sequenze di caratteri per determinare la lingua sorgente",
        "l elaborazione del linguaggio naturale ha fatto notevoli progressi negli ultimi anni",
        "l apprendimento profondo ha trasformato il nostro approccio alla comprensione del linguaggio",
        "il modello apprende schemi statistici da grandi corpus di testi",
        "le incorporazioni di parole catturano le relazioni semantiche tra gli elementi del vocabolario",
    ],
}

Now we implement the n-gram language model. For each language, we count all character trigrams and compute Laplace-smoothed probabilities. The smoothing parameter adds 1 to every trigram count before normalization, which ensures that every possible trigram receives a non-zero probability and prevents log-probability from being undefined for unseen n-grams.

In[6]:
Code
import math
from collections import Counter


def get_char_ngrams(text: str, n: int = 3) -> list[str]:
    """Extract all character n-grams from text."""
    text = text.lower()
    # Add padding characters to handle word boundaries
    padded = f"  {text}  "
    return [padded[i : i + n] for i in range(len(padded) - n + 1)]


def build_ngram_model(sentences: list[str], n: int = 3) -> dict[str, float]:
    """Build a smoothed character n-gram language model."""
    counts: Counter = Counter()
    prefix_counts: Counter = Counter()

    for sentence in sentences:
        ngrams = get_char_ngrams(sentence, n)
        for gram in ngrams:
            counts[gram] += 1
            prefix_counts[gram[:-1]] += 1

    # Laplace smoothed probabilities
    vocab_size = len(counts)
    log_probs = {}
    for gram, count in counts.items():
        prefix = gram[:-1]
        log_probs[gram] = math.log(
            (count + 1) / (prefix_counts[prefix] + vocab_size)
        )

    # Store the smoothing parameters for unseen n-grams
    log_probs["__unk_log_prob__"] = math.log(
        1 / (sum(counts.values()) + vocab_size)
    )

    return log_probs


def compute_log_prob(text: str, model: dict[str, float], n: int = 3) -> float:
    """Compute the log-probability of text under a language model."""
    ngrams = get_char_ngrams(text, n)
    unk_log_prob = model.get("__unk_log_prob__", -10.0)
    total = sum(model.get(gram, unk_log_prob) for gram in ngrams)
    # Normalize by document length to avoid length bias
    return total / len(ngrams) if ngrams else float("-inf")


def predict_language(
    text: str, models: dict[str, dict[str, float]]
) -> tuple[str, dict[str, float]]:
    """Predict language using the n-gram models. Returns predicted language and all scores."""
    scores = {
        lang: compute_log_prob(text, model) for lang, model in models.items()
    }
    predicted = max(scores, key=scores.__getitem__)
    return predicted, scores


# Build models for all five languages
ngram_models = {
    lang: build_ngram_model(sentences)
    for lang, sentences in TRAINING_DATA.items()
}
In[7]:
Code
# Test the model on held-out sentences
TEST_SENTENCES = [
    ("en", "the neural network was trained on a large dataset"),
    (
        "fr",
        "le reseau de neurones a ete entraine sur un grand ensemble de donnees",
    ),
    ("de", "das neuronale netz wurde auf einem grossen datensatz trainiert"),
    ("es", "la red neuronal fue entrenada en un gran conjunto de datos"),
    ("it", "la rete neurale e stata addestrata su un grande insieme di dati"),
    ("en", "hello world"),  # Short text challenge
    ("fr", "bonjour monde"),
    ("de", "hallo welt"),
]

results = []
for true_lang, sentence in TEST_SENTENCES:
    predicted, scores = predict_language(sentence, ngram_models)
    correct = predicted == true_lang
    results.append(
        {
            "sentence": sentence[:45] + "..."
            if len(sentence) > 45
            else sentence,
            "true": true_lang,
            "predicted": predicted,
            "correct": correct,
            "confidence": scores[predicted],
        }
    )

n_correct = sum(r["correct"] for r in results)
Out[8]:
Console
Character Trigram LangID Results (7/8 correct)

Sentence                                          True  Pred    Score
------------------------------------------------------------------------
the neural network was trained on a large dat...    en    en   -5.899  OK
le reseau de neurones a ete entraine sur un g...    fr    fr   -5.406  OK
das neuronale netz wurde auf einem grossen da...    de    de   -5.904  OK
la red neuronal fue entrenada en un gran conj...    es    es   -5.768  OK
la rete neurale e stata addestrata su un gran...    it    it   -5.601  OK
hello world                                         en    de   -6.317  WRONG
bonjour monde                                       fr    fr   -5.941  OK
hallo welt                                          de    de   -6.157  OK

The character trigram model gets most predictions right on full sentences. Notice that short inputs like "hello world" are harder: the model has less n-gram evidence to work with, and common short words may overlap between related languages. The score column shows the average log-probability per trigram, which is negative because all probabilities are less than one; a value closer to zero means the document fits the language model better.

Using fastText for Production LangID

For real data curation pipelines, we use FastText's LID.176 model. This model was trained discriminatively on text from Wikipedia, Tatoeba, and SETimes across 176 languages. It runs at millions of documents per second on a single CPU core and produces well-calibrated probability scores.

In[9]:
Code
import os
import urllib.request

# Download FastText language identification model if not present
model_path = "/tmp/lid.176.bin"
model_url = (
    "https://dl.fbaipublicfiles.com/fasttext/supervised-models/lid.176.bin"
)

if not os.path.exists(model_path):
    print("Downloading FastText LID model (126 MB)...")
    urllib.request.urlretrieve(model_url, model_path)
    print("Downloaded successfully.")
In[10]:
Code
try:
    import fasttext

    ft_model = fasttext.load_model(model_path)
    FASTTEXT_AVAILABLE = True
except Exception:
    FASTTEXT_AVAILABLE = False
In[11]:
Code
def detect_language_fasttext(
    text: str, model, top_k: int = 3
) -> list[tuple[str, float]]:
    """Detect language using FastText. Returns top_k (language, probability) pairs."""
    # FastText expects single-line input; use internal f.predict to avoid NumPy 2.x
    # compatibility issue with np.array(probs, copy=False) in fasttext's Python wrapper.
    text_nl = text.replace("\n", " ").strip() + "\n"
    predictions = model.f.predict(text_nl, top_k, 0.0, "strict")
    # predictions is a list of (prob, label) tuples
    result = [
        (label.replace("__label__", ""), prob) for prob, label in predictions
    ]
    return result


# Test documents including a code-switched example
test_documents = [
    "The attention mechanism allows the model to focus on relevant parts of the input sequence.",
    "Le mecanisme d attention permet au modele de se concentrer sur les parties pertinentes de l entree.",
    "Der Aufmerksamkeitsmechanismus ermoglicht es dem Modell, sich auf relevante Teile der Eingabe zu konzentrieren.",
    "No puedo creer que this actually worked, increible!",  # Code-switching
    "안녕하세요, 오늘 날씨가 정말 좋네요.",  # Korean
    "这是一个关于自然语言处理的例子。",  # Chinese
    "مرحبا بالعالم",  # Arabic
    "Hi",  # Ambiguous short text
]

if FASTTEXT_AVAILABLE:
    detection_results = []
    for doc in test_documents:
        results_ft = detect_language_fasttext(doc, ft_model, top_k=3)
        detection_results.append(
            {
                "doc": doc[:50] + "..." if len(doc) > 50 else doc,
                "results": results_ft,
            }
        )
Out[12]:
Console
FastText Language Detection Results

Text: The attention mechanism allows the model to focus ...
   en: 0.865 |=========================
   hi: 0.007 |
   fa: 0.005 |

Text: Le mecanisme d attention permet au modele de se co...
   fr: 0.730 |=====================
   ca: 0.057 |=
   es: 0.035 |=

Text: Der Aufmerksamkeitsmechanismus ermoglicht es dem M...
   de: 0.994 |=============================
   es: 0.001 |
   en: 0.001 |

Text: No puedo creer que this actually worked, increible...
   es: 0.770 |=======================
   en: 0.115 |===
   ca: 0.018 |

Text: 안녕하세요, 오늘 날씨가 정말 좋네요.
   ko: 1.000 |=============================
   zh: 0.000 |
   en: 0.000 |

Text: 这是一个关于自然语言处理的例子。
   zh: 1.000 |==============================
  wuu: 0.000 |
   ja: 0.000 |

Text: مرحبا بالعالم
   ar: 0.922 |===========================
   fa: 0.046 |=
  arz: 0.014 |

Text: Hi
   de: 0.516 |===============
   en: 0.319 |=========
   ga: 0.033 |

The FastText results reveal several important behaviors. High-confidence predictions (close to 1.0) appear for full sentences in distinct scripts or language families. The code-switched Spanish-English example shows Spanish as the dominant language but with lower confidence, showing the mixed content. The single word "Hi" produces uncertain predictions, with probability spread across multiple languages: this is not a model failure but an accurate reflection of real ambiguity.

Notice how the top-3 predictions for ambiguous inputs are informative. For "Hi", seeing English, German, and Dutch as the top candidates makes linguistic sense: "Hi" appears naturally in all three languages. A pipeline can use this distribution to decide whether to accept, reject, or queue for manual review.

Implementing a Language Filter for Data Curation

In a real data curation pipeline, LangID is used to filter documents by target language. We need a filter that supports configurable confidence thresholds and tracks statistics about rejection reasons. Understanding why documents are rejected (wrong language vs. insufficient confidence) helps tune the pipeline.

In[13]:
Code
class LanguageFilter:
    """Filter documents by detected language with configurable confidence threshold."""

    def __init__(
        self,
        target_languages: list[str],
        min_confidence: float = 0.8,
        model=None,
    ):
        self.target_languages = set(target_languages)
        self.min_confidence = min_confidence
        self.model = model
        self.stats = {
            "total": 0,
            "accepted": 0,
            "rejected_language": 0,
            "rejected_confidence": 0,
        }

    def filter_document(self, text: str) -> tuple[bool, str, float]:
        """
        Check whether a document passes the language filter.
        Returns (passes, detected_language, confidence).
        """
        self.stats["total"] += 1

        if self.model is not None and FASTTEXT_AVAILABLE:
            results_ft = detect_language_fasttext(text, self.model, top_k=1)
            detected_lang, confidence = results_ft[0]
        else:
            # Fall back to n-gram model (covers only en/fr/de/es/it)
            predicted, scores = predict_language(text, ngram_models)
            detected_lang = predicted
            sorted_scores = sorted(scores.values(), reverse=True)
            if len(sorted_scores) >= 2:
                confidence = math.exp(sorted_scores[0] - sorted_scores[1])
                confidence = min(confidence, 1.0)
            else:
                confidence = 1.0

        if detected_lang not in self.target_languages:
            self.stats["rejected_language"] += 1
            return False, detected_lang, confidence

        if confidence < self.min_confidence:
            self.stats["rejected_confidence"] += 1
            return False, detected_lang, confidence

        self.stats["accepted"] += 1
        return True, detected_lang, confidence

    def filter_corpus(self, documents: list[str]) -> list[str]:
        """Filter a list of documents, returning only those that pass."""
        return [doc for doc in documents if self.filter_document(doc)[0]]

    def report(self) -> dict:
        """Return filtering statistics."""
        total = self.stats["total"]
        return {
            "total": total,
            "accepted": self.stats["accepted"],
            "rejected_language": self.stats["rejected_language"],
            "rejected_confidence": self.stats["rejected_confidence"],
            "acceptance_rate": self.stats["accepted"] / total
            if total > 0
            else 0.0,
        }
In[14]:
Code
# Simulate a mixed-language corpus
mixed_corpus = [
    "The transformer architecture uses self-attention to process sequences in parallel.",
    "L architecture transformer utilise l auto-attention pour traiter les sequences en parallele.",
    "Die Transformer-Architektur verwendet Selbst-Aufmerksamkeit zur parallelen Sequenzverarbeitung.",
    "La arquitectura transformer usa auto-atencion para procesar secuencias en paralelo.",
    "Tokenization splits text into subword units before model processing.",
    "La tokenisation divise le texte en unites de sous-mots avant le traitement par le modele.",
    "For more information, visit our website.",
    "Besuchen Sie unsere Website fur weitere Informationen.",
    "Hi",  # Short ambiguous text
    "ok",  # Very short
    "Natural language processing has many practical applications.",
    "La langue naturelle est complexe et riche en nuances semantiques.",
    "Diese Methode verbessert die Sprachverarbeitung erheblich.",
]

# Filter for English only with 70% confidence threshold
en_filter = LanguageFilter(
    target_languages=["en"],
    min_confidence=0.7,
    model=ft_model if FASTTEXT_AVAILABLE else None,
)

accepted_docs = en_filter.filter_corpus(mixed_corpus)
filter_report = en_filter.report()
Out[15]:
Console
Language Filter Report
  Total documents:      13
  Accepted (English):   5
  Rejected (language):  8
  Rejected (confidence):0
  Acceptance rate:      38.5%

Accepted documents:
  - The transformer architecture uses self-attention to process sequences 
  - Tokenization splits text into subword units before model processing.
  - For more information, visit our website.
  - ok
  - Natural language processing has many practical applications.

The filter correctly accepts English documents and rejects non-English content. Documents rejected for confidence rather than wrong language are often the short, ambiguous inputs like "Hi" and "ok" where even correct English text falls below the confidence threshold. Whether to keep or discard these depends on your pipeline goals: for a pretraining corpus where quality matters, discarding low-confidence short inputs is sensible; for a dataset focused on short-form content, you might apply lower thresholds or bypass LangID for very short documents entirely.

The rejection statistics are informative for pipeline tuning. A high ratio of "rejected for confidence" to "rejected for language" suggests that your minimum confidence threshold may be too aggressive, or that your corpus contains many short documents. Monitoring these ratios in production helps detect drift.

Visualizations

Let's visualize key aspects of language identification: how character trigram distributions differ across languages, how confidence varies with document length, and the tradeoff between coverage and precision at different confidence thresholds.

Character N-gram Profiles

Out[16]:
Visualization
Horizontal bar charts showing top 20 character trigram frequencies for five European languages side by side.
Top 20 character trigrams by frequency for English, French, German, Spanish, and Italian in the chapter's small training corpus. Each panel shows the complete ranked trigram list and its relative frequencies. The distributions differ across languages, providing the statistical fingerprint that character n-gram LangID systems use for discrimination.

The trigram profiles make the LangID signal visually tangible. Each language has a different ranked pattern, but the individual entries also reflect the vocabulary and subject matter in this deliberately small training corpus. A production system therefore learns from far more varied text and compares the complete distribution across hundreds or thousands of trigrams rather than relying on a few prominent fragments.

Confidence vs. Document Length

A critical property of LangID systems is that short-text scores are often unstable. The raw score margin used by this toy model is not a calibrated probability, so plotting it by document length also reveals where a seemingly intuitive confidence proxy can mislead.

In[17]:
Code
# Generate test sentences of varying lengths by taking prefixes of longer texts
full_sentences = {
    "en": "the transformer model processes sequences by computing attention weights between all pairs of tokens in the input enabling parallel computation and long range dependencies",
    "fr": "le modele transformateur traite les sequences en calculant des poids d attention entre toutes les paires de tokens dans l entree permettant un calcul parallele et des dependances a longue portee",
    "de": "das transformer modell verarbeitet sequenzen durch berechnung von aufmerksamkeitsgewichten zwischen allen tokenpaaren in der eingabe was parallele berechnung und weitreichende abhaengigkeiten ermoeglicht",
}

# Test at different word counts
word_counts = [1, 2, 3, 5, 7, 10, 15, 20, 25, 30]
confidence_by_length = {lang: [] for lang in full_sentences}

for lang, full_text in full_sentences.items():
    words = full_text.split()
    for wc in word_counts:
        snippet = " ".join(words[: min(wc, len(words))])
        predicted, scores = predict_language(snippet, ngram_models)
        sorted_scores = sorted(scores.values(), reverse=True)
        if len(sorted_scores) >= 2:
            # Margin between top-1 and top-2 as a proxy for confidence
            margin = sorted_scores[0] - sorted_scores[1]
            conf = min(
                1.0, max(0.0, (margin + 2.0) / 4.0)
            )  # normalize to [0,1]
        else:
            conf = 1.0
        confidence_by_length[lang].append(conf)
Out[18]:
Visualization
Line chart showing irregular normalized LangID score margins by document word count for English, French, and German, with a dashed reference line at 0.8.
Normalized top-two LangID score margin as a function of document length for English, French, and German. The curves are irregular rather than monotonic: the one-word English prefix produces an anomalously high margin, while longer prefixes settle into language-specific bands. This demonstrates that a raw score margin is not calibrated confidence. The dashed line at 0.8 provides a reference threshold.

The plot shows why the score margin should not be interpreted directly as confidence. The model averages log probabilities by the number of trigrams, so longer prefixes do not automatically accumulate a larger margin. A short but distinctive fragment can appear spuriously decisive, while longer prefixes converge toward different language-specific levels.

A minimum document length can still be a useful operational rule, but it must be validated against a labeled set. Production pipelines typically calibrate model scores and measure accuracy by length bucket before choosing separate thresholds for short and long documents.

Confusion Matrix Across Languages

Let's evaluate the n-gram model on all five languages and visualize the confusion matrix to understand which languages are most often confused with each other.

In[19]:
Code
# Build held-out test set: take one sentence from each language not in training
HELD_OUT = {
    "en": [
        "the model generates text by predicting the next token at each step",
        "pretraining on large corpora gives models broad language understanding",
        "fine tuning adapts the pretrained model to specific downstream tasks",
        "the vocabulary is built using byte pair encoding to handle rare words",
    ],
    "fr": [
        "le modele genere du texte en predisant le prochain token a chaque etape",
        "le preentrainement sur de grands corpus donne aux modeles une large comprehension du langage",
        "l ajustement fin adapte le modele preentraine a des taches specifiques en aval",
        "le vocabulaire est construit en utilisant le codage par paires d octets pour gerer les mots rares",
    ],
    "de": [
        "das modell generiert text indem es bei jedem schritt das naechste token vorhersagt",
        "das vortraining auf grossen korpora gibt modellen ein breites sprachverstaendnis",
        "die feinabstimmung passt das vortrainierte modell an spezifische nachgelagerte aufgaben an",
        "das vokabular wird mit byte pair encoding erstellt um seltene woerter zu behandeln",
    ],
    "es": [
        "el modelo genera texto prediciendo el siguiente token en cada paso",
        "el preentrenamiento en grandes corpus le da a los modelos una amplia comprension del lenguaje",
        "el ajuste fino adapta el modelo preentrenado a tareas especificas posteriores",
        "el vocabulario se construye utilizando codificacion de pares de bytes para manejar palabras raras",
    ],
    "it": [
        "il modello genera testo prevedendo il prossimo token ad ogni passo",
        "il preaddestramento su grandi corpus fornisce ai modelli una comprensione ampia del linguaggio",
        "la messa a punto adatta il modello preaddestrato a compiti specifici a valle",
        "il vocabolario viene costruito utilizzando la codifica a coppie di byte per gestire le parole rare",
    ],
}

LANG_LIST = ["en", "fr", "de", "es", "it"]
confusion = np.zeros((5, 5), dtype=int)

for i, true_lang in enumerate(LANG_LIST):
    for sentence in HELD_OUT[true_lang]:
        predicted, _ = predict_language(sentence, ngram_models)
        j = LANG_LIST.index(predicted)
        confusion[i][j] += 1

# Accuracy per language
per_lang_acc = [
    confusion[i][i] / confusion[i].sum() if confusion[i].sum() > 0 else 0.0
    for i in range(5)
]
overall_acc = confusion.diagonal().sum() / confusion.sum()
Out[20]:
Visualization
Five-language confusion-matrix heatmap with values of 1.00 on every diagonal cell, 0.00 off the diagonal, and overall accuracy of 100 percent on 20 held-out sentences.
Confusion matrix for character trigram language identification across five European languages on 20 held-out test sentences. Rows represent the true language and columns the predicted language. All four examples per language are classified correctly, producing a 100% diagonal matrix. The result is a sanity check for this small sample, not evidence of production-level accuracy.
Out[21]:
Console
Overall accuracy: 100.0%

Per-language accuracy:
  en: 100.0%
  fr: 100.0%
  de: 100.0%
  es: 100.0%
  it: 100.0%

The matrix is perfectly diagonal because this held-out set contains only four clean, topic-matched sentences per language. That makes it useful as an execution sanity check, but far too small and homogeneous for estimating real error rates or identifying the hardest language pairs.

A meaningful evaluation would include many more documents, varied domains, short and noisy inputs, and closely related language varieties. Larger benchmarks often expose weaker separation between related languages, but this particular matrix does not show such errors and should not be used to claim that it does.

Confidence Threshold vs. Precision and Coverage

In data curation, you must choose a confidence threshold that balances two competing goals: maximizing the precision of your language-filtered corpus (keeping only correctly identified documents) while maximizing coverage (retaining as many documents as possible). A higher threshold gives you more reliable language predictions but discards more documents. Let's visualize this tradeoff directly.

In[22]:
Code
# Simulate scores for a large set of documents
# We model two populations: correct predictions (high confidence) and wrong predictions (lower confidence)
np.random.seed(42)
n_docs = 2000

# True positive scores: correctly classified docs tend to have high confidence
tp_scores = np.clip(np.random.beta(8, 2, size=int(n_docs * 0.85)), 0.01, 0.999)
# False positive scores: incorrectly classified docs tend to have lower but sometimes high confidence
fp_scores = np.clip(np.random.beta(3, 5, size=int(n_docs * 0.15)), 0.01, 0.999)

all_scores = np.concatenate([tp_scores, fp_scores])
all_labels = np.array([True] * len(tp_scores) + [False] * len(fp_scores))

thresholds = np.linspace(0.3, 0.99, 50)
precision_vals = []
coverage_vals = []

for thresh in thresholds:
    accepted = all_scores >= thresh
    if accepted.sum() == 0:
        precision_vals.append(1.0)
        coverage_vals.append(0.0)
    else:
        precision_vals.append(all_labels[accepted].mean())
        coverage_vals.append(accepted.sum() / len(all_scores))
Out[23]:
Visualization
Line chart showing precision increasing and coverage decreasing as the LangID confidence threshold rises, with markers at 0.7, 0.8, and 0.9.
Precision and coverage as a function of the language identification confidence threshold. Higher thresholds improve the precision of accepted documents (fewer non-target-language documents slip through) but reduce coverage (more total documents are discarded). The three vertical dotted lines mark commonly used threshold values at 0.7, 0.8, and 0.9. Practical deployments often choose values in the 0.7 to 0.9 range, with the exact choice depending on whether the pipeline prioritizes corpus purity or volume.

The useful operating zone is better identified by the bend in the coverage curve than by the curves' crossing point, which occurs near the low end of this simulation. Precision is already high at moderate thresholds, while coverage begins falling much faster beyond roughly 0.7. The 0.7 to 0.9 range therefore illustrates progressively stronger corpus-purity choices with increasingly large losses in retained documents.

In practice, the optimal threshold depends on your use case. For pretraining data where you have 500 billion tokens to work with, you can afford aggressive filtering and might choose 0.9. For a low-resource language where every correct document is precious, you might choose 0.5 and accept more noise in exchange for higher recall.

Handling Multilingual Documents

Real web pages frequently contain text in multiple languages. A single Wikipedia article may have inline quotations in a foreign language. A news article may quote speakers in their native language. An e-commerce product page might have a description in English and user reviews in a mix of languages. Technical documentation might mix English code identifiers with prose in another language. These multilingual documents require more sophisticated handling than simple binary language classification.

Document-Level vs. Sentence-Level LangID

Most LangID systems operate at the document level: they return a single language label for an entire document. This works well when a document is predominantly monolingual, but fails for multilingual documents where neither the majority-language label nor any single label captures the full picture.

Sentence-level LangID applies language detection to each sentence independently. This is more computationally expensive but provides finer granularity. For a data curation pipeline, sentence-level detection allows you to extract only the sentences in your target language from otherwise mixed documents, build a multilingual corpus with per-sentence language labels, and detect and study code-switching patterns at the sentence boundary level.

The tradeoff is that sentence-level detection is more affected by the short-text problem: individual sentences are shorter than full documents, so confidence tends to be lower. A sentence like "Thank you for your support" embedded in a French article will correctly classify as English, but a five-word French sentence at the beginning of a mostly-English document may not classify confidently enough to trust the label.

There is also the question of sentence tokenization quality. Sentence boundaries are not always obvious, especially in noisy web text. A sentence splitter may combine two sentences from different languages into one "sentence," creating an artificial code-switching scenario. The pipeline quality of document-level operations cascades into sentence-level ones.

Byte Fraction Estimation

CLD2's approach to multilingual documents estimates the fraction of bytes in each detected language. If a 1,000-character document has 600 characters classified as English and 400 as French, the tool reports English at 60% and French at 40%. This allows downstream filtering based on dominance thresholds: accept documents where the target language exceeds 70% of content, and reject those with more evenly split language mixes.

The byte fraction approach is more informative than a single label for mixed documents. A data pipeline building an English corpus can use byte fraction to filter documents where at least 80% of the content is English, retaining pages like English technical documentation that contains occasional French quotations while rejecting pages where the content is truly split.

The limitation is that byte fraction is an approximation. CLD2 classifies fixed-length chunks of text and aggregates the results, which means that language boundaries within a paragraph may not align with chunk boundaries. For fine-grained analysis of multilingual documents, more sophisticated segmentation approaches are needed.

Language-Aware Segmentation

For some downstream tasks, you want to segment a multilingual document into monolingual spans. The procedure is as follows: first, run sentence tokenization to split the document into sentences; second, run LangID on each sentence; third, merge consecutive sentences with the same language label into spans; finally, apply a minimum span length filter to discard isolated sentences in foreign languages.

This pipeline is more robust than naive per-sentence filtering because it handles the common case where a foreign-language sentence is surrounded by target-language content. A single French quotation embedded in an English article will form a short, isolated French span. If the pipeline requires spans of at least three sentences in the same language, that isolated quotation will be attributed to the surrounding English context rather than creating a separate French span.

Language-aware segmentation is also the foundation for building code-switching corpora, where you specifically want to identify and preserve the points at which language boundaries occur. Research on multilingual models, cross-lingual transfer, and code-switching phenomena all benefit from this finer-grained view.

Language Filtering in LLM Data Curation

Large language model training corpora are assembled from web crawls that span hundreds of languages. The data curation pipeline must decide which languages to include, at what quantities, and with what quality filters. Language identification is the mechanism that makes language-specific decisions possible, but the policy decisions around how to use those identifications are equally important and have large downstream effects on model capabilities.

Single-Language Pipelines

Models like GPT-2, GPT-3, and most early LLMs were predominantly English. Their data pipelines applied a simple filter: keep only documents identified as English with high confidence. The threshold is typically 0.5 to 0.9 depending on the tool and the desired strictness.

This filter dramatically reduces corpus size relative to the raw web crawl. CommonCrawl data, after language filtering for English, retains roughly 50-60% of the original document count (varying by time period and filtering aggressiveness), but with much higher quality for downstream English language modeling. The discarded documents are either non-English or are mixed-language pages where the model cannot confidently assign English as the dominant language.

For an English-focused model, this is the right tradeoff. The model does not need Mandarin Wikipedia, French news articles, or German academic papers, and including them would dilute the signal from English text at the same number of training tokens. However, this policy implicitly assumes that the pipeline wants a monolingual model. Many modern deployment contexts require multilingual capabilities, which demands a different approach.

Multilingual Pipelines

Multilingual LLMs like mBERT, XLM-R, mT5, and BLOOM handle dozens to hundreds of languages. Their curation pipelines face additional challenges that single-language pipelines never encounter.

Language balance is a central question. Web content is deeply skewed toward high-resource languages: English accounts for a large fraction of indexed web content, with a long tail of languages that have minimal web presence. If you train on data proportional to web prevalence, your model will be predominantly English with poor capabilities in most other languages. Most multilingual model developers upsample low-resource languages, typically using a temperature-based sampling scheme where the probability of sampling from language LL is proportional to pL1/Tp_L^{1/T} where pLp_L is the language's fraction of available data and T>1T > 1 is a temperature parameter that flattens the distribution.

Cross-language contamination is another concern. A document misclassified as one language and included in that language's training split introduces noise. For high-resource languages, occasional contamination is tolerable. For a low-resource language where the training split might have only millions of tokens, contamination from a closely related language could represent a significant fraction of the training signal.

Language-specific quality filters add another layer of complexity. Quality heuristics developed for English, such as perplexity filtering against a reference language model, minimum word count requirements, and duplicate removal, must be calibrated separately for each language. The average sentence length differs across languages. The distribution of common words differs. A quality filter tuned on English may over-filter or under-filter text in other languages if applied naively.

The BLOOM model (Scao et al., 2022) addressed these challenges through careful language-specific data sourcing. Rather than relying entirely on web crawls, the team curated dedicated datasets for low-resource languages from organizations like Masakhane (African languages) and AmericasNLP (Indigenous American languages). This required language identification as both a filter and a verification step to confirm that sourced data matched its claimed language labels. Even professionally curated datasets sometimes contain mislabeled content, and applying LangID as a verification step on supposedly monolingual sources catches these errors before they enter the training pipeline.

The Misclassification Problem at Scale

At the scale of a CommonCrawl-based corpus, even a 1% misclassification rate introduces tens of millions of documents in the wrong language bucket. For a model training on 100 billion English tokens, a 1% contamination means 1 billion tokens of non-English content mixed into the English training split. This has a measurable effect on English language modeling quality, particularly on tasks involving rare vocabulary, idioms, or cultural knowledge that depends on consistent English exposure.

The effect on low-resource languages is potentially much larger in relative terms. A small language with only 100 million tokens of training data in the corpus, where 20% consists of misclassified documents from related languages, receives 80 million tokens of target-language signal and 20 million tokens of noise. The noise is not random: it comes from the most closely related languages, which means the model may learn to confuse the target language with its relatives, producing outputs that mix the languages in unnatural ways.

This is why production pipelines often use multiple LangID systems in ensemble, or use LangID as an initial filter followed by secondary quality checks like perplexity filtering against a reference language model. If a document passes the LangID filter but scores poorly under a language-specific perplexity model, it is likely either misclassified or low-quality content within the target language. Both cases benefit from additional scrutiny.

Language Identification in Practice: The C4 Dataset

The C4 dataset (Raffel et al., 2020), used to train T5, applied langdetect with a threshold of 0.99 for English classification. This very high threshold was chosen to maximize English purity at the cost of some recall, discarding documents where the model was uncertain. C4 started from 156 billion tokens of CommonCrawl data and retained roughly 156 billion English tokens after all filtering steps, suggesting that the vast majority of web crawl data passing other quality filters is correctly identified as English. The 0.99 threshold is stricter than the 0.7 to 0.9 range commonly used in other pipelines; C4's designers prioritized purity, accepting the risk of discarding some valid English documents.

Ensemble and Secondary Validation

A robust production pipeline does not rely on a single LangID tool. Using two or more independent systems and requiring agreement before accepting a document reduces the false positive rate at the cost of some additional computation and potentially lower recall. The systems to combine depend on the target languages: FastText for broad coverage and speed, CLD2 for multilingual byte-fraction estimation, and lingua for high-precision short-text identification.

Secondary validation with a language-specific perplexity model is particularly effective at catching misclassified documents. A character-level or subword language model trained on high-quality text in the target language assigns low perplexity to grammatically coherent text in that language and high perplexity to text in other languages or low-quality content. Running this filter after LangID catches the documents that tricked the classifier but cannot fool a language model trained on the target language itself.

The computational cost of secondary validation is significant at scale. A perplexity-based filter requires a forward pass through a language model for every document, which is orders of magnitude more expensive than a FastText classification. In practice, this filter is applied only to documents that pass the primary LangID filter and are destined for lower-resource language buckets where quality is more critical.

Limitations and Practical Implications

Language identification has several important limitations that practitioners must understand when designing data curation pipelines. These are not edge cases: they affect a significant fraction of real-world web content.

Coverage and Script Support

No LangID system covers all 7,000-plus living languages. FastText's LID.176 model covers 176 languages, which is excellent by industry standards but still leaves thousands of languages without support. For languages outside the model's vocabulary, all text will be misclassified to whatever language the model finds most plausible, typically a related language with more training data. This is a serious issue for anyone working with low-resource languages: the tool that is supposed to identify your language of interest simply does not know it exists.

The situation is worse for scripts with multiple closely related languages. Arabic, Dari, and Pashto all use Arabic script and are often confused. Chinese (Standard Mandarin), Cantonese written in Chinese characters, and Classical Chinese look nearly identical at the character level but are very different linguistically and culturally. Misclassification between these pairs is frequent and difficult to avoid without additional context such as metadata, source URL, or geolocation information.

Some writing systems also support multiple unrelated languages. Latin script covers everything from English to Yoruba to Vietnamese. A LangID system that recognizes the Latin script has not yet done the hard work: it still needs to distinguish among the hundreds of languages that use it.

The Short Text Problem in Production

LangID systems reliably identify languages in long, clean documents. They become much less reliable on short text: tweets, titles, captions, code comments, and metadata fields. A practical rule of thumb is that 20 or more words are needed for high-confidence predictions; below 5 words, accuracy drops substantially. Even for longer documents, short documents account for a significant fraction of web-crawled text, since many pages contain primarily navigation, metadata, or short-form content.

Pipelines should consider minimum document length requirements before language filtering, or should accept lower confidence scores on short documents rather than discarding them entirely. An alternative is to concatenate short documents that appear to come from the same source, apply LangID to the concatenated text, and then split them back out. This reduces the short-text problem at the cost of introducing a dependency between documents.

Code-Switching is Not a Defect

Code-switching is a natural feature of multilingual communication, not a noise artifact. When a Hindi-English speaker switches between languages mid-sentence, that text is linguistically valid and may be exactly what you want in a multilingual model's training data. A pipeline that routes such text to neither the Hindi nor the English bucket loses this valuable signal.

The problem is compounded by the fact that code-switching patterns are language-pair specific. Spanglish (Spanish-English) has well-documented grammatical patterns. Hinglish (Hindi-English) has its own conventions. Taglish (Tagalog-English) is different again. A model trained only on monolingual data may not naturally acquire code-switching competence, which affects its utility for multilingual communities where code-switching is the norm rather than the exception.

Dedicated code-switching datasets and language identification systems that support mixed-language labels are an active research area. The CALCS (Computational Approaches to Linguistic Code Switching) community has developed benchmarks and models for this task. For now, most production pipelines label code-switched text with the dominant language and accept the imprecision as a known limitation rather than an unsolved problem.

Domain Shift

LangID models trained on Wikipedia and news text may underperform on domain-specific web content: product reviews, social media, legal documents, medical records, and technical forums. Each domain has its own vocabulary distribution, and a model calibrated on standard prose may struggle with informal writing, heavy use of domain jargon, or non-standard orthography.

This is particularly acute for social media and informal web text. Internet abbreviations, intentional misspellings, emoji, and non-standard punctuation all distort the character n-gram distributions that LangID systems rely on. A social media filter should validate its LangID tool on domain-matched text before deploying it at scale. Without this validation, you may be applying a threshold calibrated on formal text to informal text, with unknown consequences for precision and coverage.

Domain shift also interacts with language-specific conventions in non-obvious ways. German internet slang frequently uses English loanwords without inflection, making German social media text look more English-like than standard German. French internet slang uses phonetic spelling (verlan, word reversal) that breaks standard trigram patterns. Any LangID system applied to social media data should be evaluated on a sample of that data before being trusted.

Temporal Drift

Web language evolves. New words, new domains, new communities, and shifts in which languages are most active online all change the statistical properties of web-crawled text over time. A LangID model trained on a 2015 CommonCrawl snapshot may not perform as well on a 2025 snapshot, particularly for languages that have grown or changed their web presence in the intervening decade.

Practitioners should periodically re-evaluate their LangID tools on recent samples from their target crawl, especially when switching to newer CommonCrawl snapshots. If recall on target languages has dropped or contamination from related languages has increased, it may be time to retrain or update the LangID component.

Summary

Language identification is the foundation of multilingual data curation. Before any language model can learn from a document, the pipeline must know what language that document is in. The concepts from this chapter:

  • Character n-gram models exploit the distinctive statistical fingerprints of each language's character sequences. They assign a document to the language whose model assigns the highest average per-n-gram log-probability. Laplace smoothing handles unseen n-grams, and length normalization makes scores comparable across documents of different sizes.
  • Rank-based methods (Cavnar and Trenkle, 1994) build language profiles from the most frequent n-grams and measure the out-of-place distance between a document's n-gram ranks and each language profile. This approach is robust to short text and avoids probability smoothing.
  • FastText-based systems like LID.176 use discriminative training to sharpen language boundaries and cover 176 languages with extremely high throughput, making them the standard choice for large-scale web filtering.
  • CLD2 and CLD3 explicitly handle multilingual documents by estimating byte fractions per language, which is valuable when you need to identify and extract content from pages that mix languages.
  • Confidence thresholds control the precision-coverage tradeoff: higher thresholds give cleaner language buckets at the cost of discarding more documents. Practical deployments use thresholds between 0.7 and 0.9 depending on whether corpus purity or volume is the priority.
  • Short text degrades accuracy substantially below 10 to 20 words. Data curation pipelines should impose minimum document length requirements or apply separate handling for short-form content.
  • Code-switching produces text that belongs to multiple languages. Document-level LangID assigns a single dominant language label; sentence-level approaches provide finer granularity at higher computational cost. Code-switching is a natural linguistic phenomenon, not a noise artifact, and should be treated accordingly.
  • Multilingual pipelines face additional challenges around language balance, per-language quality calibration, and the compounding effect of misclassification at scale. Language upsampling with temperature-based sampling is a standard technique for giving low-resource languages equitable representation.
  • Real-world curation often combines multiple LangID tools in ensemble, applies minimum document length requirements, and uses secondary quality checks like perplexity filtering to catch misclassifications that pass the primary filter.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Language Identification.

Language Identification Quiz

Question 1 of 100 of 10 completed
What is the core statistical signal that character n-gram language models exploit for language identification?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026languageidentification, author = {Michael Brenndoerfer}, title = {Language Identification: Models, Multilingual Handling}, year = {2026}, url = {https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Language Identification: Models, Multilingual Handling. Retrieved from https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching
MLAAcademic
Michael Brenndoerfer. "Language Identification: Models, Multilingual Handling." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching>.
CHICAGOAcademic
Michael Brenndoerfer. "Language Identification: Models, Multilingual Handling." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Language Identification: Models, Multilingual Handling'. Available at: https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Language Identification: Models, Multilingual Handling. https://mbrenndoerfer.com/writing/language-identification-models-multilingual-code-switching

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.