Quality Filtering: Heuristics, Perplexity, and Classifiers

Michael BrenndoerferJanuary 10, 202650 min read

Part of Language AI Handbook

Filter low-quality text from web corpora using heuristic rules, perplexity scoring, and classifier-based methods with tunable thresholds.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Quality Filtering

When you scrape the web for training data, you collect far more than clean, coherent text. You get spam, SEO-stuffed keyword lists, machine-translated garbage, duplicate boilerplate, pornographic content, malware descriptions, and page after page of raw HTML fragments that escaped the parser. Even after deduplication (covered in the previous chapter), a substantial fraction of what remains is low-quality text that would hurt model performance if included. Quality filtering is the stage where you decide what belongs in a training corpus.

The core challenge is defining "quality" for a language model. It is not the same as "grammatically correct" or "well-written by human standards." Training data quality means text that teaches the model useful patterns: coherent sentence structure, meaningful word co-occurrences, factual statements, domain-appropriate vocabulary, and the kind of reasoning chains you want the model to reproduce. A Reddit comment with spelling errors can be higher quality than a perfectly formatted but semantically empty spam article. This distinction separates data curation from proofreading.

Quality filtering encompasses several approaches that operate at different levels of precision and cost. Heuristic filters apply simple, cheap rules that catch obvious garbage. Perplexity filtering uses a reference language model to score how natural the text is relative to known-good content. Classifier-based filtering trains a model to distinguish quality text from junk. Each method has its place in a production pipeline, and serious efforts like C4, The Pile, and RefinedWeb use combinations of all three. This chapter covers each approach in depth, including how to set filter thresholds without accidentally discarding too much good data or biasing the corpus against minority linguistic communities.

Understanding quality filtering matters because it can strongly affect any pretraining pipeline. A 10% improvement in filtering precision can translate to models that need substantially less compute to reach a given performance level. The "less is more" phenomenon, where a smaller but cleaner corpus outperforms a much larger noisier one, has been replicated across many independent training runs.

Why Quality Matters for LLM Training

Before diving into the methods, it helps to understand the stakes. Language models learn by predicting the next token in a sequence. Every piece of text in the training corpus implicitly teaches the model patterns. If that text is incoherent, the model learns incoherent patterns. If it consists of keyword repetitions, the model learns to repeat keywords. The relationship between data quality and model quality is tight, and the mechanisms are instructive.

Consider what happens when a model trains on a spam article that repeats the phrase "cheap insurance online" forty times. The model updates its weights to assign high probability to the word "online" following "cheap insurance", and high probability to "cheap" following "insurance". These updates compete with updates from legitimate text that use the same words in different, more meaningful contexts. Garbage data is not neutral: it actively corrupts the statistical patterns the model is trying to learn.

The Gopher paper from DeepMind demonstrated this clearly by ablating their quality filters and measuring perplexity on held-out benchmarks. Models trained on unfiltered Common Crawl data performed substantially worse than models trained on filtered versions of the same underlying source. More dramatically, training on a smaller but higher-quality corpus often outperforms training on a much larger but lower-quality one. The Phi-1 paper pushed this principle to an extreme by showing that a 1.3B parameter model trained on "textbook-quality" filtered and synthetic data achieved performance competitive with models trained on ten times more conventional web data.

Practically, the pipeline looks like this: web scraping and HTML extraction produce raw text. Deduplication removes redundant content. Quality filtering removes low-value content from what remains. What survives goes into the training corpus. The filtering stage is where you make the hardest judgment calls, because every filter has a precision-recall tradeoff. A strict filter removes more garbage but also more legitimate text. A permissive filter preserves edge cases but lets through noise. There is no setting that eliminates this tradeoff: you are always choosing a point on a curve, not escaping it.

The stakes scale with corpus size. When you process a trillion tokens, even a 1% false negative rate means ten billion tokens of garbage in the training set. When you process ten petabytes of raw web text, even a 0.1% false positive rate discards ten terabytes of legitimate content. At these scales, decisions about filter design that seem like engineering choices have material effects on what the resulting model knows, how well it handles different communities and dialects, and what failure modes it exhibits in production.

Heuristic Filters

Heuristic filtering is the first line of defense. These filters encode explicit rules about what bad text looks like. They run fast (often in microseconds per document), scale to petabytes, and catch the most obvious problems without requiring any model inference. The downside is that heuristics are brittle: they fail on text that looks superficially clean but is semantically poor, and they may incorrectly discard legitimate text that happens to match the wrong pattern.

The insight behind heuristic filters is that low-quality web content has characteristic surface properties that distinguish it from natural prose. Spam articles repeat keywords. Navigation menus have very short lines. Scraped template content has identical repeated blocks. Pages that failed HTML parsing have high ratios of angle brackets and escaped entities. Rather than modeling these problems semantically, heuristics directly measure the surface properties that correlate with low quality.

The heuristic filters used in production pipelines like C4, MassiveText, and RefinedWeb fall into several categories:

  • Length-based filters: Remove documents that are too short (fewer than 200 words, for instance) or too long. Short documents provide too little context for the model to learn useful patterns. Extremely long documents are sometimes artifacts of parsing failures where many separate pages got concatenated.
  • Punctuation and symbol ratio filters: Flag documents where a large fraction of characters are symbols, numbers, or non-letter characters. A document that is 30% curly braces is almost certainly code or template junk, not natural language.
  • Alphabet coverage filters: Remove documents where most content is non-alphabetic or where the text comes from an unexpected script. This is also used for language identification when language-specific filtering is required.
  • Bullet point or line-length filters: Documents composed almost entirely of short, newline-delimited items (like site navigation menus or keyword lists) tend to have poor narrative structure. The ratio of lines to words can identify these.
  • Boilerplate detection: Many web pages share identical footers, cookie banners, and privacy policy text. Hard-coded string matching can remove pages where more than a fixed percentage of the text matches known boilerplate phrases. This complements deduplication, which operates at the document level, by catching pages where boilerplate comprises a large fraction of otherwise-unique content.
  • Bad word lists: Some pipelines maintain lists of profanity, spam trigger words, or adult content markers. Any document containing words from these lists is flagged or removed. The C4 pipeline famously uses the "List of Dirty, Naughty, Obscene, or Otherwise Bad Words" maintained on GitHub.
  • Repetition filters: A document that repeats the same sentence or paragraph is almost certainly low-quality. You can detect this by checking whether any n-gram above a certain length occurs more than once in the document, or by measuring the character-level entropy of the content. Low entropy means low informational diversity.

Each rule has an explicit failure mode. Length filters discard short but high-quality documents like dictionary definitions or aphorisms. Symbol ratio filters may incorrectly remove technical documentation with code snippets. Bad word lists may disproportionately remove text from communities where that vocabulary appears in legitimate cultural or literary contexts. Understanding these failure modes is not academic: it determines which heuristics to apply, which to skip for a given domain, and how to calibrate thresholds.

Worked Example: Repetition and Symbol Filters

Let's look at what these filters detect and why they work.

In[3]:
Code
import re
from collections import Counter


def symbol_ratio(text):
    """Fraction of characters that are non-alphabetic and non-space."""
    if not text:
        return 1.0
    non_alpha = sum(1 for c in text if not c.isalpha() and not c.isspace())
    return non_alpha / len(text)


def max_word_repetition(text, top_n=5):
    """Fraction of words accounted for by the top-N most common words (excluding stopwords)."""
    stopwords = {
        "the",
        "a",
        "an",
        "and",
        "or",
        "in",
        "on",
        "at",
        "to",
        "of",
        "is",
        "it",
        "this",
        "that",
    }
    words = [
        w.lower()
        for w in text.split()
        if w.isalpha() and w.lower() not in stopwords
    ]
    if not words:
        return 1.0
    counts = Counter(words)
    top_count = sum(v for _, v in counts.most_common(top_n))
    return top_count / len(words)


def duplicate_line_ratio(text):
    """Fraction of lines that are duplicates."""
    lines = [l.strip() for l in text.splitlines() if l.strip()]
    if not lines:
        return 0.0
    duplicates = len(lines) - len(set(lines))
    return duplicates / len(lines)


# Example documents
docs = {
    "Clean news article": (
        "Scientists at MIT have developed a new approach to carbon capture that "
        "could reduce costs by up to 40 percent. The technique uses a novel "
        "polymer membrane that selectively absorbs CO2 from industrial exhaust. "
        "Initial tests suggest the system could be deployed at scale within five years."
    ),
    "Keyword spam": (
        "buy cheap insurance online get insurance quotes best insurance rates "
        "cheap car insurance home insurance life insurance online insurance quotes "
        "buy insurance now insurance deals insurance comparison insurance calculator"
    ),
    "Template boilerplate": (
        "Home | About | Contact | Privacy Policy | Terms of Service\n"
        "Home | About | Contact | Privacy Policy | Terms of Service\n"
        "Home | About | Contact | Privacy Policy | Terms of Service\n"
        "Copyright 2023. All rights reserved."
    ),
    "Symbol-heavy": (
        ">>> import sys; sys.path.append('/usr/local/lib') "
        "{{config.key}}: [val1, val2, val3] "
        "SELECT * FROM table WHERE id=42 AND status='active';"
    ),
}
Out[4]:
Console
Document                        Symbol%   WordRep%   DupLine%
--------------------------------------------------------------
Clean news article                2.09%     18.75%      0.00%
Keyword spam                      0.00%     65.52%      0.00%
Template boilerplate              8.45%     62.50%     50.00%
Symbol-heavy                     24.82%    100.00%      0.00%

The clean news article scores low on all three metrics. The keyword spam shows a very high word repetition ratio because the same words ("insurance") dominate the vocabulary after stopword removal. The template document is flagged by the duplicate line ratio. The symbol-heavy code fragment is caught by the symbol ratio.

These thresholds are tunable. If you set the symbol ratio cutoff at 0.3, you reject documents where more than 30% of characters are symbols. If you set the word repetition threshold at 0.4, you reject documents where the top 5 words account for more than 40% of non-stopword content. The right thresholds depend on your domain: code corpora need higher symbol ratio thresholds, while technical documentation might legitimately use more repetitive terminology than general prose.

The C4 Heuristic Suite

The C4 dataset (used to train T5 and mT5) is one of the most influential filtered corpora, and its heuristic rules have become a baseline for many subsequent efforts. The C4 processing pipeline starts with Common Crawl's WET files (pre-extracted text), then applies the following filters. Lines are discarded if they:

  • Do not end with terminal punctuation (a period, question mark, or exclamation point)
  • Contain fewer than 3 words or more than 100,000 characters
  • Contain the string "javascript" (to catch pages where the parser captured script blocks)
  • Contain the string "lorem ipsum" (to catch placeholder content)
  • Contain any words from the "List of Dirty, Naughty, Obscene, or Otherwise Bad Words"

Documents are then removed if fewer than 5 sentences remain after line filtering, or if they contain the phrase "lorem ipsum". The entire pipeline removed roughly 70% of raw Common Crawl documents, producing a cleaner but still large 750 GB corpus.

Later analyses questioned some of these choices and their consequences. The terminal punctuation filter disproportionately removes non-English text where punctuation conventions differ: many East Asian languages do not end sentences with a period. The JavaScript filter turned out to remove legitimate pages that happened to discuss JavaScript programming. The bad word filter significantly underrepresented content from communities that use that vocabulary in legitimate cultural, literary, or conversational contexts.

These are typical failure modes of heuristic filters: rules designed around one language or content type, or calibrated on a particular demographic's writing, can be inappropriate for others. Every rule encodes an assumption about what normal text looks like, and those assumptions tend to reflect the dominant culture and language of whoever designed the rules.

Character-Level Entropy as a Quality Signal

One less-obvious but useful heuristic is character-level entropy. Informative text has high character diversity: it uses a wide variety of words and transitions. Repetitive spam or boilerplate has low entropy because the same short n-grams repeat constantly.

Given a document with NN characters, the empirical character unigram distribution is pc=count(c)/Np_c = \text{count}(c) / N for each character cc. The character-level entropy is:

H=−∑cpclog⁡2pcH = -\sum_{c} p_c \log_2 p_c

where the sum runs over all distinct characters appearing in the document. Well-written prose in English typically has character entropy between 4.5 and 5.2 bits. Repetitive spam with limited vocabulary falls below 4.0. Compressed or encoded binary data exceeds 6.0. Setting a window of acceptable entropy values gives you a fast filter that catches a class of garbage that other rules miss.

The entropy measure is also language-aware in a useful way: it naturally adapts to the expected character distribution of the target language. Chinese text will have different typical entropy values than English, but the threshold "reject documents below the 5th percentile of entropy for this language" works without needing separate calibration.

Perplexity Filtering

Heuristic filters catch obvious garbage, but they cannot assess whether text is coherent and natural in a linguistic sense. A document can pass all heuristic checks and still be semantically nonsensical, machine-translated in a stilted way, or composed of grammatically valid but meaningless phrases. Perplexity filtering addresses this by scoring how well a reference language model can predict the text.

Perplexity

Perplexity measures how surprised a language model is by a sequence of text. Given a trained language model with probabilities PP, the perplexity of a sequence w1,w2,…,wNw_1, w_2, \ldots, w_N is:

PPL(w1,…,wN)=exp⁡(−1N∑i=1Nlog⁡P(wi∣w1,…,wi−1))\text{PPL}(w_1, \ldots, w_N) = \exp\left(-\frac{1}{N} \sum_{i=1}^{N} \log P(w_i \mid w_1, \ldots, w_{i-1})\right)

where:

  • wiw_i: the ii-th token in the sequence
  • P(wi∣w1,…,wi−1)P(w_i \mid w_1, \ldots, w_{i-1}): the conditional probability the model assigns to token wiw_i given all preceding tokens
  • NN: total number of tokens in the sequence
  • exp⁡(⋅)\exp(\cdot): the exponential function, which converts the average negative log-probability back to a probability-like scale

Lower perplexity means the model assigns higher probability to the sequence. The sequence is more predictable given what the model has learned.

The key insight behind perplexity filtering is this: a language model trained on high-quality text assigns low perplexity to text that resembles its training distribution and high perplexity to text that does not. Coherent English prose gets low perplexity from a model trained on books or Wikipedia. Garbled machine translation, keyword soup, or foreign language text gets high perplexity.

This is powerful because it captures something heuristics cannot: the sequential coherence of language. A document can have perfect punctuation, normal symbol ratios, and varied vocabulary, but if the word sequences are unnatural (as happens with low-quality machine translation or automated content generation), a language model will find it surprising. The perplexity signal captures that surprise.

How Perplexity Filtering Works

You train or obtain a small reference language model on known-good text. Wikipedia, books corpora, or curated news articles work well. Then you score each candidate document by running it through the reference model and measuring the average perplexity. Documents above a threshold are discarded.

The reference model does not need to be large. In practice, character-level n-gram models (KenLM) or small neural models work well for this purpose. The Gopher paper and CCNet pipeline both use KenLM models trained on Wikipedia for perplexity scoring. KenLM implements a modified Kneser-Ney smoothed n-gram language model stored in a compact hash structure. It is extremely fast: scoring billions of words takes hours, not days. For a trillion-token pipeline, KenLM's throughput is essential; running a multi-billion-parameter neural model at scoring time would be prohibitively expensive.

The choice of reference corpus matters enormously. A KenLM model trained on Wikipedia learns the patterns of encyclopedic English: neutral tone, third-person perspective, factual claims, formal vocabulary. It will assign high perplexity to legal documents, poetry, informal dialogue, scientific papers, and code, not because these are low quality, but because they differ from encyclopedic style. If you want all these content types in your training corpus, you need either multiple domain-specific reference models or a permissive threshold that only removes the most extreme outliers.

Computing Perplexity with KenLM

Let's demonstrate the concept using a simple n-gram language model, which captures the same core intuition as KenLM-based filtering.

In[5]:
Code
import math
from collections import Counter, defaultdict


class NgramLM:
    """Simple n-gram language model for perplexity scoring."""

    def __init__(self, n=3, smoothing=0.1):
        self.n = n
        self.smoothing = smoothing
        self.ngram_counts = defaultdict(Counter)
        self.vocab = set()

    def tokenize(self, text):
        return text.lower().split()

    def train(self, texts):
        for text in texts:
            tokens = ["<s>"] * (self.n - 1) + self.tokenize(text) + ["</s>"]
            self.vocab.update(tokens)
            for i in range(self.n - 1, len(tokens)):
                context = tuple(tokens[i - (self.n - 1) : i])
                word = tokens[i]
                self.ngram_counts[context][word] += 1
        self.vocab_size = len(self.vocab)

    def log_prob(self, context, word):
        context_count = sum(self.ngram_counts[context].values())
        word_count = self.ngram_counts[context][word]
        # Laplace smoothing
        return math.log(
            (word_count + self.smoothing)
            / (context_count + self.smoothing * self.vocab_size)
        )

    def perplexity(self, text):
        tokens = ["<s>"] * (self.n - 1) + self.tokenize(text) + ["</s>"]
        log_probs = []
        for i in range(self.n - 1, len(tokens)):
            context = tuple(tokens[i - (self.n - 1) : i])
            word = tokens[i]
            log_probs.append(self.log_prob(context, word))
        avg_log_prob = sum(log_probs) / len(log_probs)
        return math.exp(-avg_log_prob)


# Train on high-quality text (representative Wikipedia-style sentences)
quality_training = [
    "the researchers developed a new method for protein folding prediction",
    "climate change has accelerated glacial melting in polar regions",
    "the study examined correlations between diet and cardiovascular disease",
    "neural networks have transformed image recognition and speech processing",
    "economic growth depends on productivity gains and capital investment",
    "the algorithm achieves better performance with larger training datasets",
    "natural language processing enables computers to understand human text",
    "the experiment measured reaction times under different stimulus conditions",
    "photosynthesis converts light energy into chemical energy stored in glucose",
    "machine learning models require careful validation to avoid overfitting",
]

lm = NgramLM(n=3)
lm.train(quality_training)
In[6]:
Code
# Score candidate documents
candidates = {
    "Science news": "the scientists published their findings on neural plasticity in developing brains",
    "Coherent prose": "economic research suggests that education investment yields long term productivity gains",
    "Keyword spam": "buy cheap loans loan rates best loans cheap mortgage refinance loan calculator loans",
    "Garbled text": "the of from into systems processing language neural have been used for",
    "HTML artifact": "span class footer nav li ul div href src img alt copyright reserved",
}
Out[7]:
Console
Document                    Perplexity
----------------------------------------
Science news                      68.7
Coherent prose                    75.7
Keyword spam                      89.9
Garbled text                      69.9
HTML artifact                     89.9

The perplexity values separate the candidates along the expected lines. Coherent prose that matches the training distribution (science news, economic research sentences) gets low perplexity. Keyword spam and garbled text get high perplexity because their word sequences are unpredictable given the model's learned patterns. The HTML artifact scores high because isolated tag names and attributes have no sequential predictability from the reference model's perspective.

Notice that this works without any explicit rules about what constitutes "good" text. The n-gram model does not know that "buy cheap loans" is spam. It knows that those three words appearing together is unexpected given the contexts it has seen during training on science and technology sentences. Perplexity filtering is powerful precisely because it generalizes beyond any fixed rule set.

Perplexity Thresholds in Practice

The CCNet pipeline (used for XLM-R and LLaMA pretraining data) filters documents by perplexity into three buckets:

  • Head: Bottom 30th percentile by perplexity (highest quality text, closest to the reference distribution)
  • Middle: 30th to 65th percentile
  • Tail: Top 35th percentile by perplexity (lowest quality, often excluded entirely)

Rather than a fixed absolute threshold, CCNet uses percentile-based filtering. This is more robust because absolute perplexity values depend heavily on the reference model and the tokenization. A threshold of "reject documents with perplexity above 500" that works for English may be completely wrong for another language if you use the same model. Absolute thresholds also shift if you retrain the reference model with different hyperparameters. Percentile thresholds are stable across these variations because they are defined relative to the corpus distribution.

One important caveat is that perplexity filters can be too aggressive about domain specificity. A language model trained on Wikipedia will assign high perplexity to legal documents, poetry, and code, even though these are perfectly valid forms of text. If you want to include diverse domains in your training corpus, you need either domain-specific reference models or a more permissive threshold that only removes the truly extreme outliers. CCNet handles this partially by training separate language models for each language, but within a language the domain bias persists.

A second caveat is that perplexity filtering is circular in a subtle way. If you train your reference model on text that excludes certain dialects or styles, those dialects get high perplexity scores and are removed from the corpus. The resulting training corpus underrepresents those styles, producing a final model that performs poorly on them. The reference model's biases become the corpus's biases. This is distinct from heuristic bias (where a rule is explicitly wrong) but equally consequential.

Visualizing Perplexity Distributions

In[8]:
Code
import matplotlib
import matplotlib.pyplot as plt
import numpy as np

rng = np.random.default_rng(42)

# Simulate perplexity distributions for quality tiers
quality_ppl = rng.lognormal(mean=3.2, sigma=0.4, size=5000)  # high quality
medium_ppl = rng.lognormal(mean=4.2, sigma=0.5, size=5000)  # medium quality
low_ppl = rng.lognormal(mean=5.4, sigma=0.6, size=5000)  # low quality

# Clip to reasonable range for visualization
quality_ppl = np.clip(quality_ppl, 5, 800)
medium_ppl = np.clip(medium_ppl, 5, 800)
low_ppl = np.clip(low_ppl, 5, 800)
Out[9]:
Visualization
Overlapping density histograms showing perplexity distributions for three quality tiers with threshold lines.
Simulated perplexity distributions for three text quality tiers. High-quality text (solid blue) concentrates at low perplexity values, medium-quality text (dashed orange) spreads across the middle range, and low-quality text (dotted green) extends to higher values. Dashed vertical lines mark the 65th percentile thresholds used in CCNet-style filtering for each tier, illustrating why fixed absolute thresholds are less reliable than percentile-based ones.

The overlapping distributions reveal why perplexity thresholds are imperfect. The tails of the high-quality and medium-quality distributions overlap substantially. A threshold that catches most low-quality documents will inevitably discard some legitimate high-quality text. This overlap is the core precision-recall tradeoff in filtering. It is not a failure of the method: the task itself is ambiguous. Some text lies on the quality boundary, and no scalar metric will cleanly separate it. This is why production pipelines do not rely on perplexity filtering alone.

Perplexity and Tokenization Sensitivity

One practical subtlety is that perplexity scores depend on tokenization. A character-level model, a word-level model, and a BPE model trained on the same corpus will assign different perplexity values to the same document, even if all three models have similar quality. This makes it difficult to compare perplexity thresholds across different reference model implementations.

The CCNet approach of using percentile thresholds partially mitigates this: because the percentile is computed over the same corpus with the same model, relative rankings are stable even if absolute values vary. If you change the reference model (for example, updating from a 3-gram to a 5-gram KenLM model), you need to recompute the threshold percentiles on a representative sample rather than reusing old absolute cutoffs.

Classifier-Based Filtering

Classifier-based filtering takes a more direct approach: train a binary classifier to predict whether a given document is high-quality or not. Rather than encoding rules or using indirect signals like perplexity, the classifier learns from examples of good and bad text. This makes it more flexible and often more accurate than heuristic or perplexity-based approaches, at the cost of requiring labeled training data and higher inference cost.

The classic implementation from OpenAI (used for the WebText dataset, which GPT-2 trained on) trains a logistic regression classifier on n-gram features extracted from two classes of text:

  • Positive examples: Text from outbound links on Reddit that received at least 3 upvotes. The assumption is that Reddit users curate quality content by voting, so pages linked from Reddit and upvoted represent human-approved text across a wide range of topics and styles.
  • Negative examples: Random Common Crawl pages.

The classifier then scores every Common Crawl document by its predicted probability of being a positive example. Documents below a threshold are discarded. This approach uses the crowdsourced curation signal embedded in Reddit voting without requiring any manual labeling by the research team.

The Reddit upvote signal is not perfect. Reddit's user base skews toward certain demographics, interests, and languages. Content that Reddit users find interesting may not be the same as content that makes the best language model training data. Despite this, the resulting WebText dataset and models trained on it showed that the Reddit-curated filtering dramatically improved quality over unfiltered Common Crawl.

The Logistic Regression Approach

The classifier trained on TF-IDF features over character n-grams is surprisingly effective for this task. Character n-grams capture low-level patterns (spelling, punctuation, character repetition) that distinguish boilerplate and spam from natural prose without requiring linguistic analysis. The character-level representation is language-agnostic and robust to vocabulary differences between the training examples and the documents being scored.

The TF-IDF weighting (term frequency times inverse document frequency) ensures that features that appear uniformly across all documents (common characters, punctuation) get down-weighted, while features that are distinctive to certain document types get up-weighted. The logistic regression then learns to combine these features into a scalar quality score.

In[10]:
Code
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split

# Positive examples: high-quality web text
positive_examples = [
    "The study found that regular exercise reduces the risk of cardiovascular disease by 30 percent.",
    "Scientists have discovered a new species of deep-sea fish in the Mariana Trench.",
    "The Federal Reserve raised interest rates by 25 basis points in response to inflation data.",
    "Python 3.12 introduces several performance improvements including faster startup time.",
    "The novel explores themes of identity and belonging through the lens of immigrant experience.",
    "Researchers at Stanford developed an algorithm that predicts protein structures with high accuracy.",
    "Climate models suggest global temperatures could rise by 2 degrees by 2100 under current policies.",
    "The Supreme Court issued a ruling on Section 230 liability for online platforms.",
    "Astronomers detected gravitational waves from the merger of two neutron stars.",
    "The company reported quarterly earnings that exceeded analyst expectations by 15 percent.",
    "New research suggests that sleep deprivation impairs cognitive function significantly.",
    "The documentary examines the impact of social media on adolescent mental health.",
]

# Negative examples: low-quality/spam web text
negative_examples = [
    "Buy cheap insurance online get insurance quotes best rates insurance deals today.",
    "Click here now!!!! FREE MONEY LIMITED OFFER ACT NOW BEFORE ITS TOO LATE!!!",
    "Home About Contact Privacy Terms Copyright 2023 All Rights Reserved Footer Nav.",
    "var i=0;for(i=0;i<arr.length;i++){document.getElementById('div'+i).style.display='none';}",
    "best seo services cheap seo packages buy backlinks google rank first page guaranteed",
    "Lorem ipsum dolor sit amet consectetur adipiscing elit sed do eiusmod tempor incididunt",
    "!!! AMAZING DEAL !!! buy now pay later no credit check free shipping returns accepted !!!",
    "Copyright reserved terms privacy home about contact us sitemap footer links header nav",
    "casino online poker slots gambling win money jackpot bonus free spins play now sign up",
    "download free mp3 songs download movies free download software crack serial key keygen",
    "online pharmacy drugs buy prescription without prescription cheap medications fast delivery",
    "make money online work from home earn 500 daily passive income guaranteed no experience",
]

texts = positive_examples + negative_examples
labels = [1] * len(positive_examples) + [0] * len(negative_examples)

# TF-IDF on character n-grams (2-6 grams)
vectorizer = TfidfVectorizer(
    analyzer="char_wb",
    ngram_range=(2, 6),
    max_features=50000,
    sublinear_tf=True,
)

X = vectorizer.fit_transform(texts)
X_train, X_test, y_train, y_test = train_test_split(
    X, labels, test_size=0.3, random_state=42, stratify=labels
)

clf = LogisticRegression(C=1.0, max_iter=1000, random_state=42)
clf.fit(X_train, y_train)
y_pred = clf.predict(X_test)
Out[11]:
Console
Classification Report:
              precision    recall  f1-score   support

 Low quality       0.75      0.75      0.75         4
High quality       0.75      0.75      0.75         4

    accuracy                           0.75         8
   macro avg       0.75      0.75      0.75         8
weighted avg       0.75      0.75      0.75         8
In[12]:
Code
# Score new documents using classifier probabilities
new_docs = {
    "Research finding": "A study published in Nature found that mRNA vaccines produce durable immune responses.",
    "Boilerplate nav": "Home Products Services Blog About Contact Privacy Terms Footer Copyright",
    "Spam": "earn money fast online 1000 per day guaranteed no experience needed click now",
    "Coherent opinion": "The debate over AI regulation centers on balancing innovation with safety concerns.",
    "Social post": "just got the new iphone and honestly its pretty good but the camera app is confusing",
}

new_texts = list(new_docs.values())
X_new = vectorizer.transform(new_texts)
quality_scores = clf.predict_proba(X_new)[:, 1]
Out[13]:
Console
Document                   Quality Score
------------------------------------------
Research finding                 0.539  (KEEP)
Boilerplate nav                  0.430  (FILTER)
Spam                             0.435  (FILTER)
Coherent opinion                 0.537  (KEEP)
Social post                      0.505  (KEEP)

The quality scores reflect the classifier's learned pattern: text that resembles curated web content scores high, while boilerplate, spam, and keyword-stuffed text scores low. The casual social media post sits near the boundary, where its classification is ambiguous. Whether to include informal social text depends on your training objectives: if you want the model to handle conversational queries, discarding all informal text is counterproductive.

This example uses a small training set for illustration. In production, these classifiers are trained on millions of positive and negative examples, which substantially improves their accuracy and generalization. The WebText classifier reportedly used millions of Reddit-linked pages as positives and millions of random crawl pages as negatives.

What Character N-Grams Capture

It is worth understanding why character-level n-gram features are so effective for quality classification. Character n-grams at the 2-6 gram range capture patterns like:

  • Repetition of exclamation marks ("!!!!") that signal spam
  • Camel-case identifiers ("getElementById") that signal code
  • Copyright boilerplate sequences ("rights reserved", "© 20")
  • Structural patterns of URLs (".html", "href=")
  • Distinctive spacing and punctuation patterns of natural prose

These patterns are robust because they occur at a finer granularity than words. A new spam campaign can use completely different vocabulary but will still use the same exclamation-mark repetition patterns. A new programming language will still produce code with distinctive symbol combinations. The character-level representation generalizes better than word-level features across the long tail of web content.

The char_wb analyzer in scikit-learn adds word-boundary tokens to the character n-grams, which captures the shapes of individual words (short words, long words, words with unusual character combinations) without building an explicit vocabulary. This makes it much harder to game: an adversarial spam author cannot simply replace flagged words with synonyms, because the character-level patterns of their writing style remain detectable.

Fasttext-Based Filtering

For production pipelines that process billions of documents, even logistic regression with sparse features can become a bottleneck. Facebook's fastText classifier offers a faster alternative. Trained on subword n-gram embeddings, fastText achieves near-linear-time inference and can be trained on large datasets in minutes.

The CCNet pipeline uses fastText for language identification, assigning each document to a language with a confidence score, and the same architecture works for quality classification. FastText learns a shallow embedding for each subword n-gram and sums them to produce document representations, which are then used for classification. The model is shallow (no deep neural layers) but captures rich character-level patterns through the n-gram vocabulary.

The RefinedWeb dataset (used to train Falcon) extends this approach with a quality classifier that incorporates multiple signals: URL patterns, document length distribution, and language model scores, combined with a neural classifier that learns to weight these features. This ensemble approach is more robust than any single signal because different signals catch different failure modes, and the neural combination learns which signal is most reliable for which type of document.

Self-Supervised Quality Signals

A more recent approach avoids the need for manual labeling entirely. Instead of asking "does this look like Reddit-approved content?", you can train on implicit quality signals extracted automatically:

  • Educational content detection: The Phi-1 paper introduced the idea of training classifiers to detect "textbook quality" text. A GPT-4 model rated a sample of documents on educational richness, and those ratings were used to train a smaller fastText classifier that then scored the full corpus. This transfers expensive LLM judgment to cheap classifier inference.
  • Model-as-teacher at scale: Use a strong language model (like GPT-4) to rate a sample of documents, then train a smaller classifier on those ratings. This allows you to scale expensive LLM ratings across billions of documents that would be too costly to rate directly.
  • Consistency with known facts: Score documents by how many verifiable factual claims they contain versus unsupported assertions. Pages with many specific, checkable claims (dates, measurements, named entities) tend to be more informative than vague or opinion-heavy content.
  • Citation and evidence quality: Academic-style text that cites sources and qualifies claims with evidence tends to be more reliable training data than unsupported assertions. Some pipelines use structural markers (citation patterns, hedging language like "according to" or "researchers found") as lightweight quality signals.

The "textbook quality" approach from the Phi-1 paper showed that filtering for educationally rich text and synthetically generating high-quality examples could produce capable small models from surprisingly little data. A 1.3B parameter model trained on 6B tokens of textbook-quality data matched models trained on 300B tokens of conventional web data on several coding benchmarks. This finding reinforced that the marginal value of data quality exceeds the marginal value of data quantity in many training regimes.

Visualizing Filter Precision and Recall

The fundamental constraint of quality filtering is the precision-recall tradeoff. Precision measures what fraction of the documents you keep are high quality. Recall measures what fraction of the high-quality documents you successfully retain. A perfect filter would have both precision and recall equal to 1.0, but in practice the two values trade off against each other through the threshold parameter.

When you raise the threshold (require a higher quality score to pass), precision goes up (fewer garbage documents pass) but recall goes down (more legitimate documents fail). When you lower the threshold, recall goes up but precision goes down. The right operating point on this curve depends on the relative costs of the two errors.

In data curation, the cost of low recall depends on how much data you have. If you have an enormous raw corpus, you can afford to be strict and discard ambiguous documents. If you are targeting a niche domain with limited available text, discarding borderline documents could leave you with a corpus that is too small for effective training. The cost of low precision depends on how much the garbage hurts. For general-purpose pretraining, some noise may be tolerable. For safety-critical fine-tuning, even a small fraction of problematic content can have outsized effects.

In[14]:
Code
import numpy as np
from sklearn.metrics import precision_recall_curve

# Generate synthetic quality scores for a large corpus
rng = np.random.default_rng(42)
n_docs = 10000
true_quality = rng.binomial(1, 0.35, size=n_docs)  # 35% high quality

# Simulate classifier scores: high-quality docs tend to score higher
scores_hq = rng.beta(a=4.0, b=2.0, size=true_quality.sum())
scores_lq = rng.beta(a=2.0, b=5.0, size=(1 - true_quality).sum())

all_scores = np.empty(n_docs)
all_scores[true_quality == 1] = scores_hq
all_scores[true_quality == 0] = scores_lq

precision, recall, thresholds = precision_recall_curve(true_quality, all_scores)
Out[15]:
Visualization
Precision-recall curve showing the tradeoff between retaining good documents and excluding bad ones.
Precision-recall curve for a quality classifier across all possible score thresholds. Three operating points are labeled: high recall (lenient threshold, keeps most good documents but admits more garbage), balanced (moderate tradeoff), and high precision (strict threshold, keeps only the most confidently good documents). The shape of the curve reflects how well-separated the classifier''s score distributions are for high-quality and low-quality documents.

The curve illustrates the fundamental tradeoff. Requiring high precision means setting a strict threshold, which means filtering out more documents including some legitimate ones (low recall). Requiring high recall means setting a lenient threshold, accepting more low-quality documents (lower precision). The right operating point depends on how much high-quality data you have available. If you have an abundance of data, you can afford to be strict. If your target domain is data-sparse, you may need to accept more noise.

The "balanced" operating point shown at roughly equal precision and recall is not always the right choice, even when the costs of both errors seem symmetric. For language model pretraining, most practitioners find that a recall-biased operating point (say, 0.85 recall, 0.70 precision) is preferable to a precision-biased one at the same area under the curve, because the model can partially compensate for noisy training data through its own averaging during pretraining, but it cannot compensate for data it never saw.

Setting Filter Thresholds

Choosing the right threshold is as important as choosing the right filtering method. Set it too strict and you discard legitimate training data, reducing diversity and potentially introducing biases. Set it too lenient and you let through too much noise. There is no universally correct answer, but several strategies help.

Threshold Selection Strategies

The most principled approach is downstream evaluation. You vary the threshold, train (or fine-tune) a model on each filtered corpus, and measure performance on held-out benchmarks. The threshold that maximizes downstream task performance is the right one. This is expensive (it requires training multiple models) but gives ground truth. For large-scale pretraining, this approach is often impractical, but it is feasible for domain-specific fine-tuning where training is cheap.

A cheaper proxy is to monitor the acceptance rate at different thresholds. If your filter accepts 15% of documents, you might be too strict unless you have an enormous raw corpus. If it accepts 85%, you are probably not filtering enough for a general-purpose quality pipeline. Most production pipelines aim for an acceptance rate in the 20-60% range for quality-focused filtering, though this depends heavily on the quality of the source crawl. Fresh, high-quality crawls from curated seeds may have 60% acceptance rates; broad Common Crawl data might have 20-30% acceptance rates.

Percentile-based thresholds avoid the sensitivity to reference model calibration. Instead of "reject documents with perplexity > 500", you set "reject the top 30% of documents by perplexity." The percentile threshold is stable across different reference models and corpora, making it easier to reproduce and compare across pipeline iterations.

Visual inspection remains important at every stage. After setting a threshold, sample 100 accepted documents and 100 rejected documents and review them manually. You will almost always find surprises: legitimate text being rejected for unexpected reasons, or garbage that somehow passes the filter. Manual inspection catches failure modes that quantitative metrics miss. In practice, teams at Google, OpenAI, and other organizations report spending significant time reading samples from their filtered corpora before finalizing thresholds.

Stratified threshold analysis is an advanced strategy for pipelines that care about corpus diversity. Rather than setting a single acceptance rate across all domains, you set domain-specific acceptance rates. You might accept 70% of Wikipedia-style documents (because they are almost all clean) while accepting only 25% of forum content (because it has a higher noise rate). The domain labels can come from a separate classifier, URL patterns, or the language identification stage. This ensures that stricter filtering on noisy sources does not eliminate the noisier-but-legitimate domains entirely.

Threshold Effects on Corpus Statistics

Let's look at how different threshold choices affect corpus composition in a concrete way.

In[16]:
Code
import numpy as np

rng = np.random.default_rng(0)
n_docs = 100_000

# Simulate document corpus with quality scores and text properties
quality_scores = rng.beta(2.5, 4.0, size=n_docs)  # Most docs are low quality
doc_lengths = rng.lognormal(mean=6.5, sigma=1.0, size=n_docs).astype(int)
domain_labels = rng.choice(
    ["news", "wiki", "forums", "spam", "boilerplate"],
    size=n_docs,
    p=[0.15, 0.10, 0.25, 0.30, 0.20],
)

# Higher quality scores correlate with certain domains
domain_score_boost = {
    "news": 0.25,
    "wiki": 0.30,
    "forums": 0.05,
    "spam": -0.20,
    "boilerplate": -0.25,
}
for i, domain in enumerate(domain_labels):
    quality_scores[i] = np.clip(
        quality_scores[i] + domain_score_boost[domain], 0, 1
    )

thresholds = [0.2, 0.35, 0.5, 0.65, 0.8]
Out[17]:
Console
 Threshold   Accept%  Docs Kept  Mean Length Domain Distribution (news/wiki/forums/spam/bplate)
----------------------------------------------------------------------------------------------------
      0.20     68.2%     68,174         1103   22.1% / 14.5% / 33.6% / 19.6% / 10.3%
      0.35     48.8%     48,828         1103   29.8% / 20.1% / 33.2% / 11.6% / 5.3%
      0.50     30.4%     30,379         1103   37.3% / 27.1% / 29.0% / 4.8% / 1.7%
      0.65     15.5%     15,505         1091   42.9% / 35.2% / 21.0% / 0.9% / 0.1%
      0.80      6.2%      6,175         1101   45.8% / 43.7% / 10.5% / 0.0% / 0.0%

As the threshold increases, the acceptance rate drops and the domain composition shifts away from spam and boilerplate toward news and Wikipedia content. This is exactly the desired behavior. However, note that even at the strictest threshold, some spam still passes through, and some legitimate forum content is excluded. No threshold completely separates the classes.

Notice also how the accepted corpus becomes less diverse as the threshold rises. High-threshold filtering produces a corpus concentrated in formal, structured text (news, encyclopedia entries) and underrepresents informal but legitimate content (forums, personal blogs). This can affect the model's downstream capabilities: models trained only on formal text may struggle with casual conversational queries, colloquial language, or the writing styles of communities whose text disproportionately appears in lower-scored domains.

Visualizing Threshold Effects on Domain Composition

In[18]:
Code
import matplotlib
import matplotlib.pyplot as plt
import numpy as np

domains = ["news", "wiki", "forums", "spam", "boilerplate"]
domain_colors = {
    "news": theme_color("#4c72b0"),
    "wiki": theme_color("#55a868"),
    "forums": theme_color("#c44e52"),
    "spam": theme_color("#dd8452"),
    "boilerplate": theme_color("#8172b3"),
}

threshold_vals = np.linspace(0.1, 0.9, 50)
domain_fracs_by_thresh = {d: [] for d in domains}
accept_rates = []

for thresh in threshold_vals:
    mask = quality_scores >= thresh
    kept = mask.sum()
    accept_rates.append(100 * kept / n_docs)
    for d in domains:
        frac = (domain_labels[mask] == d).sum() / kept if kept > 0 else 0
        domain_fracs_by_thresh[d].append(frac)
Out[19]:
Visualization
Stacked area chart of domain fractions and line chart of acceptance rate vs. quality threshold.
Domain composition and corpus size as functions of quality score threshold. The left panel shows how each domain''s share of the accepted corpus shifts with threshold: spam (orange) and boilerplate (purple) fractions decline as expected, but forum content (red) also shrinks, illustrating how aggressive filtering reduces linguistic diversity. The right panel shows the sharp decline in total corpus size as the threshold increases, revealing the data efficiency cost of strict quality requirements.

The right panel makes the cost of aggressive filtering concrete: moving from a threshold of 0.3 to 0.7 cuts the accepted corpus to a fraction of its former size. The left panel shows what you gain and lose in detail. The gains are obvious: less spam and boilerplate. The losses are subtler: you also lose forum content that includes useful informal language patterns, colloquialisms, and practical knowledge that does not appear in formal sources. A model trained on the high-threshold corpus will produce more formal text by default and will have worse calibration on casual queries.

This visualization also reveals an important property of corpus composition: the relative fractions of domains shift even when you do not explicitly target them. A quality filter calibrated entirely on spam versus clean text will still disproportionately remove forum content because forum text happens to share surface properties with lower-quality content. If forum coverage matters for your downstream use case, you need either a domain-stratified filter or a post-filtering resampling step that restores target domain proportions.

Combining Filtering Methods

No single filtering method is complete on its own. Production pipelines chain multiple filters in sequence, with each stage removing different failure modes. The order of stages reflects a cost-accuracy tradeoff: cheap methods run first to reduce the volume reaching expensive methods, while more sophisticated methods handle the harder cases that simple rules cannot resolve.

The typical ordering of a production filtering pipeline proceeds through four stages. Language identification runs first. Documents classified as not being in the target language are removed immediately, and this alone eliminates a large fraction of raw web data for any target language. Even for a multilingual corpus, this stage assigns language confidence scores and routes documents to language-specific downstream processing.

Heuristic filters run next because they are fast (microseconds per document) and remove obvious garbage (symbol-heavy text, duplicated lines, too-short documents) before more expensive methods are applied. A heuristic stage that processes a trillion-token corpus in a few hours might prevent the perplexity stage from having to score 30-40% of the raw documents that would have been discarded anyway.

Perplexity filtering runs after heuristics. It is slower than heuristics (requiring n-gram model inference) but faster than neural classifiers. Documents surviving heuristic filtering are scored against a KenLM model and those above a percentile threshold are removed. This stage is particularly effective at catching coherently formatted but linguistically unnatural text: machine-translated boilerplate, auto-generated content, and documents in the wrong register for the target domain.

Classifier-based filtering is the final and most expensive stage. Only documents that pass the previous stages reach the classifier. This keeps inference cost manageable while getting the benefit of a high-capacity classifier for the hard borderline cases that heuristics and perplexity cannot cleanly resolve. The classifier has seen enough of the data distribution from the positive and negative examples to recognize quality patterns that are invisible to simpler methods.

The pipeline design reflects a deliberate cost-accuracy tradeoff. Each stage is more accurate but slower than the previous one. By placing cheap methods first, you ensure that expensive methods only process documents that need careful evaluation.

A Practical Filtering Pipeline

In[20]:
Code
from dataclasses import dataclass
from typing import Callable, List


@dataclass
class FilterResult:
    kept: bool
    reason: str


class FilterPipeline:
    """Chainable document quality filter."""

    def __init__(self):
        self.filters: List[tuple[str, Callable[[str], bool]]] = []

    def add_filter(self, name: str, fn: Callable[[str], bool]):
        """Add a filter. Function returns True if document should be KEPT."""
        self.filters.append((name, fn))
        return self

    def evaluate(self, text: str) -> FilterResult:
        for name, fn in self.filters:
            if not fn(text):
                return FilterResult(kept=False, reason=name)
        return FilterResult(kept=True, reason="passed_all")

    def batch_filter(self, texts: list) -> dict:
        results = {"kept": [], "rejected": []}
        rejection_counts = {}
        for text in texts:
            result = self.evaluate(text)
            if result.kept:
                results["kept"].append(text)
            else:
                results["rejected"].append(text)
                rejection_counts[result.reason] = (
                    rejection_counts.get(result.reason, 0) + 1
                )
        results["rejection_breakdown"] = rejection_counts
        return results


# Build a pipeline with increasing specificity
pipeline = FilterPipeline()

pipeline.add_filter("min_length", lambda t: len(t.split()) >= 30)
pipeline.add_filter("symbol_ratio", lambda t: symbol_ratio(t) < 0.25)
pipeline.add_filter("word_repetition", lambda t: max_word_repetition(t) < 0.5)
pipeline.add_filter("duplicate_lines", lambda t: duplicate_line_ratio(t) < 0.3)
pipeline.add_filter("perplexity_threshold", lambda t: lm.perplexity(t) < 400)

# Test documents of varying quality
test_corpus = [
    "Scientists have developed a new battery technology that could double the range of electric vehicles. The lithium-sulfur chemistry offers higher energy density than conventional lithium-ion cells and uses cheaper materials. Field trials are expected to begin next year.",
    "buy cheap insurance online best insurance deals lowest rates guaranteed compare insurance quotes now save money on auto home life insurance click here for free quote today",
    "The committee reviewed proposals for urban infrastructure development including transportation networks water management and housing density regulations. Several competing approaches were evaluated against long-term sustainability criteria.",
    "!!! FREE !!! click now limited time 50% off everything sale ends midnight tonight dont miss this amazing deal buy now pay later no interest",
    "Home | Products | Services | About | Contact | Blog | Privacy Policy | Terms | Sitemap\nHome | Products | Services | About | Contact | Blog | Privacy Policy | Terms | Sitemap",
    "var x = document.getElementById('main'); if (x) { x.style.display = 'block'; } else { console.log('not found'); } window.onload = function() { init(); };",
    "The philosophical tradition stretching from Kant through Hegel to contemporary analytic philosophy grapples with persistent questions about the nature of knowledge, justification, and the relationship between mind and world. These debates inform how we understand scientific practice.",
    "weather weather weather forecast forecast today tonight tomorrow weekend extended forecast high low temperatures precipitation probability humidity wind speed direction",
]
Out[21]:
Console
Pipeline results: 2/8 documents passed all filters (25%)

Rejection breakdown:
  min_length               : 5 document(s)
  word_repetition          : 1 document(s)

Kept documents:
  - Scientists have developed a new battery technology that could double the range of electric...
  - The philosophical tradition stretching from Kant through Hegel to contemporary analytic ph...

The pipeline successfully routes different failure modes to different filters. The word repetition filter catches keyword spam. The duplicate line filter catches navigation boilerplate. The perplexity filter catches language model-detectable incoherence. Documents that are legitimate prose pass through all stages. The rejection breakdown tells you which filter is doing the most work, which helps calibrate the pipeline: if the perplexity filter is removing very few documents after the heuristics, you might be able to drop it and save computation; if it is removing many documents, it is earning its cost.

In a real pipeline at scale, the rejection breakdown is also a diagnostic tool. If your "symbol_ratio" filter is suddenly rejecting 40% of documents in a new crawl batch, that is a signal that something changed in the upstream HTML extraction: perhaps a new template is generating more symbol-heavy artifacts. Monitoring rejection rates by stage over time turns the filtering pipeline into a data quality sensor.

Soft Scoring Versus Hard Filtering

The pipeline above uses hard binary decisions: each filter either keeps or discards a document. An alternative is soft scoring, where each stage assigns a real-valued quality estimate and the final inclusion decision is based on a combined score.

Soft scoring has several advantages. It allows a document that barely fails one filter but strongly passes all others to potentially be included, which a hard pipeline would miss. It enables downstream users of the corpus to apply different quality cuts without reprocessing the raw data. And it provides gradient-like signals that help calibrate thresholds by examining the score distribution.

The practical implementation assigns each document a vector of per-filter scores (heuristic score, perplexity score, classifier score) and combines them with a learned or hand-tuned weighting. The Dolma toolkit (Allen AI's open-source data curation pipeline) supports this soft-scoring approach, storing per-document quality attributes alongside the filtered text so that users can apply different thresholds for different downstream applications.

Limitations and Practical Implications

Quality filtering is powerful but has well-documented failure modes that practitioners need to understand. These are not edge cases to be noted and dismissed: they have real consequences for the models produced and the communities affected.

Demographic bias is the most thoroughly documented concern. Filters trained on "high quality" examples that are predominantly formal English text from mainstream sources will systematically score African American Vernacular English (AAVE), code-switching, and other non-standard dialects as low quality. The C4 dataset was shown to significantly underrepresent text associated with Black, Hispanic, and other minority communities relative to their representation in the raw web. Measurements by researchers at the University of Washington found that documents containing dialect features associated with minority communities were disproportionately removed by the C4 filtering pipeline. If a model is trained on such filtered data, it performs worse on text from those communities, generating more errors and less fluent completions. This is not hypothetical: it has been measured in multiple independent studies. Any filtering pipeline that uses a model trained on a biased reference corpus will inherit and potentially amplify that bias.

Domain exclusion is a subtler problem. Quality filters calibrated on news and encyclopedia text will deprioritize legal documents, scientific papers, programming documentation, and other specialized domains that have different linguistic patterns. This can hurt the model's performance on specialized tasks, even tasks that are highly valuable (legal AI assistants, scientific literature search, code generation). The solution is domain-stratified filtering: separate quality models for different content types, or domain-specific thresholds that account for expected differences in register and style. Implementing this requires a domain classifier that runs before the quality classifier, adding another component to the pipeline but substantially improving coverage.

Filtering for scale introduces a different pressure. When your raw corpus is in the petabyte range, even a 1% false negative rate (failing to remove low-quality documents) translates to hundreds of terabytes of garbage in the training set. This pushes practitioners toward conservative thresholds, which in turn makes demographic bias and domain exclusion worse. The tension between scale-appropriate strictness and representational fairness has no clean engineering resolution. It requires explicit policy decisions about acceptable tradeoffs.

Temporal consistency is another challenge. Web text quality distribution shifts over time. A quality model trained in 2020 may perform poorly on 2024 web crawl data because the patterns of spam, machine-generated content, and boilerplate have changed. In particular, the massive increase in LLM-generated content on the web since 2022 has introduced a new category of technically clean but potentially low-value text: content that passes all surface-level quality checks but was generated by a model rather than a human. Production pipelines need regular re-evaluation and potentially retraining on fresh samples. This is operationally non-trivial: it requires maintaining labeled datasets, retraining classifiers, and re-filtering previously processed crawl data.

Over-filtering and the homogeneity problem affect downstream model behavior in ways that are difficult to measure but consequential. Models trained on aggressively filtered data tend to produce more uniform, formal-sounding text. They handle formal queries well but may sound stilted in casual conversation or fail to match the register of informal queries. They may underperform on tasks involving non-standard writing (social media analysis, dialect translation, informal summarization). The filtering pipeline's threshold settings are effectively hyperparameters of the final model, and the effects of those hyperparameters on model behavior are poorly characterized in most published work.

For researchers and practitioners, quality filtering is both a technical and a policy decision. The thresholds you choose, the reference model you train on, and the heuristics you encode all embed assumptions about what constitutes good text and whose language is considered standard. Documenting these choices, measuring their demographic effects, and auditing samples from filtered corpora for unexpected biases is increasingly expected in responsible data curation practice. Several recent data papers (Dolma, RedPajama-V2, ROOTS) explicitly include quality scores as separate metadata rather than pre-applying hard thresholds, precisely to allow downstream users to apply different filtering policies appropriate to their use cases.

Summary

Quality filtering removes low-value text from web-scraped corpora to improve model training efficiency and output quality. The three main approaches differ in cost, accuracy, and failure mode:

  • Heuristic filters: Fast rule-based filters that catch obvious garbage including symbol-heavy text, keyword repetition, boilerplate lines, and too-short documents. They cost microseconds per document and scale to petabytes. The limitations are brittleness (rules designed for one domain or language may fail on others) and potential demographic bias in the rules themselves.
  • Perplexity filtering: Scores documents against a reference language model trained on known-good text. Documents with high perplexity (surprising to the model) are flagged for removal. More accurate than heuristics for detecting incoherent or non-target-domain text. Uses percentile-based thresholds for robustness across different reference models. The main limitations are domain sensitivity (the reference model's training distribution determines what gets high perplexity) and the circular bias risk when the reference corpus is itself demographically skewed.
  • Classifier-based filtering: Trains a binary quality classifier from labeled examples. More flexible and often more accurate than perplexity-based approaches. Can use logistic regression on character n-gram features (fast, interpretable) or neural classifiers with multiple input signals (more accurate). The main limitations are the cost of labeled data and inference, and the risk of inheriting whatever biases exist in the labeling process.

Threshold selection involves tradeoffs between precision (fewer false passes) and recall (fewer false rejections). Downstream evaluation, percentile-based thresholds, acceptance rate monitoring, and manual inspection of samples are the key tools. Domain-stratified filtering and soft scoring extend the basic pipeline to handle heterogeneous corpora.

Real pipelines chain all three approaches in order of increasing cost. Each stage removes different failure modes, and careful threshold calibration determines the final quality-diversity tradeoff. The key limitation is that "quality" is defined relative to a reference corpus, making filters potentially biased against non-standard language and specialized domains. Documenting these choices, measuring their demographic effects, and auditing filtered corpora for unexpected biases is part of responsible data curation practice. The models that result from this pipeline will reflect the filtering decisions made here: every threshold, every heuristic, and every reference corpus choice shapes what the model learns and what communities it serves well.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about quality filtering for language model training data.

Quality Filtering Quiz

Question 1 of 80 of 8 completed
Why is perplexity filtering effective at identifying low-quality text?

Comments

1 comment

  1. ADITYA KumarMember

    Till now, this is the best resource on quality filter. Very easy and crisp information and it doesn't feel boring also.

    1. Michael BrenndoerferMember

      Thanks Aditya - glad you found it helpful and clear!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026qualityfiltering, author = {Michael Brenndoerfer}, title = {Quality Filtering: Heuristics, Perplexity, and Classifiers}, year = {2026}, url = {https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Quality Filtering: Heuristics, Perplexity, and Classifiers. Retrieved from https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds
MLAAcademic
Michael Brenndoerfer. "Quality Filtering: Heuristics, Perplexity, and Classifiers." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds>.
CHICAGOAcademic
Michael Brenndoerfer. "Quality Filtering: Heuristics, Perplexity, and Classifiers." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Quality Filtering: Heuristics, Perplexity, and Classifiers'. Available at: https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Quality Filtering: Heuristics, Perplexity, and Classifiers. https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.