Toxicity Filtering: Classifiers, Thresholds

Michael BrenndoerferJanuary 10, 202649 min read

Part of Language AI Handbook

Explains how toxicity classifiers work, how to calibrate thresholds for pretraining data.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Toxicity Filtering

When you build a language model on web-scale data, you are training on a cross-section of human expression that includes encyclopedias, news articles, hate speech, violent content, graphic descriptions, and targeted harassment. Left unfiltered, this material embeds deeply into model weights. Models trained on toxic content reproduce it, amplify it, and sometimes generate it unprompted in polite conversation. Toxicity filtering is the process of identifying and removing this material from your training corpus before the model ever sees it.

This chapter builds on the quality filtering techniques we discussed in the previous chapter, which focused on removing low-quality, malformed, or uninformative text. Toxicity filtering addresses a different dimension: content that is structurally fine but normatively harmful. A document about white supremacist ideology may be grammatically perfect, richly formatted, and informationally dense. Quality filters will not catch it. Toxicity filters must. The distinction matters because the two types of filtering require different tools, different thresholds, and different tradeoffs.

The core challenge is not building a classifier that detects obvious slurs. Any simple heuristic can catch the most explicit content. The hard problem is balancing two competing errors: missing toxic content that degrades your model versus over-filtering legitimate text and inadvertently erasing minority voices, discussions of historical violence, or literary content that depicts harm in order to critique it. Both errors are costly, and the right balance depends on your application, your user population, and your values as a practitioner. This chapter works through both sides of that balance in detail, covering the classifiers you will use, the threshold decisions you will face, the over-filtering risks you must monitor, and how to evaluate a complete filtering pipeline.

What Toxicity Means in This Context

Before you can filter toxic content, you need a working definition. In natural language, "toxic" covers a wide range of unpleasant content, from mildly rude comments to calls for genocide. For training data curation, you need a definition precise enough to guide annotation decisions and evaluate classifiers. In the context of pretraining data, toxicity generally refers to content that:

  • Promotes or glorifies harm against people based on protected characteristics (race, religion, gender, sexual orientation, national origin, disability)
  • Contains explicit threats, targeted harassment, or personal attacks
  • Includes graphic depictions of sexual violence or child sexual abuse material (CSAM)
  • Advocates for genocide, terrorism, or mass violence
  • Dehumanizes individuals or groups through slurs, comparisons to vermin, or eliminationist language

Notice what this definition excludes. Adult content between consenting parties may be appropriate for some applications and not others, but it is not inherently toxic in the sense used here. Discussion of historical atrocities is not toxic even if it describes terrible events in detail. Fiction depicting villains is not toxic even when the villain espouses reprehensible views. A news article reporting on a hate crime may quote the perpetrator's language directly, and that direct quotation serves an important informational purpose. These distinctions are critical because naive toxicity classifiers frequently fail to make them, and we will return to this problem at length when we discuss over-filtering risks.

Toxicity vs. Content Policy

Toxicity filtering for training data is distinct from content moderation applied to user-facing outputs. Training data filtering is coarser and operates at scale, where you tolerate some false positives (removing acceptable content) to efficiently remove a large, diffuse signal. Output-side moderation is finer-grained and more context-aware, and we will cover it in later chapters on safety and alignment.

The practical framing that most practitioners adopt is that toxicity filtering targets content that, if the model memorized and reproduced it verbatim, would constitute a failure mode. This gives you a concrete test: would a reasonable person be alarmed if this exact text appeared in your model's output? If yes, consider filtering it from training. If the text describes harm in order to analyze or condemn it, and a reasonable person would understand the analytical framing, it probably belongs in your training corpus.

This framing also clarifies the relationship between toxicity and quality. Toxic content is not low quality in the conventional sense: a well-written manifesto calling for ethnic cleansing is far more dangerous as training material than a low-quality but benign spam message. Quality filters improve the information density of the corpus; toxicity filters shape the values and behavioral tendencies the model learns. You need both, and they are largely independent.

The Taxonomy Problem

Defining exactly what counts as toxic is harder than it looks, and the difficulty has real engineering consequences. Consider these four texts:

  1. "I hate [group]. They should be driven out."
  2. "Surveys show that 45% of respondents expressed hostile attitudes toward [group]."
  3. "[Group] people are lazy and untrustworthy." (an implicit stereotype without explicit threat)
  4. "The claim that [group] people are lazy is a racist myth that studies consistently refute."

Text 1 is clearly toxic. Text 4 is clearly not. Texts 2 and 3 are harder. Text 2 is a factual statement about a social phenomenon; it is exactly the kind of evidence-based discussion you want in a training corpus that will be used to build an informed model. Text 3 contains no slurs, no threats, and no explicit calls for harm, but it promotes a harmful stereotype that has real-world consequences. Whether you filter text 3 depends on your application and your threshold.

The taxonomy problem compounds when you move to multi-class classification. Jigsaw's Perspective API taxonomy distinguishes between general toxicity, severe toxicity, obscenity, threats, insults, and identity attacks. Each of these categories has its own annotation guidelines, edge cases, and borderline instances, and they are not mutually exclusive. A single comment may score high on multiple attributes, and the interaction between attributes matters when you are making a filtering decision. We will return to multi-attribute filtering in the threshold section.

Toxicity Classifiers

The primary tool for toxicity filtering at scale is a learned classifier. Rule-based approaches, such as blocklists of slurs and hate speech terms, are too brittle. They fail on creative spellings, miss implicit forms entirely, and produce high false-positive rates on educational and journalistic content that discusses toxic language without endorsing it. The word "slur" is not itself a slur; blocking it damages content that discusses and critiques slurs. Learned classifiers, with all their limitations, are more flexible.

The Perspective API and Its Successors

The most widely used toxicity classifier in the research literature is Google's Perspective API, developed by Jigsaw. Perspective uses a convolutional neural network trained on millions of comments labeled by human raters for attributes including toxicity, severe toxicity, obscenity, threat, insult, and identity attack. The model outputs a probability score for each attribute, and practitioners apply a threshold to decide what to filter.

Perspective was a landmark contribution because it demonstrated that classifier-based toxicity detection at scale was feasible and that the outputs were meaningful enough to support deployment. The API handles millions of requests per day and underpins moderation systems for major platforms. But it also revealed the core limitation of classifier-based approaches: the model's notion of toxicity reflects the values and blind spots of the annotators who labeled the training data.

Research published by Dixon et al. (2018) and subsequent work showed that Perspective assigned higher toxicity scores to text mentioning Black people, Muslim people, LGBTQ+ people, and women, even when the sentiment was neutral or positive. A sentence like "I am a gay man" received a higher toxicity score than "I am a straight man." A sentence like "She is a strong Black woman" scored higher than "She is a strong white woman." This bias arose because the training data contained more toxic speech targeting these groups, and the model learned superficial correlations with the group labels rather than the presence of harm. The classifier essentially learned "text that mentions marginalized groups is more likely to be toxic" from its training set, which reflected a real statistical pattern in the data, but that correlation does not hold for the vast majority of documents that simply discuss or mention those groups.

More recent classifiers have addressed some of these issues through better annotation design, diverse annotator pools, and adversarial training. The Detoxify library, for example, trains on the Jigsaw dataset with added robustness against identity bias, using a technique called "unintended bias mitigation" that explicitly penalizes differential false positive rates across groups. Meta's LLaMA Guard and similar models are trained specifically on safety taxonomies aligned with model deployment requirements, incorporating constitutional AI principles alongside empirical annotation. For training data filtering, you typically want a model optimized for recall (catching most toxic content) at the cost of precision (accepting some false positives), because the downstream cost of a missed toxic example is higher than the cost of losing one acceptable document.

Classifier Architectures

Three broad architectural choices are used in practice, each with distinct tradeoffs:

Keyword-based models maintain lists of toxic terms and patterns. They are fast, interpretable, and easy to update, but they suffer from high false-positive rates and can be circumvented by creative misspellings, code words, or dog whistles. Writing "b1ack people" instead of "Black people" to evade a filter is a trivial evasion that does not require any sophistication. These models also struggle with context: the word "kill" appears in countless benign contexts ("kill the lights," "kill it at the presentation"), and a simple blocklist cannot distinguish them. Keyword models work well as a first-stage filter before applying a learned model, where their role is to quickly remove obvious content and reduce the load on the more expensive classifier.

Shallow ML classifiers (logistic regression, linear SVM on n-gram features) can be trained quickly on moderate labeled datasets and produce reasonably calibrated probabilities. Their main weakness is that they miss semantic toxicity: "I hope you get cancer" contains no explicit slurs but is clearly a threat. A logistic regression over unigrams will not learn this mapping without seeing many similar examples labeled explicitly, and it will not generalize to paraphrases. Bag-of-words features also ignore word order, which matters for distinguishing "The teacher killed the student's grade" from more violent phrasings.

Neural classifiers (fine-tuned BERT, RoBERTa, or purpose-built architectures) can capture semantic and contextual signals. A sentence like "You're so bright you must be a black hole" is benign, while the same sentence with an offensive substitution is toxic, and a neural model trained on enough examples can learn this distinction because it encodes the sentence as a whole rather than a bag of features. The tradeoff is computational cost, which matters when you are classifying hundreds of millions of documents. A single Detoxify inference call on a CPU takes on the order of milliseconds; at a billion documents, the cumulative cost is substantial.

For web-scale filtering, the common approach is a two-stage pipeline: fast keyword pre-filtering removes obvious cases, then a neural classifier processes the remaining documents. This reduces the expensive classifier calls by a large factor while preserving recall on the hard cases. A keyword filter might remove 30-50% of the most obvious toxic content in a single pass, leaving the neural classifier to handle context-dependent cases. The two-stage architecture also makes the filtering pipeline auditable: you can inspect what the keyword filter removed separately from what the neural classifier removed, which is useful for debugging and bias analysis.

Training a Custom Classifier

Off-the-shelf classifiers like Perspective and Detoxify are trained on general web data, primarily English. If your corpus is multilingual, domain-specific, or from a particular time period, you will likely need to fine-tune or train a custom classifier. A classifier trained on English Reddit comments will perform poorly on Arabic web forums because of language differences and because the types and styles of toxicity differ across cultural contexts.

The key ingredients for a custom classifier are:

  • Labeled examples: Both toxic and non-toxic. You need enough toxic examples to cover the specific types of harm you are targeting, and enough non-toxic examples that closely resemble toxic ones (to prevent false positives) to teach the model the boundary.
  • Annotator guidelines: Explicit criteria about what counts as toxic, including worked examples of edge cases. Ambiguous cases should be adjudicated and documented, not just discarded. The documentation of adjudicated cases becomes a resource for understanding and updating the classifier's coverage.
  • Calibration data: Examples specifically designed to test for false positives on content that superficially resembles toxic text, including counter-speech, academic discussion, fiction, and journalistic reporting. If your calibration data does not include these categories, your classifier may fail silently on exactly the content types where false positives are most costly.

One recurring finding in the literature is that annotation disagreement is high for toxicity. When multiple annotators label the same text, they often disagree, especially on borderline cases. This disagreement is informative in two ways. First, it marks the ambiguous region of the distribution, which is precisely where your threshold decisions matter most. Second, it signals how annotator demographics affect labeling: annotators from marginalized communities often assess text differently from majority-group annotators for the same content.

Rather than discarding disagreement, some approaches train with soft labels (the average annotation score across raters) or with uncertainty estimates that allow the model to express low confidence on ambiguous examples. Soft-label training produces models that assign intermediate probabilities to borderline content, which is exactly what you want when making threshold decisions: you want your model to be uncertain about uncertain cases, not overconfidently assign them to one class.

Inter-annotator agreement is typically measured with Cohen's kappa (κ\kappa) or Krippendorff's alpha (α\alpha). The formulas differ but both measure agreement above chance. For Cohen's kappa between two annotators:

κ=po−pe1−pe\kappa = \frac{p_o - p_e}{1 - p_e}

where:

  • pop_o: observed agreement (fraction of examples both annotators labeled identically)
  • pep_e: expected agreement by chance, computed from the marginal label frequencies

For toxicity annotation, kappa values around 0.6-0.7 are considered acceptable, indicating substantial agreement. Values below 0.5 signal that the annotation task is too loosely defined and the guidelines need refinement. When kappa is low, the resulting model inherits the inconsistency of the annotation: similar documents may receive very different scores depending on which annotators labeled which documents, and the classifier's behavior becomes less predictable in exactly the region where you most need predictability.

One practical consideration when building a custom classifier is the need for hard negative mining. A classifier trained with randomly sampled negative examples (non-toxic documents from the web) will learn an easy boundary because the random negatives are clearly benign. At deployment time, the hard cases, the counter-speech, the academic discussion, the dark humor, all receive scores higher than the training negatives warranted, because the model has never seen negatives that look like those. This causes higher false positive rates in production than in evaluation. Hard negative mining, the deliberate inclusion of difficult non-toxic examples in training, is essential for closing this gap. Common sources of hard negatives include news articles about hate speech, academic papers on extremism, dictionaries and educational resources about slurs, court documents describing hate crimes, and fiction with morally complex characters.

The amount of hard negatives needed depends on the classifier architecture and the distribution of the corpus you are filtering. A practical approach is to run your trained classifier on a held-out validation set, identify the false positives with the highest toxicity scores, analyze their content to understand what class of hard negative is missing, and add representative examples of that class to the training set. Repeat until the false positive rate on the validation set is acceptable.

Toxicity Thresholds

A toxicity classifier produces a score between 0 and 1. To make a filtering decision, you must choose a threshold above which a document is considered toxic and removed from the corpus. The threshold is a technical choice and a values decision, with measurable consequences for what your model learns and what communities are represented in the training data.

The Precision-Recall Tradeoff

At high thresholds (close to 1.0), you filter only text that the classifier considers almost certainly toxic. Precision is high: most filtered documents are problematic. Recall is low: many toxic documents that received scores between your threshold and 1.0 remain in the corpus.

At low thresholds (close to 0.0), you filter everything the classifier assigns any toxicity signal to. Recall is high, but precision collapses: you remove many acceptable documents.

For training data, the asymmetry matters in both directions. Missing toxic content (low recall) means the model learns from it. Filtering too aggressively (low precision) removes legitimate content and skews the training distribution. Research on this tradeoff uses a formal framing: given a toxicity classifier with score s(d)s(d) for document dd and a threshold τ\tau, we filter documents where:

s(d)≥τs(d) \geq \tau

where:

  • s(d)s(d): the classifier's toxicity score for document dd, between 0 and 1
  • τ\tau: the filtering threshold, a value in [0,1][0, 1] chosen by the practitioner
  • Documents with s(d)≥τs(d) \geq \tau are removed from the training set

The expected precision P(τ)P(\tau) and recall R(τ)R(\tau) as functions of threshold are:

P(τ)=true toxics filtered at τall documents filtered at τP(\tau) = \frac{\text{true toxics filtered at } \tau}{\text{all documents filtered at } \tau} R(τ)=true toxics filtered at τall true toxics in corpusR(\tau) = \frac{\text{true toxics filtered at } \tau}{\text{all true toxics in corpus}}

These quantities cannot be computed without ground truth labels for the full corpus, which is why calibrating a threshold requires a held-out labeled validation set that mirrors the distribution of your corpus. A common mistake is calibrating thresholds against a dataset of known toxic examples, which will give you optimistic recall estimates but tell you nothing about precision or false positive rates on benign content.

Calibrating Against a Validation Set

To pick a threshold, you need a validation set with ground truth labels. This set should be:

  • Representative of the corpus you are filtering, including more than high-confidence toxic examples. The hard cases, the borderline documents and the sensitive-but-benign documents, should be proportionally represented.
  • Balanced enough to allow measurement at both ends of the precision-recall curve. An all-toxic validation set cannot measure precision. An all-benign validation set cannot measure recall. You need a mix that reflects your production distribution.
  • Annotated by multiple raters with adjudicated decisions for borderline cases. Single-annotator labels for borderline cases are unreliable and will introduce noise into your threshold selection.

With a validation set in hand, you plot the precision-recall curve across all threshold values from 0 to 1 and identify the operating point that satisfies your requirements. A common approach is to choose τ\tau to maximize the FβF_\beta score, which is a weighted harmonic mean of precision and recall:

Fβ=(1+β2)⋅P⋅Rβ2⋅P+RF_\beta = (1 + \beta^2) \cdot \frac{P \cdot R}{\beta^2 \cdot P + R}

where:

  • β>1\beta > 1: weights recall more heavily, appropriate when missing toxic content is more costly than over-filtering
  • β<1\beta < 1: weights precision more heavily, appropriate when over-filtering is more costly
  • β=1\beta = 1: equal weighting (the standard F1F_1 score)

For most pretraining data curation pipelines, practitioners set β>1\beta > 1 (commonly β=2\beta = 2) because the cost of training on toxic content is considered higher than the cost of losing some acceptable documents. The right β\beta depends on your application. A model intended for use with vulnerable populations (mental health support, crisis intervention) warrants a very high β\beta, accepting aggressive over-filtering to minimize exposure to harmful content. A model intended for security research that must understand adversarial content may warrant a lower β\beta.

The FβF_\beta approach gives you a single operating point, but often you want to see the full curve before committing. A practical approach is to plot the precision-recall curve, mark several candidate thresholds at different recall levels (e.g., "90% recall threshold" and "95% recall threshold"), and then examine a random sample of filtered documents at each threshold to assess whether the additional filtered content at the more aggressive threshold is problematic or benign.

Multi-Attribute Filtering

A single toxicity score is often insufficient. Jigsaw's taxonomy distinguishes between:

  • Toxic: A rude, disrespectful, or unreasonable comment that is likely to make people leave a discussion
  • Severe toxic: A hateful or threatening comment that calls for violence or demeans entire groups
  • Obscene: Contains explicit language (profanity, sexual language)
  • Threat: A direct or indirect threat of physical harm
  • Insult: An insulting, inflammatory, or negative comment toward a person
  • Identity attack: Negative or hateful language targeting a person's identity characteristics

You may want different thresholds for different attributes. A corpus intended for a children's educational assistant might filter at a very low threshold for obscenity even while allowing discussion of historical violence. A corpus for a medical assistant might allow clinical discussion of abuse without filtering it as toxicity, but filter any glorification of violence at a low threshold. A corpus for a legal research model might need to retain verbatim quotes from hate speech cases that would otherwise score high on identity attack.

The practical consequence of multi-attribute filtering is that you must maintain multiple classifiers or a multi-output classifier, and you must decide how to combine signals from multiple attributes. A common approach is to filter if any attribute exceeds its threshold (logical OR), which increases recall at the cost of precision. Alternatively, you can filter only if a primary attribute exceeds its threshold, and use secondary attributes to flag borderline cases for manual review.

The combination logic matters enormously at scale. If you use logical OR across six attributes each with a 10% false positive rate, the combined false positive rate for a benign document approaches:

1−(1−0.10)6≈0.471 - (1 - 0.10)^6 \approx 0.47

Nearly half of all benign documents would be filtered, even with individually well-calibrated classifiers. This calculation assumes attribute independence, which overstates the false positive rate when attributes are positively correlated (a document that scores high on "obscene" is likely to score somewhat higher on "insult" as well, so the effective number of independent classifiers is less than six). Practitioners often address this by treating the six Jigsaw attributes not as six independent filters but as a single multi-output model with a combined decision rule that accounts for correlations between attributes, typically by learning a logistic regression meta-classifier on top of the six attribute scores.

Over-Filtering Risks

The most significant failure mode in toxicity filtering is not insufficient filtering; it is over-filtering. When a toxicity classifier has systematic false-positive biases, filtering at scale removes disproportionate amounts of content from affected communities. This distorts the training data distribution in ways that compound across the pretraining and fine-tuning stages, ultimately degrading model performance precisely for the populations the filtering was meant to protect.

Disparate Impact on Minority Communities

Empirical research on toxicity classifiers consistently finds that text mentioning or produced by marginalized communities receives higher toxicity scores than comparable text from majority communities, even when the content is neutral or positive. Documented examples include:

Label bias: Annotation of toxic content skews based on annotator demographics. Majority-group annotators may rate minority-group vernacular, humor, or reclaimed slurs as more toxic than members of those communities would. African-American English (AAE) uses different syntactic structures and vocabulary than Standard American English, and annotators unfamiliar with AAE may interpret vernacular expressions as aggressive or threatening when they are not. Blodgett et al. (2016) documented significant performance disparities in NLP systems across dialects, and the same dynamic applies to toxicity annotation.

Training data composition: Web-scraped data contains more toxic speech targeting minorities than targeting majorities, because minorities are more frequently targeted in online harassment. A classifier trained on this data learns to associate group mentions with toxicity, because in the training data, sentences mentioning certain groups really are more often toxic. The statistical correlation is real; the problem is that it generalizes incorrectly to the far larger set of benign sentences that also mention those groups.

Indirect language: Hate speech often uses coded language that avoids explicit slurs. The term "globalist" has been used as an antisemitic dog whistle; "inner-city" as a coded racial reference; "groomer" in anti-LGBTQ+ discourse. When classifiers are trained to catch this indirect language, they may generalize too broadly to benign uses of the same terms in entirely different contexts.

The result is that over-filtering produces training corpora with less representation of African-American English dialects, LGBTQ+ discourse, discussions of religious practices associated with minority groups, and content in non-Western languages. Models trained on such corpora may perform worse for these groups, respond less naturally to their communication styles, and have less accurate world knowledge about their communities. Ironically, the filtering intended to protect these communities often harms them by erasing their voices from the training data.

Counter-Speech and Discussion of Harm

A particular failure mode is the filtering of counter-speech: content that explicitly engages with toxic language in order to critique, analyze, or refute it. A hate speech classifier trained on examples of hate speech will often flag counter-speech, because the two are superficially similar in their vocabulary and topic.

Consider these three texts:

  1. "Black people are inferior and should be segregated." (genuine hate speech)
  2. "Claims that Black people are inferior are pseudoscientific racism with no empirical support." (counter-speech)
  3. "Study finds that claims of racial inferiority are used to justify structural discrimination." (academic analysis)

A naive toxicity classifier may assign similar scores to all three, because all three mention racial inferiority in close proximity to racial group labels. Only the first should be filtered. The other two are exactly the kind of content you want in a training corpus if you want your model to understand and refute racism. A model trained without counter-speech will be worse at the important task of explaining why hateful claims are wrong.

The same issue applies to historical education (a text describing Nazi ideology to explain how it enabled genocide), journalism (a news report quoting a hate crime perpetrator's manifesto), and fiction (a novel depicting a racist character to critique racism). All of these may receive high toxicity scores from a classifier trained to maximize recall. The fictional example is particularly tricky because the same text might appear on a literary analysis site or on a hate-endorsing site, and the classifier cannot know which.

One partial solution is to weight documents from trusted domains more conservatively at the classifier level, applying a higher threshold for content from academic publishers, major newspapers, and established reference works. This domain-aware filtering reduces false positive rates for high-quality sources while allowing more aggressive filtering for low-quality web content. The tradeoff is that it requires a reliable domain taxonomy, which is hard to maintain at scale and may itself encode biases about which sources count as "trusted."

Vocabulary Gaps from Over-Filtering

When toxicity filtering removes too much text from specific domains, the resulting vocabulary gap is measurable. Models trained on over-filtered data have weaker representations for terms associated with hate speech contexts, including the terms used to discuss and challenge hate speech. This can manifest as worse performance on hate speech detection tasks (the model has not seen enough examples of hate speech to recognize it), lower accuracy on questions about historical atrocities, or unusual avoidance of topics where the training data lacked coverage.

Careful practitioners track vocabulary coverage before and after filtering. If you are removing more than a few percent of your corpus, you should audit the removed content for systematic patterns. Tools for this audit include confusion matrices by demographic mention (computing false positive rates separately for documents that mention different groups), analysis of removed vocabulary (computing the drop in term frequency for specific terms before and after filtering), and manual review samples from filtered content. A ten-minute manual review of a random sample of filtered documents often reveals patterns that quantitative analysis misses.

One quantitative approach is to compute the Kullback-Leibler divergence between the vocabulary distributions before and after filtering:

DKL(Pbefore∥Pafter)=∑w∈VPbefore(w)log⁡Pbefore(w)Pafter(w)D_{\text{KL}}(P_{\text{before}} \| P_{\text{after}}) = \sum_{w \in V} P_{\text{before}}(w) \log \frac{P_{\text{before}}(w)}{P_{\text{after}}(w)}

where:

  • Pbefore(w)P_{\text{before}}(w): the normalized frequency of word ww in the unfiltered corpus
  • Pafter(w)P_{\text{after}}(w): the normalized frequency of word ww after filtering
  • VV: the vocabulary

A low overall KL divergence means filtering has not substantially changed the vocabulary distribution. But you should also compute this separately for demographic term subsets: if the KL divergence is high for terms associated with specific communities, those communities are underrepresented in the filtered corpus relative to the unfiltered one.

Toxicity Filter Evaluation

Evaluating a toxicity filter requires more than measuring classifier performance on a test set. You need to measure the downstream effect on the training data distribution and, ideally, on model behavior. A filter that achieves 95% AUC on a held-out test set may still cause serious problems if the test set is not representative of the corpus or if the filter's false positive biases are concentrated in important content categories.

Classifier Evaluation Metrics

At the classifier level, standard evaluation metrics apply:

  • AUC-ROC: Measures classifier discrimination across all thresholds, appropriate for comparing classifiers without committing to a threshold choice. A random classifier has AUC 0.5; a perfect classifier has AUC 1.0.
  • Precision at fixed recall: Choose a target recall level (e.g., 90% of toxic documents filtered) and measure what fraction of filtered documents are truly toxic. This directly measures how much collateral filtering you accept to achieve your recall target.
  • Recall at fixed precision: Choose a target precision level (e.g., 80% of filtered documents are truly toxic) and measure what fraction of toxic documents are caught at that precision level.
  • Average Precision (AP): The area under the precision-recall curve, which summarizes classifier quality across all thresholds. AP is more informative than AUC-ROC for highly imbalanced datasets where true positives are rare.
  • Calibration (Expected Calibration Error): The degree to which classifier scores reflect true probabilities of toxicity. A score of 0.7 should correspond to approximately 70% probability of being toxic. Calibration matters for threshold selection: if a classifier is poorly calibrated, your threshold choices will be systematically off.

For calibration evaluation, reliability diagrams plot classifier output probability on the x-axis against the fraction of truly toxic examples in each probability bin on the y-axis. A well-calibrated classifier produces points near the diagonal. The Expected Calibration Error (ECE) is the weighted mean absolute deviation from the diagonal:

ECE=∑b=1B∣Bb∣n∣acc(Bb)−conf(Bb)∣\text{ECE} = \sum_{b=1}^{B} \frac{|B_b|}{n} \left| \text{acc}(B_b) - \text{conf}(B_b) \right|

where:

  • BbB_b: the set of documents whose predicted probability falls in bin bb
  • ∣Bb∣|B_b|: the number of documents in bin bb
  • nn: total number of documents
  • acc(Bb)\text{acc}(B_b): the true fraction of toxic documents in bin bb (the "accuracy" in the calibration sense)
  • conf(Bb)\text{conf}(B_b): the mean predicted probability for documents in bin bb

Low ECE means predicted probabilities are trustworthy as probability estimates. High ECE means the classifier's scores are poorly calibrated and should not be treated as probabilities when choosing thresholds.

Bias Evaluation

Beyond standard accuracy metrics, toxicity classifiers should be evaluated for disparate impact across demographic groups. The key metrics are:

Subgroup AUC: Measure classifier AUC separately on subsets of examples that explicitly mention a specific group. Significant drops in subgroup AUC indicate the classifier performs differently for different groups. For example, if overall AUC is 0.92 but AUC on examples mentioning LGBTQ+ people is 0.78, the classifier is substantially less reliable for that category.

False positive rate parity: The fraction of non-toxic documents that are incorrectly flagged as toxic should be approximately equal across groups. If the false positive rate for documents mentioning Muslim people is 15% while the false positive rate for documents mentioning Christian people is 3%, the filter will disproportionately remove content about Muslim people from your training corpus, regardless of whether that content is benign.

Counterfactual consistency: Pairs of sentences that differ only in a demographic identifier (e.g., "I am a [group] person" versus "I am a [other group] person") should receive similar toxicity scores. Significant differences indicate group-specific sensitivity rather than content-based toxicity detection. The Jigsaw Unintended Bias in Toxicity Classification dataset provides labeled pairs of this type, and it remains the standard benchmark for this evaluation.

Pinned AUC (BPSN and BNSP): The Jigsaw competition introduced two specialized metrics. BPSN (Background Positive, Subgroup Negative) measures AUC on examples that are toxic overall but do not mention a specific group, paired with examples that mention the group but are not toxic. BNSP (Background Negative, Subgroup Positive) measures the reverse. Together they diagnose whether the classifier is biased toward associating a group with toxicity (high BPSN, low BNSP) or toward treating group-mentioning toxic content as benign (low BPSN, high BNSP).

Corpus-Level Evaluation

After applying a toxicity filter to a training corpus, measure:

  • Filtered fraction by domain: What percentage of web pages, books, Wikipedia articles, and other sources are removed? Large variation by domain may indicate miscalibration or domain mismatch between the classifier's training data and the corpus.
  • Filtered fraction by language: Multilingual corpora often have language-specific false positive rates. A classifier trained primarily on English may be poorly calibrated for other languages, assigning higher toxicity scores to non-English text simply because it is unusual relative to the training distribution.
  • Manual audit of filtered content: A random sample of filtered documents should be reviewed by human annotators to estimate the empirical precision of the filter. If your estimated precision is below 70-80%, you are over-filtering substantially and should raise the threshold.
  • Vocabulary drift: Compare token frequencies before and after filtering. Unusual drops in frequency for terms associated with specific communities indicate disparate impact. This is most efficiently computed by computing the ratio of pre-filter to post-filter frequency for every token and flagging the tokens with the largest drops.

Model-Level Evaluation

The ultimate test of toxicity filtering is its effect on model behavior after training. Corpus-level metrics tell you what was removed; model-level metrics tell you whether the removal had the intended effect. Required evaluations include:

  • Toxicity benchmarks: Held-out test sets of known toxic and non-toxic examples that the model never saw during training. Measure generation toxicity by prompting the model with topic openings and evaluating the completions.
  • Targeted probe sets: Prompts designed to elicit toxic completions. A model trained on insufficiently filtered data will complete "People of [group] are..." with negative stereotypes more often than a model trained on filtered data. The RealToxicityPrompts dataset provides a large collection of such probes.
  • Fairness probes: Measure whether model toxicity levels differ by demographic group mentioned in the prompt. A model that generates more toxic content about some groups than others has learned biased associations that the training data filter did not fully address.
  • Counter-speech performance: Test whether the model can accurately refute hate speech and discuss hate-related topics in an educational context. Over-filtering degrades this capability, and a model that cannot engage with hate speech topics educationally has limited utility for content moderation, historical education, and safety research applications.

Implementation Walkthrough

Let's implement a toxicity filtering pipeline. We will work with a simulated dataset that captures the statistical properties of a realistic pretraining corpus: a mix of toxic content, borderline content, and benign content that includes the sensitive categories (counter-speech, group mentions) where over-filtering is most common. This allows us to measure classifier performance, calibrate a threshold, and audit the filter's behavior across demographic categories.

Setup and Installation

In[3]:
Code
# Install required packages
# uv pip install detoxify scikit-learn numpy pandas matplotlib

import warnings

warnings.filterwarnings("ignore")

Synthetic Toxicity Dataset

We will create a realistic synthetic dataset that simulates the key challenge: a mix of toxic content, borderline content, and benign content that a naive classifier might falsely flag. The category distributions reflect a realistic pretraining corpus, where toxic content is a minority of the total and the challenging cases make up a substantial fraction of the benign documents.

In[4]:
Code
import numpy as np

np.random.seed(42)

# Simulate toxicity classifier scores and ground truth labels
# This represents a corpus of 2000 documents with ground truth labels
# The classifier scores are calibrated to match realistic Detoxify output distributions

n_docs = 2000

# Simulate different document categories
categories = {
    "clearly_toxic": {"n": 200, "mean_score": 0.85, "std": 0.10, "label": 1},
    "mildly_toxic": {"n": 300, "mean_score": 0.60, "std": 0.15, "label": 1},
    "borderline": {"n": 200, "mean_score": 0.40, "std": 0.15, "label": 0},
    "group_mention": {
        "n": 300,
        "mean_score": 0.30,
        "std": 0.20,
        "label": 0,
    },  # Benign but mentions minority groups
    "counter_speech": {
        "n": 200,
        "mean_score": 0.35,
        "std": 0.18,
        "label": 0,
    },  # Benign counter-speech
    "clearly_benign": {"n": 800, "mean_score": 0.08, "std": 0.07, "label": 0},
}

scores = []
labels = []
category_tags = []

for cat_name, cat in categories.items():
    n = cat["n"]
    cat_scores = np.clip(
        np.random.normal(cat["mean_score"], cat["std"], n), 0.001, 0.999
    )
    scores.extend(cat_scores)
    labels.extend([cat["label"]] * n)
    category_tags.extend([cat_name] * n)

scores = np.array(scores)
labels = np.array(labels)
category_tags = np.array(category_tags)

# Calculate ground truth metrics
n_toxic = labels.sum()
n_total = len(labels)
prevalence = n_toxic / n_total
Out[5]:
Console
Dataset: 2000 documents
Toxic documents: 500 (25.0% of corpus)
Non-toxic documents: 1500 (75.0% of corpus)

Documents by category:
  clearly_toxic         200 docs  [TOXIC]
  mildly_toxic          300 docs  [TOXIC]
  borderline            200 docs  [benign]
  group_mention         300 docs  [benign]
  counter_speech        200 docs  [benign]
  clearly_benign        800 docs  [benign]

This distribution reflects a realistic pretraining corpus: toxic content is a minority of the total (about 25%), and the challenging cases (borderline, group mentions, counter-speech) make up a substantial fraction of the non-toxic documents. The group-mention category represents documents discussing minority communities in neutral or positive terms; the counter-speech category represents documents that engage directly with harmful claims in order to refute them.

Classifier Performance Analysis

With our dataset in place, we compute the standard classifier evaluation metrics. These give a threshold-independent view of classifier quality before we commit to a specific operating point.

In[6]:
Code
from sklearn.metrics import (
    average_precision_score,
    precision_recall_curve,
    roc_auc_score,
    roc_curve,
)

# Calculate ROC curve and AUC
fpr, tpr, roc_thresholds = roc_curve(labels, scores)
auc_roc = roc_auc_score(labels, scores)

# Calculate Precision-Recall curve and Average Precision
precision, recall, pr_thresholds = precision_recall_curve(labels, scores)
avg_precision = average_precision_score(labels, scores)

# Find threshold that maximizes F-beta for beta=2 (recall-weighted)
beta = 2
f_beta_scores = []
for p, r in zip(precision, recall):
    if p + r > 0:
        f_beta = (1 + beta**2) * p * r / (beta**2 * p + r)
    else:
        f_beta = 0.0
    f_beta_scores.append(f_beta)

f_beta_scores = np.array(f_beta_scores)
best_idx = np.argmax(f_beta_scores)
best_threshold = (
    pr_thresholds[best_idx]
    if best_idx < len(pr_thresholds)
    else pr_thresholds[-1]
)
best_precision = precision[best_idx]
best_recall = recall[best_idx]
best_f_beta = f_beta_scores[best_idx]
Out[7]:
Console
Classifier Performance Summary
  AUC-ROC:           0.958
  Average Precision: 0.891

Optimal threshold (F-beta=2):
  Threshold:   0.383
  Precision:   0.610
  Recall:      0.966
  F-2 score: 0.865

The AUC-ROC and average precision give us a threshold-independent view of classifier quality. The optimal threshold for our recall-weighted objective sits below 0.5, showing the asymmetric cost: we accept lower precision to achieve higher recall on toxic content. The threshold is selected by maximizing F2F_2, which weights recall twice as heavily as precision.

Precision-Recall Tradeoff Visualization

Out[8]:
Visualization
Precision-recall curve with the F2-optimal threshold marked as a red dot and baseline shown.
Precision-recall curve for the toxicity classifier across all threshold values. The operating point maximizing the F2 score (which weights recall twice as heavily as precision) is marked as a red dot. Achieving high recall requires accepting lower precision. The horizontal dashed line shows the baseline precision for a random classifier, equal to the prevalence of toxic documents in the corpus. The area under the curve (average precision) summarizes classifier quality independent of threshold choice.

False Positive Rate by Category

A critical diagnostic is how false positive rates differ across document categories. This directly measures the disparate impact of a given threshold: which categories of benign content are disproportionately removed? A threshold that is appropriate overall may still cause serious over-filtering for specific content types.

In[9]:
Code
# Apply threshold and compute false positive rates by category
applied_threshold = best_threshold
filtered = scores >= applied_threshold

# Compute per-category false positive rates (among non-toxic docs)
fp_rates = {}
for cat_name in categories:
    if categories[cat_name]["label"] == 0:  # Only non-toxic categories
        mask = category_tags == cat_name
        cat_n = mask.sum()
        cat_fp = (filtered[mask]).sum()
        fp_rates[cat_name] = {
            "total": cat_n,
            "filtered": int(cat_fp),
            "fp_rate": cat_fp / cat_n if cat_n > 0 else 0.0,
        }

# Overall stats
n_filtered = filtered.sum()
tp = (filtered & (labels == 1)).sum()
fp = (filtered & (labels == 0)).sum()
fn = (~filtered & (labels == 1)).sum()
tn = (~filtered & (labels == 0)).sum()
overall_precision = tp / (tp + fp) if (tp + fp) > 0 else 0
overall_recall = tp / (tp + fn) if (tp + fn) > 0 else 0
Out[10]:
Console
Filtering results at threshold tau=0.383
  Documents filtered: 792 / 2000 (39.6%)
  True positives:     483  (toxic docs filtered)
  False positives:    309  (benign docs filtered)
  False negatives:    17  (toxic docs missed)
  Precision:          0.610
  Recall:             0.966

False positive rates by category:
  borderline           51.0% (102/200 docs filtered)
  group_mention        35.3% (106/300 docs filtered)
  counter_speech       50.5% (101/200 docs filtered)
  clearly_benign       0.0% (0/800 docs filtered)

The false positive rates reveal the disparate impact directly. Categories like group_mention and counter_speech have substantially higher false positive rates than clearly_benign, even though none of these categories contain toxic content. This is the over-filtering risk in concrete numbers: at a threshold calibrated for high recall, you lose substantially more content from documents that discuss or critique hate speech than from neutral benign content. If those categories represent important community voices or educational content, the filter is causing harm.

Calibration Analysis

A well-calibrated classifier is essential for threshold selection. If the classifier outputs 0.7 for documents that are toxic only 40% of the time, your threshold selection will be systematically miscalibrated: you will believe you are operating at one precision-recall point when you are at another. The Expected Calibration Error (ECE) measures how far the classifier's probability estimates deviate from true frequencies.

In[11]:
Code
from sklearn.calibration import calibration_curve

# Compute calibration curve
n_bins = 10
fraction_of_positives, mean_predicted_value = calibration_curve(
    labels, scores, n_bins=n_bins, strategy="uniform"
)

# Compute Expected Calibration Error (ECE)
# ECE = weighted mean absolute difference between predicted probability and fraction positive
bin_counts = []
bin_width = 1.0 / n_bins
for i in range(n_bins):
    lower = i * bin_width
    upper = (i + 1) * bin_width
    in_bin = (scores >= lower) & (scores < upper)
    bin_counts.append(in_bin.sum())

bin_counts = np.array(bin_counts)
valid = bin_counts > 0
ece = (
    np.sum(
        bin_counts[valid]
        / n_total
        * np.abs(fraction_of_positives - mean_predicted_value)
    )
    if valid.any()
    else 0.0
)
Out[12]:
Console
Calibration Analysis
  Expected Calibration Error (ECE): 0.1155
  (Lower is better; 0.0 = perfect calibration)

Calibration by probability bin:
     Predicted   Observed
         0.045      0.000
         0.144      0.003
         0.250      0.034
         0.351      0.105
         0.448      0.286
         0.545      0.435
         0.648      0.654
         0.749      0.862
         0.850      0.973
         0.952      1.000

Calibration Reliability Diagram

The reliability diagram makes calibration quality visually obvious. Well-calibrated classifiers produce points near the diagonal; systematic deviations indicate that the classifier is overconfident (points below the diagonal) or underconfident (points above).

Out[13]:
Visualization
Reliability diagram comparing predicted toxicity probability to observed toxicity fraction across 10 bins.
Calibration reliability diagram for the toxicity classifier. The dashed diagonal represents perfect calibration, where the predicted probability exactly matches the observed fraction of truly toxic documents in each bin. Points above the diagonal indicate underconfidence (the true toxicity rate is higher than predicted), while points below indicate overconfidence. The Expected Calibration Error (ECE) summarizes overall calibration quality as a weighted average deviation from perfect calibration.

Threshold Sensitivity Analysis

Before committing to a threshold for a production pipeline, you should understand how key outcomes change as the threshold varies. This analysis reveals the relationship between threshold, filtering aggressiveness, and disparate impact, allowing you to make an informed choice rather than simply accepting the FβF_\beta-optimal threshold.

In[14]:
Code
import pandas as pd

thresholds = np.linspace(0.1, 0.9, 50)
results = []

for tau in thresholds:
    filtered_mask = scores >= tau
    n_filt = filtered_mask.sum()
    tp_t = (filtered_mask & (labels == 1)).sum()
    fp_t = (filtered_mask & (labels == 0)).sum()
    fn_t = (~filtered_mask & (labels == 1)).sum()
    prec = tp_t / (tp_t + fp_t) if (tp_t + fp_t) > 0 else 1.0
    rec = tp_t / (tp_t + fn_t) if (tp_t + fn_t) > 0 else 0.0
    pct_filtered = n_filt / n_total

    # False positive rates for sensitive categories
    counter_speech_mask = category_tags == "counter_speech"
    group_mention_mask = category_tags == "group_mention"
    cs_fp = (
        filtered_mask & counter_speech_mask
    ).sum() / counter_speech_mask.sum()
    gm_fp = (
        filtered_mask & group_mention_mask
    ).sum() / group_mention_mask.sum()

    results.append(
        {
            "threshold": tau,
            "pct_filtered": pct_filtered,
            "precision": prec,
            "recall": rec,
            "counter_speech_fp": cs_fp,
            "group_mention_fp": gm_fp,
        }
    )

results_df = pd.DataFrame(results)
Out[15]:
Visualization
Line plot of precision and recall vs threshold with F2 optimal threshold marked as dashed vertical line.
Precision and recall as the toxicity threshold varies from 0.1 to 0.9. Higher thresholds improve precision but reduce recall, while lower thresholds catch more toxic content at the cost of more false positives. The vertical dashed line marks the F2-optimal operating point, which sits below 0.5 because the recall-weighted objective accepts lower precision to maximize toxic content removal.
Line plot of false positive rates for counter-speech and group-mention categories vs threshold with F2 optimal marked.
False positive rates for two sensitive non-toxic categories as the threshold varies. Counter-speech and group-mention documents are substantially over-filtered at low thresholds relative to clearly benign content, showing how aggressive recall-oriented filtering causes disparate impact on content that discusses or mentions harmful topics without endorsing them.

The left panel shows the precision-recall tradeoff that any filtering practitioner must manage. The right panel is equally important: it shows that the false positive rates for counter-speech and group-mention documents are substantially higher than for clearly benign content across the entire threshold range. This gap does not close at higher thresholds; it narrows, but both sensitive categories remain over-filtered relative to neutral benign content at any threshold that achieves meaningful recall. The pattern is structural and cannot be solved by threshold calibration alone; it reflects the underlying classifier's learned associations.

Applying the Filter and Auditing Results

In[16]:
Code
# Simulate applying the filter to the corpus
final_threshold = best_threshold

corpus_filtered = scores >= final_threshold
corpus_kept = ~corpus_filtered

# Audit: compute fraction removed by category
audit_results = []
for cat_name in categories:
    mask = category_tags == cat_name
    n_cat = mask.sum()
    n_removed = (corpus_filtered & mask).sum()
    pct_removed = n_removed / n_cat if n_cat > 0 else 0.0
    is_toxic = categories[cat_name]["label"] == 1
    audit_results.append(
        {
            "category": cat_name,
            "total": n_cat,
            "removed": int(n_removed),
            "pct_removed": pct_removed,
            "expected_removal": "Yes" if is_toxic else "No (FP)",
        }
    )

audit_df = pd.DataFrame(audit_results)
total_removed = corpus_filtered.sum()
Out[17]:
Console
Filter Audit at threshold=0.383
Total removed: 792 / 2000 (39.6%)

Category                Total  Removed  % Removed  Expected?
------------------------------------------------------------
clearly_toxic             200      200    100.0%        Yes
mildly_toxic              300      283     94.3%        Yes
borderline                200      102     51.0%    No (FP)
group_mention             300      106     35.3%    No (FP)
counter_speech            200      101     50.5%    No (FP)
clearly_benign            800        0      0.0%    No (FP)

This audit table is what every practitioner should examine before deploying a filter at scale. The expected toxic categories should have high removal rates; the non-toxic categories should have low ones. When categories like counter_speech and group_mention show higher removal rates, that is the concrete signal to investigate whether the filter is properly calibrated for your corpus, or whether the threshold needs to be raised to reduce disparate impact at the cost of some recall.

Out[18]:
Visualization
Horizontal bar chart showing removal rate per document category, color-coded by whether removal is expected or a false positive.
Removal rate by document category at the F2-optimal threshold. Toxic categories (clearly toxic, mildly toxic) show high removal rates as expected. Among non-toxic categories, counter-speech and group-mention documents exhibit substantially elevated false positive rates compared to clearly benign content. This pattern demonstrates the disparate impact problem directly: categories that discuss or critique harmful content are over-filtered relative to neutral content at any threshold that achieves high recall on genuinely toxic material.

Key Parameters

The key parameters for a toxicity filtering pipeline are:

  • Classifier architecture: BERT-based classifiers like Detoxify or LLaMA Guard capture semantic toxicity better than shallow models, at higher computational cost. For very large corpora, a two-stage pipeline (keyword pre-filter followed by neural classifier) reduces cost while preserving recall.
  • Filtering threshold (τ\tau): Values around 0.5-0.7 are typical starting points; calibrate to your precision-recall requirements using a labeled validation set. The FβF_\beta-optimal threshold gives a systematic starting point, but manual review of filtered content should always follow.
  • Beta for FβF_\beta threshold selection: β=2\beta = 2 weights recall twice as heavily as precision, appropriate when missing toxic content is the primary concern. Adjust β\beta based on your application's sensitivity to over-filtering versus under-filtering.
  • Attribute selection: Decide which toxicity attributes (severe toxicity, identity attacks, threats, obscenity) to filter independently versus jointly. Combined OR filtering compounds false positive rates; a meta-classifier on attribute scores is more precise.
  • Audit frequency: Re-audit filter performance whenever the corpus changes substantially, when the filter is applied to a new language or domain, or when community norms shift in ways that may affect classifier calibration.

Limitations and Practical Considerations

Toxicity filtering is a hard problem without clean solutions. Several limitations deserve explicit acknowledgment before you deploy a filter in production.

The definition of toxic content is culturally and contextually contingent. What counts as acceptable discourse varies by country, community, time period, and platform. A toxicity classifier trained on US English web content encodes US English norms. When you apply it to content from other cultures, the filter's behavior may be systematically miscalibrated in ways that are difficult to detect without local expertise. In some cultures, forms of expression that appear highly confrontational to US annotators are conventional and not intended as hostile. In others, coded language for discrimination may not appear in the classifier's training data at all.

The adversarial nature of toxic content means classifiers quickly become outdated. New slurs emerge, existing slurs are reclaimed by the communities they target, coded language evolves to evade detection, and community norms shift over time. A classifier trained in 2020 may miss substantial toxic content that emerged in 2023, and it may flag reclaimed terms that were unambiguously toxic in 2020 but are now used with positive valence within those communities. This means toxicity filtering requires ongoing maintenance, not just initial deployment. Budget for periodic retraining and recalibration as part of your data pipeline infrastructure.

Classifier-based toxicity filtering operates at the document level, which is a coarse granularity. A long document that is overwhelmingly high-quality but contains a single toxic sentence will either be fully retained (if its overall score is low) or fully discarded (if its overall score is high). Sentence-level or span-level toxicity detection is more precise but dramatically more expensive and can miss patterns that only become apparent at document scale (such as a document that builds toward incitement through individually benign statements). Span-level filtering also requires span-level annotation, which is more expensive and harder to obtain at the volumes needed for pretraining data.

The interaction between toxicity filtering and model capabilities is not fully understood. Research suggests that filtering toxic content from pretraining data reduces model toxicity in generation tasks, but the magnitude of the effect depends on the filtering method, the threshold, and the downstream fine-tuning procedure. A model fine-tuned with RLHF or constitutional AI methods may effectively suppress toxic outputs regardless of pretraining data quality, while a model deployed without fine-tuning may exhibit strong sensitivity to pretraining data toxicity. This means the importance of pretraining-level toxicity filtering varies substantially by your deployment pipeline, and you should not assume that pretraining filtering alone is sufficient for a deployed model.

Finally, no toxicity filter catches everything. Implicit toxicity, coordinated harassment that is benign in isolation, and sophisticated adversarial content require human review that automated filtering cannot replace. Toxicity filtering is one layer of a defense-in-depth approach. It should be combined with output-side moderation, human review processes for high-stakes decisions, and ongoing monitoring of model outputs in production. We will cover output-side moderation in detail in later chapters on safety and alignment.

In the next chapter on PII Removal, we will see a complementary filtering challenge: identifying and removing personally identifiable information from training data. Like toxicity filtering, PII removal requires balancing recall (catching sensitive information) against precision (preserving useful context), and it faces similar challenges of classifier bias, over-filtering, and culturally contingent definitions of what constitutes sensitive information.

Summary

Toxicity filtering is a necessary step in pretraining data curation, aimed at removing content that would degrade model behavior if memorized and reproduced. The central tension is that the tools needed to catch subtle toxicity (recall-oriented classifiers with aggressive thresholds) also remove legitimate content from sensitive topics, creating vocabulary gaps and representation imbalances that can harm the very communities the filtering intends to protect.

The key concepts from this chapter:

  • Toxicity definition: Content that promotes harm against protected groups, contains direct threats, or includes graphic abuse material. Distinct from adult content, historical discussion, or fiction depicting harm. The boundary requires explicit annotation guidelines and adjudicated edge cases, not just intuition.

  • Classifier approaches: Keyword filters provide fast first-stage filtering; neural classifiers (Detoxify, Perspective, LLaMA Guard) capture semantic toxicity but are computationally expensive and reflect annotator biases. The two-stage pipeline, keyword filtering followed by neural classification, is the standard approach for web-scale corpora.

  • Threshold selection: The choice of threshold τ\tau determines the precision-recall tradeoff. The FβF_\beta score framework allows you to express the relative cost of false negatives versus false positives. For pretraining data, recall-weighted objectives (β>1\beta > 1) are common, but the right β\beta depends on your application and user population.

  • Over-filtering risks: Toxicity classifiers systematically assign higher scores to content mentioning or produced by minority groups, counter-speech, and academic discussion of harm. At aggressive recall-oriented thresholds, these categories are over-filtered relative to neutral content. The effect is measurable as higher false positive rates for sensitive categories and is structural rather than a threshold calibration issue.

  • Evaluation framework: Assess classifiers on AUC-ROC and calibration (ECE), measure bias through per-group false positive rates and counterfactual consistency, and audit corpus-level filtering fractions by domain, language, and content category. Model-level evaluation (generation toxicity on held-out prompts, counter-speech performance) is the ultimate test of whether filtering achieved its goal.

  • Limitations: Toxicity definitions are culturally contingent, classifiers require ongoing maintenance as language evolves, document-level filtering is coarse, and no filter achieves perfect recall without unacceptable false positives. Toxicity filtering is one layer of defense-in-depth, not a complete solution.

Toxicity filtering is a values-laden engineering decision as much as a technical one. The choice of threshold, classifier, and evaluation criteria reflects judgments about which errors are more costly, and different applications will make these tradeoffs differently. Understanding those tradeoffs, and measuring them explicitly, is what separates a thoughtful filtering pipeline from one that causes harm while trying to prevent it.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about toxicity filtering in pretraining data curation.

Toxicity Filtering Quiz

Question 1 of 80 of 8 completed
Which of the following best describes why toxicity filtering for pretraining data differs from output-side content moderation?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026toxicityfiltering, author = {Michael Brenndoerfer}, title = {Toxicity Filtering: Classifiers, Thresholds}, year = {2026}, url = {https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Toxicity Filtering: Classifiers, Thresholds. Retrieved from https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks
MLAAcademic
Michael Brenndoerfer. "Toxicity Filtering: Classifiers, Thresholds." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks>.
CHICAGOAcademic
Michael Brenndoerfer. "Toxicity Filtering: Classifiers, Thresholds." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Toxicity Filtering: Classifiers, Thresholds'. Available at: https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Toxicity Filtering: Classifiers, Thresholds. https://mbrenndoerfer.com/writing/toxicity-filtering-classifiers-thresholds-over-filtering-risks

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.