Document Extraction with Trafilatura and HTML Parsing

Michael BrenndoerferJanuary 9, 202655 min read

Part of Language AI Handbook

Extract clean text from HTML and PDFs for LLM training data. Topics include boilerplate removal, text density algorithms, Trafilatura.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Document Extraction

Training a language model at scale starts long before any gradient update. Before text can enter a training pipeline, it must first be extracted from the raw formats in which it exists: HTML pages downloaded by web crawlers, PDFs, Word documents, and other structured formats. The process of extracting clean, readable text from these formats is called document extraction, and it is one of the most consequential steps in the entire data curation pipeline. Get this step wrong and every downstream stage, from tokenization to model training, inherits the damage. Get it right and you create a foundation of clean, coherent prose that a model can learn from.

Building on the Web Crawling chapter, which explained how to collect raw HTML at scale, this chapter focuses on the next step: converting that raw HTML (and other document formats) into clean text suitable for model training. The challenge is harder than it sounds. A typical web page bundles navigation menus, sidebars, cookie consent banners, footer links, advertisements, and tracking scripts alongside the actual article or blog post you care about. Naive extraction pulls everything indiscriminately. The result is text poisoned with "Click here to subscribe" and "Copyright 2023 XYZ Corp" repeated thousands of times, patterns a language model will dutifully memorize and regurgitate at the worst possible moments.

This chapter covers the technical mechanics of HTML parsing, the problem of boilerplate removal, the main algorithmic approaches to content extraction, and production-ready tools like Trafilatura that implement these ideas. We close by looking at PDF extraction, a separate and harder problem, and by examining how to build a robust multi-format extraction pipeline with proper quality filtering.

Why Document Extraction Is Hard

The web was designed for human browsers, not text extractors. HTML is a display language: its semantic structure is often weak or misleading, and publishers have no incentive to make their markup machine-readable. The result is a gap between what a page visually presents to a reader and what a naive text extractor pulls from the raw HTML bytes.

The Boilerplate Problem

Every web page contains two kinds of text. The first is the main content: the article, the forum post, the product description, the research paper that a reader came to read. The second is boilerplate: navigation links, advertisements, footers, sidebars, related-article carousels, subscription prompts, and legal disclaimers. For training data quality, boilerplate is toxic.

The mechanism of harm is subtle. When you train a language model on text that contains thousands of repetitions of "Privacy Policy | Terms of Service | Cookie Settings", the model assigns high probability to generating exactly that kind of text in contexts where it does not belong. A model trained on a corpus with poorly removed boilerplate learns to interpolate navigation-like fragments into generated prose. It may produce text that reads naturally for a few sentences and then inexplicably inserts a list of menu items. These artifacts are difficult to detect at training time and frustrating to encounter at inference time.

Quantifying the problem: studies of Common Crawl, the most widely used web corpus for LLM pretraining, have found that after full boilerplate removal, the fraction of retained text can drop dramatically from the raw HTML byte count. A page that weighs 80 KB as HTML may yield only 3 KB of real content. The ratio varies enormously by site type. News articles have relatively high content ratios: perhaps 20-40% of the HTML bytes represent prose worth keeping. Link aggregators, forums with short posts, and shopping pages can be almost entirely boilerplate, with meaningful text occupying less than 5% of the raw bytes. At the scale of Common Crawl, which archives tens of billions of pages, even a small improvement in boilerplate removal translates to billions of higher-quality training tokens.

HTML Structure Is Unreliable

Semantically, HTML has tags like <article>, <main>, and <p> that should indicate content. The HTML5 specification describes <article> as "a self-contained composition in a document, page, application, or site, which is intended to be independently distributable or reusable." In theory, finding <article> tags should locate the content. In practice, many sites use <div> for everything and assign meaning only through CSS classes or JavaScript behavior. A navigation bar is often <div class="nav">. Whether that class name signals navigation depends entirely on the site's conventions, which vary infinitely across the web.

The situation is compounded by the fact that HTML semantics are rarely enforced. There is nothing stopping a developer from placing an advertisement inside an <article> tag or using <p> tags for navigation items. The semantic tags provide hints, not guarantees. An extractor that relies entirely on semantic tags will perform well on carefully constructed sites like Wikipedia or the New York Times and poorly on the long tail of hobbyist pages, small businesses, and older web content.

Additionally, HTML in the wild is frequently malformed. Studies of large web crawls consistently find that a substantial fraction of pages contain unclosed tags, incorrectly nested elements, duplicate IDs, and other violations of HTML syntax. Browsers implement extensive error-recovery heuristics to display broken HTML gracefully: they can infer where a missing closing tag should go and reconstruct a valid tree from broken markup. A parser that expects well-formed markup will fail on a significant fraction of real-world pages, silently returning empty or incorrect results.

Encoding and Character Set Issues

Web pages arrive in a range of character encodings: UTF-8, Latin-1 (ISO-8859-1), Windows-1252, Shift-JIS, GB2312, Big5, and many others. UTF-8 has become dominant for new content, but a large fraction of older pages and non-English content uses other encodings. The declared encoding in the HTTP Content-Type header or the HTML <meta charset> tag is sometimes wrong, inconsistent with the actual bytes, or entirely missing.

Mishandling encoding produces mojibake: text where multi-byte sequences are decoded with the wrong charset, resulting in garbled characters. The French word "café" stored as UTF-8 (bytes: 63 61 66 c3 a9) read as Latin-1 becomes "café" with a visible artifact. At corpus scale, mojibake-corrupted text quietly poisons training data. A model trained on it will learn to associate those garbled character sequences with surrounding words, producing subtle generation artifacts for affected languages.

The correct approach is a layered encoding detection strategy: check the HTTP header first, then the HTML meta charset tag, and finally fall back to statistical encoding detection using a library like chardet or charset-normalizer. After decoding, normalize to Unicode NFC form to ensure consistent representation of characters that have multiple valid Unicode representations.

Dynamic Content and JavaScript Rendering

A growing fraction of the modern web delivers content via JavaScript. Single-page applications (SPAs) built with React, Vue, or Angular send a nearly empty HTML skeleton to the browser. All meaningful content is injected into the DOM by JavaScript running after page load. An HTML parser that receives the raw HTTP response from such a page sees almost nothing: a <div id="root"> or <div id="app"> with no child content.

This is increasingly common for news sites, social media platforms, and documentation portals. Standard static HTML parsing cannot handle it. The only solution is to render the page with a headless browser (Playwright, Puppeteer, or Selenium), wait for JavaScript to execute, and extract text from the rendered DOM. Headless rendering is five to fifty times slower than static parsing and consumes significantly more memory and CPU. At the scale of a web crawl, rendering every page is impractical. Production pipelines typically skip JavaScript-heavy pages entirely, or maintain a separate rendering tier for high-value domains where JavaScript content is worth the cost.

HTML Parsing Fundamentals

Before any content extraction algorithm can run, the raw HTML bytes must be parsed into a structured representation that the algorithm can traverse and analyze. Two dominant approaches exist: tree-based parsing and streaming parsing.

DOM Tree Parsing

Tree-based parsers convert the entire HTML document into a Document Object Model (DOM): a hierarchical tree of nodes where each element, attribute, and text fragment is a node. The root of the tree is the <html> element. Its children are <head> and <body>. The body contains all the visible page elements as nested descendants.

Once the tree is built, you can traverse it with XPath queries, CSS selectors, or recursive traversals. You can walk from any node to its parent, children, or siblings. You can select all <p> elements within an <article> element, or all elements whose class attribute contains a particular substring. This rich navigation model is what makes DOM-based extraction powerful.

The primary Python library for DOM-based parsing is Beautiful Soup, which wraps underlying parsers (lxml or the built-in html.parser) and provides a high-level API for tree navigation:

In[4]:
Code
from bs4 import BeautifulSoup

html = """
<html>
<head><title>Example</title></head>
<body>
  <nav><a href="/">Home</a><a href="/about">About</a></nav>
  <main>
    <h1>Article Title</h1>
    <p>This is the first paragraph of the article.</p>
    <p>This is the second paragraph with more content.</p>
  </main>
  <footer>Copyright 2024. All rights reserved.</footer>
</body>
</html>
"""

soup = BeautifulSoup(html, "lxml")
Out[5]:
Console
Title: Example
Main content:
Article Title
This is the first paragraph of the article.
This is the second paragraph with more content.

The get_text() method extracts all text within a node, joining inline elements and stripping tags. The separator argument controls how block-level elements are joined, which matters when you need to preserve paragraph breaks for downstream processing.

Beautiful Soup's tree-building step involves two phases: lexing (breaking the HTML byte stream into tokens: open tags, close tags, attributes, text nodes, and comments) and parsing (assembling the token stream into a tree according to HTML grammar rules). The lxml backend handles this more efficiently than Python's built-in html.parser and applies browser-compatible error recovery for malformed HTML.

lxml and XPath

For high-performance parsing at scale, lxml provides direct bindings to the libxml2 C library. It parses faster than Python-native parsers and supports XPath, a powerful query language for selecting nodes in a tree:

In[6]:
Code
from lxml import html as lhtml

tree = lhtml.fromstring(html.encode("utf-8"))

# XPath to select all paragraph text within main
paragraphs = tree.xpath("//main//p/text()")
Out[7]:
Console
Paragraphs found:
  - This is the first paragraph of the article.
  - This is the second paragraph with more content.

XPath expressions describe paths through the tree: //main//p/text() means "find all <p> elements anywhere inside a <main> element, and return their text content." XPath supports predicates (filtering by attribute values), positional indexing, and aggregate functions, making it possible to write highly targeted queries.

XPath is powerful when you know the structure in advance: for example, parsing a specific news site with a stable template. It becomes unwieldy when dealing with arbitrary pages at crawl scale, where the structure is unpredictable. An XPath rule that works perfectly on one site may extract nothing or the wrong content on a structurally different site.

Handling Malformed HTML

The lxml HTML parser implements the same error recovery heuristics as browsers, making it robust to the broken markup common in real web data. Beautiful Soup's lxml backend inherits this robustness:

In[8]:
Code
broken_html = "<p>Unclosed paragraph <b>bold text <p>Second paragraph"
soup2 = BeautifulSoup(broken_html, "lxml")
Out[9]:
Console
<html>
 <body>
  <p>
   Unclosed paragraph
   <b>
    bold text
   </b>
  </p>
  <p>
   Second paragraph
  </p>
 </body>
</html>

The parser infers closing tags and produces a valid tree from the broken input. In this case, lxml recognizes that <p> elements cannot be nested (in the HTML content model, a <p> element is implicitly closed when another block-level element begins), so it creates two sibling <p> elements, each properly closed. This error recovery is essential for processing real-world web data: a pipeline that crashes or returns empty results on malformed HTML will lose a significant fraction of crawled pages.

The Tree as a Data Structure

Understanding the DOM tree as a data structure clarifies why DOM-based parsing is so useful for extraction algorithms. Each node in the tree has a type (element, text, comment, CDATA), a tag name (for element nodes), a set of attributes, and a list of children. The tree preserves the full containment hierarchy of the page.

This hierarchy is the key signal for content extraction. If a block of text is contained within a <nav> element, it is almost certainly boilerplate, regardless of its textual content. If it is contained within an <article> element that is itself a child of <main>, it is almost certainly content. The extraction problem reduces, in part, to learning which subtrees of the DOM contain content and which contain boilerplate, using the tree structure itself as a feature.

Content Extraction Approaches

Parsing gives you a tree. The harder problem is deciding which parts of the tree contain main content and which are boilerplate. Three main algorithmic families address this problem: rule-based extraction, density-based extraction, and machine learning classification.

Rule-Based Extraction

The simplest approach uses hand-crafted rules based on HTML semantics and structural patterns. The intuition is that content-bearing elements tend to have certain properties: they appear inside <article> or <main> tags, they consist of <p> elements with substantial text, they have high text-to-HTML ratios, and they avoid link-heavy regions.

The first step in rule-based extraction is removing known boilerplate containers before attempting to extract content from whatever remains. Navigation bars, headers, footers, sidebars, advertisements, and scripts are reliably non-content across almost all pages:

In[10]:
Code
from bs4 import BeautifulSoup


def remove_boilerplate_tags(soup: BeautifulSoup) -> BeautifulSoup:
    """Remove elements that reliably contain boilerplate."""
    boilerplate_selectors = [
        "nav",
        "header",
        "footer",
        "aside",
        "[class*='menu']",
        "[class*='sidebar']",
        "[class*='ad']",
        "[class*='banner']",
        "[class*='cookie']",
        "[class*='popup']",
        "script",
        "style",
        "noscript",
    ]
    for selector in boilerplate_selectors:
        for tag in soup.select(selector):
            tag.decompose()
    return soup


sample_html = """
<html><body>
<nav><a href="/">Home</a></nav>
<article>
  <h1>Understanding Neural Networks</h1>
  <p>Neural networks are computational models inspired by biological neurons.</p>
  <p>They consist of layers of interconnected nodes that process information.</p>
</article>
<aside class="sidebar-widget">Related articles</aside>
<footer>Privacy Policy | Terms of Service</footer>
</body></html>
"""

soup = BeautifulSoup(sample_html, "lxml")
cleaned = remove_boilerplate_tags(soup)
Out[11]:
Console
Understanding Neural Networks
Neural networks are computational models inspired by biological neurons.
They consist of layers of interconnected nodes that process information.

After removing the known-boilerplate containers, the extractor looks for the best remaining content candidate. A common heuristic is to find the element with the highest density of <p> tags, on the theory that paragraphs are the natural unit of prose content. Another heuristic is to find the deepest subtree that contains the most text characters, preferring deeply nested content over shallow wrappers.

Rule-based approaches work well on sites with consistent, semantic markup and on news sites where the article is consistently inside a <article> or a <div class="article-body"> element. They fail on sites that use <div> for everything and when boilerplate classes have unpredictable or obfuscated names (some sites deliberately obfuscate class names to hinder scrapers).

The deeper problem with rule-based approaches is brittleness. Rules that work well on the 1,000 sites you tested may fail silently on the 10,000 sites you have not tested. There is no principled way to know in advance which rules will generalize. This motivates the density-based approaches that follow.

Text Density Analysis

A more robust approach exploits the observation that content regions are text-dense, while boilerplate regions are link-dense. This observation holds across a wide variety of sites and does not require knowing anything about the specific site's HTML structure or class naming conventions.

The text density of an HTML node is the ratio of characters in actual text to the total number of characters in the HTML markup for that node. Consider the arithmetic: a paragraph with 200 characters of prose text and 40 characters of wrapping tags has text density:

density(b)=200200+40≈0.83\text{density}(b) = \frac{200}{200 + 40} \approx 0.83

A navigation bar with 10 characters of link text and 300 characters of anchor tags has text density:

density(b)=1010+300≈0.03\text{density}(b) = \frac{10}{10 + 300} \approx 0.03

The contrast is dramatic and consistent across site types. Content paragraphs almost always have densities above 0.5. Navigation, advertisement, and footer regions almost always have densities below 0.2. The threshold between them provides a reliable content/boilerplate signal.

The classic algorithm implementing this idea is Content Extraction via Text-to-Link Ratio (CETR), originally proposed by YinLei Hu et al. A simplified version computes, for each block bb in the page:

density(b)=char_count(b)tag_count(b)+1\text{density}(b) = \frac{\text{char\_count}(b)}{\text{tag\_count}(b) + 1}

where:

  • char_count(b)\text{char\_count}(b): the number of text characters in block bb
  • tag_count(b)\text{tag\_count}(b): the number of HTML tags in block bb
  • The +1+1 in the denominator prevents division by zero for tag-only nodes

Blocks with density above a threshold are classified as content; those below are classified as boilerplate.

A more detailed version also factors in the link density within each block. Navigation and footer regions are not only tag-dense, they are specifically anchor-tag-dense: a high fraction of their characters are inside <a href="..."> elements. The link density of a block is:

link_density(b)=link_chars(b)char_count(b)+1\text{link\_density}(b) = \frac{\text{link\_chars}(b)}{\text{char\_count}(b) + 1}

where:

  • link_chars(b)\text{link\_chars}(b): characters inside anchor tags in block bb
  • char_count(b)\text{char\_count}(b): total text characters in block bb

Content blocks have low link density (prose rarely consists primarily of hyperlinks). Navigation and footer blocks have high link density (they are essentially lists of links). Combining both signals gives a more robust classifier than either signal alone.

Let's implement a basic text density extractor to see these signals in practice:

In[12]:
Code
import re

from bs4 import BeautifulSoup


def compute_block_features(soup: BeautifulSoup) -> list[dict]:
    """Compute text density and link density for each block-level element."""
    block_tags = {
        "p",
        "div",
        "article",
        "section",
        "li",
        "td",
        "h1",
        "h2",
        "h3",
    }
    results = []

    for tag in soup.find_all(block_tags):
        full_html = str(tag)
        text = tag.get_text(strip=True)
        char_count = len(text)
        tag_count = len(re.findall(r"<[^>]+>", full_html))

        # Count characters inside anchor tags
        link_chars = sum(len(a.get_text(strip=True)) for a in tag.find_all("a"))

        density = char_count / (tag_count + 1)
        link_density = link_chars / (char_count + 1)

        results.append(
            {
                "tag": tag.name,
                "text_preview": text[:60] + "..." if len(text) > 60 else text,
                "char_count": char_count,
                "tag_count": tag_count,
                "density": round(density, 2),
                "link_density": round(link_density, 2),
            }
        )

    return results


complex_html = """
<html><body>
<div class="nav">
  <a href="/">Home</a> | <a href="/about">About</a> | <a href="/contact">Contact</a>
</div>
<article>
  <h1>The Science of Deep Learning</h1>
  <p>Deep learning has revolutionized computer vision, natural language processing,
  and speech recognition over the past decade. The key insight is that hierarchical
  representations learned from data can outperform hand-crafted features.</p>
  <p>Convolutional neural networks (CNNs) learn spatial hierarchies of features,
  making them particularly effective for image recognition tasks.</p>
</article>
<div class="footer-links">
  <a href="/privacy">Privacy</a> | <a href="/terms">Terms</a> |
  <a href="/cookies">Cookies</a> | <a href="/sitemap">Sitemap</a>
</div>
</body></html>
"""

soup3 = BeautifulSoup(complex_html, "lxml")
features = compute_block_features(soup3)
Out[13]:
Console
Tag         Density  Link Density  Text Preview
----------------------------------------------------------------------
div            2.00          0.84  Home|About|Contact
article       45.00          0.00  The Science of Deep LearningDeep learning has revolutionized...
h1             9.33          0.00  The Science of Deep Learning
p             78.67          0.00  Deep learning has revolutionized computer vision, natural la...
p             47.00          0.00  Convolutional neural networks (CNNs) learn spatial hierarchi...
div            2.64          0.87  Privacy|Terms|Cookies|Sitemap

The output demonstrates the core principle: content paragraphs show high density and low link density, while navigation and footer divs show the opposite pattern. This signal is stable enough to generalize across sites without requiring any site-specific configuration.

The density-based approach handles the <div>-for-everything problem that defeats pure rule-based methods. Even if a navigation menu is wrapped in a generic <div> with no class names, its text density will be low because the anchor tags dominate the markup. The algorithm does not need to know that the div contains navigation; the structure of the markup reveals it implicitly.

Machine Learning-Based Extraction

The text density heuristics are hand-tuned and may not generalize perfectly across all site types. A more principled approach trains a classifier on annotated examples where human reviewers have labeled each paragraph as content or boilerplate.

The Web-Boilerplate Dataset (CLEANEVAL) provided human annotations of content and boilerplate at sentence or paragraph level, enabling supervised learning for this task. Features used in such classifiers include:

  • Text density and link density (as computed above)
  • Position in the document (content tends to appear near the middle; navigation at the top, footers at the bottom)
  • Surrounding tag context (what parent, sibling, and child elements surround this block)
  • N-gram features of the text itself (boilerplate phrases like "all rights reserved", "click here to subscribe", and "terms of service" are distinctive)
  • DOM depth: navigation tends to be shallow in the tree (close to the root), while content is typically nested several levels deep
  • Paragraph character count: very short paragraphs are disproportionately boilerplate
  • Ratio of capital letters: navigation items and legal text often use unusual capitalization patterns

A well-trained classifier can learn non-obvious combinations of these features. For instance, a block that has medium text density but zero link density and appears directly inside an <article> element is almost certainly content, even if a density-only threshold would mark it as ambiguous. The classifier combines all available signals into a single decision.

Modern tools like Trafilatura, discussed in detail below, use a combination of rule-based signals, tree structure analysis, and machine learning heuristics, achieving strong precision-recall tradeoffs across diverse web content without requiring site-specific configuration.

The CETR Algorithm in Detail

CETR (Content Extraction via Tag Ratio) deserves a deeper look because it introduced the density-based paradigm that most modern extractors build on. The algorithm works as follows.

First, the HTML is converted to a line-by-line representation where each line corresponds to a rendered text fragment. Lines that contain only HTML tags (no visible text) are assigned a density of zero. Lines that contain substantial prose are assigned high density. This representation treats the HTML as a signal stream rather than a tree.

The density values across lines form a histogram. Content regions appear as peaks in this histogram: consecutive lines all with high density. Boilerplate regions appear as valleys. The algorithm applies a smoothing step (a moving average over adjacent lines) to reduce the impact of single noisy lines, then identifies content regions as contiguous spans where the smoothed density exceeds a threshold.

The key insight is that content appears in contiguous blocks. A paragraph consists of many consecutive text-rich lines, not isolated fragments scattered across the page. This contiguity constraint allows CETR to handle cases where individual lines have ambiguous density by looking at whether neighboring lines are also text-rich.

The weakness of CETR is that it makes a hard binary decision at a fixed threshold. Pages with unusual structures, such as academic papers with dense citations or e-commerce pages with paragraph-length product descriptions in navigation-like containers, can confuse the algorithm. More recent algorithms replace the fixed threshold with adaptive thresholds estimated per-page based on the overall density distribution.

Trafilatura: Industrial-Strength Extraction

Trafilatura is an open-source Python library specifically designed for web content extraction at scale. It was developed with LLM pretraining data quality in mind and has become a standard tool in the field. It handles HTML parsing, boilerplate removal, encoding detection, and output formatting in a single, well-tested pipeline.

How Trafilatura Works

Trafilatura's extraction pipeline has several stages that combine multiple signals to make robust content/boilerplate decisions.

The first stage attempts to locate the content candidate using XPath rules targeting semantic HTML elements like <article>, <main>, <section>, and high-density <p>-containing <div> elements. Trafilatura also uses class and ID heuristics to identify content regions even when semantic tags are absent: class names containing "article", "content", "post-body", or "entry-content" are strong positive signals. Class names containing "nav", "menu", "sidebar", "footer", "ad", "banner", or "cookie" are strong negative signals.

The second stage computes tree-level features for candidate blocks, including text density, link density, and paragraph length distribution. Blocks are scored and filtered based on these features. Short blocks with high link density are discarded. Blocks consisting primarily of boilerplate indicators (short copyright notices, "read more" links, metadata fragments) are also removed.

The third stage reassembles the surviving blocks into a coherent document, preserving heading structure and paragraph breaks. Trafilatura is careful to maintain the logical reading order of the original document rather than simply concatenating extracted blocks. It can output plain text, Markdown, or structured XML depending on the use case.

Trafilatura also handles encoding detection automatically, using a cascade of strategies: HTTP headers (if metadata is provided), HTML meta tags, and statistical detection as a fallback. It normalizes all output to UTF-8.

Basic Usage

In[14]:
Code
import subprocess
import sys

# Install trafilatura if not present
subprocess.run(
    [sys.executable, "-m", "pip", "install", "trafilatura", "-q"],
    capture_output=True,
)
In[15]:
Code
import trafilatura

# Simulate fetching and extracting a page
# In practice: downloaded_html = trafilatura.fetch_url(url)
sample_page = """
<!DOCTYPE html>
<html>
<head><meta charset="utf-8"><title>Introduction to Transformers</title></head>
<body>
<header>
  <nav><ul><li><a href="/">Home</a></li><li><a href="/blog">Blog</a></li></ul></nav>
</header>
<main>
  <article>
    <h1>Introduction to Transformers</h1>
    <p>Published on January 15, 2024 by Jane Smith</p>
    <p>The transformer architecture, introduced in the landmark paper
    "Attention Is All You Need" (Vaswani et al., 2017), fundamentally
    changed natural language processing. Unlike recurrent neural networks
    that process sequences step by step, transformers process all tokens
    in parallel using a mechanism called self-attention.</p>
    <p>Self-attention allows each token to directly attend to every other
    token in the sequence, regardless of distance. This enables transformers
    to capture long-range dependencies that RNNs struggle with, and to
    train much more efficiently on modern hardware due to parallelism.</p>
    <h2>The Attention Mechanism</h2>
    <p>At the core of the transformer is scaled dot-product attention.
    For each query vector, attention computes dot products against all
    key vectors, scales by the square root of the dimension, applies
    softmax to obtain weights, and takes a weighted sum of value vectors.</p>
  </article>
</main>
<aside>
  <h3>Related Posts</h3>
  <ul>
    <li><a href="/bert">Understanding BERT</a></li>
    <li><a href="/gpt">GPT Architecture Explained</a></li>
  </ul>
</aside>
<footer>
  <p>© 2024 AI Blog. All rights reserved.</p>
  <nav><a href="/privacy">Privacy</a> | <a href="/terms">Terms</a></nav>
</footer>
</body>
</html>
"""

extracted = trafilatura.extract(sample_page)
extracted_with_meta = trafilatura.extract(
    sample_page,
    include_comments=False,
    include_tables=True,
    favor_recall=False,
    output_format="markdown",
)
Out[16]:
Console
=== Plain text extraction ===
Introduction to Transformers
Published on January 15, 2024 by Jane Smith
The transformer architecture, introduced in the landmark paper "Attention Is All You Need" (Vaswani et al., 2017), fundamentally changed natural language processing. Unlike recurrent neural networks that process sequences step by step, transformers process all tokens in parallel using a mechanism called self-attention.
Self-attention allows each token to directly attend to every other token in the sequence, regardless of distance. This enables transformers to capture long-range dependencies that RNNs struggle with, and to train much more efficiently on modern hardware due to parallelism.
The Attention Mechanism
At the core of the transformer is scaled dot-product attention. For each query vector, attention computes dot products against all key vectors, scales by the square root of the dimension, applies softmax to obtain weights, and takes a weighted sum of value vectors.

=== Markdown extraction ===
# Introduction to Transformers

Published on January 15, 2024 by Jane Smith

The transformer architecture, introduced in the landmark paper "Attention Is All You Need" (Vaswani et al., 2017), fundamentally changed natural language processing. Unlike recurrent neural networks that process sequences step by step, transformers process all tokens in parallel using a mechanism called self-attention.

Self-attention allows each token to directly attend to every other token in the sequence, regardless of distance. This enables transformers to capture long-range dependencies that RNNs struggle with, and to train much more efficiently on modern hardware due to parallelism.

## The Attention Mechanism

At the core of the transformer is scaled dot-product attention. For each query vector, attention computes dot products against all key vectors, scales by the square root of the dimension, applies softmax to obtain weights, and takes a weighted sum of value vectors.

Trafilatura correctly isolates the article body, discarding the navigation, sidebar, and footer. The Markdown output preserves the heading hierarchy, which is valuable when the extracted text will be used in a retrieval system or a training corpus where structural signals matter.

Metadata Extraction

Beyond the main text, Trafilatura can extract structured metadata from pages: titles, author names, publication dates, and canonical URLs. This metadata is valuable for filtering and provenance tracking.

Publication dates allow you to filter content by recency, which matters when training data quality depends on the time period. Author information can help identify and deduplicate content from the same source appearing under different URLs. Canonical URLs help resolve the common situation where the same article is accessible at multiple URLs (with and without trailing slashes, with different UTM parameters, syndicated to partner sites).

In[17]:
Code
metadata = extract_metadata(sample_page)
Out[18]:
Console
Title:  Introduction to Transformers
Author: None
Date:   2024-01-15
URL:    None

Trafilatura extracts metadata from multiple sources in the page: the <title> tag, <meta> tags with properties like og:title and og:author, JSON-LD structured data blocks (a machine-readable format increasingly used by modern websites), and microdata annotations. This multi-source approach increases robustness: if the <title> tag is missing or generic, the og:title property often contains the real article title.

Extraction Quality Settings

Trafilatura exposes a favor_recall versus precision tradeoff that you can control explicitly:

  • When favor_recall=True, Trafilatura includes more candidate text, accepting some boilerplate in order to avoid discarding real content. This is preferable when building training corpora where missing content is more costly than having some noise, because downstream quality filtering can remove boilerplate that slips through, but cannot recover content that was never extracted.
  • When favor_recall=False (the default), the extractor is more conservative, keeping only high-confidence content blocks. This yields cleaner output but may discard short but valid paragraphs, particularly on pages with sparse content.
In[19]:
Code
short_content_page = """
<html><body>
<nav><a href="/">Home</a><a href="/news">News</a></nav>
<article>
  <h1>Quick Update</h1>
  <p>We released version 2.0 today.</p>
</article>
<footer>Contact us at info@example.com</footer>
</body></html>
"""

precise = trafilatura.extract(short_content_page, favor_recall=False)
recalled = trafilatura.extract(short_content_page, favor_recall=True)
Out[20]:
Console
Precision mode: 'Home\nNews\nQuick Update\nWe released version 2.0 today.'
Recall mode:    'Home\nNews\nQuick Update\nWe released version 2.0 today.'

For LLM pretraining, recall-favoring settings are often preferable because the downstream deduplication and quality filtering steps can remove remaining noise, whereas content that is never extracted cannot be recovered. The optimal choice depends on how sophisticated your downstream filtering is. A pipeline with perplexity-based quality scoring and n-gram deduplication can tolerate more extraction noise than a pipeline that relies on extraction quality alone.

Performance at Scale

Trafilatura is designed to handle the throughput demands of processing a full web crawl. On a single CPU core, it can extract text from thousands of pages per second when processing pre-downloaded HTML. The library includes built-in parallelism support for batch processing:

In[21]:
Code
# Batch processing pattern for large corpora
# trafilatura.process_xmltei() handles batches efficiently
# For maximum throughput, process HTML in parallel using multiprocessing

from typing import Optional


def safe_extract(html: str) -> Optional[str]:
    """Extract text from HTML with error handling."""
    try:
        return trafilatura.extract(html, favor_recall=True)
    except Exception:
        return None


# Example of batch processing structure
html_documents = [sample_page, short_content_page]
results = list(map(safe_extract, html_documents))
Out[22]:
Console
Processed: 2 documents
Successful extractions: 2
Total characters extracted: 1010

At LLM pretraining scale, where you might process hundreds of terabytes of HTML, the throughput characteristics of your extraction library matter enormously. Trafilatura's C-library backend (via lxml) keeps per-page extraction time in the single-millisecond range for typical pages, making it feasible to process billions of pages on a cluster in reasonable time.

Other Extraction Tools

Trafilatura is not the only option. Several other tools occupy different points in the precision/speed/flexibility tradeoff space, and knowing their characteristics helps you choose the right tool for your specific use case.

Newspaper3k and Newspaper4k

The newspaper3k library (and its actively maintained fork newspaper4k) targets news article extraction specifically. It uses a combination of heuristics tuned for news site conventions: the assumption that the page title is in the largest heading, that the article body follows the headline, and that bylines typically appear near the top of the article.

In[23]:
Code
subprocess.run(
    [sys.executable, "-m", "pip", "install", "newspaper4k", "-q"],
    capture_output=True,
)
In[24]:
Code
from newspaper import Article
from newspaper.article import ArticleDownloadState

# Demonstrate with HTML input (avoiding live network requests)
article = Article(url="https://example.com/article")
article.html = sample_page
article.is_downloaded = True
article.download_state = ArticleDownloadState.SUCCESS
article.parse()
Out[25]:
Console
Title:   Introduction to Transformers
Authors: []
Text preview: Introduction to Transformers

Published on January 15, 2024 by Jane Smith

The transformer architecture, introduced in the landmark paper "Attention Is All You Need" (Vaswani et al., 2017), fundamenta...

Newspaper is fast to configure and extracts author and date information reliably on standard news sites. It also performs natural language processing on the article to extract summary sentences and keywords, which can be useful for building search indexes or topic-based filtering.

Its weakness is poor generalization to non-news content. Forum posts, documentation pages, e-commerce product descriptions, and academic papers do not follow the news article conventions that Newspaper's heuristics assume. For a general-purpose web corpus, Newspaper should not be the primary extractor, but it can be used as a high-quality specialized extractor for news-domain content.

Readability-lxml

The readability-lxml library is a Python port of Mozilla's Readability algorithm, which powers Firefox's "Reader Mode" feature. When you click the reader icon in Firefox, this algorithm is what strips the page down to its core content.

In[26]:
Code
subprocess.run(
    [sys.executable, "-m", "pip", "install", "readability-lxml", "-q"],
    capture_output=True,
)
In[27]:
Code
from readability import Document

doc = Document(sample_page)
Out[28]:
Console
Title: Introduction to Transformers
Extracted text preview:

Introduction to Transformers
Published on January 15, 2024 by Jane Smith
The transformer architecture, introduced in the landmark paper
    "Attention Is All You Need" (Vaswani et al., 2017), fundamentally
    changed natural language processing. Unlike recurrent neural networks
    that process se

Readability-lxml returns cleaned HTML rather than plain text. This is a meaningful distinction: the returned HTML preserves heading tags, bold and italic emphasis, table structures, and list formatting. For applications where you need structured content rather than a raw text dump, this is valuable. The algorithm works by scoring candidate blocks based on their class names, tag types, and content length, then selecting the highest-scoring subtree as the main content container.

The scoring heuristic is simple but effective: elements receive positive scores for being <div>, <p>, <td>, or <article> elements, and negative scores for class or ID names matching patterns like "comment", "meta", "footer", "nav", "sidebar", or "ad". The element with the highest cumulative score in its subtree becomes the content root.

Because it returns HTML, Readability-lxml pairs well with a secondary processing step that converts the cleaned HTML to Markdown or plain text while preserving structural signals. The markdownify library performs this conversion.

JusText

JusText implements a paragraph-level boilerplate detection algorithm that classifies each paragraph as "good" (content), "near-good", "short", or "bad" (boilerplate). It uses a combination of stop-word frequency and link density as its primary signals.

The key insight behind JusText is that stop words (high-frequency grammatical words like "the", "is", "and", "of") are distributed differently in content versus boilerplate. Ordinary prose has a natural density of stop words because language has grammatical structure. Navigation menus, copyright notices, and button labels consist primarily of content words and lack the grammatical connective tissue of real prose. JusText counts the stop-word ratio in each paragraph and uses it as a content signal.

In[29]:
Code
subprocess.run(
    [sys.executable, "-m", "pip", "install", "justext", "-q"],
    capture_output=True,
)
In[30]:
Code
import justext

paragraphs = justext.justext(
    sample_page.encode("utf-8"), justext.get_stoplist("English")
)
Out[31]:
Console
Class        Text Preview
------------------------------------------------------------
bad          Home
bad          Blog
short        Introduction to Transformers
short        Published on January 15, 2024 by Jane Smith
neargood     The transformer architecture, introduced in the la
good         Self-attention allows each token to directly atten
short        The Attention Mechanism
good         At the core of the transformer is scaled dot-produ
short        Related Posts
bad          Understanding BERT
bad          GPT Architecture Explained
bad          © 2024 AI Blog. All rights reserved.
bad          Privacy | Terms

JusText's classification gives you more granular control than a binary content/boilerplate decision. The "near-good" class is particularly interesting: it captures paragraphs that are probably content but have some boilerplate-like characteristics (perhaps a short paragraph that happens to have low stop-word density). You can choose to include "near-good" paragraphs for recall-heavy pipelines or restrict to "good" only for high-precision pipelines.

JusText's reliance on stop words also means it handles multilingual content differently from English-centric heuristics. Stop word lists exist for many languages, and the algorithm's behavior is consistent across languages as long as the appropriate stop word list is used. This makes JusText a good choice for multilingual corpora where English-specific heuristics may fail.

Comparing the Tools

Each tool reflects a different design philosophy and set of tradeoffs:

  • Trafilatura is the most general-purpose and is best suited for processing diverse web corpora at scale. Its combination of rule-based, density-based, and machine-learning signals makes it robust across site types.
  • Newspaper4k is the best choice for news content specifically. Its domain-specific heuristics achieve excellent accuracy on news sites while being simpler to configure than Trafilatura.
  • Readability-lxml is best when you need structured output (preserved HTML with headings, lists, and tables) rather than plain text, and when your content is articles or documentation.
  • JusText is best for multilingual corpora and when granular classification (good/near-good/bad) is useful for your downstream filtering logic.

PDF and Other Document Formats

Web crawls capture HTML, but LLM training corpora increasingly incorporate PDFs (academic papers, technical reports, books, government documents) and other formats. PDF extraction presents a distinct set of challenges that are fundamentally different from HTML extraction.

PDF Extraction Challenges

PDFs do not store text as a linear stream that mirrors the reading order. Instead, the PDF format describes text placement using absolute coordinates: "draw glyph 'H' at position (72, 720), draw glyph 'e' at position (78, 720)..." The PDF viewer assembles these glyphs into readable text visually. An extractor must reconstruct the reading order from raw positional data, which requires spatial reasoning that HTML parsing does not.

Multi-column academic papers are a common challenge. A two-column paper has text in the left column that should be read before the text at the same vertical position in the right column. A naive extractor that reads glyphs in order of their y-coordinate (top to bottom) will interleave text from both columns, producing incoherent output. The correct reading order requires first identifying column boundaries, then sorting text within each column independently.

Additional complications include:

  • Scanned PDFs: documents scanned from paper contain no text at all, only images of text. These require Optical Character Recognition (OCR) before any text extraction can occur, adding significant computational cost and reducing accuracy compared to born-digital PDFs.
  • Mathematical notation: equations in PDFs are often stored as sequences of specially positioned characters from mathematical symbol fonts, or as embedded images. Accurate reconstruction of mathematical formulas from raw PDF character streams is a research-grade problem; most extractors either skip equations or produce approximations.
  • Headers and footers: page numbers, running chapter titles, and copyright notices appear on every page. They must be identified as repeating elements and removed, but this requires recognizing that the same text pattern appears across multiple pages.
  • Tables: tabular data in PDFs loses its grid structure during extraction unless the extractor applies spatial analysis to group cells into rows and columns based on their bounding box coordinates.
  • Font encoding: some PDFs use custom font encodings where the Unicode code points of characters in the PDF do not match standard Unicode assignments. The character displayed as 'A' may be stored as code point 0x41 in one font and 0x01 in another. Extractors must map through font encoding tables to recover correct characters.

PyMuPDF (fitz)

PyMuPDF (imported as fitz in Python) is among the fastest and most accurate PDF text extractors. It wraps the MuPDF library, which is also used in PDF viewers like Evince and the Chrome PDF viewer:

In[32]:
Code
subprocess.run(
    [sys.executable, "-m", "pip", "install", "pymupdf", "-q"],
    capture_output=True,
)
In[33]:
Code
import fitz  # PyMuPDF

# Create a minimal PDF in memory for demonstration
# In practice: doc = fitz.open("path/to/document.pdf")
# We demonstrate the API structure


def extract_text_from_pdf_bytes(pdf_bytes: bytes) -> str:
    """Extract text from PDF bytes, removing common boilerplate."""
    doc = fitz.open(stream=pdf_bytes, filetype="pdf")
    pages_text = []
    for page_num in range(len(doc)):
        page = doc[page_num]
        text = page.get_text("text")
        pages_text.append(text)
    doc.close()
    return "\n\n".join(pages_text)


# Show the API usage pattern
print("PyMuPDF API demonstrated. Key methods:")
print("  doc = fitz.open('file.pdf')")
print("  page = doc[0]")
print("  text = page.get_text('text')  # plain text")
print("  blocks = page.get_text('blocks')  # bounding-box blocks")
print("  words = page.get_text('words')   # individual words with coords")
Out[34]:
Console
PyMuPDF extraction modes:
  'text': Raw text, preserving line breaks as in PDF
  'blocks': List of (x0, y0, x1, y1, text, block_no, block_type) tuples
  'words': List of (x0, y0, x1, y1, word, block_no, line_no, word_no)
  'dict': Full structured representation with spans and fonts
  'html': HTML reconstruction of the page layout

The "blocks" mode is particularly useful for multi-column PDFs. Each block includes its bounding box coordinates: the (x0,y0)(x_0, y_0) top-left corner and (x1,y1)(x_1, y_1) bottom-right corner. Given these bounding boxes, you can sort blocks first by column (x-coordinate) and then by vertical position within each column, reconstructing the correct reading order. A simple heuristic: if the midpoint of a block's x-range falls in the left half of the page, it belongs to the left column; if in the right half, to the right column.

The "dict" mode gives the most complete structural information, including font names, font sizes, and character-level spans. This is valuable when you want to identify headings (which typically use larger font sizes or bold weight), separate body text from captions, or detect figure and table labels.

pdfminer.six

An alternative with more fine-grained control over the extraction process is pdfminer.six, which exposes the PDF internal rendering model:

In[35]:
Code
subprocess.run(
    [sys.executable, "-m", "pip", "install", "pdfminer.six", "-q"],
    capture_output=True,
)
In[36]:
Code
from pdfminer.layout import LAParams

# Demonstrate the LAParams (Layout Analysis Parameters)
laparams = LAParams(
    line_overlap=0.5,  # Fraction of overlap to consider lines the same
    char_margin=2.0,  # Max distance between chars to join into word
    word_margin=0.1,  # Min whitespace fraction between words
    line_margin=0.5,  # Max vertical distance to consider same line
    boxes_flow=0.5,  # Balance between column vs reading order (0=col, 1=reading)
    detect_vertical=False,
    all_texts=False,
)
print("LAParams configured for standard single-column documents")
print(f"  char_margin:  {laparams.char_margin}")
print(f"  line_margin:  {laparams.line_margin}")
print(f"  boxes_flow:   {laparams.boxes_flow}")
Out[37]:
Console
For two-column academic papers, adjust boxes_flow=0.3 to prioritize column order.
For single-column content, boxes_flow=0.5 gives good results.

The boxes_flow parameter controls the balance between two competing strategies for reconstructing reading order. A value of 0 treats each text box independently and sorts by horizontal position (column-first order). A value of 1 uses flow analysis that tries to follow the visual reading order, even across column boundaries. Values between 0 and 1 interpolate between these strategies. For academic papers, boxes_flow=0.3 typically works well; for single-column content, boxes_flow=0.5 is appropriate.

Handling Scanned Documents: OCR

For scanned PDFs and image-based documents, OCR is necessary before any text processing can begin. The two dominant open-source OCR engines are Tesseract and EasyOCR. Tesseract (developed at Hewlett-Packard and now maintained by Google) is a C++ library with Python bindings via pytesseract. EasyOCR is a more recent Python-native engine that uses deep learning models and handles a wider variety of scripts and fonts.

In[38]:
Code
# pytesseract wraps the Tesseract OCR engine
# Requires: brew install tesseract (macOS) or apt install tesseract-ocr (Linux)
# subprocess.run([sys.executable, "-m", "pip", "install", "pytesseract", "-q"], capture_output=True)

# OCR pipeline example (requires Tesseract installed):
# from PIL import Image
# import pytesseract
#
# img = Image.open("scanned_page.png")
# text = pytesseract.image_to_string(img, lang="eng")

print("OCR pipeline: Image -> Tesseract -> Text")
print("Key considerations for OCR in LLM data pipelines:")
print("  - OCR confidence scores can filter low-quality extractions")
print("  - Language detection should run after OCR to identify language")
print("  - Layout analysis before OCR improves accuracy on multi-column docs")

OCR introduces errors that are qualitatively different from boilerplate noise. A misrecognized character ("0" read as "O", "l" read as "1") produces text that is locally incorrect but syntactically similar to real text. These errors are difficult to detect and filter because they look plausible to statistical filters. At LLM pretraining scale, OCR errors accumulate across millions of scanned pages and can subtly degrade model accuracy on tasks involving numbers, proper names, and technical terminology.

At LLM pretraining scale, OCR is computationally expensive. Pipelines like those used to build Dolma and The Pile selectively apply OCR to PDFs that do not already contain embedded text (detected by attempting text extraction first, which is fast), avoiding unnecessary compute on the majority of documents.

Emerging Tools: Marker and Nougat

Two newer tools address specific limitations of traditional PDF extraction.

Marker is a pipeline that combines PyMuPDF for text-layer extraction with vision models for understanding complex layouts, tables, and figures. It produces Markdown output with proper heading hierarchy, preserving the logical structure of the document rather than just the raw text.

Nougat (Neural Optical Understanding for Academic Documents), developed by Meta AI, uses a vision-language model trained specifically on academic papers to convert PDFs to a structured Markdown format. It handles mathematical equations in LaTeX syntax, which is a significant improvement over existing methods. The tradeoff is that Nougat is substantially slower than traditional extractors, making it suitable for high-value academic documents where quality matters more than throughput.

Building a Reliable Extraction Pipeline

Individual tools handle individual document types. A production extraction pipeline must orchestrate multiple tools, handle errors gracefully, track quality metrics, and make decisions about which documents are worth the extraction cost.

Document Type Detection

Before invoking any extractor, you need to know what type of document you are processing. The raw bytes provide several reliable signals. PDF files begin with the magic bytes %PDF (hex: 25 50 44 46). HTML files typically begin with a DOCTYPE declaration or an <html> tag. The URL extension and the HTTP Content-Type header are additional signals, though both can be unreliable (a server might serve HTML with a .htm extension or set the wrong content type).

A layered detection strategy is more reliable than any single signal:

In[39]:
Code
from dataclasses import dataclass, field
from enum import Enum
from typing import Optional


class DocumentType(Enum):
    HTML = "html"
    PDF = "pdf"
    UNKNOWN = "unknown"


@dataclass
class ExtractionResult:
    text: Optional[str]
    title: Optional[str]
    language: Optional[str]
    doc_type: DocumentType
    char_count: int
    success: bool
    error: Optional[str] = None
    warnings: list[str] = field(default_factory=list)


def detect_doc_type(content: bytes, url: str = "") -> DocumentType:
    """Detect document type from content signature or URL."""
    if content[:4] == b"%PDF":
        return DocumentType.PDF
    if (
        b"<html" in content[:1000].lower()
        or b"<!doctype" in content[:100].lower()
    ):
        return DocumentType.HTML
    if url.endswith(".pdf"):
        return DocumentType.PDF
    return DocumentType.UNKNOWN


def extract_html(content: bytes) -> ExtractionResult:
    """Extract text from HTML using Trafilatura."""
    try:
        html_str = content.decode("utf-8", errors="replace")
        text = trafilatura.extract(html_str, favor_recall=True)
        metadata = extract_metadata(html_str)
        title = metadata.title if metadata else None
        return ExtractionResult(
            text=text,
            title=title,
            language=None,
            doc_type=DocumentType.HTML,
            char_count=len(text) if text else 0,
            success=text is not None and len(text) > 50,
        )
    except Exception as e:
        return ExtractionResult(
            text=None,
            title=None,
            language=None,
            doc_type=DocumentType.HTML,
            char_count=0,
            success=False,
            error=str(e),
        )


def extract_document(content: bytes, url: str = "") -> ExtractionResult:
    """Unified document extraction dispatcher."""
    doc_type = detect_doc_type(content, url)
    if doc_type == DocumentType.HTML:
        return extract_html(content)
    # PDF extraction would follow similar pattern using fitz
    return ExtractionResult(
        text=None,
        title=None,
        language=None,
        doc_type=doc_type,
        char_count=0,
        success=False,
        error=f"Unsupported type: {doc_type.value}",
    )
Out[40]:
Console
Document type: html
Extraction success: True
Title: Introduction to Transformers
Character count: 957
Text preview: Introduction to Transformers
Published on January 15, 2024 by Jane Smith
The transformer architecture, introduced in the landmark paper "Attention Is ...

Quality Filtering After Extraction

Not every extracted document is worth keeping. After extraction, a filtering pass removes documents that are too short, too repetitive, or structured in ways that suggest extraction failure rather than valid content.

Minimum length thresholds are the simplest filter. A document with fewer than 200 characters almost certainly represents a failed extraction (an empty page, a redirect, a captcha gate) rather than a legitimately short piece of content. Most production pipelines set higher thresholds: 1,000 characters or 200 words is more typical. The right threshold depends on the document types in your corpus.

Repetition filters catch a different class of problems. Certain pages consist primarily of repeated content: product listing pages where the same price and review count template repeats hundreds of times, paginated index pages where the same navigation repeats across dozens of pages, or spam pages that repeat the same keyword phrase to manipulate search engines. A simple check on the ratio of unique lines to total lines catches the most egregious cases.

In[41]:
Code
def quality_filter(result: ExtractionResult) -> tuple[bool, str]:
    """Apply quality filters to an extraction result."""
    if not result.success or not result.text:
        return False, "extraction_failed"

    text = result.text
    char_count = len(text)

    # Minimum content length
    if char_count < 200:
        return False, "too_short"

    # Check for excessive repetition (lorem ipsum, test pages)
    lines = [l.strip() for l in text.splitlines() if l.strip()]
    if lines:
        unique_line_ratio = len(set(lines)) / len(lines)
        if unique_line_ratio < 0.5 and len(lines) > 5:
            return False, "too_repetitive"

    # Check content-to-whitespace ratio
    non_space_chars = len(text.replace(" ", "").replace("\n", ""))
    if non_space_chars / (char_count + 1) < 0.5:
        return False, "excessive_whitespace"

    return True, "passed"


# Test with our extraction result
passed, reason = quality_filter(result)
Out[42]:
Console
Quality filter result: PASS (passed)
  Characters: 957
  Words: 140
  Unique line ratio: 1.00

More sophisticated quality filters include language identification (keeping only documents in the target language or languages), perplexity scoring against a small reference model (documents with unusually high perplexity relative to the reference language model are likely garbled or non-prose), and classifier-based quality scoring trained on human judgments of document quality.

Logging and Monitoring

A production extraction pipeline processes millions of documents per day. Without logging and monitoring, it is impossible to detect when the pipeline starts failing silently: when a site changes its HTML structure and extraction starts returning empty strings, when encoding detection starts failing for a language, or when a particular document type starts causing crashes that are silently caught by exception handlers.

The key metrics to track are: success rate (fraction of documents that return non-empty, non-error results), mean character count of successful extractions (a sudden drop signals that extraction is returning short, low-quality results), and the distribution of failure modes (which filters are catching which documents). Alerting on these metrics allows you to catch extraction quality regressions before they contaminate a large fraction of the corpus.

Comparing Extraction Tools

The right tool depends on the use case. Evaluating tools on the same document reveals how their approaches differ in practice:

In[43]:
Code
test_document = """
<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>Backpropagation Through Time: A Complete Guide</title>
</head>
<body>
<header>
  <nav class="main-nav">
    <a href="/">Machine Learning Blog</a>
    <a href="/tutorials">Tutorials</a>
    <a href="/subscribe">Subscribe</a>
  </nav>
</header>
<div class="breadcrumb">Home > Tutorials > Deep Learning > BPTT</div>
<main class="content">
  <article>
    <h1>Backpropagation Through Time: A Complete Guide</h1>
    <div class="meta">By Alex Johnson | March 15, 2024 | 8 min read</div>
    <p>Backpropagation through time (BPTT) is the algorithm used to train
    recurrent neural networks (RNNs). It extends the standard backpropagation
    algorithm to handle sequences by "unrolling" the network across time steps
    and applying the chain rule across the resulting computation graph.</p>
    <p>Understanding BPTT is essential for grasping why RNNs are difficult to
    train on long sequences. The gradient must flow backward through every time
    step, and repeated multiplication of gradient matrices causes either
    vanishing or exploding gradients.</p>
    <h2>The Vanishing Gradient Problem</h2>
    <p>When gradients flow through many time steps, repeated matrix multiplication
    causes them to shrink exponentially. A gradient that starts at 1.0 multiplied
    by 0.9 at each of 50 steps becomes less than 0.005. The weights at early time
    steps receive nearly zero gradient signal and fail to learn long-range dependencies.</p>
  </article>
</main>
<aside class="sidebar">
  <div class="widget">
    <h4>Popular Articles</h4>
    <ul>
      <li><a href="/lstm">LSTM Networks Explained</a></li>
      <li><a href="/gru">GRU: Gated Recurrent Units</a></li>
      <li><a href="/transformers">Why Transformers Beat RNNs</a></li>
    </ul>
  </div>
  <div class="ad-banner">Advertisement</div>
</aside>
<footer>
  <p>© 2024 ML Blog. Privacy Policy | Cookie Settings | Contact</p>
</footer>
</body>
</html>
"""

# Extract using each tool
trafilatura_result = trafilatura.extract(test_document) or ""
bs4_result = remove_boilerplate_tags(
    BeautifulSoup(test_document, "lxml")
).get_text(separator="\n", strip=True)
Out[44]:
Console
Trafilatura:
  Words extracted: 136
  Boilerplate markers retained: 0
  Preview: Backpropagation Through Time: A Complete Guide Backpropagation through time (BPTT) is the algorithm used to train recurr

BS4 + Rules:
  Words extracted: 153
  Boilerplate markers retained: 0
  Preview: Backpropagation Through Time: A Complete Guide Backpropagation Through Time: A Complete Guide By Alex Johnson | March 15

The comparison reveals that rule-based removal using Beautiful Soup, while straightforward to implement, often leaves residual boilerplate that Trafilatura's more sophisticated pipeline eliminates. The rule-based approach removes known boilerplate containers but cannot identify boilerplate that appears inside ambiguous <div> elements or that uses non-standard class names.

Visualizations

Let's visualize how text density varies across page regions, which is the core signal used by most extraction algorithms:

Out[45]:
Visualization
Bar chart comparing text density and link density across nav, heading, content, sidebar, and footer regions.
Text density and link density for different regions of a typical web page. Content paragraphs have high text density (many characters per tag) and low link density (few anchor characters). Navigation and footer regions show the opposite pattern, giving a clear signal for boilerplate detection algorithms.
Out[46]:
Visualization
Two-panel bar chart comparing three extraction approaches: retained word counts on the left and tracked boilerplate marker counts on the right.
Word count retained and six keyword-based boilerplate markers across Trafilatura, rule-based BS4, and naive extraction applied to the same test document. Both structured extractors remove all six tracked markers in this example, while Trafilatura retains fewer words. Naive extraction retains the most text and all tracked markers. The marker count is a targeted diagnostic rather than a general extraction-quality score.
Out[47]:
Visualization
Line chart showing percentage of documents retained after each stage of the extraction pipeline from raw crawl to final corpus.
Cumulative document retention through successive stages of a document extraction and quality filtering pipeline. Each stage reduces the corpus size while increasing quality. The largest drops occur at content extraction (many pages yield no usable text) and the length filter (short pages are discarded), with smaller but consistent drops at each subsequent quality stage.

The pipeline retention chart illustrates a pattern common to all large-scale data curation efforts: the vast majority of the filtering work happens early, in the extraction and basic quality filtering stages. The remaining stages apply progressively finer-grained selection criteria to the already-filtered subset.

Limitations and Practical Considerations

Document extraction is reliable on well-structured HTML and standard document formats, but it has limits that practitioners must account for at scale.

The Long Tail of Web Formats

The web contains an enormous variety of specialized formats: JavaScript-rendered single-page applications (SPAs) where content only appears after JavaScript executes, Flash-based content (largely defunct but still present in archived crawls), XML feeds, custom CMS markup, and more. Standard HTML extractors cannot process SPAs because they receive essentially blank HTML before JavaScript runs. Handling SPAs requires a headless browser like Playwright or Puppeteer, which is orders of magnitude slower and more resource-intensive than static parsing.

For LLM pretraining corpora, SPAs are often simply excluded: the added infrastructure cost is not justified when billions of static HTML pages are available. But for specialized domains (financial dashboards, real-time data applications, social media platforms), SPA rendering becomes necessary if those domains are important to the target use case.

The long tail also includes content that is technically accessible as HTML but deliberately obfuscated to prevent extraction: encrypted or encoded content bodies, dynamically assembled content via <script> injection, and anti-scraping techniques that detect non-browser user agents and return empty or misleading content. These techniques are increasingly common on sites that are concerned about their content being used for AI training.

Accuracy Is Site-Dependent

No extractor achieves uniform accuracy across all site types. Trafilatura's creators report precision and recall figures in the 85-95% range on news corpora, but accuracy drops significantly on forums, Q&A sites, and non-English pages with unusual character distributions. The benchmark numbers reported in research papers are measured on curated test sets that may not represent the distribution of content in a real web crawl.

Pipeline builders should evaluate their chosen extractor on a representative sample of their target corpus before committing to it, rather than assuming published benchmarks generalize. A practical approach is to manually inspect 100-200 randomly sampled extraction results, noting cases where content was dropped (false negatives) or boilerplate was retained (false positives), then adjusting extractor settings or adding post-processing steps to address the most common failure modes.

Encoding Errors at Scale

Even with careful encoding detection, at corpus scale you will encounter documents where the declared encoding is wrong, the encoding is a rare dialect with no detection support, or the file is simply corrupted. Standard practice is to apply a final UTF-8 encoding pass with error replacement (errors='replace'), which substitutes the Unicode replacement character (U+FFFD, displayed as ) for undecodable bytes. Documents with high replacement-character density can then be filtered out.

The threshold for filtering depends on the target languages in your corpus. Scripts that encode characters as multi-byte sequences in UTF-8 (Chinese, Japanese, Korean, Arabic, and many others) will naturally have higher byte counts per character, but the replacement-character density should still be near zero for correctly encoded documents. A replacement rate above 1-2% suggests encoding problems worth filtering on.

The Recall-Precision Tradeoff in Practice

For model pretraining, recall matters more than it might for other applications. Every valuable paragraph that an extractor misclassifies as boilerplate is permanently lost from the training data. Practitioners frequently find that a recall-biased extractor combined with aggressive downstream quality filtering outperforms a precision-biased extractor on final model quality, because the quality filtering can remove noise that slipped through, but cannot recover content that was never extracted.

The right balance depends on the downstream pipeline. If you have a robust multi-stage quality pipeline (deduplication, perplexity filtering, classifier-based quality scoring), favor recall at the extraction stage. If extraction is your only quality gate, favor precision.

This tradeoff has an asymmetry worth emphasizing: boilerplate that passes through extraction will be seen repeatedly across many documents (a given footer template might appear on every page of a large site), and the repetition signal allows downstream deduplication to identify and remove it. But content that is filtered at extraction time disappears from the corpus entirely and permanently. Erring on the side of inclusion at the extraction stage, with careful filtering downstream, is generally the safer strategy.

Document extraction from web crawls raises questions about copyright and terms of service that the technical literature often underweights. Many websites prohibit automated scraping in their terms of service. The legality of using scraped content for AI training is an area of active litigation in multiple jurisdictions, with outcomes that will shape industry practice for years to come. The robots.txt crawl politeness standard described in the Web Crawling chapter applies equally here: content that was crawled against a site's expressed wishes remains ethically questionable even after successful extraction.

The appropriate response to this legal uncertainty varies by organization and use case. Academic research pipelines often operate under fair use or equivalent doctrines. Commercial AI training at scale faces more scrutiny. The safest approaches involve obtaining explicit permission, using licensed datasets (such as Common Crawl, which provides data under its terms of use), or restricting training data to content with permissive licensing.

Multimodal Content Loss

Current text extraction pipelines discard everything that is not text: images, charts, diagrams, and tables. For a purely text-focused language model, this is acceptable. But much of the knowledge in academic papers, technical documentation, and even news articles is conveyed through visual elements. A scientific paper describing a new algorithm might have a diagram that makes the algorithm immediately clear but whose meaning is lost when the paper is reduced to its text.

Multimodal training, where models learn from text and images together, requires preserving and linking these visual elements during extraction. This is an active area of development: tools like Nougat and Marker are beginning to handle figures and tables in PDF extraction, and some HTML extractors preserve <img> tags with their alt text to provide at least a textual representation of visual content.

Summary

Document extraction bridges raw web data and usable training text. The key ideas to carry forward:

  • Boilerplate removal is as important as content extraction. A naive text dump of HTML is polluted with navigation, advertising, and legal text that degrades model training by introducing repetitive, context-free fragments.
  • Text density and link density are the core signals that most algorithms rely on. Content regions are character-dense and link-sparse; boilerplate regions are the opposite. This signal holds across diverse site types without requiring site-specific configuration.
  • The DOM tree is the central data structure for HTML extraction. Understanding the tree enables principled feature engineering and makes it possible to exploit the containment hierarchy as a content signal.
  • Trafilatura is the state-of-the-practice tool for HTML extraction at LLM scale, giving strong precision-recall balance, metadata extraction, and reliable encoding handling. Its recall-favoring mode is appropriate for pretraining pipelines with downstream quality filtering.
  • PDF extraction is a separate and harder problem, requiring spatial reconstruction of reading order from absolute glyph coordinates. PyMuPDF handles most cases well; multi-column layouts require additional spatial reasoning. Scanned PDFs require OCR, which is computationally expensive and introduces a different class of errors.
  • Favor recall at extraction time when a downstream quality pipeline will further filter the corpus. Missed content cannot be recovered; boilerplate that slips through can be removed by deduplication.
  • Evaluate on your specific data. Published benchmark numbers do not substitute for empirical testing on representative samples of your target crawl. Extraction quality varies significantly across site types and languages.
  • The legal and ethical rules around web extraction for AI training is actively evolving. Be aware of the terms of service of sources you extract from and the licensing of datasets you use.

The next chapters in this section cover deduplication and quality filtering: the downstream stages that refine the text this extraction pipeline produces into a training corpus of sufficient quality for language model pretraining.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about document extraction.

Document Extraction Quiz

Question 1 of 80 of 8 completed
Why is boilerplate removal important when building LLM training data from web crawls?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026documentextraction, author = {Michael Brenndoerfer}, title = {Document Extraction with Trafilatura and HTML Parsing}, year = {2026}, url = {https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Document Extraction with Trafilatura and HTML Parsing. Retrieved from https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura
MLAAcademic
Michael Brenndoerfer. "Document Extraction with Trafilatura and HTML Parsing." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura>.
CHICAGOAcademic
Michael Brenndoerfer. "Document Extraction with Trafilatura and HTML Parsing." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Document Extraction with Trafilatura and HTML Parsing'. Available at: https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Document Extraction with Trafilatura and HTML Parsing. https://mbrenndoerfer.com/writing/document-extraction-html-parsing-boilerplate-removal-trafilatura

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.