Web Crawling: Common Crawl, Robots.txt, and Freshness

Michael BrenndoerferJanuary 8, 202652 min read

Part of Language AI Handbook

Explains how web crawlers discover and download internet content at scale, covering Common Crawl architecture, robots.txt compliance, politeness policies.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Web Crawling

The internet is the largest corpus of human knowledge ever assembled. Hundreds of billions of web pages document everything from scientific research to casual conversation, from formal legislation to informal social media posts. When researchers at Common Crawl or AI labs want to train a language model on this knowledge, they face a fundamental engineering challenge: how do you systematically discover, download, and store that content at a scale that spans petabytes of data?

Web crawling is the answer. A web crawler, also called a spider or bot, is a program that automatically traverses the web by following hyperlinks, downloading page content as it goes. The same technology that powers Google's search index also underpins the training data pipelines for virtually every large language model trained today. GPT-3 used filtered Common Crawl data for roughly 60% of its training corpus. LLaMA, Falcon, and Mistral all draw heavily from crawled web text. Understanding how web crawling works, where its data comes from, and how to respect its social contract is foundational knowledge for anyone building modern language AI systems.

This chapter covers the mechanics of web crawling, from seed URL discovery through frontier management and politeness policies. We examine Common Crawl in depth because it is the single most important public data source in language AI. We then explore how freshness, depth, and breadth tradeoffs shape the character of crawled datasets, and we work through a hands-on implementation of a focused crawler. By the end, you will understand how crawling works mechanically and why every design decision, from seed selection to rate limiting, has downstream consequences for the quality and character of the training data you produce.

How Web Crawlers Work

Web crawlers operate a simple loop: take a URL, download the page at that URL, extract all links from the page, add new links to a queue, and repeat. The complexity comes from doing this at scale, doing it without overwhelming servers, and doing it in a way that produces high-quality, well-organized data.

The conceptual simplicity of this loop is deceptive. A production crawler must handle hundreds of millions of URLs daily, deal with malformed HTML, follow redirects, manage connection pooling, cache DNS results, respect per-host rate limits, parse and obey robots.txt rules, detect and skip duplicate pages, and write output in formats that downstream pipelines can process efficiently. Every one of these concerns interacts with the others. A DNS cache that is too aggressive causes stale IP addresses that lead to connection failures. A rate limiter that is too loose causes site operators to block the crawler's IP. An HTML parser that fails silently on malformed markup drops entire pages from the corpus without warning.

This section walks through the core components of the crawl process: seed selection, the main loop, URL normalization, DNS and HTTP mechanics, and frontier management at scale.

Seed URL Selection

Before any crawling begins, a crawler needs a starting point. Seed URLs are the initial set of pages from which the crawler begins following links. The choice of seeds has an outsized influence on what the final crawl contains, because breadth-first traversal explores the neighborhood of seeds first and most thoroughly.

For general-purpose crawls like Common Crawl, seeds typically come from several sources. Authoritative web directories, such as the DMOZ Open Directory (now archived) or curated lists of major news sites, government portals, and educational institutions, provide high-quality starting points. URLs from previous crawls that were productive are re-seeded to maintain continuity. XML sitemaps published by individual sites enumerate pages that owners want indexed, giving a direct invitation to crawl. Domain registrar data reveals newly registered domains that may contain fresh content.

The seed strategy creates a quality bias from the start. A crawl seeded from Wikipedia will produce data richer in factual reference content than a crawl seeded from social media domains. A crawl seeded from curated academic directories will skew toward formal writing and technical vocabulary. This quality inheritance, where crawled content mirrors the editorial standards of the seeded domains, is a feature that language AI practitioners exploit deliberately when building specialized corpora.

The number of seeds also matters. A few hundred seed URLs in a breadth-first crawler will naturally cluster the crawl around those seeds' neighborhoods. A million seeds, carefully selected to span diverse languages, topics, and regions, produces much more balanced coverage. Common Crawl uses tens of millions of seeds across a wide variety of domains and languages, which is part of why its coverage is unusually broad compared to purpose-built enterprise crawlers.

For a focused crawl building a domain-specific corpus, seeds are chosen differently. If you want a dataset of Python programming content, you might seed from the official Python documentation, PyPI project pages, Stack Overflow Python questions, major Python blogs, and key GitHub repositories. These seeds will guide the link-following traversal toward the neighborhood of Python content. You might also feed the crawler a keyword list and have it score discovered URLs by the presence of those keywords, making sure that pages added to the frontier are relevant to your target domain.

The Crawl Loop

The fundamental crawl algorithm proceeds through a series of well-defined steps. The crawler begins with a set of seed URLs, typically highly authoritative pages like major news sites, Wikipedia, or web directories. For each URL in the queue, the crawler issues an HTTP GET request and stores the response. It then parses the response to extract hyperlinks using the href attributes in anchor tags. Each new URL that has not already been visited and passes acceptability filters gets added to the frontier, the data structure holding URLs yet to visit. This loop continues until either the frontier is exhausted or some stopping condition, like a document count or time limit, is reached.

Web Crawler

A web crawler is an automated program that systematically browses the World Wide Web, following hyperlinks to discover and download web pages. Crawlers index web content for search engines or build datasets for training language models. They are also called spiders or bots.

The simplest frontier is a FIFO queue, which produces a breadth-first traversal. Start from seeds, visit all pages linked from seeds, then all pages linked from those pages, and so on. Breadth-first crawling naturally covers the most highly connected pages first, since popular pages appear as targets in many early crawls.

Depth-first traversal, using a stack instead of a queue, follows link chains deep into individual sites before moving on. This strategy is more useful for focused crawls that want to exhaust a specific domain's content, but it risks getting stuck in link traps: cycles of dynamically generated pages that extend indefinitely. A product catalog with 10,000 items and a sorting parameter, a calendar system generating every possible date URL, or a search results page with query string variations can all trap a depth-first crawler in a technically infinite loop if the crawler does not implement loop detection.

Production crawlers typically use priority queues rather than simple FIFO structures. URLs are scored based on predicted importance, recency, or content type, and higher-scored URLs are downloaded first. This allows the crawler to be more selective about what it visits when resources are limited. A priority crawler might assign higher scores to URLs from domains that have historically produced high-quality content, URLs that contain specific keywords relevant to the crawl's focus, or URLs that appear as link targets on many different pages, a signal of importance analogous to PageRank.

URL Normalization and Deduplication

The web is full of URLs that point to the same content. Consider these URLs:

  • http://example.com/page
  • HTTP://EXAMPLE.COM/PAGE
  • http://example.com/page?utm_source=twitter
  • http://example.com/page#section2
  • http://www.example.com/page

All of these may return identical or nearly identical content. Without normalization, a crawler visits the same page multiple times under different URLs, wasting bandwidth and cluttering the dataset with duplicates. URL normalization applies a sequence of transformations to produce a canonical form: lowercasing the scheme and host, sorting query parameters, removing fragment identifiers, stripping known tracking parameters like utm_source and fbclid, and resolving relative URLs against the page's base URL.

URL normalization is more subtle than it appears. Some query parameters are semantically significant, like ?lang=en or ?page=2, while others are pure tracking noise. Stripping the wrong parameters produces false merges where two different pages get collapsed to the same canonical URL. Keeping too many tracking parameters produces false splits where many crawl attempts hit the same actual page. Production crawlers maintain lists of known tracking parameters to strip and use heuristics for unknown parameters based on their structure and context.

Even after normalization, URL deduplication alone is insufficient. Many websites serve the same content at multiple canonical URLs through redirects, mirrors, or content management system artifacts. A news article might be accessible at /2024/03/article-title, /article/12345, and /amp/article-12345, all returning essentially the same text. For this reason, content-level deduplication, which we will explore in the Deduplication chapter, is equally important downstream.

Link traps deserve special attention. Some websites generate infinite URLs intentionally or accidentally. A calendar application might produce URLs for every possible date. A dynamic filter interface might allow arbitrary combinations of filtering parameters, each yielding a unique URL. E-commerce sites with faceted search can generate billions of unique-looking URLs from a few dozen actual products. Crawlers defend against link traps with several heuristics: URL length limits, maximum path depth, pattern detection for parametric URL explosions, and maximum pages-per-domain caps.

DNS Resolution and HTTP Fetching

For each URL, the crawler must resolve the hostname to an IP address using the Domain Name System, then open a TCP connection and issue an HTTP request. This two-step process sounds straightforward but becomes a significant bottleneck at scale.

DNS resolution is inherently network-dependent. Resolving a single hostname may take anywhere from 1 millisecond for a cached entry to several hundred milliseconds for a cold lookup against distant authoritative name servers. At web scale, with millions of unique hostnames being crawled, DNS resolution latency accumulates to a substantial fraction of total crawl time. Production crawlers cache DNS results aggressively, typically with TTLs of several hours, and batch DNS lookups for new domains using asynchronous DNS libraries that can resolve thousands of hostnames concurrently. Some large-scale crawlers maintain their own authoritative DNS resolvers to reduce latency further.

HTTP connections carry significant overhead from TCP handshake and TLS negotiation. A full TLS 1.3 handshake requires two round trips before any application data can be sent. Crawlers reuse connections via HTTP keep-alive within the same host and batch requests to the same IP to amortize connection costs. HTTP/2 support provides multiplexed connections where multiple requests share a single TCP connection, further reducing per-request overhead. At the scale of millions of pages per day, even small per-request latency improvements translate to meaningful increases in crawl throughput.

Response handling requires checking status codes, and each category demands a different response. A 200 OK means success; the content can be stored. A 301 Moved Permanently or 302 Found means the page has moved; the crawler follows the redirect URL and records the mapping between old and new URL. A 404 Not Found means the page no longer exists; the crawler records this and removes the URL from future crawl schedules. A 429 Too Many Requests means the server is throttling requests; the crawler must back off exponentially and try again later. A 503 Service Unavailable means temporary unavailability; the crawler queues the URL for a retry after a delay. A 5xx error from a server generally indicates a transient problem, and well-designed crawlers implement exponential backoff with jitter for such failures to avoid thundering-herd effects when many crawl workers hit the same server simultaneously.

Content-type negotiation matters too. Crawlers should request text/html and handle other content types appropriately. A URL that returns a PDF needs different processing than one that returns HTML. Some crawlers include PDF and EPUB extraction pipelines for high-quality documents, while others skip non-HTML content types entirely for simplicity.

Robots.txt files tell crawlers to avoid certain paths entirely, and well-behaved crawlers check and respect these instructions before fetching any URL on a domain. We discuss robots.txt in detail in the next section.

The Frontier at Scale

For small crawls, an in-memory queue is sufficient. At web scale, the frontier can contain billions of URLs that must be managed persistently across machine restarts. Production crawlers like those operated by Google or Common Crawl use distributed frontier managers backed by databases or specialized queue systems.

A key design challenge is balancing crawl politeness against crawl speed. You cannot send a thousand requests per second to a single server. Instead, the frontier must ensure that no server receives requests faster than its politeness policy allows. This typically means grouping URLs by host, and for each host, maintaining a separate subqueue with delays between requests. The classic approach partitions the frontier into per-host buckets and draws one URL from each bucket in round-robin fashion, inserting a mandatory wait before re-visiting any bucket.

Crawl Frontier

The crawl frontier is the data structure that holds URLs discovered but not yet fetched by a web crawler. Managing the frontier involves deciding which URLs to visit next (scheduling), how frequently to revisit known pages (freshness policy), and how to avoid overwhelming any single server (politeness).

Distributed crawlers add further complexity. When many worker machines are fetching pages in parallel, the shared frontier must coordinate so that no two workers attempt to fetch the same URL simultaneously. This requires either a centralized frontier service with high availability, or a partitioned frontier where URL assignments are sharded by host and each shard is owned by exactly one worker. Common Crawl's crawler infrastructure, for example, uses a distributed queue system that ensures consistent URL assignment across hundreds of parallel workers while maintaining per-host politeness.

The frontier also maintains metadata for each URL beyond the URL string itself. Crawlers track the URL's discovery timestamp, its link depth from the nearest seed, its estimated relevance score, its crawl priority, and any known metadata from sitemaps. This metadata drives scheduling decisions that determine which pages get crawled and when.

Robots.txt and Ethical Crawling

Web crawlers have enormous power. A single aggressive crawler can bring down a web server by flooding it with requests. Even well-intentioned crawlers can cause problems: generating excessive server load, consuming bandwidth that costs site owners money, and accessing private or sensitive content that was never meant to be indexed.

The robots exclusion protocol, formalized as robots.txt, is the web's social contract between site owners and crawlers. When a crawler first visits a domain, it should fetch https://example.com/robots.txt and parse the access rules before crawling any other URL on that domain. This file tells the crawler which paths are off limits, which crawlers are welcome, and how fast requests should be made.

Parsing robots.txt

The robots.txt format specifies rules for different user agents using User-agent directives, followed by Disallow and Allow rules. The wildcard * matches any user agent not explicitly named.

User-agent: * Disallow: /private/ Disallow: /admin/ Allow: /public/ User-agent: Googlebot Disallow: /staging/ Crawl-delay: 1 Sitemap: https://example.com/sitemap.xml

In this example, all crawlers are blocked from /private/ and /admin/, but explicitly allowed to access /public/. Googlebot has an additional exclusion for /staging/ and must wait at least 1 second between requests. The sitemap declaration points to a structured list of URLs the site owner wants crawled.

Robots.txt parsing has several edge cases that matter for correct implementation. The Allow directive takes precedence over Disallow when both match a URL, but the more specific rule wins when there is no explicit precedence conflict. An Allow: /public/ rule overrides Disallow: / for any URL under /public/. The order of rules within a group matters in some implementations, while others apply longest-match semantics. Wildcards (*) and end-of-string anchors ($) add expressiveness: Disallow: /*.pdf$ blocks all PDF files regardless of directory.

The Crawl-delay directive specifies the minimum number of seconds to wait between successive requests to the server. Respecting this directive is polite and often legally required in jurisdictions where unauthorized automated access may violate terms of service. When both the robots.txt crawl delay and your own configured delay exist, you should honor whichever is more conservative.

The Sitemap directive points to XML sitemap files that enumerate the site's publicly available pages. Smart crawlers fetch these sitemaps to seed their frontier with known good URLs rather than relying entirely on link following. A well-maintained sitemap tells the crawler exactly which URLs the site owner considers canonical and worth indexing, eliminating guesswork about URL normalization and reducing the chance of crawling duplicate or unwanted pages.

Why Respecting robots.txt Matters

Respecting robots.txt is both an ethical obligation and a practical necessity. Many websites disallow crawling of user-generated content areas, login-protected pages, or dynamically generated search results that would produce near-infinite URLs. Crawling these areas wastes resources and produces low-quality data. Pages behind authentication walls return redirect responses or login forms rather than actual content. Search result pages produce query-dependent content that changes with every request and may contain little informational value as training data.

The legal dimension is significant. Some jurisdictions treat ignoring robots.txt as unauthorized computer access. In the United States, the Computer Fraud and Abuse Act has been used to argue against unauthorized scraping, though legal interpretations vary. In Europe, the GDPR creates obligations around processing personal data, which web pages may contain. Beyond legality, violating robots.txt damages trust between the open-web ecosystem and the research community. Sites that feel violated by crawlers often block entire IP ranges, penalizing legitimate crawlers along with bad actors.

Some data pipelines used to train language models have been criticized for ignoring robots.txt at scale. When Common Crawl began operations, some site owners found their content in crawl data despite having disallowed bots. This tension between the value of open training data and the rights of content creators has become a central legal and ethical debate in AI development. If you are building a crawler, respecting robots.txt is the minimum acceptable standard.

Robots Exclusion Protocol

The robots exclusion protocol is a standard for websites to communicate which parts of their content crawlers are permitted to access. A file named robots.txt at the root of a domain contains rules specifying which user agents are allowed or disallowed from fetching which URL paths.

Politeness Policies Beyond robots.txt

Even when robots.txt grants access, good crawlers implement additional politeness measures. The most important is rate limiting. Sending too many requests too quickly stresses servers and degrades the experience for human visitors. A common heuristic is to wait a multiple of the server's response time before issuing the next request, adapting automatically to server load: if a server takes 2 seconds to respond, wait an additional 4 seconds before sending the next request. This adaptive delay naturally backs off when servers are under load without requiring you to hard-code domain-specific delays.

IP-level rate limiting matters too. A single website may be hosted behind a shared IP that serves thousands of domains. Hitting that IP with aggregate traffic from all those domains can cause collateral damage to unrelated sites. Production crawlers track per-host request rates and per-IP request rates to avoid this kind of collateral congestion.

Identifying yourself through the User-Agent HTTP header is another courtesy. Well-behaved crawlers provide a user agent string that includes a URL or email address where the site owner can contact you if your crawler causes problems. This allows site operators to reach out about issues rather than simply blocking your IP. Compare Mozilla/5.0 (compatible; MyCrawler/1.0) with MyCrawler/1.0 (+https://myproject.org/crawler; contact@myproject.org). The second form gives site operators actionable information.

Crawlers should also implement exponential backoff when they encounter server errors. A 503 Service Unavailable response means the server is temporarily overwhelmed. Immediately retrying adds to that overwhelm. Instead, wait an exponentially increasing delay: 30 seconds, then 60, then 120, then give up. This pattern is respectful and also more likely to succeed, since transient overloads typically resolve within minutes.

Common Crawl

Common Crawl is a nonprofit organization that has been crawling the web since 2007 and making its data freely available. It has become the single most important source of training data for large language models. Understanding Common Crawl's structure, scale, and quality characteristics is essential for anyone building language AI systems.

Scale and Coverage

Common Crawl releases new crawls roughly monthly, each containing 3 to 5 billion web pages. The archive now holds over 250 billion pages accumulated over more than 15 years. The raw data is stored in WARC (Web ARChive) format, a standard format for web archives that preserves both HTTP response headers and page content. Each monthly crawl is approximately 70 to 100 terabytes of WARC files, hosted on Amazon S3 where they are accessible at no cost for data transfer within AWS and at very low cost for external downloads.

The coverage of Common Crawl is deliberately broad rather than deep. It does not attempt to crawl every page on every site exhaustively. Instead, it prioritizes popular, linked-to pages across a wide variety of domains and languages. The result is good coverage of the most important content on the internet, with sparser coverage of niche or low-traffic sites.

This breadth-over-depth strategy has clear implications for language AI. Common Crawl data captures mainstream web writing well: news articles, Wikipedia, Stack Overflow, Reddit, GitHub, and major blogs are all well represented. Niche academic forums, regional news in underrepresented languages, and community wikis may appear rarely or not at all. Understanding these coverage gaps is important when evaluating what knowledge a model trained on Common Crawl data will and will not possess.

WARC Format

WARC (Web ARChive) is a file format for storing web crawl data. Each WARC record contains the full HTTP response for a single URL, including headers and body content, along with metadata like the crawl timestamp and target URL. WARC is an ISO standard (ISO 28500) commonly used by web archives and Common Crawl.

Common Crawl Data Products

Common Crawl publishes three main data products, each serving a different use case in the data processing pipeline.

WARC files are the primary artifact from which all other products are derived. Each WARC file is a compressed sequence of WARC records, where each record contains the full HTTP response for one URL: status line, response headers, and body content. WARC files also contain request records showing what the crawler sent, and metadata records with crawl-specific information. A WARC record for an HTML page includes the raw HTML bytes, while a record for an image includes the raw image bytes. WARC files are the authoritative source and can be reprocessed to produce alternative text extractions if the WET extraction quality is inadequate.

WAT files (Web Archive Transformation) contain JSON metadata extracted from WARC records. Each WAT record includes HTTP headers, detected content type, page title, outgoing link structure, and language detection results. WAT files are much smaller than WARC files, typically 5-10x smaller, making them practical for metadata analysis without downloading full page content. They are useful for analyzing link graph structure, tracking domain coverage statistics, and filtering which WARC segments are worth downloading for a specific application.

WET files (WARC Encapsulated Text) contain only the plain text extracted from HTML pages, with all markup stripped. Each WET record includes the URL, timestamp, and plain text content. WET files are the most commonly used starting point for text-based NLP applications because they require no HTML parsing, have already extracted text content, and are substantially smaller than WARC files. A single monthly crawl's WET files total around 5 to 10 terabytes, compared to 70+ terabytes for the WARC files.

For language model pretraining, WET files are often the entry point. However, WET extraction quality varies significantly. Some pages produce clean, readable text, while others produce garbled outputs from malformed HTML or aggressive content extraction. Navigation menus, cookie consent dialogs, advertisement text, and JavaScript-rendered content that the simple HTML parser cannot access all end up either missing from WET files or appearing as noise. Research groups building high-quality training corpora often prefer to start from WARC files and apply their own text extraction pipeline to achieve better quality and consistency.

Accessing Common Crawl Data

Common Crawl data lives on Amazon S3 in the commoncrawl bucket. You can list available crawls and their indices via the Common Crawl index API or by browsing the S3 bucket directly. Each monthly crawl is identified by a name like CC-MAIN-2024-10 that encodes the year and approximate week number.

The Common Crawl index is a key tool for selective access. Rather than downloading entire WARC files, which are enormous, you can query the index to find specific URLs and retrieve only the relevant byte ranges using HTTP range requests. The Columnar Index format stores the index in Parquet files, allowing efficient filtered scans using tools like Apache Spark or DuckDB. You can filter by URL prefix to find all pages from a given domain, by content type to find only HTML pages, or by timestamp to find pages crawled during a specific time window, and then download only the relevant WARC byte ranges.

Let us look at what a WET record looks like in practice:

In[3]:
Code
# The warcio library provides Python bindings for reading WARC/WET files
# uv pip install warcio requests

# A single WET segment file path from a Common Crawl release
# We simulate the structure by working with a minimal example
# In practice: download from s3://commoncrawl/crawl-data/CC-MAIN-2024-10/...
sample_wet_record = {
    "uri": "https://en.wikipedia.org/wiki/Natural_language_processing",
    "content_type": "text/plain",
    "content_length": 45231,
    "timestamp": "2024-03-15T14:23:11Z",
    "text_preview": (
        "Natural language processing (NLP) is a subfield of computer science "
        "and artificial intelligence that uses machine learning to enable "
        "computers to understand and communicate with human language in a "
        "valuable way. NLP enables computers to read, hear, and understand "
        "human language..."
    ),
}
Out[4]:
Console
WET Record Structure:
  URI:          https://en.wikipedia.org/wiki/Natural_language_processing
  Content-Type: text/plain
  Length:       45,231 bytes
  Timestamp:    2024-03-15T14:23:11Z

Text preview:
  Natural language processing (NLP) is a subfield of computer science and artificial intelligence that uses machine learning to enable computers to understand and communicate with human language in a va...

Each WET record stores the URI of the crawled page alongside the extracted plain text. The content length gives us a sense of document size. For a Wikipedia article on NLP, nearly 45 KB of text reflects the depth of coverage typical for encyclopedic content. In contrast, many spam pages or thin-content sites produce only a few hundred bytes of text, a signal used by quality filters.

The practical workflow for using Common Crawl data in a training pipeline typically follows several steps. First, you identify which monthly crawl releases to use, often selecting multiple crawls across different months to ensure temporal diversity. Next, you download WAT files to analyze the metadata and identify which WARC segments contain content relevant to your goals. Then you download the WET files for those segments and apply language identification, quality filtering, and deduplication. Finally, you post-process the filtered text into the format your training pipeline expects.

Crawl Strategies

Not all crawls are created equal. The strategy a crawler uses shapes the character of the resulting dataset in fundamental ways. The tradeoffs between breadth, depth, freshness, and topical focus determine what kinds of knowledge end up in the data.

Breadth-First vs. Focused Crawling

A breadth-first crawler makes no assumptions about what topics are valuable. It follows all links equally, expanding outward from seeds in all directions. This produces broad coverage of the web graph but shallow coverage of any individual site or topic. Common Crawl uses a broad, general-purpose strategy: it wants to cover as much of the web as possible across all languages and topics. The resulting dataset reflects the web's actual distribution of content, which means a lot of English-language commercial and social media content.

A focused crawler, by contrast, prioritizes pages that match a specific topic or domain. It uses classifiers or keyword heuristics to score each URL before adding it to the frontier, then visits high-scoring URLs first. A focused crawl on scientific literature would prioritize domains like arxiv.org, pubmed.ncbi.nlm.nih.gov, and academic publishers. A focused crawl on programming would prioritize Stack Overflow, GitHub, and documentation sites. A focused crawl on legal text would target court opinion databases, regulatory archives, and law review journals.

The advantage of focused crawling is efficiency: you download far fewer irrelevant pages to assemble a domain-specific corpus. If you want a high-quality programming dataset, starting from Common Crawl and filtering is one approach, but directly crawling programming-relevant domains is often more efficient and produces cleaner results. The disadvantage of focused crawling is that topic classifiers can miss relevant content that uses unexpected terminology or lives on unexpected domains. A blog post about machine learning hosted on a personal site that the crawler has never seen might score poorly on a classifier trained on known ML domains.

Hybrid strategies combine the two approaches. Some crawlers use broad crawling to discover new domains and focused policies to determine crawl priority within each domain. Common Crawl itself applies some form of quality and relevance scoring to prioritize crawl bandwidth toward higher-quality pages.

Link distance from seed pages is a strong proxy for content quality. Pages that are one click from a trusted seed, like Wikipedia or a major news site, tend to be more relevant and authoritative than pages buried six hops deep in the web graph. Common Crawl seeds from web directories and past crawl data, which means its coverage is richer for well-connected sites and sparser for isolated niche domains.

Crawl depth, the maximum number of links followed from a seed URL, directly controls how far into the long tail the crawler reaches. A shallow crawl at depth 2 or 3 captures mainstream content efficiently. A deep crawl can surface rare or specialized content that shallow crawls miss, but at the cost of downloading much more spam and low-quality content.

The relationship between crawl depth and content quality reflects how the web's link structure works. High-quality, authoritative content tends to receive many inbound links from other authoritative pages. This is exactly what PageRank, the algorithm underlying Google's original search ranking, measures: the probability that a random web surfer lands on a page by following links. Pages with high PageRank appear early in breadth-first traversals. Training data composed primarily of high-PageRank pages tends to be more coherent and authoritative than data from the long tail.

The practical implication is that very deep crawls, while theoretically yielding more content, often produce diminishing returns for training data quality. Each additional hop into the web graph brings in more spam, low-quality content, and link-farm pages that are only there to manipulate search rankings. Most high-quality training data pipelines cap crawl depth and apply quality filters to compensate, rather than trying to recover quality from a very deep crawl.

Crawl Freshness

The web changes constantly. Pages are updated, deleted, and created at rates that vary enormously by site type. News sites publish dozens of articles per day. Academic papers are rarely updated after publication. Social media posts are ephemeral. A static snapshot of the web taken today is already partially stale tomorrow.

Freshness policies govern how often a crawler revisits pages it has previously seen. There are several common approaches. The simplest is a fixed-interval policy: revisit every page every NN days regardless of content. A more sophisticated approach is change-rate estimation: track how often a page's content changes and revisit it proportionally more often if it changes frequently. A site that publishes multiple articles per day warrants daily revisits, while a site whose content changes once a year warrants annual revisits.

Crawl Freshness

Crawl freshness refers to how up-to-date the content in a crawl is relative to the current state of the web. A fresh crawl contains recent versions of pages, while a stale crawl may contain outdated content. Freshness is especially important for news, financial data, and other time-sensitive information.

For language model training data, freshness has subtle implications. Models trained on more recent data know about events and terminology that did not exist when older data was collected. However, very recent data may include lower-quality content, as the web accumulates more ephemeral, low-effort content over time. Common Crawl's monthly release cadence provides a reasonable freshness tradeoff for training data pipelines that can select crawls by date.

The freshness-versus-bandwidth tradeoff is unavoidable. Revisiting all pages frequently requires enormous crawler bandwidth. In practice, production crawlers use adaptive scheduling that prioritizes fresh revisits to high-value, frequently-changing pages and less frequent revisits to stable content. The optimal revisit interval for a page is proportional to the page's change rate. Pages that change daily should be revisited daily; pages that change annually need only annual revisits. Some crawlers model this explicitly using exponential smoothing of observed change rates, while others use simpler heuristics based on content type.

Implementing a Focused Crawler

Let us build a functional focused crawler that demonstrates the core concepts. Our crawler will target a specific topic, respect robots.txt, implement rate limiting, and extract structured text from crawled pages. This is deliberately simplified for clarity; production crawlers add distributed execution, persistent frontiers, and much more sophisticated scheduling.

Setting Up the Environment

We need a few Python libraries for HTTP fetching, HTML parsing, and robots.txt handling:

In[5]:
Code
# uv pip install requests beautifulsoup4 lxml

Core Data Structures

A clean crawler separates concerns between the frontier (URL management), the fetcher (HTTP I/O), and the extractor (content parsing). Let us define the key data structures first:

In[6]:
Code
from dataclasses import dataclass, field


@dataclass
class CrawledPage:
    url: str
    title: str
    text: str
    links: list[str]
    status_code: int
    content_length: int
    crawl_timestamp: float


@dataclass
class CrawlConfig:
    seed_urls: list[str]
    max_pages: int = 100
    max_depth: int = 3
    crawl_delay: float = 1.0  # Seconds between requests to same host
    request_timeout: int = 10  # Seconds before HTTP timeout
    user_agent: str = "LanguageAIHandbookCrawler/1.0 (+https://example.com)"
    respect_robots: bool = True
    keywords: list[str] = field(default_factory=list)  # For focused crawling

The CrawlConfig bundles all tunable parameters. The user_agent string identifies our crawler so site operators can contact us. The keywords list enables focused crawling: we will score pages and URLs by relevance to these keywords and prioritize higher-scoring URLs. By separating configuration from crawler state, we can reuse the same crawler logic with different configurations for different crawl targets.

Robots.txt Cache

Checking robots.txt for every URL would be expensive. We cache the parsed robots.txt for each host, fetching it once per domain and reusing the parsed rules for all subsequent URLs on that domain:

In[7]:
Code
import urllib.parse
import urllib.robotparser
from typing import Optional


class RobotsCache:
    """Cache robots.txt rules for each domain to avoid repeated fetches."""

    def __init__(self, user_agent: str, timeout: int = 5):
        self._cache: dict[str, urllib.robotparser.RobotFileParser] = {}
        self.user_agent = user_agent
        self.timeout = timeout

    def can_fetch(self, url: str) -> bool:
        """Return True if robots.txt permits fetching this URL."""
        parsed = urllib.parse.urlparse(url)
        base = f"{parsed.scheme}://{parsed.netloc}"

        if base not in self._cache:
            rp = urllib.robotparser.RobotFileParser()
            rp.set_url(f"{base}/robots.txt")
            try:
                rp.read()
            except Exception:
                # If we cannot fetch robots.txt, assume access is allowed
                rp.allow_all = True
            self._cache[base] = rp

        return self._cache[base].can_fetch(self.user_agent, url)

    def get_crawl_delay(self, url: str) -> Optional[float]:
        """Return the crawl delay specified in robots.txt, if any."""
        parsed = urllib.parse.urlparse(url)
        base = f"{parsed.scheme}://{parsed.netloc}"
        if base in self._cache:
            return self._cache[base].crawl_delay(self.user_agent)
        return None

The urllib.robotparser.RobotFileParser class handles parsing robots.txt syntax, including User-agent, Disallow, Allow, and Crawl-delay directives. We extend it with a per-host cache so each domain's rules are fetched only once. The fallback when robots.txt is unreachable, to assume access is allowed, follows the convention adopted by most major crawlers: a missing robots.txt is not equivalent to a disallow-all directive.

The Focused Crawler

Now let us build the main crawler class. The scoring function is the heart of focused crawling: it assigns higher priority to URLs whose path and surrounding anchor text suggest relevance to our target keywords:

In[8]:
Code
import time
import urllib.parse
from collections import defaultdict

import requests
from bs4 import BeautifulSoup


class FocusedCrawler:
    """A focused web crawler that prioritizes pages matching specified keywords."""

    def __init__(self, config: CrawlConfig):
        self.config = config
        self.robots = RobotsCache(config.user_agent)
        self.visited: set[str] = set()
        self.results: list[CrawledPage] = []
        # Track last fetch time per host for rate limiting
        self._last_fetch: dict[str, float] = defaultdict(float)
        # Frontier: list of (score, depth, url) tuples
        self.frontier: list[tuple[float, int, str]] = []

    def _score_url(self, url: str, text_context: str = "") -> float:
        """Score a URL by keyword relevance. Higher = more relevant."""
        if not self.config.keywords:
            return 1.0
        url_lower = url.lower()
        context_lower = text_context.lower()
        score = sum(
            2.0 if kw in url_lower else (1.0 if kw in context_lower else 0.0)
            for kw in self.config.keywords
        )
        return score

    def _normalize_url(self, url: str, base_url: str) -> Optional[str]:
        """Normalize and validate a URL relative to its page."""
        try:
            abs_url = urllib.parse.urljoin(base_url, url)
            parsed = urllib.parse.urlparse(abs_url)
            # Only HTTP/HTTPS
            if parsed.scheme not in ("http", "https"):
                return None
            # Remove fragment
            normalized = parsed._replace(fragment="").geturl()
            return normalized
        except Exception:
            return None

    def _wait_for_host(self, url: str):
        """Enforce rate limiting per host."""
        parsed = urllib.parse.urlparse(url)
        host = parsed.netloc
        elapsed = time.time() - self._last_fetch[host]
        # Check for crawl-delay from robots.txt
        robots_delay = self.robots.get_crawl_delay(url) or 0.0
        delay = max(self.config.crawl_delay, robots_delay)
        if elapsed < delay:
            time.sleep(delay - elapsed)
        self._last_fetch[host] = time.time()

    def _fetch(self, url: str) -> Optional[requests.Response]:
        """Fetch a URL, returning the response or None on failure."""
        self._wait_for_host(url)
        try:
            headers = {"User-Agent": self.config.user_agent}
            response = requests.get(
                url,
                headers=headers,
                timeout=self.config.request_timeout,
                allow_redirects=True,
            )
            return response
        except requests.RequestException:
            return None

    def _extract(self, url: str, response: requests.Response) -> CrawledPage:
        """Extract title, text, and links from an HTML response."""
        soup = BeautifulSoup(response.content, "lxml")

        # Extract title
        title_tag = soup.find("title")
        title = title_tag.get_text(strip=True) if title_tag else ""

        # Remove script and style elements
        for tag in soup(["script", "style", "nav", "footer", "header"]):
            tag.decompose()

        # Extract main text
        text = soup.get_text(separator=" ", strip=True)
        # Clean excessive whitespace
        text = re.sub(r"\s+", " ", text).strip()

        # Extract all links
        links = []
        for anchor in soup.find_all("a", href=True):
            normalized = self._normalize_url(anchor["href"], url)
            if normalized and normalized not in self.visited:
                links.append(normalized)

        return CrawledPage(
            url=url,
            title=title,
            text=text,
            links=links,
            status_code=response.status_code,
            content_length=len(response.content),
            crawl_timestamp=time.time(),
        )

    def crawl(self) -> list[CrawledPage]:
        """Run the crawl and return all successfully fetched pages."""
        # Seed the frontier
        for url in self.config.seed_urls:
            score = self._score_url(url)
            self.frontier.append((score, 0, url))

        while self.frontier and len(self.results) < self.config.max_pages:
            # Sort by score descending to visit best URLs first
            self.frontier.sort(key=lambda x: x[0], reverse=True)
            score, depth, url = self.frontier.pop(0)

            if url in self.visited:
                continue
            self.visited.add(url)

            # Check robots.txt
            if self.config.respect_robots and not self.robots.can_fetch(url):
                continue

            # Fetch the page
            response = self._fetch(url)
            if response is None or response.status_code != 200:
                continue

            # Extract content
            content_type = response.headers.get("Content-Type", "")
            if "text/html" not in content_type:
                continue

            page = self._extract(url, response)
            self.results.append(page)

            # Expand frontier if within depth limit
            if depth < self.config.max_depth:
                for link in page.links:
                    link_score = self._score_url(link, page.text)
                    self.frontier.append((link_score, depth + 1, link))

        return self.results

The scoring function rewards URLs that contain relevant keywords in the URL path itself (score 2.0 per keyword match) or in the surrounding anchor text context (score 1.0 per match). This distinction matters because a URL like /python-debugging-guide strongly suggests relevant content, while an anchor link that says "click here" but happens to be on a Python page provides weaker evidence. By sorting the frontier by score before each fetch, the crawler naturally gravitates toward the highest-relevance content, spending its bandwidth budget where it is most likely to find useful pages.

Running a Simulated Crawl

For demonstration purposes, we simulate what a crawl would produce without hitting live servers. The simulation uses realistic distributions for content lengths, domain distributions, and page link structures:

In[9]:
Code
import random
import string
import time

# Simulate crawl results to show statistics
random.seed(42)


def simulate_crawl_results(n_pages: int = 150) -> list[CrawledPage]:
    """Simulate realistic crawl output for analysis."""
    domains = [
        "nlp-papers.org",
        "machinelearning.wiki",
        "ai-research.io",
        "deep-learning-blog.com",
        "transformer-models.net",
        "linguistics.edu",
        "corpus-research.org",
        "text-mining.io",
        "spam-site-123.xyz",
        "thin-content-456.click",
    ]

    # Power-law weights: top domains get disproportionately more pages
    # Reflects web's link-following concentration effect
    domain_weights = [30, 22, 14, 9, 6, 4, 7, 3, 3, 2]

    pages = []
    for i in range(n_pages):
        domain = random.choices(domains, weights=domain_weights)[0]
        url = f"https://{domain}/page-{i:04d}"

        # Realistic content length distribution (log-normal)
        mean_log = 8.5  # log(~4900 chars)
        std_log = 1.2
        content_len = int(random.lognormvariate(mean_log, std_log))
        content_len = max(50, min(content_len, 200_000))

        # Generate text with varying quality
        is_quality = domain not in [
            "spam-site-123.xyz",
            "thin-content-456.click",
        ]
        word_count = content_len // 6
        text = " ".join(
            random.choice(string.ascii_lowercase * 3)
            for _ in range(min(word_count, 500))
        )

        # Vary status codes (most succeed)
        status = random.choices(
            [200, 200, 200, 200, 200, 301, 404, 429, 503],
            weights=[70, 5, 5, 5, 5, 4, 3, 2, 1],
        )[0]

        pages.append(
            CrawledPage(
                url=url,
                title=f"Page {i} on {domain}",
                text=text,
                links=[
                    f"https://{domain}/page-{random.randint(0, n_pages)}"
                    for _ in range(random.randint(2, 15))
                ],
                status_code=status if status == 200 else status,
                content_length=content_len,
                crawl_timestamp=time.time() - random.uniform(0, 3600),
            )
        )

    # Only keep successful fetches (status 200)
    return [p for p in pages if p.status_code == 200]


crawl_results = simulate_crawl_results(150)

# Compute statistics
total_pages = len(crawl_results)
total_bytes = sum(p.content_length for p in crawl_results)
avg_content_len = total_bytes / total_pages if total_pages > 0 else 0

domain_counts: dict[str, int] = defaultdict(int)
for page in crawl_results:
    domain = urllib.parse.urlparse(page.url).netloc
    domain_counts[domain] += 1

top_domains = sorted(domain_counts.items(), key=lambda x: x[1], reverse=True)[
    :5
]
Out[10]:
Console
Crawl Summary
========================================
Pages successfully fetched: 136
Total content downloaded:   1323.1 KB
Average content length:     9,962 bytes

Top domains by page count:
  nlp-papers.org                       40 pages  (29.4%)
  machinelearning.wiki                 32 pages  (23.5%)
  ai-research.io                       21 pages  (15.4%)
  deep-learning-blog.com               17 pages  (12.5%)
  transformer-models.net                7 pages  (5.1%)

The distribution of pages across domains illustrates a fundamental characteristic of web crawls: content is not uniformly distributed. Some domains contribute disproportionately many pages, especially if the crawler follows many internal links. This domain imbalance is one reason deduplication and domain-level filtering are critical steps in the data processing pipeline.

Analyzing Content Length Distribution

Content length is a simple but powerful proxy for content quality. Very short pages, under 200 bytes of text, are often spam, error pages, or thin content with no informational value. Very long pages occasionally contain high-quality reference material but can also be terms-of-service documents, cookie consent dialogs, or concatenated low-quality content. The relationship between length and quality is real but imperfect: a 200-word blog post may be more useful than a 5,000-word spun article that repeats the same idea ten different ways.

In[11]:
Code
import numpy as np

content_lengths = np.array([p.content_length for p in crawl_results])

# Compute percentile statistics
p10 = np.percentile(content_lengths, 10)
p25 = np.percentile(content_lengths, 25)
p50 = np.percentile(content_lengths, 50)
p75 = np.percentile(content_lengths, 75)
p90 = np.percentile(content_lengths, 90)

# Filter thresholds used in practice
min_threshold = 200  # bytes; below this is almost certainly junk
max_threshold = 100_000  # bytes; above this is often boilerplate-heavy

filtered_pages = [
    p
    for p in crawl_results
    if min_threshold <= p.content_length <= max_threshold
]
filter_rate = 1 - len(filtered_pages) / len(crawl_results)
Out[12]:
Console
Content Length Distribution
========================================
10th percentile:       1,292 bytes
25th percentile:       2,288 bytes
Median:                6,118 bytes
75th percentile:      12,633 bytes
90th percentile:      24,074 bytes

After length filtering (200–100,000 bytes):
  Pages retained:  136 (100.0%)
  Pages filtered:  0 (0.0%)

The log-normal distribution of content lengths means most pages cluster around a moderate size, with long tails at both extremes. Quality filtering based on length typically removes 10 to 30 percent of crawled pages in practice, depending on how aggressively the crawler followed low-quality domains. This filtering step is fast and cheap, making it an excellent first pass before the more expensive operations like language detection and quality classification.

Key Parameters

The main parameters governing a web crawler's behavior are:

  • crawl_delay: Minimum seconds between requests to the same host. Typical values range from 0.5 to 5 seconds. Always honor the Crawl-delay directive in robots.txt when it specifies a higher value.
  • max_depth: Maximum link distance from seed URLs to follow. Shallow crawls (depth 2-3) are efficient and produce higher average quality; deep crawls (depth 5+) surface long-tail content but bring in more noise.
  • max_pages: A stopping condition based on page count. Combined with max_depth, it prevents runaway crawls.
  • keywords (focused crawling): Relevance keywords used to score and prioritize URLs. More specific keywords produce higher-precision focused crawls at the cost of recall.
  • user_agent: The crawler's identifier string. Always include contact information so site operators can reach you.

Visualizing Crawl Dynamics

The following visualizations illustrate how content length distribution and domain diversity evolve during a crawl, along with how relevance scoring concentrates focused crawl results.

Out[13]:
Visualization
Histogram of content lengths in bytes on a log x-axis with filter threshold lines marked.
Log-scale histogram of page content lengths from a simulated web crawl. The distribution follows an approximately log-normal shape, with most pages falling between 1,000 and 50,000 bytes. Vertical dashed lines mark the lower threshold (200 bytes, below which pages are almost certainly spam or error responses) and the upper threshold (100 KB, above which pages tend to be boilerplate-heavy). Both thresholds are commonly applied in production data pipelines as the cheapest first-pass quality filter.
Out[14]:
Visualization
Horizontal bar chart showing page counts per domain, with the highest-count domain at top.
Distribution of crawled pages across domains in a simulated focused crawl. The uneven distribution reflects the web''s power-law link structure: a small number of domains receive disproportionate crawler attention when following links, mirroring the real-world concentration of web traffic. In production pipelines, domain capping limits any single domain''s share of the training corpus to prevent the final dataset from being dominated by a handful of sites.
Out[15]:
Visualization
Bar chart showing average content length decreasing at higher link depths.
Average content length by link depth from seed URLs in a simulated crawl. Pages closer to seeds tend to be longer and more substantive, showing the editorial quality of well-linked authoritative pages. The decline at greater depths reflects the increasing prevalence of thin content and spam in the web graph's periphery. Error bars show one standard deviation across simulated pages at each depth level.
Horizontal bar chart comparing change rates across page types on a log scale.
Simulated page change rate by content type on a log scale, showing how frequently pages are updated relative to one another. News articles change orders of magnitude more frequently than academic papers, which justifies very different revisit scheduling for each content type. Freshness-aware crawlers allocate revisit bandwidth proportionally to each category's observed change rate.

Crawl Freshness in Depth

Understanding freshness matters operationally and epistemologically. A model does not simply know facts; it knows the snapshot of the web that existed during its training data collection. The date range of the crawl data shapes what the model considers current, recent, and historically established.

Freshness Estimation

Estimating how fresh a crawl is for a given URL requires comparing the crawl timestamp to the page's last-modification time. HTTP provides two mechanisms for this. The Last-Modified response header indicates when the server believes the content was last changed. The ETag header provides a content fingerprint that changes when the content changes. Crawlers that support conditional requests can send If-Modified-Since or If-None-Match headers with a subsequent request, and the server responds with a 304 Not Modified status if the content has not changed, saving download bandwidth.

In practice, many servers do not implement these headers correctly. Pages may have Last-Modified headers that reflect the page template's modification date rather than the actual content update time. ETags may be generated randomly on each request due to server misconfiguration. Crawlers that want accurate freshness information often fall back to content change detection: hash the extracted text from the previous crawl, compare against the current hash, and flag pages as changed only when the hashes differ. This approach is accurate but requires storing a hash for every URL ever crawled, which adds significant storage overhead at web scale.

Temporal Bias in Training Data

An important subtlety is that Common Crawl data is not a uniform temporal sample. Newer crawls contain more pages from recently created websites and fewer pages from sites that have since gone offline. The web has grown substantially over Common Crawl's history, so earlier crawls have sparser coverage than recent ones even for domains that have existed for many years.

This temporal non-uniformity can create training artifacts. A model trained on data from a narrow time window may underrepresent historical knowledge and overrepresent contemporary internet culture. A model trained only on 2022 crawls will have strong beliefs about what is "current" as of 2022 and little knowledge of events after that date. Mixing crawls from different time periods helps address this, and several recent training data pipelines explicitly stratify their Common Crawl samples by crawl date to ensure temporal diversity.

The model's knowledge cutoff, the date after which it has no information, arises directly from the crawl dates of its training data. When users notice a model claiming not to know about a recent event, the root cause is almost always in the crawl dates. Understanding the connection between crawl freshness and model knowledge cutoffs helps practitioners reason clearly about what their models can and cannot be expected to know.

HTTP Status Code Patterns in Crawls

Understanding how status codes distribute across a real crawl reveals important characteristics of the web that affect training data quality. Let us visualize what a typical status code distribution looks like:

In[16]:
Code
# Build a simulated status code distribution based on empirical crawl data
# Distributions reflect typical web crawl observations from Common Crawl reports
status_categories = {
    "200 OK": 0.72,
    "301/302 Redirect": 0.11,
    "404 Not Found": 0.09,
    "403 Forbidden": 0.04,
    "429 Rate Limited": 0.02,
    "5xx Server Error": 0.015,
    "Other": 0.005,
}

# Simulate counts for a crawl of 10000 URLs
total_attempted = 10_000
status_counts = {
    k: int(v * total_attempted) for k, v in status_categories.items()
}

# Compute pages actually usable for training
usable = status_counts["200 OK"]
usable_pct = usable / total_attempted * 100
Out[17]:
Visualization
Horizontal bar chart showing HTTP status code categories by percentage of total crawl attempts.
Simulated HTTP status code distribution for a web crawl of 10,000 URLs, based on patterns observed in large-scale crawls. Successful responses (200 OK) account for roughly 72% of attempts. Redirects account for an additional 11%, many of which ultimately resolve to successful pages. Together, the non-2xx responses reveal that nearly 30% of crawl bandwidth is spent on pages that cannot directly contribute usable training data, reinforcing the importance of efficient URL prioritization.

The status code breakdown exposes the real efficiency of a web crawl. Only pages with 200 OK status codes directly contribute training data, and even among those, quality filtering will remove a significant fraction. The 11% redirect rate represents URLs where the content has moved; a crawler that follows redirects correctly can recover some of this content, but chains of more than two or three redirects often lead to low-quality destinations or infinite redirect loops. The 9% 404 rate represents dead links in the web graph, which is unavoidable: the web changes constantly and old links break.

Limitations and Practical Challenges

Web crawling at scale poses both engineering and sociotechnical challenges. Several important limitations shape what crawled data can and cannot represent.

Coverage Bias

The web is not uniformly accessible. Content behind login walls, paywalls, or JavaScript-rendered single-page applications is largely invisible to simple crawlers. Academic journals, enterprise intranets, subscription newsletters, and private social media posts collectively represent an enormous body of knowledge that web crawling cannot access. Common Crawl's text is therefore biased toward publicly accessible content, which skews toward English-language, commercially oriented, and technically accessible writing.

Dynamic content generation poses a separate challenge. A growing fraction of web pages are rendered entirely client-side using JavaScript frameworks. A basic HTTP crawler receives only the empty HTML skeleton, missing all the content that JavaScript would populate. Crawling JavaScript-rendered pages requires a full browser automation layer using tools like Playwright or Puppeteer, which adds significant complexity and resource cost. Running a headless browser for every page is 10 to 50 times slower and more resource-intensive than simple HTTP fetching. For this reason, most large-scale crawls still use HTTP-only fetching and accept the coverage gap for JavaScript-heavy sites.

The language distribution of the web also creates bias. English dominates the web in terms of both page count and content density, but the actual population of potential language AI users is globally distributed. Common Crawl contains content in dozens of languages, but English pages vastly outnumber pages in any other language. Models trained on unfiltered Common Crawl data inherit this English-language dominance. Building multilingual training data requires deliberate effort: explicitly identifying and proportionally including content from underrepresented languages rather than relying on raw crawl proportions.

Domain Imbalance

Without domain capping, link-following crawlers naturally accumulate enormous numbers of pages from a few highly interconnected domains. A crawler that follows all links would eventually download billions of Reddit comments and Stack Overflow answers, which, while potentially useful, would swamp other content types. This concentration effect arises from the web's power-law link structure: the most popular domains receive many times more inbound links than less popular ones, so a link-following crawler visits them many times more often.

Production data pipelines typically apply domain caps, limiting any single domain to a fixed fraction of the total corpus, to ensure diversity. If Reddit alone contributes 20% of a training corpus, the model will be heavily biased toward casual conversational writing and the specific cultural norms of Reddit communities. A cap of 1% per domain across thousands of domains produces a much more balanced representation of writing styles and topics.

Copyright law creates significant uncertainty around web crawling for training data purposes. While robots.txt provides a social contract for access permission, it does not confer copyright permission. Content creators increasingly argue that training language models on their writing without compensation constitutes copyright infringement. Several lawsuits in the United States and Europe have challenged major AI companies on exactly these grounds, with ongoing legal proceedings that have not yet produced clear settled law.

The tension remains unresolved. Open research into language AI depends on large-scale crawled data. But the people who created that data, journalists, authors, coders, educators, often received no compensation and gave no explicit consent. Responsible use of crawled data should at minimum respect robots.txt, provide opt-out mechanisms, and acknowledge the human labor that produced the training corpus. Some organizations have begun offering opt-out registries where content creators can request removal from training datasets, though the practical effectiveness of these mechanisms remains limited.

The situation is evolving rapidly. The European Union's AI Act and related regulations are establishing new requirements around training data transparency. Some jurisdictions are exploring data rights frameworks that would give content creators enforceable claims over uses of their work. Anyone building a production crawling pipeline should monitor legal developments in their target jurisdictions and consult legal counsel before deploying at scale.

Data Quality Across the Web

The web is not a curated library; it is the sum of everything humans have chosen to publish. This includes rigorous scholarship alongside spam, misinformation, hate speech, and low-effort content. A crawl that follows links aggressively will collect substantial quantities of all of these. Spam farms generate thousands of keyword-stuffed pages designed to manipulate search rankings rather than inform readers. Content mills produce formulaic articles that repeat the same information in slightly varied phrasing, contributing enormous volume with minimal unique information. Deliberately misleading content, conspiracy theories, and factually incorrect claims appear throughout the web at substantial scale.

The downstream filtering pipeline is as important as the crawl itself. A thoughtfully filtered medium-sized crawl often produces better training data than a massive unfiltered one. We discuss quality filtering, toxicity filtering, and deduplication in the following chapters in this section. For now, the key insight is that web crawling produces raw material that requires substantial refinement before it is suitable as training data. The crawl is the beginning of the pipeline, not the end.

Summary

Web crawling provides the raw material for the largest language model training corpora. The core loop is simple: fetch URLs, extract links, repeat. The complexity comes from doing this at scale while respecting servers, managing quality, and complying with legal constraints.

Key takeaways from this chapter:

  • Common Crawl is the foundational public data source for language AI, giving petabyte-scale WARC, WAT, and WET data with monthly crawl releases. WET files provide pre-extracted plain text for NLP use cases, while WARC files preserve the complete raw crawl data for custom text extraction.
  • Robots.txt defines the social contract between sites and crawlers. Respecting it is both an ethical obligation and practically essential for sustainable crawling. The protocol includes allow/disallow rules, crawl delay directives, and sitemap pointers.
  • Politeness policies prevent a single crawler from overwhelming servers. Per-host rate limiting, crawl delay respect, adaptive backoff on server errors, and responsible user-agent identification are minimum requirements.
  • Crawl strategy shapes data character. Breadth-first crawls produce broad coverage; focused crawls produce topically relevant but narrower datasets. Depth and freshness tradeoffs further determine what ends up in the corpus.
  • Coverage bias is inherent: web crawls miss paywalled, dynamic, and login-required content, skewing the resulting data toward publicly accessible, often English-language material. JavaScript rendering, domain imbalance, and temporal non-uniformity create additional biases that practitioners must understand and manage.
  • Data quality varies enormously across the web. Content length filtering is a fast first pass, but the full pipeline requires deduplication, language identification, quality filtering, and toxicity filtering, topics we turn to next.

The next chapter on Document Extraction explores how to reliably convert raw HTML from crawl data into clean text, a step that proves more difficult than it first appears given the diversity of web page structures and encodings.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about web crawling.

Web Crawling Quiz

Question 1 of 80 of 8 completed
What is the robots.txt file used for?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026webcrawling, author = {Michael Brenndoerfer}, title = {Web Crawling: Common Crawl, Robots.txt, and Freshness}, year = {2026}, url = {https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Web Crawling: Common Crawl, Robots.txt, and Freshness. Retrieved from https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt
MLAAcademic
Michael Brenndoerfer. "Web Crawling: Common Crawl, Robots.txt, and Freshness." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt>.
CHICAGOAcademic
Michael Brenndoerfer. "Web Crawling: Common Crawl, Robots.txt, and Freshness." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Web Crawling: Common Crawl, Robots.txt, and Freshness'. Available at: https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Web Crawling: Common Crawl, Robots.txt, and Freshness. https://mbrenndoerfer.com/writing/web-crawling-common-crawl-strategies-robots-txt

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.