Part of Language AI Handbook
Explains how web crawlers discover and download internet content at scale, covering Common Crawl architecture, robots.txt compliance, politeness policies.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Web Crawling
The internet is the largest corpus of human knowledge ever assembled. Hundreds of billions of web pages document everything from scientific research to casual conversation, from formal legislation to informal social media posts. When researchers at Common Crawl or AI labs want to train a language model on this knowledge, they face a fundamental engineering challenge: how do you systematically discover, download, and store that content at a scale that spans petabytes of data?
Web crawling is the answer. A web crawler, also called a spider or bot, is a program that automatically traverses the web by following hyperlinks, downloading page content as it goes. The same technology that powers Google's search index also underpins the training data pipelines for virtually every large language model trained today. GPT-3 used filtered Common Crawl data for roughly 60% of its training corpus. LLaMA, Falcon, and Mistral all draw heavily from crawled web text. Understanding how web crawling works, where its data comes from, and how to respect its social contract is foundational knowledge for anyone building modern language AI systems.
This chapter covers the mechanics of web crawling, from seed URL discovery through frontier management and politeness policies. We examine Common Crawl in depth because it is the single most important public data source in language AI. We then explore how freshness, depth, and breadth tradeoffs shape the character of crawled datasets, and we work through a hands-on implementation of a focused crawler. By the end, you will understand how crawling works mechanically and why every design decision, from seed selection to rate limiting, has downstream consequences for the quality and character of the training data you produce.
How Web Crawlers Work
Web crawlers operate a simple loop: take a URL, download the page at that URL, extract all links from the page, add new links to a queue, and repeat. The complexity comes from doing this at scale, doing it without overwhelming servers, and doing it in a way that produces high-quality, well-organized data.
The conceptual simplicity of this loop is deceptive. A production crawler must handle hundreds of millions of URLs daily, deal with malformed HTML, follow redirects, manage connection pooling, cache DNS results, respect per-host rate limits, parse and obey robots.txt rules, detect and skip duplicate pages, and write output in formats that downstream pipelines can process efficiently. Every one of these concerns interacts with the others. A DNS cache that is too aggressive causes stale IP addresses that lead to connection failures. A rate limiter that is too loose causes site operators to block the crawler's IP. An HTML parser that fails silently on malformed markup drops entire pages from the corpus without warning.
This section walks through the core components of the crawl process: seed selection, the main loop, URL normalization, DNS and HTTP mechanics, and frontier management at scale.
Seed URL Selection
Before any crawling begins, a crawler needs a starting point. Seed URLs are the initial set of pages from which the crawler begins following links. The choice of seeds has an outsized influence on what the final crawl contains, because breadth-first traversal explores the neighborhood of seeds first and most thoroughly.
For general-purpose crawls like Common Crawl, seeds typically come from several sources. Authoritative web directories, such as the DMOZ Open Directory (now archived) or curated lists of major news sites, government portals, and educational institutions, provide high-quality starting points. URLs from previous crawls that were productive are re-seeded to maintain continuity. XML sitemaps published by individual sites enumerate pages that owners want indexed, giving a direct invitation to crawl. Domain registrar data reveals newly registered domains that may contain fresh content.
The seed strategy creates a quality bias from the start. A crawl seeded from Wikipedia will produce data richer in factual reference content than a crawl seeded from social media domains. A crawl seeded from curated academic directories will skew toward formal writing and technical vocabulary. This quality inheritance, where crawled content mirrors the editorial standards of the seeded domains, is a feature that language AI practitioners exploit deliberately when building specialized corpora.
The number of seeds also matters. A few hundred seed URLs in a breadth-first crawler will naturally cluster the crawl around those seeds' neighborhoods. A million seeds, carefully selected to span diverse languages, topics, and regions, produces much more balanced coverage. Common Crawl uses tens of millions of seeds across a wide variety of domains and languages, which is part of why its coverage is unusually broad compared to purpose-built enterprise crawlers.
For a focused crawl building a domain-specific corpus, seeds are chosen differently. If you want a dataset of Python programming content, you might seed from the official Python documentation, PyPI project pages, Stack Overflow Python questions, major Python blogs, and key GitHub repositories. These seeds will guide the link-following traversal toward the neighborhood of Python content. You might also feed the crawler a keyword list and have it score discovered URLs by the presence of those keywords, making sure that pages added to the frontier are relevant to your target domain.
The Crawl Loop
The fundamental crawl algorithm proceeds through a series of well-defined steps. The crawler begins with a set of seed URLs, typically highly authoritative pages like major news sites, Wikipedia, or web directories. For each URL in the queue, the crawler issues an HTTP GET request and stores the response. It then parses the response to extract hyperlinks using the href attributes in anchor tags. Each new URL that has not already been visited and passes acceptability filters gets added to the frontier, the data structure holding URLs yet to visit. This loop continues until either the frontier is exhausted or some stopping condition, like a document count or time limit, is reached.
A web crawler is an automated program that systematically browses the World Wide Web, following hyperlinks to discover and download web pages. Crawlers index web content for search engines or build datasets for training language models. They are also called spiders or bots.
The simplest frontier is a FIFO queue, which produces a breadth-first traversal. Start from seeds, visit all pages linked from seeds, then all pages linked from those pages, and so on. Breadth-first crawling naturally covers the most highly connected pages first, since popular pages appear as targets in many early crawls.
Depth-first traversal, using a stack instead of a queue, follows link chains deep into individual sites before moving on. This strategy is more useful for focused crawls that want to exhaust a specific domain's content, but it risks getting stuck in link traps: cycles of dynamically generated pages that extend indefinitely. A product catalog with 10,000 items and a sorting parameter, a calendar system generating every possible date URL, or a search results page with query string variations can all trap a depth-first crawler in a technically infinite loop if the crawler does not implement loop detection.
Production crawlers typically use priority queues rather than simple FIFO structures. URLs are scored based on predicted importance, recency, or content type, and higher-scored URLs are downloaded first. This allows the crawler to be more selective about what it visits when resources are limited. A priority crawler might assign higher scores to URLs from domains that have historically produced high-quality content, URLs that contain specific keywords relevant to the crawl's focus, or URLs that appear as link targets on many different pages, a signal of importance analogous to PageRank.
URL Normalization and Deduplication
The web is full of URLs that point to the same content. Consider these URLs:
http://example.com/pageHTTP://EXAMPLE.COM/PAGEhttp://example.com/page?utm_source=twitterhttp://example.com/page#section2http://www.example.com/page
All of these may return identical or nearly identical content. Without normalization, a crawler visits the same page multiple times under different URLs, wasting bandwidth and cluttering the dataset with duplicates. URL normalization applies a sequence of transformations to produce a canonical form: lowercasing the scheme and host, sorting query parameters, removing fragment identifiers, stripping known tracking parameters like utm_source and fbclid, and resolving relative URLs against the page's base URL.
URL normalization is more subtle than it appears. Some query parameters are semantically significant, like ?lang=en or ?page=2, while others are pure tracking noise. Stripping the wrong parameters produces false merges where two different pages get collapsed to the same canonical URL. Keeping too many tracking parameters produces false splits where many crawl attempts hit the same actual page. Production crawlers maintain lists of known tracking parameters to strip and use heuristics for unknown parameters based on their structure and context.
Even after normalization, URL deduplication alone is insufficient. Many websites serve the same content at multiple canonical URLs through redirects, mirrors, or content management system artifacts. A news article might be accessible at /2024/03/article-title, /article/12345, and /amp/article-12345, all returning essentially the same text. For this reason, content-level deduplication, which we will explore in the Deduplication chapter, is equally important downstream.
Link traps deserve special attention. Some websites generate infinite URLs intentionally or accidentally. A calendar application might produce URLs for every possible date. A dynamic filter interface might allow arbitrary combinations of filtering parameters, each yielding a unique URL. E-commerce sites with faceted search can generate billions of unique-looking URLs from a few dozen actual products. Crawlers defend against link traps with several heuristics: URL length limits, maximum path depth, pattern detection for parametric URL explosions, and maximum pages-per-domain caps.
DNS Resolution and HTTP Fetching
For each URL, the crawler must resolve the hostname to an IP address using the Domain Name System, then open a TCP connection and issue an HTTP request. This two-step process sounds straightforward but becomes a significant bottleneck at scale.
DNS resolution is inherently network-dependent. Resolving a single hostname may take anywhere from 1 millisecond for a cached entry to several hundred milliseconds for a cold lookup against distant authoritative name servers. At web scale, with millions of unique hostnames being crawled, DNS resolution latency accumulates to a substantial fraction of total crawl time. Production crawlers cache DNS results aggressively, typically with TTLs of several hours, and batch DNS lookups for new domains using asynchronous DNS libraries that can resolve thousands of hostnames concurrently. Some large-scale crawlers maintain their own authoritative DNS resolvers to reduce latency further.
HTTP connections carry significant overhead from TCP handshake and TLS negotiation. A full TLS 1.3 handshake requires two round trips before any application data can be sent. Crawlers reuse connections via HTTP keep-alive within the same host and batch requests to the same IP to amortize connection costs. HTTP/2 support provides multiplexed connections where multiple requests share a single TCP connection, further reducing per-request overhead. At the scale of millions of pages per day, even small per-request latency improvements translate to meaningful increases in crawl throughput.
Response handling requires checking status codes, and each category demands a different response. A 200 OK means success; the content can be stored. A 301 Moved Permanently or 302 Found means the page has moved; the crawler follows the redirect URL and records the mapping between old and new URL. A 404 Not Found means the page no longer exists; the crawler records this and removes the URL from future crawl schedules. A 429 Too Many Requests means the server is throttling requests; the crawler must back off exponentially and try again later. A 503 Service Unavailable means temporary unavailability; the crawler queues the URL for a retry after a delay. A 5xx error from a server generally indicates a transient problem, and well-designed crawlers implement exponential backoff with jitter for such failures to avoid thundering-herd effects when many crawl workers hit the same server simultaneously.
Content-type negotiation matters too. Crawlers should request text/html and handle other content types appropriately. A URL that returns a PDF needs different processing than one that returns HTML. Some crawlers include PDF and EPUB extraction pipelines for high-quality documents, while others skip non-HTML content types entirely for simplicity.
Robots.txt files tell crawlers to avoid certain paths entirely, and well-behaved crawlers check and respect these instructions before fetching any URL on a domain. We discuss robots.txt in detail in the next section.
The Frontier at Scale
For small crawls, an in-memory queue is sufficient. At web scale, the frontier can contain billions of URLs that must be managed persistently across machine restarts. Production crawlers like those operated by Google or Common Crawl use distributed frontier managers backed by databases or specialized queue systems.
A key design challenge is balancing crawl politeness against crawl speed. You cannot send a thousand requests per second to a single server. Instead, the frontier must ensure that no server receives requests faster than its politeness policy allows. This typically means grouping URLs by host, and for each host, maintaining a separate subqueue with delays between requests. The classic approach partitions the frontier into per-host buckets and draws one URL from each bucket in round-robin fashion, inserting a mandatory wait before re-visiting any bucket.
The crawl frontier is the data structure that holds URLs discovered but not yet fetched by a web crawler. Managing the frontier involves deciding which URLs to visit next (scheduling), how frequently to revisit known pages (freshness policy), and how to avoid overwhelming any single server (politeness).
Distributed crawlers add further complexity. When many worker machines are fetching pages in parallel, the shared frontier must coordinate so that no two workers attempt to fetch the same URL simultaneously. This requires either a centralized frontier service with high availability, or a partitioned frontier where URL assignments are sharded by host and each shard is owned by exactly one worker. Common Crawl's crawler infrastructure, for example, uses a distributed queue system that ensures consistent URL assignment across hundreds of parallel workers while maintaining per-host politeness.
The frontier also maintains metadata for each URL beyond the URL string itself. Crawlers track the URL's discovery timestamp, its link depth from the nearest seed, its estimated relevance score, its crawl priority, and any known metadata from sitemaps. This metadata drives scheduling decisions that determine which pages get crawled and when.
Robots.txt and Ethical Crawling
Web crawlers have enormous power. A single aggressive crawler can bring down a web server by flooding it with requests. Even well-intentioned crawlers can cause problems: generating excessive server load, consuming bandwidth that costs site owners money, and accessing private or sensitive content that was never meant to be indexed.
The robots exclusion protocol, formalized as robots.txt, is the web's social contract between site owners and crawlers. When a crawler first visits a domain, it should fetch https://example.com/robots.txt and parse the access rules before crawling any other URL on that domain. This file tells the crawler which paths are off limits, which crawlers are welcome, and how fast requests should be made.
Parsing robots.txt
The robots.txt format specifies rules for different user agents using User-agent directives, followed by Disallow and Allow rules. The wildcard * matches any user agent not explicitly named.
User-agent: *
Disallow: /private/
Disallow: /admin/
Allow: /public/
User-agent: Googlebot
Disallow: /staging/
Crawl-delay: 1
Sitemap: https://example.com/sitemap.xml
In this example, all crawlers are blocked from /private/ and /admin/, but explicitly allowed to access /public/. Googlebot has an additional exclusion for /staging/ and must wait at least 1 second between requests. The sitemap declaration points to a structured list of URLs the site owner wants crawled.
Robots.txt parsing has several edge cases that matter for correct implementation. The Allow directive takes precedence over Disallow when both match a URL, but the more specific rule wins when there is no explicit precedence conflict. An Allow: /public/ rule overrides Disallow: / for any URL under /public/. The order of rules within a group matters in some implementations, while others apply longest-match semantics. Wildcards (*) and end-of-string anchors ($) add expressiveness: Disallow: /*.pdf$ blocks all PDF files regardless of directory.
The Crawl-delay directive specifies the minimum number of seconds to wait between successive requests to the server. Respecting this directive is polite and often legally required in jurisdictions where unauthorized automated access may violate terms of service. When both the robots.txt crawl delay and your own configured delay exist, you should honor whichever is more conservative.
The Sitemap directive points to XML sitemap files that enumerate the site's publicly available pages. Smart crawlers fetch these sitemaps to seed their frontier with known good URLs rather than relying entirely on link following. A well-maintained sitemap tells the crawler exactly which URLs the site owner considers canonical and worth indexing, eliminating guesswork about URL normalization and reducing the chance of crawling duplicate or unwanted pages.
Why Respecting robots.txt Matters
Respecting robots.txt is both an ethical obligation and a practical necessity. Many websites disallow crawling of user-generated content areas, login-protected pages, or dynamically generated search results that would produce near-infinite URLs. Crawling these areas wastes resources and produces low-quality data. Pages behind authentication walls return redirect responses or login forms rather than actual content. Search result pages produce query-dependent content that changes with every request and may contain little informational value as training data.
The legal dimension is significant. Some jurisdictions treat ignoring robots.txt as unauthorized computer access. In the United States, the Computer Fraud and Abuse Act has been used to argue against unauthorized scraping, though legal interpretations vary. In Europe, the GDPR creates obligations around processing personal data, which web pages may contain. Beyond legality, violating robots.txt damages trust between the open-web ecosystem and the research community. Sites that feel violated by crawlers often block entire IP ranges, penalizing legitimate crawlers along with bad actors.
Some data pipelines used to train language models have been criticized for ignoring robots.txt at scale. When Common Crawl began operations, some site owners found their content in crawl data despite having disallowed bots. This tension between the value of open training data and the rights of content creators has become a central legal and ethical debate in AI development. If you are building a crawler, respecting robots.txt is the minimum acceptable standard.
The robots exclusion protocol is a standard for websites to communicate which parts of their content crawlers are permitted to access. A file named robots.txt at the root of a domain contains rules specifying which user agents are allowed or disallowed from fetching which URL paths.
Politeness Policies Beyond robots.txt
Even when robots.txt grants access, good crawlers implement additional politeness measures. The most important is rate limiting. Sending too many requests too quickly stresses servers and degrades the experience for human visitors. A common heuristic is to wait a multiple of the server's response time before issuing the next request, adapting automatically to server load: if a server takes 2 seconds to respond, wait an additional 4 seconds before sending the next request. This adaptive delay naturally backs off when servers are under load without requiring you to hard-code domain-specific delays.
IP-level rate limiting matters too. A single website may be hosted behind a shared IP that serves thousands of domains. Hitting that IP with aggregate traffic from all those domains can cause collateral damage to unrelated sites. Production crawlers track per-host request rates and per-IP request rates to avoid this kind of collateral congestion.
Identifying yourself through the User-Agent HTTP header is another courtesy. Well-behaved crawlers provide a user agent string that includes a URL or email address where the site owner can contact you if your crawler causes problems. This allows site operators to reach out about issues rather than simply blocking your IP. Compare Mozilla/5.0 (compatible; MyCrawler/1.0) with MyCrawler/1.0 (+https://myproject.org/crawler; contact@myproject.org). The second form gives site operators actionable information.
Crawlers should also implement exponential backoff when they encounter server errors. A 503 Service Unavailable response means the server is temporarily overwhelmed. Immediately retrying adds to that overwhelm. Instead, wait an exponentially increasing delay: 30 seconds, then 60, then 120, then give up. This pattern is respectful and also more likely to succeed, since transient overloads typically resolve within minutes.
Common Crawl
Common Crawl is a nonprofit organization that has been crawling the web since 2007 and making its data freely available. It has become the single most important source of training data for large language models. Understanding Common Crawl's structure, scale, and quality characteristics is essential for anyone building language AI systems.
Scale and Coverage
Common Crawl releases new crawls roughly monthly, each containing 3 to 5 billion web pages. The archive now holds over 250 billion pages accumulated over more than 15 years. The raw data is stored in WARC (Web ARChive) format, a standard format for web archives that preserves both HTTP response headers and page content. Each monthly crawl is approximately 70 to 100 terabytes of WARC files, hosted on Amazon S3 where they are accessible at no cost for data transfer within AWS and at very low cost for external downloads.
The coverage of Common Crawl is deliberately broad rather than deep. It does not attempt to crawl every page on every site exhaustively. Instead, it prioritizes popular, linked-to pages across a wide variety of domains and languages. The result is good coverage of the most important content on the internet, with sparser coverage of niche or low-traffic sites.
This breadth-over-depth strategy has clear implications for language AI. Common Crawl data captures mainstream web writing well: news articles, Wikipedia, Stack Overflow, Reddit, GitHub, and major blogs are all well represented. Niche academic forums, regional news in underrepresented languages, and community wikis may appear rarely or not at all. Understanding these coverage gaps is important when evaluating what knowledge a model trained on Common Crawl data will and will not possess.
WARC (Web ARChive) is a file format for storing web crawl data. Each WARC record contains the full HTTP response for a single URL, including headers and body content, along with metadata like the crawl timestamp and target URL. WARC is an ISO standard (ISO 28500) commonly used by web archives and Common Crawl.
Common Crawl Data Products
Common Crawl publishes three main data products, each serving a different use case in the data processing pipeline.
WARC files are the primary artifact from which all other products are derived. Each WARC file is a compressed sequence of WARC records, where each record contains the full HTTP response for one URL: status line, response headers, and body content. WARC files also contain request records showing what the crawler sent, and metadata records with crawl-specific information. A WARC record for an HTML page includes the raw HTML bytes, while a record for an image includes the raw image bytes. WARC files are the authoritative source and can be reprocessed to produce alternative text extractions if the WET extraction quality is inadequate.
WAT files (Web Archive Transformation) contain JSON metadata extracted from WARC records. Each WAT record includes HTTP headers, detected content type, page title, outgoing link structure, and language detection results. WAT files are much smaller than WARC files, typically 5-10x smaller, making them practical for metadata analysis without downloading full page content. They are useful for analyzing link graph structure, tracking domain coverage statistics, and filtering which WARC segments are worth downloading for a specific application.
WET files (WARC Encapsulated Text) contain only the plain text extracted from HTML pages, with all markup stripped. Each WET record includes the URL, timestamp, and plain text content. WET files are the most commonly used starting point for text-based NLP applications because they require no HTML parsing, have already extracted text content, and are substantially smaller than WARC files. A single monthly crawl's WET files total around 5 to 10 terabytes, compared to 70+ terabytes for the WARC files.
For language model pretraining, WET files are often the entry point. However, WET extraction quality varies significantly. Some pages produce clean, readable text, while others produce garbled outputs from malformed HTML or aggressive content extraction. Navigation menus, cookie consent dialogs, advertisement text, and JavaScript-rendered content that the simple HTML parser cannot access all end up either missing from WET files or appearing as noise. Research groups building high-quality training corpora often prefer to start from WARC files and apply their own text extraction pipeline to achieve better quality and consistency.
Accessing Common Crawl Data
Common Crawl data lives on Amazon S3 in the commoncrawl bucket. You can list available crawls and their indices via the Common Crawl index API or by browsing the S3 bucket directly. Each monthly crawl is identified by a name like CC-MAIN-2024-10 that encodes the year and approximate week number.
The Common Crawl index is a key tool for selective access. Rather than downloading entire WARC files, which are enormous, you can query the index to find specific URLs and retrieve only the relevant byte ranges using HTTP range requests. The Columnar Index format stores the index in Parquet files, allowing efficient filtered scans using tools like Apache Spark or DuckDB. You can filter by URL prefix to find all pages from a given domain, by content type to find only HTML pages, or by timestamp to find pages crawled during a specific time window, and then download only the relevant WARC byte ranges.
Let us look at what a WET record looks like in practice:
# The warcio library provides Python bindings for reading WARC/WET files
# uv pip install warcio requests
# A single WET segment file path from a Common Crawl release
# We simulate the structure by working with a minimal example
# In practice: download from s3://commoncrawl/crawl-data/CC-MAIN-2024-10/...
sample_wet_record = {
"uri": "https://en.wikipedia.org/wiki/Natural_language_processing",
"content_type": "text/plain",
"content_length": 45231,
"timestamp": "2024-03-15T14:23:11Z",
"text_preview": (
"Natural language processing (NLP) is a subfield of computer science "
"and artificial intelligence that uses machine learning to enable "
"computers to understand and communicate with human language in a "
"valuable way. NLP enables computers to read, hear, and understand "
"human language..."
),
}WET Record Structure: URI: https://en.wikipedia.org/wiki/Natural_language_processing Content-Type: text/plain Length: 45,231 bytes Timestamp: 2024-03-15T14:23:11Z Text preview: Natural language processing (NLP) is a subfield of computer science and artificial intelligence that uses machine learning to enable computers to understand and communicate with human language in a va...
Each WET record stores the URI of the crawled page alongside the extracted plain text. The content length gives us a sense of document size. For a Wikipedia article on NLP, nearly 45 KB of text reflects the depth of coverage typical for encyclopedic content. In contrast, many spam pages or thin-content sites produce only a few hundred bytes of text, a signal used by quality filters.
The practical workflow for using Common Crawl data in a training pipeline typically follows several steps. First, you identify which monthly crawl releases to use, often selecting multiple crawls across different months to ensure temporal diversity. Next, you download WAT files to analyze the metadata and identify which WARC segments contain content relevant to your goals. Then you download the WET files for those segments and apply language identification, quality filtering, and deduplication. Finally, you post-process the filtered text into the format your training pipeline expects.
Crawl Strategies
Not all crawls are created equal. The strategy a crawler uses shapes the character of the resulting dataset in fundamental ways. The tradeoffs between breadth, depth, freshness, and topical focus determine what kinds of knowledge end up in the data.
Breadth-First vs. Focused Crawling
A breadth-first crawler makes no assumptions about what topics are valuable. It follows all links equally, expanding outward from seeds in all directions. This produces broad coverage of the web graph but shallow coverage of any individual site or topic. Common Crawl uses a broad, general-purpose strategy: it wants to cover as much of the web as possible across all languages and topics. The resulting dataset reflects the web's actual distribution of content, which means a lot of English-language commercial and social media content.
A focused crawler, by contrast, prioritizes pages that match a specific topic or domain. It uses classifiers or keyword heuristics to score each URL before adding it to the frontier, then visits high-scoring URLs first. A focused crawl on scientific literature would prioritize domains like arxiv.org, pubmed.ncbi.nlm.nih.gov, and academic publishers. A focused crawl on programming would prioritize Stack Overflow, GitHub, and documentation sites. A focused crawl on legal text would target court opinion databases, regulatory archives, and law review journals.
The advantage of focused crawling is efficiency: you download far fewer irrelevant pages to assemble a domain-specific corpus. If you want a high-quality programming dataset, starting from Common Crawl and filtering is one approach, but directly crawling programming-relevant domains is often more efficient and produces cleaner results. The disadvantage of focused crawling is that topic classifiers can miss relevant content that uses unexpected terminology or lives on unexpected domains. A blog post about machine learning hosted on a personal site that the crawler has never seen might score poorly on a classifier trained on known ML domains.
Hybrid strategies combine the two approaches. Some crawlers use broad crawling to discover new domains and focused policies to determine crawl priority within each domain. Common Crawl itself applies some form of quality and relevance scoring to prioritize crawl bandwidth toward higher-quality pages.
Crawl Depth and Link Distance
Link distance from seed pages is a strong proxy for content quality. Pages that are one click from a trusted seed, like Wikipedia or a major news site, tend to be more relevant and authoritative than pages buried six hops deep in the web graph. Common Crawl seeds from web directories and past crawl data, which means its coverage is richer for well-connected sites and sparser for isolated niche domains.
Crawl depth, the maximum number of links followed from a seed URL, directly controls how far into the long tail the crawler reaches. A shallow crawl at depth 2 or 3 captures mainstream content efficiently. A deep crawl can surface rare or specialized content that shallow crawls miss, but at the cost of downloading much more spam and low-quality content.
The relationship between crawl depth and content quality reflects how the web's link structure works. High-quality, authoritative content tends to receive many inbound links from other authoritative pages. This is exactly what PageRank, the algorithm underlying Google's original search ranking, measures: the probability that a random web surfer lands on a page by following links. Pages with high PageRank appear early in breadth-first traversals. Training data composed primarily of high-PageRank pages tends to be more coherent and authoritative than data from the long tail.
The practical implication is that very deep crawls, while theoretically yielding more content, often produce diminishing returns for training data quality. Each additional hop into the web graph brings in more spam, low-quality content, and link-farm pages that are only there to manipulate search rankings. Most high-quality training data pipelines cap crawl depth and apply quality filters to compensate, rather than trying to recover quality from a very deep crawl.
Crawl Freshness
The web changes constantly. Pages are updated, deleted, and created at rates that vary enormously by site type. News sites publish dozens of articles per day. Academic papers are rarely updated after publication. Social media posts are ephemeral. A static snapshot of the web taken today is already partially stale tomorrow.
Freshness policies govern how often a crawler revisits pages it has previously seen. There are several common approaches. The simplest is a fixed-interval policy: revisit every page every days regardless of content. A more sophisticated approach is change-rate estimation: track how often a page's content changes and revisit it proportionally more often if it changes frequently. A site that publishes multiple articles per day warrants daily revisits, while a site whose content changes once a year warrants annual revisits.
Crawl freshness refers to how up-to-date the content in a crawl is relative to the current state of the web. A fresh crawl contains recent versions of pages, while a stale crawl may contain outdated content. Freshness is especially important for news, financial data, and other time-sensitive information.
For language model training data, freshness has subtle implications. Models trained on more recent data know about events and terminology that did not exist when older data was collected. However, very recent data may include lower-quality content, as the web accumulates more ephemeral, low-effort content over time. Common Crawl's monthly release cadence provides a reasonable freshness tradeoff for training data pipelines that can select crawls by date.
The freshness-versus-bandwidth tradeoff is unavoidable. Revisiting all pages frequently requires enormous crawler bandwidth. In practice, production crawlers use adaptive scheduling that prioritizes fresh revisits to high-value, frequently-changing pages and less frequent revisits to stable content. The optimal revisit interval for a page is proportional to the page's change rate. Pages that change daily should be revisited daily; pages that change annually need only annual revisits. Some crawlers model this explicitly using exponential smoothing of observed change rates, while others use simpler heuristics based on content type.
Implementing a Focused Crawler
Let us build a functional focused crawler that demonstrates the core concepts. Our crawler will target a specific topic, respect robots.txt, implement rate limiting, and extract structured text from crawled pages. This is deliberately simplified for clarity; production crawlers add distributed execution, persistent frontiers, and much more sophisticated scheduling.
Setting Up the Environment
We need a few Python libraries for HTTP fetching, HTML parsing, and robots.txt handling:
# uv pip install requests beautifulsoup4 lxmlCore Data Structures
A clean crawler separates concerns between the frontier (URL management), the fetcher (HTTP I/O), and the extractor (content parsing). Let us define the key data structures first:
from dataclasses import dataclass, field
@dataclass
class CrawledPage:
url: str
title: str
text: str
links: list[str]
status_code: int
content_length: int
crawl_timestamp: float
@dataclass
class CrawlConfig:
seed_urls: list[str]
max_pages: int = 100
max_depth: int = 3
crawl_delay: float = 1.0 # Seconds between requests to same host
request_timeout: int = 10 # Seconds before HTTP timeout
user_agent: str = "LanguageAIHandbookCrawler/1.0 (+https://example.com)"
respect_robots: bool = True
keywords: list[str] = field(default_factory=list) # For focused crawlingThe CrawlConfig bundles all tunable parameters. The user_agent string identifies our crawler so site operators can contact us. The keywords list enables focused crawling: we will score pages and URLs by relevance to these keywords and prioritize higher-scoring URLs. By separating configuration from crawler state, we can reuse the same crawler logic with different configurations for different crawl targets.
Robots.txt Cache
Checking robots.txt for every URL would be expensive. We cache the parsed robots.txt for each host, fetching it once per domain and reusing the parsed rules for all subsequent URLs on that domain:
import urllib.parse
import urllib.robotparser
from typing import Optional
class RobotsCache:
"""Cache robots.txt rules for each domain to avoid repeated fetches."""
def __init__(self, user_agent: str, timeout: int = 5):
self._cache: dict[str, urllib.robotparser.RobotFileParser] = {}
self.user_agent = user_agent
self.timeout = timeout
def can_fetch(self, url: str) -> bool:
"""Return True if robots.txt permits fetching this URL."""
parsed = urllib.parse.urlparse(url)
base = f"{parsed.scheme}://{parsed.netloc}"
if base not in self._cache:
rp = urllib.robotparser.RobotFileParser()
rp.set_url(f"{base}/robots.txt")
try:
rp.read()
except Exception:
# If we cannot fetch robots.txt, assume access is allowed
rp.allow_all = True
self._cache[base] = rp
return self._cache[base].can_fetch(self.user_agent, url)
def get_crawl_delay(self, url: str) -> Optional[float]:
"""Return the crawl delay specified in robots.txt, if any."""
parsed = urllib.parse.urlparse(url)
base = f"{parsed.scheme}://{parsed.netloc}"
if base in self._cache:
return self._cache[base].crawl_delay(self.user_agent)
return NoneThe urllib.robotparser.RobotFileParser class handles parsing robots.txt syntax, including User-agent, Disallow, Allow, and Crawl-delay directives. We extend it with a per-host cache so each domain's rules are fetched only once. The fallback when robots.txt is unreachable, to assume access is allowed, follows the convention adopted by most major crawlers: a missing robots.txt is not equivalent to a disallow-all directive.
The Focused Crawler
Now let us build the main crawler class. The scoring function is the heart of focused crawling: it assigns higher priority to URLs whose path and surrounding anchor text suggest relevance to our target keywords:
import time
import urllib.parse
from collections import defaultdict
import requests
from bs4 import BeautifulSoup
class FocusedCrawler:
"""A focused web crawler that prioritizes pages matching specified keywords."""
def __init__(self, config: CrawlConfig):
self.config = config
self.robots = RobotsCache(config.user_agent)
self.visited: set[str] = set()
self.results: list[CrawledPage] = []
# Track last fetch time per host for rate limiting
self._last_fetch: dict[str, float] = defaultdict(float)
# Frontier: list of (score, depth, url) tuples
self.frontier: list[tuple[float, int, str]] = []
def _score_url(self, url: str, text_context: str = "") -> float:
"""Score a URL by keyword relevance. Higher = more relevant."""
if not self.config.keywords:
return 1.0
url_lower = url.lower()
context_lower = text_context.lower()
score = sum(
2.0 if kw in url_lower else (1.0 if kw in context_lower else 0.0)
for kw in self.config.keywords
)
return score
def _normalize_url(self, url: str, base_url: str) -> Optional[str]:
"""Normalize and validate a URL relative to its page."""
try:
abs_url = urllib.parse.urljoin(base_url, url)
parsed = urllib.parse.urlparse(abs_url)
# Only HTTP/HTTPS
if parsed.scheme not in ("http", "https"):
return None
# Remove fragment
normalized = parsed._replace(fragment="").geturl()
return normalized
except Exception:
return None
def _wait_for_host(self, url: str):
"""Enforce rate limiting per host."""
parsed = urllib.parse.urlparse(url)
host = parsed.netloc
elapsed = time.time() - self._last_fetch[host]
# Check for crawl-delay from robots.txt
robots_delay = self.robots.get_crawl_delay(url) or 0.0
delay = max(self.config.crawl_delay, robots_delay)
if elapsed < delay:
time.sleep(delay - elapsed)
self._last_fetch[host] = time.time()
def _fetch(self, url: str) -> Optional[requests.Response]:
"""Fetch a URL, returning the response or None on failure."""
self._wait_for_host(url)
try:
headers = {"User-Agent": self.config.user_agent}
response = requests.get(
url,
headers=headers,
timeout=self.config.request_timeout,
allow_redirects=True,
)
return response
except requests.RequestException:
return None
def _extract(self, url: str, response: requests.Response) -> CrawledPage:
"""Extract title, text, and links from an HTML response."""
soup = BeautifulSoup(response.content, "lxml")
# Extract title
title_tag = soup.find("title")
title = title_tag.get_text(strip=True) if title_tag else ""
# Remove script and style elements
for tag in soup(["script", "style", "nav", "footer", "header"]):
tag.decompose()
# Extract main text
text = soup.get_text(separator=" ", strip=True)
# Clean excessive whitespace
text = re.sub(r"\s+", " ", text).strip()
# Extract all links
links = []
for anchor in soup.find_all("a", href=True):
normalized = self._normalize_url(anchor["href"], url)
if normalized and normalized not in self.visited:
links.append(normalized)
return CrawledPage(
url=url,
title=title,
text=text,
links=links,
status_code=response.status_code,
content_length=len(response.content),
crawl_timestamp=time.time(),
)
def crawl(self) -> list[CrawledPage]:
"""Run the crawl and return all successfully fetched pages."""
# Seed the frontier
for url in self.config.seed_urls:
score = self._score_url(url)
self.frontier.append((score, 0, url))
while self.frontier and len(self.results) < self.config.max_pages:
# Sort by score descending to visit best URLs first
self.frontier.sort(key=lambda x: x[0], reverse=True)
score, depth, url = self.frontier.pop(0)
if url in self.visited:
continue
self.visited.add(url)
# Check robots.txt
if self.config.respect_robots and not self.robots.can_fetch(url):
continue
# Fetch the page
response = self._fetch(url)
if response is None or response.status_code != 200:
continue
# Extract content
content_type = response.headers.get("Content-Type", "")
if "text/html" not in content_type:
continue
page = self._extract(url, response)
self.results.append(page)
# Expand frontier if within depth limit
if depth < self.config.max_depth:
for link in page.links:
link_score = self._score_url(link, page.text)
self.frontier.append((link_score, depth + 1, link))
return self.resultsThe scoring function rewards URLs that contain relevant keywords in the URL path itself (score 2.0 per keyword match) or in the surrounding anchor text context (score 1.0 per match). This distinction matters because a URL like /python-debugging-guide strongly suggests relevant content, while an anchor link that says "click here" but happens to be on a Python page provides weaker evidence. By sorting the frontier by score before each fetch, the crawler naturally gravitates toward the highest-relevance content, spending its bandwidth budget where it is most likely to find useful pages.
Running a Simulated Crawl
For demonstration purposes, we simulate what a crawl would produce without hitting live servers. The simulation uses realistic distributions for content lengths, domain distributions, and page link structures:
import random
import string
import time
# Simulate crawl results to show statistics
random.seed(42)
def simulate_crawl_results(n_pages: int = 150) -> list[CrawledPage]:
"""Simulate realistic crawl output for analysis."""
domains = [
"nlp-papers.org",
"machinelearning.wiki",
"ai-research.io",
"deep-learning-blog.com",
"transformer-models.net",
"linguistics.edu",
"corpus-research.org",
"text-mining.io",
"spam-site-123.xyz",
"thin-content-456.click",
]
# Power-law weights: top domains get disproportionately more pages
# Reflects web's link-following concentration effect
domain_weights = [30, 22, 14, 9, 6, 4, 7, 3, 3, 2]
pages = []
for i in range(n_pages):
domain = random.choices(domains, weights=domain_weights)[0]
url = f"https://{domain}/page-{i:04d}"
# Realistic content length distribution (log-normal)
mean_log = 8.5 # log(~4900 chars)
std_log = 1.2
content_len = int(random.lognormvariate(mean_log, std_log))
content_len = max(50, min(content_len, 200_000))
# Generate text with varying quality
is_quality = domain not in [
"spam-site-123.xyz",
"thin-content-456.click",
]
word_count = content_len // 6
text = " ".join(
random.choice(string.ascii_lowercase * 3)
for _ in range(min(word_count, 500))
)
# Vary status codes (most succeed)
status = random.choices(
[200, 200, 200, 200, 200, 301, 404, 429, 503],
weights=[70, 5, 5, 5, 5, 4, 3, 2, 1],
)[0]
pages.append(
CrawledPage(
url=url,
title=f"Page {i} on {domain}",
text=text,
links=[
f"https://{domain}/page-{random.randint(0, n_pages)}"
for _ in range(random.randint(2, 15))
],
status_code=status if status == 200 else status,
content_length=content_len,
crawl_timestamp=time.time() - random.uniform(0, 3600),
)
)
# Only keep successful fetches (status 200)
return [p for p in pages if p.status_code == 200]
crawl_results = simulate_crawl_results(150)
# Compute statistics
total_pages = len(crawl_results)
total_bytes = sum(p.content_length for p in crawl_results)
avg_content_len = total_bytes / total_pages if total_pages > 0 else 0
domain_counts: dict[str, int] = defaultdict(int)
for page in crawl_results:
domain = urllib.parse.urlparse(page.url).netloc
domain_counts[domain] += 1
top_domains = sorted(domain_counts.items(), key=lambda x: x[1], reverse=True)[
:5
]Crawl Summary ======================================== Pages successfully fetched: 136 Total content downloaded: 1323.1 KB Average content length: 9,962 bytes Top domains by page count: nlp-papers.org 40 pages (29.4%) machinelearning.wiki 32 pages (23.5%) ai-research.io 21 pages (15.4%) deep-learning-blog.com 17 pages (12.5%) transformer-models.net 7 pages (5.1%)
The distribution of pages across domains illustrates a fundamental characteristic of web crawls: content is not uniformly distributed. Some domains contribute disproportionately many pages, especially if the crawler follows many internal links. This domain imbalance is one reason deduplication and domain-level filtering are critical steps in the data processing pipeline.
Analyzing Content Length Distribution
Content length is a simple but powerful proxy for content quality. Very short pages, under 200 bytes of text, are often spam, error pages, or thin content with no informational value. Very long pages occasionally contain high-quality reference material but can also be terms-of-service documents, cookie consent dialogs, or concatenated low-quality content. The relationship between length and quality is real but imperfect: a 200-word blog post may be more useful than a 5,000-word spun article that repeats the same idea ten different ways.
import numpy as np
content_lengths = np.array([p.content_length for p in crawl_results])
# Compute percentile statistics
p10 = np.percentile(content_lengths, 10)
p25 = np.percentile(content_lengths, 25)
p50 = np.percentile(content_lengths, 50)
p75 = np.percentile(content_lengths, 75)
p90 = np.percentile(content_lengths, 90)
# Filter thresholds used in practice
min_threshold = 200 # bytes; below this is almost certainly junk
max_threshold = 100_000 # bytes; above this is often boilerplate-heavy
filtered_pages = [
p
for p in crawl_results
if min_threshold <= p.content_length <= max_threshold
]
filter_rate = 1 - len(filtered_pages) / len(crawl_results)Content Length Distribution ======================================== 10th percentile: 1,292 bytes 25th percentile: 2,288 bytes Median: 6,118 bytes 75th percentile: 12,633 bytes 90th percentile: 24,074 bytes After length filtering (200–100,000 bytes): Pages retained: 136 (100.0%) Pages filtered: 0 (0.0%)
The log-normal distribution of content lengths means most pages cluster around a moderate size, with long tails at both extremes. Quality filtering based on length typically removes 10 to 30 percent of crawled pages in practice, depending on how aggressively the crawler followed low-quality domains. This filtering step is fast and cheap, making it an excellent first pass before the more expensive operations like language detection and quality classification.
Key Parameters
The main parameters governing a web crawler's behavior are:
- crawl_delay: Minimum seconds between requests to the same host. Typical values range from 0.5 to 5 seconds. Always honor the
Crawl-delaydirective in robots.txt when it specifies a higher value. - max_depth: Maximum link distance from seed URLs to follow. Shallow crawls (depth 2-3) are efficient and produce higher average quality; deep crawls (depth 5+) surface long-tail content but bring in more noise.
- max_pages: A stopping condition based on page count. Combined with max_depth, it prevents runaway crawls.
- keywords (focused crawling): Relevance keywords used to score and prioritize URLs. More specific keywords produce higher-precision focused crawls at the cost of recall.
- user_agent: The crawler's identifier string. Always include contact information so site operators can reach you.
Visualizing Crawl Dynamics
The following visualizations illustrate how content length distribution and domain diversity evolve during a crawl, along with how relevance scoring concentrates focused crawl results.




Crawl Freshness in Depth
Understanding freshness matters operationally and epistemologically. A model does not simply know facts; it knows the snapshot of the web that existed during its training data collection. The date range of the crawl data shapes what the model considers current, recent, and historically established.
Freshness Estimation
Estimating how fresh a crawl is for a given URL requires comparing the crawl timestamp to the page's last-modification time. HTTP provides two mechanisms for this. The Last-Modified response header indicates when the server believes the content was last changed. The ETag header provides a content fingerprint that changes when the content changes. Crawlers that support conditional requests can send If-Modified-Since or If-None-Match headers with a subsequent request, and the server responds with a 304 Not Modified status if the content has not changed, saving download bandwidth.
In practice, many servers do not implement these headers correctly. Pages may have Last-Modified headers that reflect the page template's modification date rather than the actual content update time. ETags may be generated randomly on each request due to server misconfiguration. Crawlers that want accurate freshness information often fall back to content change detection: hash the extracted text from the previous crawl, compare against the current hash, and flag pages as changed only when the hashes differ. This approach is accurate but requires storing a hash for every URL ever crawled, which adds significant storage overhead at web scale.
Temporal Bias in Training Data
An important subtlety is that Common Crawl data is not a uniform temporal sample. Newer crawls contain more pages from recently created websites and fewer pages from sites that have since gone offline. The web has grown substantially over Common Crawl's history, so earlier crawls have sparser coverage than recent ones even for domains that have existed for many years.
This temporal non-uniformity can create training artifacts. A model trained on data from a narrow time window may underrepresent historical knowledge and overrepresent contemporary internet culture. A model trained only on 2022 crawls will have strong beliefs about what is "current" as of 2022 and little knowledge of events after that date. Mixing crawls from different time periods helps address this, and several recent training data pipelines explicitly stratify their Common Crawl samples by crawl date to ensure temporal diversity.
The model's knowledge cutoff, the date after which it has no information, arises directly from the crawl dates of its training data. When users notice a model claiming not to know about a recent event, the root cause is almost always in the crawl dates. Understanding the connection between crawl freshness and model knowledge cutoffs helps practitioners reason clearly about what their models can and cannot be expected to know.
HTTP Status Code Patterns in Crawls
Understanding how status codes distribute across a real crawl reveals important characteristics of the web that affect training data quality. Let us visualize what a typical status code distribution looks like:
# Build a simulated status code distribution based on empirical crawl data
# Distributions reflect typical web crawl observations from Common Crawl reports
status_categories = {
"200 OK": 0.72,
"301/302 Redirect": 0.11,
"404 Not Found": 0.09,
"403 Forbidden": 0.04,
"429 Rate Limited": 0.02,
"5xx Server Error": 0.015,
"Other": 0.005,
}
# Simulate counts for a crawl of 10000 URLs
total_attempted = 10_000
status_counts = {
k: int(v * total_attempted) for k, v in status_categories.items()
}
# Compute pages actually usable for training
usable = status_counts["200 OK"]
usable_pct = usable / total_attempted * 100
The status code breakdown exposes the real efficiency of a web crawl. Only pages with 200 OK status codes directly contribute training data, and even among those, quality filtering will remove a significant fraction. The 11% redirect rate represents URLs where the content has moved; a crawler that follows redirects correctly can recover some of this content, but chains of more than two or three redirects often lead to low-quality destinations or infinite redirect loops. The 9% 404 rate represents dead links in the web graph, which is unavoidable: the web changes constantly and old links break.
Limitations and Practical Challenges
Web crawling at scale poses both engineering and sociotechnical challenges. Several important limitations shape what crawled data can and cannot represent.
Coverage Bias
The web is not uniformly accessible. Content behind login walls, paywalls, or JavaScript-rendered single-page applications is largely invisible to simple crawlers. Academic journals, enterprise intranets, subscription newsletters, and private social media posts collectively represent an enormous body of knowledge that web crawling cannot access. Common Crawl's text is therefore biased toward publicly accessible content, which skews toward English-language, commercially oriented, and technically accessible writing.
Dynamic content generation poses a separate challenge. A growing fraction of web pages are rendered entirely client-side using JavaScript frameworks. A basic HTTP crawler receives only the empty HTML skeleton, missing all the content that JavaScript would populate. Crawling JavaScript-rendered pages requires a full browser automation layer using tools like Playwright or Puppeteer, which adds significant complexity and resource cost. Running a headless browser for every page is 10 to 50 times slower and more resource-intensive than simple HTTP fetching. For this reason, most large-scale crawls still use HTTP-only fetching and accept the coverage gap for JavaScript-heavy sites.
The language distribution of the web also creates bias. English dominates the web in terms of both page count and content density, but the actual population of potential language AI users is globally distributed. Common Crawl contains content in dozens of languages, but English pages vastly outnumber pages in any other language. Models trained on unfiltered Common Crawl data inherit this English-language dominance. Building multilingual training data requires deliberate effort: explicitly identifying and proportionally including content from underrepresented languages rather than relying on raw crawl proportions.
Domain Imbalance
Without domain capping, link-following crawlers naturally accumulate enormous numbers of pages from a few highly interconnected domains. A crawler that follows all links would eventually download billions of Reddit comments and Stack Overflow answers, which, while potentially useful, would swamp other content types. This concentration effect arises from the web's power-law link structure: the most popular domains receive many times more inbound links than less popular ones, so a link-following crawler visits them many times more often.
Production data pipelines typically apply domain caps, limiting any single domain to a fixed fraction of the total corpus, to ensure diversity. If Reddit alone contributes 20% of a training corpus, the model will be heavily biased toward casual conversational writing and the specific cultural norms of Reddit communities. A cap of 1% per domain across thousands of domains produces a much more balanced representation of writing styles and topics.
Legal and Ethical Constraints
Copyright law creates significant uncertainty around web crawling for training data purposes. While robots.txt provides a social contract for access permission, it does not confer copyright permission. Content creators increasingly argue that training language models on their writing without compensation constitutes copyright infringement. Several lawsuits in the United States and Europe have challenged major AI companies on exactly these grounds, with ongoing legal proceedings that have not yet produced clear settled law.
The tension remains unresolved. Open research into language AI depends on large-scale crawled data. But the people who created that data, journalists, authors, coders, educators, often received no compensation and gave no explicit consent. Responsible use of crawled data should at minimum respect robots.txt, provide opt-out mechanisms, and acknowledge the human labor that produced the training corpus. Some organizations have begun offering opt-out registries where content creators can request removal from training datasets, though the practical effectiveness of these mechanisms remains limited.
The situation is evolving rapidly. The European Union's AI Act and related regulations are establishing new requirements around training data transparency. Some jurisdictions are exploring data rights frameworks that would give content creators enforceable claims over uses of their work. Anyone building a production crawling pipeline should monitor legal developments in their target jurisdictions and consult legal counsel before deploying at scale.
Data Quality Across the Web
The web is not a curated library; it is the sum of everything humans have chosen to publish. This includes rigorous scholarship alongside spam, misinformation, hate speech, and low-effort content. A crawl that follows links aggressively will collect substantial quantities of all of these. Spam farms generate thousands of keyword-stuffed pages designed to manipulate search rankings rather than inform readers. Content mills produce formulaic articles that repeat the same information in slightly varied phrasing, contributing enormous volume with minimal unique information. Deliberately misleading content, conspiracy theories, and factually incorrect claims appear throughout the web at substantial scale.
The downstream filtering pipeline is as important as the crawl itself. A thoughtfully filtered medium-sized crawl often produces better training data than a massive unfiltered one. We discuss quality filtering, toxicity filtering, and deduplication in the following chapters in this section. For now, the key insight is that web crawling produces raw material that requires substantial refinement before it is suitable as training data. The crawl is the beginning of the pipeline, not the end.
Summary
Web crawling provides the raw material for the largest language model training corpora. The core loop is simple: fetch URLs, extract links, repeat. The complexity comes from doing this at scale while respecting servers, managing quality, and complying with legal constraints.
Key takeaways from this chapter:
- Common Crawl is the foundational public data source for language AI, giving petabyte-scale WARC, WAT, and WET data with monthly crawl releases. WET files provide pre-extracted plain text for NLP use cases, while WARC files preserve the complete raw crawl data for custom text extraction.
- Robots.txt defines the social contract between sites and crawlers. Respecting it is both an ethical obligation and practically essential for sustainable crawling. The protocol includes allow/disallow rules, crawl delay directives, and sitemap pointers.
- Politeness policies prevent a single crawler from overwhelming servers. Per-host rate limiting, crawl delay respect, adaptive backoff on server errors, and responsible user-agent identification are minimum requirements.
- Crawl strategy shapes data character. Breadth-first crawls produce broad coverage; focused crawls produce topically relevant but narrower datasets. Depth and freshness tradeoffs further determine what ends up in the corpus.
- Coverage bias is inherent: web crawls miss paywalled, dynamic, and login-required content, skewing the resulting data toward publicly accessible, often English-language material. JavaScript rendering, domain imbalance, and temporal non-uniformity create additional biases that practitioners must understand and manage.
- Data quality varies enormously across the web. Content length filtering is a fast first pass, but the full pipeline requires deduplication, language identification, quality filtering, and toxicity filtering, topics we turn to next.
The next chapter on Document Extraction explores how to reliably convert raw HTML from crawl data into clean text, a step that proves more difficult than it first appears given the diversity of web page structures and encodings.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about web crawling.
Web Crawling Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!