Volume 7 of 20 · PDF edition

In progress

Pre-training Data and Scaling

For teams designing defensible corpora, objectives, mixtures, and compute-optimal pre-training programs.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
24 written chapters
Planned chapters
6 planned chapters
Approximate pages
~760 pages
Price
$24 one-time
Language AI Handbook, Volume 7: Pre-training Data and Scaling cover

Author and edition details

About the author and Volume 7 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 7 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Data Curation
  • Data Governance, Provenance, and Synthetic Ecosystems
  • Pre-training Objectives
  • Scaling Laws

Audience and prerequisites

Where this volume fits

For teams designing defensible corpora, objectives, mixtures, and compute-optimal pre-training programs.

Prerequisites: Assumes Volume 5 and working knowledge of optimization.

Free online preview

Start with “Web Crawling

Covers Common Crawl, crawling strategies, robots.txt respect, crawl freshness.

Read the chapter

Exact contents

24 chapters available now

The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.

Part XX: Data Curation

  1. 01
    Web Crawling

    Covers Common Crawl, crawling strategies, robots.txt respect, crawl freshness.

  2. 02
    Document Extraction

    Covers HTML parsing, boilerplate removal, content extraction, trafilatura and similar tools.

  3. 03
    Language Identification

    Covers language ID models, multilingual document handling, code-switching, language filtering.

  4. 04
    Deduplication

    Covers exact deduplication, near-duplicate detection, document vs substring dedup, dedup at scale.

  5. 05
    MinHash

    Covers MinHash algorithm, Jaccard similarity estimation, MinHash LSH, MinHash implementation.

  6. 06
    Quality Filtering

    Covers heuristic filters, perplexity filtering, classifier-based filtering, filter thresholds.

  7. 07
    Toxicity Filtering

    Covers toxicity classifiers, toxicity thresholds, over-filtering risks, toxicity filter evaluation.

  8. 08
    PII Removal

    Covers PII detection methods, PII removal strategies, PII removal evaluation, privacy preservation.

  9. 09
    Data Mixing

    Covers domain proportions, quality weighting, data mixing experiments, optimal mixing.

  10. 10
    Synthetic Data

    Covers synthetic data generation, quality verification, synthetic data diversity, distillation.

Part XXI: Data Governance, Provenance, and Synthetic Ecosystems

  1. 01
    Dataset Lineage and ProvenancePlanned

    Tracing documents through collection, filtering, mixing, training, evaluation, and model release.

  2. 02
    Licensing, Consent, and Data-Use SignalsPlanned

    Licenses, terms, robots directives, creator consent, jurisdictional uncertainty, and defensible data decisions.

  3. 03
    Attribution and Training-Data TransparencyPlanned

    Data statements, source disclosure, influence estimation, attribution limits, and communicating uncertainty.

  4. 04
    Synthetic Data PipelinesPlanned

    Generation, filtering, diversity control, verification, curriculum design, and provenance for synthetic corpora.

  5. 05
    Recursive Training and Model CollapsePlanned

    Feedback loops from model-generated data, tail loss, distribution drift, and mitigations for synthetic ecosystems.

  6. 06
    Data Audits and Release GovernancePlanned

    Pre-training audits, sensitive-content review, documentation, approval gates, and post-release traceability.

Part XXII: Pre-training Objectives

  1. 01
    Causal Language Modeling

    Covers CLM objective formulation, autoregressive factorization, CLM loss computation, CLM for generation, CLM training data, CLM scaling properties.

  2. 02
    Masked Language Modeling

    Covers MLM objective formulation, masking strategies (15% rule), [MASK] token usage, MLM for understanding, MLM training dynamics.

  3. 03
    Whole Word Masking

    Covers subword masking problems, whole word masking procedure, WWM implementation, WWM vs random masking, WWM for different tokenizers.

  4. 04
    Span Corruption

    Covers span selection strategies, span length distribution, sentinel tokens, T5-style corruption, span corruption benefits.

  5. 05
    Prefix Language Modeling

    Covers prefix LM formulation, prefix LM attention pattern, prefix LM for generation, prefix LM training, UniLM-style objectives.

  6. 06
    Replaced Token Detection

    Covers generator-discriminator setup, replaced vs original detection, RTD efficiency advantages, ELECTRA training procedure, RTD vs MLM comparison.

  7. 07
    Denoising Objectives

    Covers token deletion, token shuffling, sentence permutation, document rotation, BART-style denoising, combining denoising tasks.

Part XXIII: Scaling Laws

  1. 01
    Power Laws in Deep Learning

    Covers power law definition, log-log linear relationships, power law fitting, power law universality, power law intuition.

  2. 02
    Kaplan Scaling Laws

    Covers loss vs parameters, loss vs data, loss vs compute, Kaplan optimal allocation, Kaplan predictions.

  3. 03
    Chinchilla Scaling Laws

    Covers Chinchilla experiments, revised scaling coefficients, optimal tokens per parameter, Chinchilla vs Kaplan, Chinchilla implications.

  4. 04
    Compute-Optimal Training

    Covers compute budget allocation, tokens vs parameters ratio, training efficiency, compute-optimal recipes, practical guidelines.

  5. 05
    Data-Constrained Scaling

    Covers data repetition effects, optimal repetition strategies, data augmentation scaling, synthetic data scaling.

  6. 06
    Inference Scaling

    Covers training vs inference compute, inference-optimal models, over-training for efficiency, deployment cost modeling.

  7. 07
    Predicting Model Performance

    Covers loss extrapolation, capability prediction, scaling law uncertainty, prediction reliability, practical forecasting.

Volume 7 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 7 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.