Volume 7 of 20 · PDF edition
In progressPre-training Data and Scaling
For teams designing defensible corpora, objectives, mixtures, and compute-optimal pre-training programs.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 24 written chapters
- Planned chapters
- 6 planned chapters
- Approximate pages
- ~760 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 7 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 7 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Data Curation
- Data Governance, Provenance, and Synthetic Ecosystems
- Pre-training Objectives
- Scaling Laws
Audience and prerequisites
Where this volume fits
For teams designing defensible corpora, objectives, mixtures, and compute-optimal pre-training programs.
Prerequisites: Assumes Volume 5 and working knowledge of optimization.
Free online preview
Start with “Web Crawling”
Covers Common Crawl, crawling strategies, robots.txt respect, crawl freshness.
Exact contents
24 chapters available now
The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.
Part XX: Data Curation
- 01Web Crawling
Covers Common Crawl, crawling strategies, robots.txt respect, crawl freshness.
- 02Document Extraction
Covers HTML parsing, boilerplate removal, content extraction, trafilatura and similar tools.
- 03Language Identification
Covers language ID models, multilingual document handling, code-switching, language filtering.
- 04Deduplication
Covers exact deduplication, near-duplicate detection, document vs substring dedup, dedup at scale.
- 05MinHash
Covers MinHash algorithm, Jaccard similarity estimation, MinHash LSH, MinHash implementation.
- 06Quality Filtering
Covers heuristic filters, perplexity filtering, classifier-based filtering, filter thresholds.
- 07Toxicity Filtering
Covers toxicity classifiers, toxicity thresholds, over-filtering risks, toxicity filter evaluation.
- 08PII Removal
Covers PII detection methods, PII removal strategies, PII removal evaluation, privacy preservation.
- 09Data Mixing
Covers domain proportions, quality weighting, data mixing experiments, optimal mixing.
- 10Synthetic Data
Covers synthetic data generation, quality verification, synthetic data diversity, distillation.
Part XXI: Data Governance, Provenance, and Synthetic Ecosystems
- 01Dataset Lineage and ProvenancePlanned
Tracing documents through collection, filtering, mixing, training, evaluation, and model release.
- 02Licensing, Consent, and Data-Use SignalsPlanned
Licenses, terms, robots directives, creator consent, jurisdictional uncertainty, and defensible data decisions.
- 03Attribution and Training-Data TransparencyPlanned
Data statements, source disclosure, influence estimation, attribution limits, and communicating uncertainty.
- 04Synthetic Data PipelinesPlanned
Generation, filtering, diversity control, verification, curriculum design, and provenance for synthetic corpora.
- 05Recursive Training and Model CollapsePlanned
Feedback loops from model-generated data, tail loss, distribution drift, and mitigations for synthetic ecosystems.
- 06Data Audits and Release GovernancePlanned
Pre-training audits, sensitive-content review, documentation, approval gates, and post-release traceability.
Part XXII: Pre-training Objectives
- 01Causal Language Modeling
Covers CLM objective formulation, autoregressive factorization, CLM loss computation, CLM for generation, CLM training data, CLM scaling properties.
- 02Masked Language Modeling
Covers MLM objective formulation, masking strategies (15% rule), [MASK] token usage, MLM for understanding, MLM training dynamics.
- 03Whole Word Masking
Covers subword masking problems, whole word masking procedure, WWM implementation, WWM vs random masking, WWM for different tokenizers.
- 04Span Corruption
Covers span selection strategies, span length distribution, sentinel tokens, T5-style corruption, span corruption benefits.
- 05Prefix Language Modeling
Covers prefix LM formulation, prefix LM attention pattern, prefix LM for generation, prefix LM training, UniLM-style objectives.
- 06Replaced Token Detection
Covers generator-discriminator setup, replaced vs original detection, RTD efficiency advantages, ELECTRA training procedure, RTD vs MLM comparison.
- 07Denoising Objectives
Covers token deletion, token shuffling, sentence permutation, document rotation, BART-style denoising, combining denoising tasks.
Part XXIII: Scaling Laws
- 01Power Laws in Deep Learning
Covers power law definition, log-log linear relationships, power law fitting, power law universality, power law intuition.
- 02Kaplan Scaling Laws
Covers loss vs parameters, loss vs data, loss vs compute, Kaplan optimal allocation, Kaplan predictions.
- 03Chinchilla Scaling Laws
Covers Chinchilla experiments, revised scaling coefficients, optimal tokens per parameter, Chinchilla vs Kaplan, Chinchilla implications.
- 04Compute-Optimal Training
Covers compute budget allocation, tokens vs parameters ratio, training efficiency, compute-optimal recipes, practical guidelines.
- 05Data-Constrained Scaling
Covers data repetition effects, optimal repetition strategies, data augmentation scaling, synthetic data scaling.
- 06Inference Scaling
Covers training vs inference compute, inference-optimal models, over-training for efficiency, deployment cost modeling.
- 07Predicting Model Performance
Covers loss extrapolation, capability prediction, scaling law uncertainty, prediction reliability, practical forecasting.
Volume 7 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 7 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.