Volume 10 of 20 · PDF edition
Available nowLarge-Scale Training and Sparse Models
For engineers training large models across accelerators and researchers studying sparse conditional computation.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 29 written chapters
- Approximate pages
- ~920 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 10 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 10 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Training Infrastructure
- Training Optimization
- Mixture of Experts
Audience and prerequisites
Where this volume fits
For engineers training large models across accelerators and researchers studying sparse conditional computation.
Prerequisites: Assumes Volumes 5, 7, and 9.
Free online preview
Start with “GPU Architecture”
Covers GPU memory hierarchy, CUDA cores, tensor cores, GPU specifications.
Exact contents
29 chapters available now
Part XXXI: Training Infrastructure
- 01GPU Architecture
Covers GPU memory hierarchy, CUDA cores, tensor cores, GPU specifications.
- 02Memory Management
Covers memory breakdown (activations, parameters, gradients, optimizer states), memory estimation, OOM debugging.
- 03Data Parallelism
Covers DDP algorithm, gradient synchronization, all-reduce operations, DDP scaling.
- 04Tensor Parallelism
Covers column parallelism, row parallelism, communication patterns, Megatron-style parallelism.
- 05Pipeline Parallelism
Covers pipeline stages, micro-batching, pipeline bubbles, pipeline schedules (GPipe, 1F1B).
- 06ZeRO Optimization
Covers ZeRO stage 1 (optimizer state partitioning), ZeRO stage 2 (gradient partitioning), ZeRO stage 3 (parameter partitioning), ZeRO memory savings.
- 07FSDP
Covers FSDP concepts, FSDP vs ZeRO, FSDP sharding strategies, FSDP usage.
- 08Activation Checkpointing
Covers checkpointing concept, checkpoint selection, checkpointing overhead, selective checkpointing.
- 09Mixed Precision Training
Covers floating point formats, loss scaling, BF16 advantages, mixed precision implementation.
- 10Communication Optimization
Covers gradient compression, communication overlap, topology-aware communication, NCCL optimization.
- 11Checkpointing and Recovery
Covers checkpoint contents, checkpoint frequency, async checkpointing, fault recovery.
Part XXXII: Training Optimization
- 01Learning Rate Warmup
Covers warmup motivation, linear warmup, warmup duration, warmup for large batches.
- 02Learning Rate Decay
Covers step decay, exponential decay, inverse square root decay, decay scheduling.
- 03Cosine Learning Rate Schedule
Covers cosine decay formula, cosine with restarts, cosine schedule parameters, cosine vs linear.
- 04Large Batch Training
Covers batch size effects, learning rate scaling, batch size limits, LAMB optimizer.
- 05Weight Decay
Covers weight decay formula, decoupled weight decay, weight decay selection, weight decay interaction with Adam.
- 06Gradient Accumulation
Covers accumulation procedure, accumulation steps, accumulation for memory, accumulation correctness.
- 07Training Stability
Covers loss spikes, gradient norm monitoring, stability techniques, training stability debugging.
- 08Hyperparameter Selection
Covers hyperparameter search, hyperparameter transfer, critical vs robust hyperparameters, default recipes.
Part XXXIII: Mixture of Experts
- 01Sparse Models
Covers dense vs sparse trade-offs, conditional computation motivation, sparse model efficiency, sparse model challenges.
- 02Expert Networks
Covers expert architecture, expert as FFN, expert capacity, expert count selection, expert placement in transformer.
- 03Gating Networks
Covers router architecture, routing score computation, router training, router learned behavior.
- 04Top-K Routing
Covers top-1 routing, top-2 routing, k selection trade-offs, routing implementation, combining expert outputs.
- 05Load Balancing
Covers expert utilization imbalance, collapse failure mode, load metrics, balanced routing importance.
- 06Auxiliary Balancing Loss
Covers load balancing loss formulation, loss coefficient tuning, balancing vs task loss, auxiliary loss implementation.
- 07Router Z-Loss
Covers router instability, z-loss formulation, z-loss benefits, z-loss coefficient, combined auxiliary losses.
- 08Expert Parallelism
Covers expert placement strategies, all-to-all communication, communication overhead, expert parallelism implementation.
- 09Switch Transformer
Covers Switch Transformer design, top-1 routing choice, capacity factor, Switch scaling results.
- 10Mixtral
Covers Mixtral architecture, Mixtral expert design, Mixtral performance, Mixtral efficiency, Mixtral vs dense models.
Volume 10 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 10 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.