Volume 10 of 20 · PDF edition

Available now

Large-Scale Training and Sparse Models

For engineers training large models across accelerators and researchers studying sparse conditional computation.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
29 written chapters
Approximate pages
~920 pages
Price
$24 one-time
Language AI Handbook, Volume 10: Large-Scale Training and Sparse Models cover

Author and edition details

About the author and Volume 10 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 10 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Training Infrastructure
  • Training Optimization
  • Mixture of Experts

Audience and prerequisites

Where this volume fits

For engineers training large models across accelerators and researchers studying sparse conditional computation.

Prerequisites: Assumes Volumes 5, 7, and 9.

Free online preview

Start with “GPU Architecture

Covers GPU memory hierarchy, CUDA cores, tensor cores, GPU specifications.

Read the chapter

Exact contents

29 chapters available now

Part XXXI: Training Infrastructure

  1. 01
    GPU Architecture

    Covers GPU memory hierarchy, CUDA cores, tensor cores, GPU specifications.

  2. 02
    Memory Management

    Covers memory breakdown (activations, parameters, gradients, optimizer states), memory estimation, OOM debugging.

  3. 03
    Data Parallelism

    Covers DDP algorithm, gradient synchronization, all-reduce operations, DDP scaling.

  4. 04
    Tensor Parallelism

    Covers column parallelism, row parallelism, communication patterns, Megatron-style parallelism.

  5. 05
    Pipeline Parallelism

    Covers pipeline stages, micro-batching, pipeline bubbles, pipeline schedules (GPipe, 1F1B).

  6. 06
    ZeRO Optimization

    Covers ZeRO stage 1 (optimizer state partitioning), ZeRO stage 2 (gradient partitioning), ZeRO stage 3 (parameter partitioning), ZeRO memory savings.

  7. 07
    FSDP

    Covers FSDP concepts, FSDP vs ZeRO, FSDP sharding strategies, FSDP usage.

  8. 08
    Activation Checkpointing

    Covers checkpointing concept, checkpoint selection, checkpointing overhead, selective checkpointing.

  9. 09
    Mixed Precision Training

    Covers floating point formats, loss scaling, BF16 advantages, mixed precision implementation.

  10. 10
    Communication Optimization

    Covers gradient compression, communication overlap, topology-aware communication, NCCL optimization.

  11. 11
    Checkpointing and Recovery

    Covers checkpoint contents, checkpoint frequency, async checkpointing, fault recovery.

Part XXXII: Training Optimization

  1. 01
    Learning Rate Warmup

    Covers warmup motivation, linear warmup, warmup duration, warmup for large batches.

  2. 02
    Learning Rate Decay

    Covers step decay, exponential decay, inverse square root decay, decay scheduling.

  3. 03
    Cosine Learning Rate Schedule

    Covers cosine decay formula, cosine with restarts, cosine schedule parameters, cosine vs linear.

  4. 04
    Large Batch Training

    Covers batch size effects, learning rate scaling, batch size limits, LAMB optimizer.

  5. 05
    Weight Decay

    Covers weight decay formula, decoupled weight decay, weight decay selection, weight decay interaction with Adam.

  6. 06
    Gradient Accumulation

    Covers accumulation procedure, accumulation steps, accumulation for memory, accumulation correctness.

  7. 07
    Training Stability

    Covers loss spikes, gradient norm monitoring, stability techniques, training stability debugging.

  8. 08
    Hyperparameter Selection

    Covers hyperparameter search, hyperparameter transfer, critical vs robust hyperparameters, default recipes.

Part XXXIII: Mixture of Experts

  1. 01
    Sparse Models

    Covers dense vs sparse trade-offs, conditional computation motivation, sparse model efficiency, sparse model challenges.

  2. 02
    Expert Networks

    Covers expert architecture, expert as FFN, expert capacity, expert count selection, expert placement in transformer.

  3. 03
    Gating Networks

    Covers router architecture, routing score computation, router training, router learned behavior.

  4. 04
    Top-K Routing

    Covers top-1 routing, top-2 routing, k selection trade-offs, routing implementation, combining expert outputs.

  5. 05
    Load Balancing

    Covers expert utilization imbalance, collapse failure mode, load metrics, balanced routing importance.

  6. 06
    Auxiliary Balancing Loss

    Covers load balancing loss formulation, loss coefficient tuning, balancing vs task loss, auxiliary loss implementation.

  7. 07
    Router Z-Loss

    Covers router instability, z-loss formulation, z-loss benefits, z-loss coefficient, combined auxiliary losses.

  8. 08
    Expert Parallelism

    Covers expert placement strategies, all-to-all communication, communication overhead, expert parallelism implementation.

  9. 09
    Switch Transformer

    Covers Switch Transformer design, top-1 routing choice, capacity factor, Switch scaling results.

  10. 10
    Mixtral

    Covers Mixtral architecture, Mixtral expert design, Mixtral performance, Mixtral efficiency, Mixtral vs dense models.

Volume 10 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 10 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.