Volume 12 of 20 · PDF edition

In progress

Alignment, Preferences, and Oversight

For teams training assistants from preference data and evaluating whether alignment survives real deployment pressure.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
16 written chapters
Planned chapters
5 planned chapters
Approximate pages
~510 pages
Price
$24 one-time
Language AI Handbook, Volume 12: Alignment, Preferences, and Oversight cover

Author and edition details

About the author and Volume 12 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 12 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Alignment and RLHF
  • Scalable Oversight and Safety Training

Audience and prerequisites

Where this volume fits

For teams training assistants from preference data and evaluating whether alignment survives real deployment pressure.

Prerequisites: Assumes Volume 11 and basic reinforcement learning concepts.

Free online preview

Start with “Alignment Problem

Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.

Read the chapter

Exact contents

16 chapters available now

The current PDF contains every linked chapter below. The remaining 5 planned chapters will be added through free volume updates.

Part XXXVII: Alignment and RLHF

  1. 01
    Alignment Problem

    Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.

  2. 02
    Human Preference Data

    Covers preference collection UI, comparison design, annotator guidelines, preference data quality.

  3. 03
    Bradley-Terry Model

    Covers pairwise comparison model, preference probability, Bradley-Terry likelihood, preference strength.

  4. 04
    Reward Modeling

    Covers reward model architecture, preference loss function, reward model training, reward model evaluation.

  5. 05
    Reward Hacking

    Covers reward hacking examples, distribution shift, over-optimization, reward hacking mitigation.

  6. 06
    Policy Gradient Methods

    Covers policy definition, REINFORCE algorithm, policy gradient derivation, variance reduction.

  7. 07
    PPO Algorithm

    Covers clipped objective, PPO derivation, trust region intuition, PPO implementation.

  8. 08
    PPO for Language Models

    Covers LLM as policy, action space (tokens), reward assignment, KL penalty importance.

  9. 09
    RLHF Pipeline

    Covers SFT stage, reward model training, PPO fine-tuning, RLHF hyperparameters, RLHF debugging.

  10. 10
    KL Divergence Penalty

    Covers KL penalty motivation, KL coefficient selection, adaptive KL, KL effects on training.

  11. 11
    DPO Concept

    Covers DPO motivation, removing reward model, DPO intuition, DPO benefits.

  12. 12
    DPO Derivation

    Covers DPO from RLHF objective, optimal policy derivation, DPO loss function, DPO as classification.

  13. 13
    DPO Implementation

    Covers DPO data format, DPO loss computation, DPO training procedure, DPO hyperparameters.

  14. 14
    DPO Variants

    Covers IPO formulation, KTO for unpaired feedback, ORPO, cDPO, comparing alignment methods.

  15. 15
    RLAIF

    Covers AI as annotator, constitutional AI principles, AI preference generation, RLAIF scalability.

  16. 16
    Iterative Alignment

    Covers iterative DPO, online preference learning, self-improvement loops, alignment stability.

Part XXXVIII: Scalable Oversight and Safety Training

  1. 01
    Constitutional and Principle-Based AlignmentPlanned

    Written principles, critique and revision, AI feedback, evaluation, and failure modes.

  2. 02
    Scalable OversightPlanned

    Decomposition, debate, recursive supervision, weak-to-strong generalization, and oversight bottlenecks.

  3. 03
    Process Supervision and Behavioral SpecificationsPlanned

    Supervising intermediate behavior, specifying policies, resolving conflicts, and testing adherence.

  4. 04
    Alignment Data Quality and Rater PopulationsPlanned

    Rater disagreement, cultural pluralism, annotator effects, preference aggregation, and auditability.

  5. 05
    Alignment Robustness and Distribution ShiftPlanned

    Sycophancy, reward tampering, jailbreak pressure, deployment drift, and adversarial evaluation.

Volume 12 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 12 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.