Volume 12 of 20 · PDF edition
In progressAlignment, Preferences, and Oversight
For teams training assistants from preference data and evaluating whether alignment survives real deployment pressure.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 16 written chapters
- Planned chapters
- 5 planned chapters
- Approximate pages
- ~510 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 12 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 12 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Alignment and RLHF
- Scalable Oversight and Safety Training
Audience and prerequisites
Where this volume fits
For teams training assistants from preference data and evaluating whether alignment survives real deployment pressure.
Prerequisites: Assumes Volume 11 and basic reinforcement learning concepts.
Free online preview
Start with “Alignment Problem”
Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.
Exact contents
16 chapters available now
The current PDF contains every linked chapter below. The remaining 5 planned chapters will be added through free volume updates.
Part XXXVII: Alignment and RLHF
- 01Alignment Problem
Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.
- 02Human Preference Data
Covers preference collection UI, comparison design, annotator guidelines, preference data quality.
- 03Bradley-Terry Model
Covers pairwise comparison model, preference probability, Bradley-Terry likelihood, preference strength.
- 04Reward Modeling
Covers reward model architecture, preference loss function, reward model training, reward model evaluation.
- 05Reward Hacking
Covers reward hacking examples, distribution shift, over-optimization, reward hacking mitigation.
- 06Policy Gradient Methods
Covers policy definition, REINFORCE algorithm, policy gradient derivation, variance reduction.
- 07PPO Algorithm
Covers clipped objective, PPO derivation, trust region intuition, PPO implementation.
- 08PPO for Language Models
Covers LLM as policy, action space (tokens), reward assignment, KL penalty importance.
- 09RLHF Pipeline
Covers SFT stage, reward model training, PPO fine-tuning, RLHF hyperparameters, RLHF debugging.
- 10KL Divergence Penalty
Covers KL penalty motivation, KL coefficient selection, adaptive KL, KL effects on training.
- 11DPO Concept
Covers DPO motivation, removing reward model, DPO intuition, DPO benefits.
- 12DPO Derivation
Covers DPO from RLHF objective, optimal policy derivation, DPO loss function, DPO as classification.
- 13DPO Implementation
Covers DPO data format, DPO loss computation, DPO training procedure, DPO hyperparameters.
- 14DPO Variants
Covers IPO formulation, KTO for unpaired feedback, ORPO, cDPO, comparing alignment methods.
- 15RLAIF
Covers AI as annotator, constitutional AI principles, AI preference generation, RLAIF scalability.
- 16Iterative Alignment
Covers iterative DPO, online preference learning, self-improvement loops, alignment stability.
Part XXXVIII: Scalable Oversight and Safety Training
- 01Constitutional and Principle-Based AlignmentPlanned
Written principles, critique and revision, AI feedback, evaluation, and failure modes.
- 02Scalable OversightPlanned
Decomposition, debate, recursive supervision, weak-to-strong generalization, and oversight bottlenecks.
- 03Process Supervision and Behavioral SpecificationsPlanned
Supervising intermediate behavior, specifying policies, resolving conflicts, and testing adherence.
- 04Alignment Data Quality and Rater PopulationsPlanned
Rater disagreement, cultural pluralism, annotator effects, preference aggregation, and auditability.
- 05Alignment Robustness and Distribution ShiftPlanned
Sycophancy, reward tampering, jailbreak pressure, deployment drift, and adversarial evaluation.
Volume 12 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 12 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.