Volume 19 of 20 · PDF edition
In progressResponsible, Secure, and Interpretable Language AI
For teams responsible for fairness, model understanding, safety, agent security, audits, and governance.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 24 written chapters
- Planned chapters
- 13 planned chapters
- Approximate pages
- ~760 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 19 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 19 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Bias and Fairness
- Interpretability
- Safety and Security
- Agentic AI and Protocol Security
- Human-AI Interaction and Governance
Audience and prerequisites
Where this volume fits
For teams responsible for fairness, model understanding, safety, agent security, audits, and governance.
Prerequisites: Assumes model fundamentals and benefits from Volume 18's evaluation framework.
Free online preview
Start with “Bias in Language Models”
Covers bias sources, bias types (demographic, cultural), bias in training data, bias amplification.
Exact contents
24 chapters available now
The current PDF contains every linked chapter below. The remaining 13 planned chapters will be added through free volume updates.
Part LVIII: Bias and Fairness
- 01Bias in Language Models
Covers bias sources, bias types (demographic, cultural), bias in training data, bias amplification.
- 02Bias Measurement
Covers embedding association tests, generation bias metrics, classification bias metrics, bias benchmarks.
- 03Bias Mitigation: Debiasing, CDA, and Fair Fine-tuning
Practical techniques for reducing demographic bias in language models: data balancing, embedding debiasing, adversarial training, and prompt-based interventions.
- 04Fairness Metrics: Demographic Parity, Equalized Odds, and Trade-offs
Learn the key mathematical definitions of algorithmic fairness, from demographic parity to equalized odds, and why satisfying all metrics simultaneously is impossible.
- 05Representation Harms
Covers stereotyping, erasure, demeaning associations, measuring representation harms.
Part LIX: Interpretability
- 01Interpretability Goals
Covers debugging, trust, safety, scientific understanding, interpretability approaches overview.
- 02Attention Visualization
Covers attention weight extraction, attention head visualization, attention interpretation caveats, attention tools.
- 03Attention Analysis Limitations
Covers attention vs importance, attention manipulation studies, gradient-based alternatives.
- 04Probing Classifiers
Covers linear probing methodology, probing task design, probing interpretation, control tasks.
- 05Probing Layers
Covers layer selection, representation evolution, task localization, layer probing patterns.
- 06Activation Patching
Covers patching methodology, locating information, patching experiments, causal tracing.
- 07Logit Lens
Covers logit lens concept, intermediate vocabulary projection, tuned lens, lens interpretation.
- 08Sparse Autoencoders
Covers SAE architecture, sparsity constraints, dictionary learning, SAE for LLMs.
- 09Feature Interpretation
Covers feature activation patterns, feature naming, automated interpretation, feature circuits.
- 10Mechanistic Interpretability
Reverse-engineer transformer networks into human-understandable algorithms by identifying circuits, induction heads, and mechanistic discoveries.
- 11Activation Steering: Steering Vectors and Representation Engineering
Learn how steering vectors and activation addition let you modify language model behavior at inference time by injecting directional nudges into the residual stream.
Part LX: Safety and Security
- 01Safety Risks
Covers harmful content generation, misuse scenarios, unintended harms, safety threat models.
- 02Red Teaming
Learn how red teams systematically probe language models for safety failures, covering attack taxonomies, ASR metrics, RL-based automated attack generation, and real-world findings.
- 03Jailbreaking
Covers jailbreak techniques, prompt injection, adversarial suffixes, jailbreak defenses.
- 04Prompt Injection
Covers direct prompt injection, indirect prompt injection, injection in RAG, injection defenses.
- 05Content Filtering
Covers classification-based filtering, rule-based filtering, filter placement, filter evaluation.
- 06Guardrails
Covers input guardrails, output guardrails, guardrail frameworks, guardrail design.
- 07Memorization and Privacy in Language Models
How language models memorize training data, methods for measuring extractable memorization, PII risks in web-scale corpora, and practical privacy mitigations.
- 08Differential Privacy
Learn how differential privacy protects training data in language models, from the mathematical guarantee to DP-SGD, privacy budgets, and practical LLM fine-tuning.
Part LXI: Agentic AI and Protocol Security
- 01Threat Modeling Compound AI SystemsPlanned
Assets, actors, trust boundaries, control flow, data flow, and failure propagation across model-centered systems.
- 02Indirect Prompt InjectionPlanned
Malicious instructions in documents, web pages, messages, tool output, and retrieved context.
- 03Tool, Retrieval, and Memory PoisoningPlanned
Schema attacks, compromised tools, poisoned indexes, persistent memory attacks, and integrity controls.
- 04Identity, Authorization, and Least PrivilegePlanned
User and agent identity, scoped credentials, delegation, expiry, consent, and policy enforcement outside the model.
- 05Secrets, Sandboxes, and Safe ExecutionPlanned
Credential isolation, capability sandboxes, network and filesystem boundaries, output validation, and containment.
- 06Agent Supply Chains and Protocol SecurityPlanned
Tool registries, dependency tampering, server discovery, trust-on-first-use, signing, and update risk.
- 07Audit, Detection, and Incident ResponsePlanned
Tamper-evident traces, anomaly detection, kill switches, investigation, rollback, and post-incident learning.
Part LXII: Human-AI Interaction and Governance
- 01Designing Human-AI CollaborationPlanned
Task allocation, mixed initiative, handoffs, feedback, shared context, and preserving human agency.
- 02Trust, Reliance, and Automation BiasPlanned
Calibrated reliance, overtrust, undertrust, explanations, uncertainty, and behavior under time pressure.
- 03Oversight, Appeals, and EscalationPlanned
Review queues, contestability, escalation paths, documentation, and accountability for consequential decisions.
- 04Accessibility and Inclusive Language AIPlanned
Disability access, literacy, language inclusion, participatory design, and evaluating who benefits or is excluded.
- 05System Cards, Audits, and Impact AssessmentsPlanned
Documenting complete systems, evaluating downstream use, independent audits, and monitoring material changes.
- 06Risk Governance Across the LifecyclePlanned
Ownership, risk tiers, review gates, incident reporting, change management, and retirement.
Volume 19 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 19 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.