Volume 19 of 20 · PDF edition

In progress

Responsible, Secure, and Interpretable Language AI

For teams responsible for fairness, model understanding, safety, agent security, audits, and governance.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
24 written chapters
Planned chapters
13 planned chapters
Approximate pages
~760 pages
Price
$24 one-time
Language AI Handbook, Volume 19: Responsible, Secure, and Interpretable Language AI cover

Author and edition details

About the author and Volume 19 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 19 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Bias and Fairness
  • Interpretability
  • Safety and Security
  • Agentic AI and Protocol Security
  • Human-AI Interaction and Governance

Audience and prerequisites

Where this volume fits

For teams responsible for fairness, model understanding, safety, agent security, audits, and governance.

Prerequisites: Assumes model fundamentals and benefits from Volume 18's evaluation framework.

Free online preview

Start with “Bias in Language Models

Covers bias sources, bias types (demographic, cultural), bias in training data, bias amplification.

Read the chapter

Exact contents

24 chapters available now

The current PDF contains every linked chapter below. The remaining 13 planned chapters will be added through free volume updates.

Part LVIII: Bias and Fairness

  1. 01
    Bias in Language Models

    Covers bias sources, bias types (demographic, cultural), bias in training data, bias amplification.

  2. 02
    Bias Measurement

    Covers embedding association tests, generation bias metrics, classification bias metrics, bias benchmarks.

  3. 03
    Bias Mitigation: Debiasing, CDA, and Fair Fine-tuning

    Practical techniques for reducing demographic bias in language models: data balancing, embedding debiasing, adversarial training, and prompt-based interventions.

  4. 04
    Fairness Metrics: Demographic Parity, Equalized Odds, and Trade-offs

    Learn the key mathematical definitions of algorithmic fairness, from demographic parity to equalized odds, and why satisfying all metrics simultaneously is impossible.

  5. 05
    Representation Harms

    Covers stereotyping, erasure, demeaning associations, measuring representation harms.

Part LIX: Interpretability

  1. 01
    Interpretability Goals

    Covers debugging, trust, safety, scientific understanding, interpretability approaches overview.

  2. 02
    Attention Visualization

    Covers attention weight extraction, attention head visualization, attention interpretation caveats, attention tools.

  3. 03
    Attention Analysis Limitations

    Covers attention vs importance, attention manipulation studies, gradient-based alternatives.

  4. 04
    Probing Classifiers

    Covers linear probing methodology, probing task design, probing interpretation, control tasks.

  5. 05
    Probing Layers

    Covers layer selection, representation evolution, task localization, layer probing patterns.

  6. 06
    Activation Patching

    Covers patching methodology, locating information, patching experiments, causal tracing.

  7. 07
    Logit Lens

    Covers logit lens concept, intermediate vocabulary projection, tuned lens, lens interpretation.

  8. 08
    Sparse Autoencoders

    Covers SAE architecture, sparsity constraints, dictionary learning, SAE for LLMs.

  9. 09
    Feature Interpretation

    Covers feature activation patterns, feature naming, automated interpretation, feature circuits.

  10. 10
    Mechanistic Interpretability

    Reverse-engineer transformer networks into human-understandable algorithms by identifying circuits, induction heads, and mechanistic discoveries.

  11. 11
    Activation Steering: Steering Vectors and Representation Engineering

    Learn how steering vectors and activation addition let you modify language model behavior at inference time by injecting directional nudges into the residual stream.

Part LX: Safety and Security

  1. 01
    Safety Risks

    Covers harmful content generation, misuse scenarios, unintended harms, safety threat models.

  2. 02
    Red Teaming

    Learn how red teams systematically probe language models for safety failures, covering attack taxonomies, ASR metrics, RL-based automated attack generation, and real-world findings.

  3. 03
    Jailbreaking

    Covers jailbreak techniques, prompt injection, adversarial suffixes, jailbreak defenses.

  4. 04
    Prompt Injection

    Covers direct prompt injection, indirect prompt injection, injection in RAG, injection defenses.

  5. 05
    Content Filtering

    Covers classification-based filtering, rule-based filtering, filter placement, filter evaluation.

  6. 06
    Guardrails

    Covers input guardrails, output guardrails, guardrail frameworks, guardrail design.

  7. 07
    Memorization and Privacy in Language Models

    How language models memorize training data, methods for measuring extractable memorization, PII risks in web-scale corpora, and practical privacy mitigations.

  8. 08
    Differential Privacy

    Learn how differential privacy protects training data in language models, from the mathematical guarantee to DP-SGD, privacy budgets, and practical LLM fine-tuning.

Part LXI: Agentic AI and Protocol Security

  1. 01
    Threat Modeling Compound AI SystemsPlanned

    Assets, actors, trust boundaries, control flow, data flow, and failure propagation across model-centered systems.

  2. 02
    Indirect Prompt InjectionPlanned

    Malicious instructions in documents, web pages, messages, tool output, and retrieved context.

  3. 03
    Tool, Retrieval, and Memory PoisoningPlanned

    Schema attacks, compromised tools, poisoned indexes, persistent memory attacks, and integrity controls.

  4. 04
    Identity, Authorization, and Least PrivilegePlanned

    User and agent identity, scoped credentials, delegation, expiry, consent, and policy enforcement outside the model.

  5. 05
    Secrets, Sandboxes, and Safe ExecutionPlanned

    Credential isolation, capability sandboxes, network and filesystem boundaries, output validation, and containment.

  6. 06
    Agent Supply Chains and Protocol SecurityPlanned

    Tool registries, dependency tampering, server discovery, trust-on-first-use, signing, and update risk.

  7. 07
    Audit, Detection, and Incident ResponsePlanned

    Tamper-evident traces, anomaly detection, kill switches, investigation, rollback, and post-incident learning.

Part LXII: Human-AI Interaction and Governance

  1. 01
    Designing Human-AI CollaborationPlanned

    Task allocation, mixed initiative, handoffs, feedback, shared context, and preserving human agency.

  2. 02
    Trust, Reliance, and Automation BiasPlanned

    Calibrated reliance, overtrust, undertrust, explanations, uncertainty, and behavior under time pressure.

  3. 03
    Oversight, Appeals, and EscalationPlanned

    Review queues, contestability, escalation paths, documentation, and accountability for consequential decisions.

  4. 04
    Accessibility and Inclusive Language AIPlanned

    Disability access, literacy, language inclusion, participatory design, and evaluating who benefits or is excluded.

  5. 05
    System Cards, Audits, and Impact AssessmentsPlanned

    Documenting complete systems, evaluating downstream use, independent audits, and monitoring material changes.

  6. 06
    Risk Governance Across the LifecyclePlanned

    Ownership, risk tiers, review gates, incident reporting, change management, and retirement.

Volume 19 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 19 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.