Volume 18 of 20 · PDF edition

In progress

Evaluation, Factuality, and Reliability

For researchers and product teams designing credible benchmarks, judges, human studies, and reliability programs.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
30 written chapters
Planned chapters
6 planned chapters
Approximate pages
~950 pages
Price
$24 one-time
Language AI Handbook, Volume 18: Evaluation, Factuality, and Reliability cover

Author and edition details

About the author and Volume 18 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 18 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Evaluation Fundamentals
  • Benchmark Evaluation
  • Human and Model Evaluation
  • Hallucination and Factuality
  • Dynamic, Interactive, and Agentic Evaluation

Audience and prerequisites

Where this volume fits

For researchers and product teams designing credible benchmarks, judges, human studies, and reliability programs.

Prerequisites: Can be read after any model-focused volume; basic statistics are helpful.

Free online preview

Start with “Perplexity Evaluation

Covers perplexity calculation, perplexity interpretation, perplexity limitations, comparing perplexities.

Read the chapter

Exact contents

30 chapters available now

The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.

Part LIII: Evaluation Fundamentals

  1. 01
    Perplexity Evaluation

    Covers perplexity calculation, perplexity interpretation, perplexity limitations, comparing perplexities.

  2. 02
    Cross-Entropy Loss

    Covers cross-entropy definition, bits-per-character, cross-entropy vs perplexity, loss curves.

  3. 03
    BLEU Score

    Covers n-gram precision, brevity penalty, BLEU formula, BLEU limitations, corpus vs sentence BLEU.

  4. 04
    ROUGE Scores

    Covers ROUGE-N, ROUGE-L, ROUGE-W, ROUGE interpretation, ROUGE limitations.

  5. 05
    BERTScore

    Covers BERTScore computation, token alignment, BERTScore variants, BERTScore vs BLEU.

  6. 06
    Exact Match and F1

    Covers exact match scoring, token-level F1, normalization for matching, metric selection.

  7. 07
    Calibration

    Covers calibration definition, expected calibration error, calibration plots, calibration methods.

  8. 08
    Evaluation Fundamentals

    Covers evaluation design principles, metric selection, evaluation pitfalls, evaluation frameworks.

  9. 09
    Benchmark Design

    Covers benchmark construction, dataset collection, annotation guidelines, benchmark validity.

  10. 10
    Evaluation Challenges

    Covers benchmark contamination, evaluation brittleness, gaming metrics, evaluation best practices.

Part LIV: Benchmark Evaluation

  1. 01
    MMLU

    Covers MMLU structure, subject coverage, MMLU evaluation protocol, MMLU limitations.

  2. 02
    HellaSwag

    Covers HellaSwag task design, adversarial filtering, HellaSwag evaluation, HellaSwag saturation.

  3. 03
    GSM8K

    Covers GSM8K problem types, chain-of-thought evaluation, GSM8K accuracy metrics, math reasoning assessment.

  4. 04
    HumanEval

    Covers HumanEval structure, functional correctness, pass@k metric, HumanEval limitations.

  5. 05
    MBPP

    Covers MBPP dataset, MBPP vs HumanEval, code evaluation challenges.

  6. 06
    TruthfulQA

    Covers TruthfulQA design, truthfulness vs informativeness, TruthfulQA evaluation methods.

  7. 07
    Benchmark Contamination

    Covers contamination problem, contamination detection methods, n-gram overlap analysis, contamination mitigation.

  8. 08
    Benchmark Saturation

    Covers ceiling effects, benchmark retirement, dynamic benchmarks, benchmark evolution.

Part LV: Human and Model Evaluation

  1. 01
    Human Evaluation Design

    Covers evaluation interface design, task instructions, annotator selection, evaluation cost.

  2. 02
    Inter-Annotator Agreement

    Covers Cohen's kappa, Fleiss' kappa, Krippendorff's alpha, handling disagreement.

  3. 03
    Preference Evaluation

    Covers A/B comparison design, Elo rating systems, preference aggregation, statistical significance.

  4. 04
    LLM-as-Judge: Scalable AI Evaluation with Language Models

    Learn how to build LLM-as-Judge evaluation pipelines: prompt design, judge model selection, calibration against human annotations, and bias mitigation.

  5. 05
    Position Bias in LLM Judges

    Covers position bias measurement, bias mitigation (swapping), verbosity bias, sycophancy.

  6. 06
    Evaluation Prompt Engineering: Designing Reliable LLM Judges

    Learn how to design reliable LLM judge prompts using explicit criteria, few-shot examples, and chain-of-thought formatting to maximize evaluation accuracy.

Part LVI: Hallucination and Factuality

  1. 01
    Hallucination Types in Language Models: A Complete Guide

    Learn how language models hallucinate: intrinsic and extrinsic hallucination, factual errors, fabrication, and inconsistency with NLI-based detection.

  2. 02
    Hallucination Detection: NLI, Self-Consistency & Learned Models

    Learn four methods for detecting LLM hallucinations: entailment-based scoring, knowledge base verification, self-consistency checks, and learned detection models.

  3. 03
    Hallucination Causes

    Covers training data issues, exposure bias, knowledge gaps, generation pressure.

  4. 04
    Hallucination Mitigation: RAG, Decoding, and Training

    Learn how to reduce LLM hallucination using retrieval augmentation, self-consistency decoding, DPO training, and calibrated uncertainty expression.

  5. 05
    Attribution and Citation: Sourcing LLM Outputs

    Learn how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.

  6. 06
    Uncertainty Quantification

    Covers confidence calibration, verbalized uncertainty, sampling-based uncertainty, uncertainty communication.

Part LVII: Dynamic, Interactive, and Agentic Evaluation

  1. 01
    Live and Continuously Updated BenchmarksPlanned

    Fresh questions, private test sets, rolling releases, objective scoring, and contamination controls.

  2. 02
    Evaluation Harnesses and ReproducibilityPlanned

    Versioned prompts, model settings, graders, confidence intervals, artifacts, and comparable reports.

  3. 03
    Interactive and Long-Horizon Task EvaluationPlanned

    Stateful environments, functional outcomes, partial credit, recovery, efficiency, and trajectory analysis.

  4. 04
    Judge Calibration and Meta-EvaluationPlanned

    Bias, self-preference, verbosity effects, rubric validity, judge ensembles, and agreement with experts.

  5. 05
    Online Experiments and Production EvaluationPlanned

    Shadow traffic, interleaving, A/B tests, guardrail metrics, drift, and linking offline scores to user outcomes.

  6. 06
    Evaluation GovernancePlanned

    Benchmark access, disclosure, privacy, gaming, deprecation, audit trails, and responsible leaderboard use.

Volume 18 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 18 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.