Volume 18 of 20 · PDF edition
In progressEvaluation, Factuality, and Reliability
For researchers and product teams designing credible benchmarks, judges, human studies, and reliability programs.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 30 written chapters
- Planned chapters
- 6 planned chapters
- Approximate pages
- ~950 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 18 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 18 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Evaluation Fundamentals
- Benchmark Evaluation
- Human and Model Evaluation
- Hallucination and Factuality
- Dynamic, Interactive, and Agentic Evaluation
Audience and prerequisites
Where this volume fits
For researchers and product teams designing credible benchmarks, judges, human studies, and reliability programs.
Prerequisites: Can be read after any model-focused volume; basic statistics are helpful.
Free online preview
Start with “Perplexity Evaluation”
Covers perplexity calculation, perplexity interpretation, perplexity limitations, comparing perplexities.
Exact contents
30 chapters available now
The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.
Part LIII: Evaluation Fundamentals
- 01Perplexity Evaluation
Covers perplexity calculation, perplexity interpretation, perplexity limitations, comparing perplexities.
- 02Cross-Entropy Loss
Covers cross-entropy definition, bits-per-character, cross-entropy vs perplexity, loss curves.
- 03BLEU Score
Covers n-gram precision, brevity penalty, BLEU formula, BLEU limitations, corpus vs sentence BLEU.
- 04ROUGE Scores
Covers ROUGE-N, ROUGE-L, ROUGE-W, ROUGE interpretation, ROUGE limitations.
- 05BERTScore
Covers BERTScore computation, token alignment, BERTScore variants, BERTScore vs BLEU.
- 06Exact Match and F1
Covers exact match scoring, token-level F1, normalization for matching, metric selection.
- 07Calibration
Covers calibration definition, expected calibration error, calibration plots, calibration methods.
- 08Evaluation Fundamentals
Covers evaluation design principles, metric selection, evaluation pitfalls, evaluation frameworks.
- 09Benchmark Design
Covers benchmark construction, dataset collection, annotation guidelines, benchmark validity.
- 10Evaluation Challenges
Covers benchmark contamination, evaluation brittleness, gaming metrics, evaluation best practices.
Part LIV: Benchmark Evaluation
- 01MMLU
Covers MMLU structure, subject coverage, MMLU evaluation protocol, MMLU limitations.
- 02HellaSwag
Covers HellaSwag task design, adversarial filtering, HellaSwag evaluation, HellaSwag saturation.
- 03GSM8K
Covers GSM8K problem types, chain-of-thought evaluation, GSM8K accuracy metrics, math reasoning assessment.
- 04HumanEval
Covers HumanEval structure, functional correctness, pass@k metric, HumanEval limitations.
- 05MBPP
Covers MBPP dataset, MBPP vs HumanEval, code evaluation challenges.
- 06TruthfulQA
Covers TruthfulQA design, truthfulness vs informativeness, TruthfulQA evaluation methods.
- 07Benchmark Contamination
Covers contamination problem, contamination detection methods, n-gram overlap analysis, contamination mitigation.
- 08Benchmark Saturation
Covers ceiling effects, benchmark retirement, dynamic benchmarks, benchmark evolution.
Part LV: Human and Model Evaluation
- 01Human Evaluation Design
Covers evaluation interface design, task instructions, annotator selection, evaluation cost.
- 02Inter-Annotator Agreement
Covers Cohen's kappa, Fleiss' kappa, Krippendorff's alpha, handling disagreement.
- 03Preference Evaluation
Covers A/B comparison design, Elo rating systems, preference aggregation, statistical significance.
- 04LLM-as-Judge: Scalable AI Evaluation with Language Models
Learn how to build LLM-as-Judge evaluation pipelines: prompt design, judge model selection, calibration against human annotations, and bias mitigation.
- 05Position Bias in LLM Judges
Covers position bias measurement, bias mitigation (swapping), verbosity bias, sycophancy.
- 06Evaluation Prompt Engineering: Designing Reliable LLM Judges
Learn how to design reliable LLM judge prompts using explicit criteria, few-shot examples, and chain-of-thought formatting to maximize evaluation accuracy.
Part LVI: Hallucination and Factuality
- 01Hallucination Types in Language Models: A Complete Guide
Learn how language models hallucinate: intrinsic and extrinsic hallucination, factual errors, fabrication, and inconsistency with NLI-based detection.
- 02Hallucination Detection: NLI, Self-Consistency & Learned Models
Learn four methods for detecting LLM hallucinations: entailment-based scoring, knowledge base verification, self-consistency checks, and learned detection models.
- 03Hallucination Causes
Covers training data issues, exposure bias, knowledge gaps, generation pressure.
- 04Hallucination Mitigation: RAG, Decoding, and Training
Learn how to reduce LLM hallucination using retrieval augmentation, self-consistency decoding, DPO training, and calibrated uncertainty expression.
- 05Attribution and Citation: Sourcing LLM Outputs
Learn how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.
- 06Uncertainty Quantification
Covers confidence calibration, verbalized uncertainty, sampling-based uncertainty, uncertainty communication.
Part LVII: Dynamic, Interactive, and Agentic Evaluation
- 01Live and Continuously Updated BenchmarksPlanned
Fresh questions, private test sets, rolling releases, objective scoring, and contamination controls.
- 02Evaluation Harnesses and ReproducibilityPlanned
Versioned prompts, model settings, graders, confidence intervals, artifacts, and comparable reports.
- 03Interactive and Long-Horizon Task EvaluationPlanned
Stateful environments, functional outcomes, partial credit, recovery, efficiency, and trajectory analysis.
- 04Judge Calibration and Meta-EvaluationPlanned
Bias, self-preference, verbosity effects, rubric validity, judge ensembles, and agreement with experts.
- 05Online Experiments and Production EvaluationPlanned
Shadow traffic, interleaving, A/B tests, guardrail metrics, drift, and linking offline scores to user outcomes.
- 06Evaluation GovernancePlanned
Benchmark access, disclosure, privacy, gaming, deprecation, audit trails, and responsible leaderboard use.
Volume 18 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 18 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.