Language AI Handbook

Start with classical NLP, then work through transformers, LLM training, evaluation, and production
357h 29m total read time
409 of 513 chapters published
Language AI Handbook Cover

Read Language AI Handbook free online

The Language AI Handbook connects the pieces of modern language AI. We begin with classical NLP, build up to transformers and LLM training, then cover retrieval, evaluation, safety, and production deployment.

Begin with the fundamentals that never go out of style: tokenization, embeddings, and the statistical foundations that inform modern approaches. Then dive deep into the transformer architecture. Learn not just how to use it, but how it actually works. Understand self-attention mathematically, grasp why positional encodings matter, and see how architectural choices like layer normalization affect training dynamics.

Author and edition details

About the author and this edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
First edition
Published
Last reviewed

Track Your Progress

Sign in to mark chapters as complete, track quiz scores, and see your reading journey

Sign in →
Part 1

Text as Data

54h 46m
Part 2

Classical Text Representations

98h 3m
Part 3

Distributional Semantics

43h 16m
Part 4

Word Embeddings

97h 24m
Part 5

Subword Tokenization

86h 52m
Part 6

Sequence Labeling

87h 10m
Part 7

Linguistic Form: Morphology and Syntax

2/61h 23m
Part 8

Semantics and Information Extraction

6 · Coming soon
Part 9

Coreference, Discourse, Pragmatics, and Dialogue

5 · Coming soon
Part 10

Neural Network Foundations

1310h 55m
Part 11

Recurrent Neural Networks

97h 18m
Part 12

Sequence-to-Sequence

75h 59m
Part 13

Self-Attention

64h 56m
Part 14

Positional Encoding

75h 40m
Part 15

Transformer Blocks

86h 38m
Part 16

Transformer Architectures

64h 36m
Part 17

Efficient Attention

97h 16m
Part 18

Long Context

76h 19m
Part 19

Alternative Sequence and Generative Architectures

6 · Coming soon
Part 20

Data Curation

108h 52m
Part 21

Data Governance, Provenance, and Synthetic Ecosystems

6 · Coming soon
Part 22

Pre-training Objectives

75h 41m
Part 23

Scaling Laws

75h 44m
Part 24

BERT and Variants

86h 31m
Part 25

Encoder-Decoder Models

64h 49m
Part 26

Multilingual Language Models and Cross-Lingual Transfer

7 · Coming soon
Part 27

Machine Translation and Speech Translation

7 · Coming soon
Part 28

GPT Architecture

109h 47m
Part 29

Modern Decoder Models

75h 9m
Part 30

Emergent Capabilities

65h 39m
Part 31

Training Infrastructure

119h 40m
Part 32

Training Optimization

87h 26m
Part 33

Mixture of Experts

109h 18m
Part 34

Fine-tuning Fundamentals

54h 55m
Part 35

Parameter-Efficient Fine-tuning

1212h 3m
Part 36

Instruction Tuning

65h 48m
Part 37

Alignment and RLHF

1616h 18m
Part 38

Scalable Oversight and Safety Training

5 · Coming soon
Part 39

Reasoning

76h 23m
Part 40

Reasoning Post-Training and Verifiable Rewards

7 · Coming soon
Part 41

Model Compression

64h 46m
Part 42

Inference Optimization

1413h 14m
Part 43

Small and On-Device Language Models

5 · Coming soon
Part 44

Retrieval-Augmented Generation

1414h 30m
Part 45

Knowledge, Memory, Editing, and Unlearning

6 · Coming soon
Part 46

Continual Learning

54h 19m
Part 47

Code Generation

65h 17m
Part 48

Tool Use and Agents

109h 9m
Part 49

Long-Horizon Agents and Interoperability

8 · Coming soon
Part 50

Multimodal Models

1210h 57m
Part 51

Speech and Audio

54h 55m
Part 52

Omni-Modal and Embodied Language Systems

7 · Coming soon
Part 53

Evaluation Fundamentals

108h 53m
Part 54

Benchmark Evaluation

86h 43m
Part 55

Human and Model Evaluation

65h 22m
Part 56

Hallucination and Factuality

65h 25m
Part 57

Dynamic, Interactive, and Agentic Evaluation

6 · Coming soon
Part 58

Bias and Fairness

54h 43m
Part 59

Interpretability

119h 20m
Part 60

Safety and Security

87h 31m
Part 61

Agentic AI and Protocol Security

7 · Coming soon
Part 62

Human-AI Interaction and Governance

6 · Coming soon
Part 63

LLM Applications

7
Part 64

Compound AI System Design

6 · Coming soon
Part 65

Production Systems

97h 47m
Part 66

Frontier Methods of 2025

87h 3m
Part 67

Frontier Outlook from 2025

65h 1m
Own the library

The whole growing library, yours forever.

Reading online is free, and always will be. The paid library unlocks every current volume as its own carefully typeset PDF, plus each future volume when it is published. There is intentionally no combined PDF.
  • All 395 written chapters across 19 available volumes

  • 19 written volume PDFs plus the Volume 3 placeholder. Files are delivered separately.

  • All future expansions, updates, and errata, free, forever

  • Every future volume included, including 104 planned chapters

  • Downloadable editions for durable offline reading

Best value, save $257 (56%)

$199

one-time

instead of $456 for the 19 available volume PDFs bought separately

Language AI Handbook Cover
5.0 from 12 readersGet all-volume access · $199Start with a single volume · $24

Secure checkout via Stripe · Individual PDF links delivered to your inbox

Reader verdict

5.0

out of 5

12 five-star reader reviews

Genuine feedback, with names abbreviated for privacy.

Meet the community
“Really appreciate the way you have explained the concepts. Simple to understand, yet builds up knowledge exponentially, all while reinforcing it with examples repeatedly. I've read the last two chapters `74.Self-attention & 75.Q,K,V` and I've to admit, my concepts have never been clearer. You really are doing an amazing job breaking down maths and complex concepts like a story!”
Muhammad H.AI/ML Engineer
“I found your book Language AI Handbook's explanations of Transformer architectures and production deployment strategies to be exceptionally clear and insightful”
San Z.Researcher
“I came across one of your books, Language AI Handbook, and man, I'm loving it. I've read many books but have never found such a comprehensive book on large models that covers everything from start to finish with such depth.”
Uon L.Reader
Read the other 9 reader reviews
“I would like to thank you for your work. Specifically, the chapter: 'LIME Explainability: Complete Guide to Local Interpretable Model-Agnostic Explanations' was extremely helpful for me in understanding LIME and communicating to colleagues. Your work is inspiring and valuable. Thank you for doing it.”
Will G.Data Scientist
“Thanks so much for establishing such an amazing community. I appreciate the opportunity to be a learner and a contributor to the community.”
Cyrille L.AI/ML Engineer
“I just came across your website where your books are available for free to read. Thank you so much for making this available to the public. This will help a lot of people.”
Vikas C.AI/ML Engineer
“I am enjoying reading your writings. I appreciate the effort you put into it. I want to drop a note of thanks. Thank you!”
Krishnan S.Researcher
“Thank you Michael. It's top!”
David X.Data Scientist
“Thanks for putting cool stuff out into the world.”
Charlie B.Applied AI
“Your writing feels like a treasure trove of great material. So, thank you for compiling everything on your website!”
Griffin P.Quantitative Researcher
“I came across your collection of online handbooks covering quantitative finance, AI, and machine learning. I wanted to reach out and thank you for making such high-quality material openly available. They are incredibly helpful for my CFA Level II review and for brushing up ahead of ML interviews.”
Karim Z.Reader
“I read a few of your blogs on NLP and oh WOW, they really are something, now I exactly know how all of the terms Entropy, Cross-entropy and how it is connected to Perplexity, will definitely read more of your blogs.”
Kanishk K.Reader

Prefer to start with a single volume?

Buy 9 volumes and all-volume access unlocks free. Already own volumes? Upgrade to all-volume access and you only pay the difference.

Table of Contents

Part I: Text as Data

5 chapters

Part II: Classical Text Representations

9 chapters
6

Bag of Words

Covers document-term matrix construction, vocabulary building from corpus, word counting and frequency vectors, sparse matrix representation (CSR/CSC formats), vocabulary pruning (min_df, max_df), binary vs count representations, limitations of word order loss.

51m
7

N-grams

Covers bigram and trigram extraction, n-gram vocabulary explosion, n-gram frequency distributions, Zipf's law in n-grams, character n-grams for robustness, skip-grams and flexible windows, n-gram indexing for search.

47m
8

N-gram Language Models

Covers Markov assumption and chain rule, maximum likelihood estimation, probability calculation for sequences, handling unseen n-grams, start and end tokens, generating text from n-gram models, model storage and lookup efficiency.

58m
9

Smoothing Techniques

Covers add-one (Laplace) smoothing, add-k smoothing and tuning, Good-Turing smoothing derivation, Kneser-Ney smoothing intuition and formula, interpolation vs backoff, modified Kneser-Ney, comparing smoothing methods empirically.

56m
10

Perplexity

Covers cross-entropy definition and derivation, perplexity as branching factor, relationship to bits-per-character, held-out evaluation methodology, perplexity vs downstream performance, comparing models with perplexity, perplexity limitations and caveats.

63m
11

Term Frequency

Covers raw term frequency, log-scaled term frequency, boolean term frequency, augmented term frequency, L2-normalized frequency vectors, term frequency sparsity patterns, efficient term frequency computation.

46m
12

Inverse Document Frequency

Covers document frequency calculation, IDF formula derivation, IDF intuition (rare words matter more), smoothed IDF variants, IDF across corpus splits, relationship to information theory, implementing IDF efficiently.

59m
13

TF-IDF

Covers TF-IDF formula and variants, TF-IDF vector computation, TF-IDF normalization options, BM25 as TF-IDF extension, document similarity with TF-IDF, TF-IDF for feature extraction, sklearn TfidfVectorizer deep dive.

51m
14

BM25

Covers BM25 derivation from probabilistic IR, saturation parameter k1, length normalization parameter b, BM25+ and BM25L variants, field-weighted BM25, implementing BM25 scoring, BM25 vs TF-IDF empirically.

52m

Part III: Distributional Semantics

4 chapters

Part IV: Word Embeddings

9 chapters
19

Skip-gram Model

Covers skip-gram architecture diagram, input/output representations, softmax over vocabulary, skip-gram objective function, training data generation, window size hyperparameter, skip-gram vs CBOW intuition.

41m
20

CBOW Model

Covers CBOW architecture, context word averaging, CBOW objective function, CBOW vs skip-gram training speed, CBOW for frequent words, implementing CBOW forward pass, CBOW gradient derivation.

52m
21

Negative Sampling

Covers softmax computational bottleneck, negative sampling objective derivation, sampling distribution (unigram^0.75), number of negatives hyperparameter, negative sampling gradient computation, NCE vs negative sampling, implementing efficient sampling.

47m
22

Hierarchical Softmax

Covers binary tree construction (Huffman coding), path probability computation, hierarchical softmax objective, gradient computation along paths, tree structure impact on learning, hierarchical softmax vs negative sampling, when to use each approach.

53m
23

Word2Vec Training

Covers data preprocessing pipeline, subsampling frequent words, learning rate scheduling, minibatch vs online training, convergence monitoring, gensim Word2Vec usage, training from scratch in PyTorch.

50m
24

Word Analogy

Covers vector arithmetic for analogies, parallelogram model, analogy evaluation datasets, 3CosAdd vs 3CosMul methods, analogy accuracy metrics, limitations of analogy evaluation, what analogies reveal about embeddings.

49m
25

GloVe

Covers GloVe objective function derivation, weighted least squares formulation, relationship to matrix factorization, weighting function design, bias terms in GloVe, GloVe vs Word2Vec comparison, training GloVe efficiently.

47m
26

FastText

Covers character n-gram representation, word vector as n-gram sum, FastText architecture, handling OOV words, morphological awareness, FastText for morphologically rich languages, training FastText models.

51m
27

Embedding Evaluation

Covers intrinsic vs extrinsic evaluation, word similarity datasets (SimLex, WordSim), analogy accuracy, embedding visualization (t-SNE, UMAP), downstream task evaluation, embedding bias detection, evaluation pitfalls.

54m

Part V: Subword Tokenization

8 chapters
28

The Vocabulary Problem

Covers OOV word problem, vocabulary size explosion, rare word representation, morphological productivity, compound words, code and technical text, the case for subword units.

48m
29

Byte Pair Encoding

Covers BPE algorithm step-by-step, merge rules learning, vocabulary size control, BPE encoding procedure, BPE decoding procedure, BPE implementation from scratch, BPE hyperparameters.

46m
30

WordPiece

Covers WordPiece vs BPE differences, likelihood objective for merges, greedy tokenization algorithm, ## prefix notation, WordPiece in BERT, training WordPiece tokenizers, handling unknown characters.

48m
31

Unigram Language Model Tokenization

Covers unigram LM formulation, EM algorithm for training, Viterbi decoding for tokenization, sampling multiple segmentations, subword regularization, unigram vs BPE comparison, SentencePiece unigram mode.

55m
32

SentencePiece

Covers treating text as raw bytes, whitespace handling (▁ prefix), BPE and unigram modes, training from raw text, pretokenization elimination, SentencePiece in production, multilingual tokenization.

44m
33

Tokenizer Training

Covers corpus preparation, vocabulary size selection, special tokens configuration, training with HuggingFace tokenizers, saving and loading tokenizers, tokenizer versioning, domain-specific tokenizers.

53m
34

Special Tokens

Covers [CLS], [SEP], [PAD], [MASK], [UNK] tokens, beginning/end of sequence tokens, custom special tokens, special token embeddings, token type IDs, handling special tokens in generation.

57m
35

Tokenization Challenges

Covers number tokenization issues, code tokenization, multilingual text mixing, emoji and Unicode edge cases, tokenization artifacts, adversarial tokenization, measuring tokenization quality.

61m

Part VI: Sequence Labeling

8 chapters
36

Part-of-Speech Tagging

Covers POS tag sets (Penn Treebank, Universal), POS tagging as classification, contextual disambiguation, POS tagging accuracy metrics, POS tagging for downstream tasks, rule-based vs statistical taggers.

64m
37

Named Entity Recognition

Covers entity types (PER, ORG, LOC, etc.), NER as sequence labeling, nested entity challenges, entity boundary detection, NER evaluation (exact vs partial match), NER datasets and benchmarks.

60m
38

BIO Tagging

Covers BIO scheme explanation, BIOES/BILOU variants, converting spans to BIO tags, BIO decoding to spans, handling tagging inconsistencies, BIO for multi-label scenarios, implementing BIO utilities.

55m
39

Chunking

Covers noun phrase chunking, chunk types (NP, VP, PP), IOB tagging for chunks, chunking vs full parsing, chunking evaluation, chunking as preprocessing, regex chunking with NLTK.

41m
40

Hidden Markov Models

Covers HMM components (states, observations, transitions), emission and transition probabilities, HMM assumptions (Markov, independence), HMM for POS tagging, HMM parameter estimation, HMM limitations for NLP.

49m
41

Viterbi Algorithm

Covers optimal path problem formulation, Viterbi recursion derivation, backpointer tracking, Viterbi complexity analysis, log-space computation, implementing Viterbi efficiently, Viterbi for beam search foundation.

47m
42

Conditional Random Fields

Covers CRF vs HMM comparison, CRF feature functions, log-linear formulation, partition function computation, CRF for NER, CRF inference complexity, neural CRF layers.

58m
43

CRF Training

Covers CRF log-likelihood objective, forward-backward algorithm, gradient computation, L-BFGS optimization, feature template design, CRF regularization, CRF training convergence.

56m

Part VII: Linguistic Form: Morphology and Syntax

6 chapters
44

Morphology and Word Formation

Morphemes, inflection, derivation, compounding, and why word structure matters for language models.

37m
45

Morphological Analysis and Generation

Finite-state and neural approaches to analyzing and generating morphologically rich language.

46m
46

Grammars and Constituency Structure

Soon

Context-free grammars, phrase structure, constituency trees, and the limits of symbolic grammars.

47

Dependency Syntax and Universal Dependencies

Soon

Head-dependent relations, dependency trees, projectivity, and cross-lingual annotation with Universal Dependencies.

48

Constituency and Dependency Parsing

Soon

Transition-based, graph-based, chart, and neural parsing methods with their decoding trade-offs.

49

Structured Prediction and Parsing Evaluation

Soon

Dynamic programming, constrained decoding, labeled attachment, span scores, and meaningful parser error analysis.

Part VIII: Semantics and Information Extraction

6 chapters
50

Lexical Semantics and Word Senses

Soon

Polysemy, synonymy, lexical resources, contextual word senses, and word-sense disambiguation.

51

Sentence Meaning and Natural Language Inference

Soon

Compositional meaning, entailment, contradiction, presupposition, and robust NLI evaluation.

52

Semantic Roles and Argument Structure

Soon

Predicates, arguments, thematic roles, frame semantics, and semantic role labeling.

53

Relation, Event, and Temporal Extraction

Soon

Extracting relations, events, arguments, and timelines from documents rather than isolated sentences.

54

Entity Linking and Knowledge Graph Construction

Soon

Linking mentions to entities, resolving ambiguity, and building provenance-aware knowledge graphs.

55

Semantic Parsing and Meaning Representations

Soon

Logical forms, AMR, text-to-SQL, executable meaning representations, and constrained generation.

Part IX: Coreference, Discourse, Pragmatics, and Dialogue

5 chapters
56

Coreference Resolution

Soon

Mention detection, entity clusters, pronoun resolution, long-document coreference, and evaluation.

57

Discourse Relations and Coherence

Soon

Discourse structure, rhetorical relations, coherence modeling, and long-form generation failures.

58

Pragmatics, Implicature, and Presupposition

Soon

Speaker intent, common ground, deixis, implicature, and what literal next-token prediction misses.

59

Dialogue Acts and Conversation Structure

Soon

Turn-taking, grounding, repair, dialogue acts, state tracking, and multi-party conversation.

60

Argumentation, Stance, and Persuasion

Soon

Claims, evidence, stance, argument structure, persuasion, and responsible evaluation.

Part X: Neural Network Foundations

13 chapters
61

Linear Classifiers

Covers linear decision boundaries, weight vectors and bias, dot product interpretation, multiclass classification (softmax), linear classifier limitations, training with gradient descent.

53m
62

Activation Functions

Covers sigmoid function and saturation, tanh properties, ReLU and dying ReLU, Leaky ReLU and PReLU, ELU and SELU, GELU derivation and properties, Swish and Mish, choosing activation functions.

50m
63

Multilayer Perceptrons

Covers hidden layers and depth, weight matrices between layers, forward pass computation, representational capacity, MLP for classification, MLP for regression, MLP architecture design.

40m
64

Loss Functions

Covers cross-entropy loss derivation, MSE for regression, binary vs multiclass cross-entropy, label smoothing, focal loss for imbalance, loss function numerical stability, custom loss functions.

56m
65

Backpropagation

Covers computational graphs, chain rule review, forward and backward pass, gradient accumulation, backprop complexity analysis, automatic differentiation, implementing backprop from scratch.

58m
66

Stochastic Gradient Descent

Covers batch vs stochastic gradient descent, minibatch gradient descent, learning rate selection, SGD convergence properties, SGD noise as regularization, learning rate schedules basics, SGD implementation.

40m
67

Momentum

Covers momentum intuition (ball rolling), momentum update equations, momentum coefficient selection, dampening oscillations, momentum vs vanilla SGD, Nesterov momentum derivation, implementing momentum.

50m
68

Adam Optimizer

Covers exponential moving averages, first moment (mean) estimation, second moment (variance) estimation, bias correction derivation, Adam update rule, Adam hyperparameters, Adam convergence properties.

51m
69

AdamW

Covers L2 regularization vs weight decay, why they differ with Adam, AdamW formulation, weight decay coefficient selection, AdamW as default optimizer, AdamW vs Adam empirically.

45m
70

Weight Initialization

Covers random initialization importance, Xavier/Glorot initialization derivation, He initialization for ReLU, initialization for different activations, layer-wise initialization, initialization debugging, modern initialization practices.

53m
71

Batch Normalization

Covers internal covariate shift, batch statistics computation, learnable scale and shift, training vs inference mode, batch norm gradient flow, batch norm placement debates, batch norm limitations.

54m
72

Dropout

Covers dropout as ensemble, dropout mask sampling, inverted dropout scaling, dropout rate selection, dropout at inference, spatial dropout for sequences, dropout in modern architectures.

55m
73

Gradient Clipping

Covers gradient explosion detection, clip by value, clip by global norm, gradient clipping implementation, when to use gradient clipping, clipping threshold selection, monitoring gradient norms.

50m

Part XI: Recurrent Neural Networks

9 chapters
74

RNN Architecture

Covers recurrent connection intuition, hidden state as memory, unrolled computation graph, parameter sharing across time, RNN for sequence classification, RNN for sequence generation, RNN equations and dimensions.

55m
75

Backpropagation Through Time

Covers BPTT derivation, gradient flow through time, truncated BPTT, BPTT memory requirements, BPTT implementation, gradient accumulation across timesteps.

52m
76

Vanishing Gradients

Covers gradient product across timesteps, vanishing gradient analysis, long-range dependency failure, gradient visualization, vanishing vs exploding trade-off, architectural solutions overview.

45m
77

LSTM Architecture

Covers cell state as information highway, gate mechanism intuition, LSTM diagram walkthrough, information flow in LSTMs, LSTM for long sequences, LSTM memory capacity.

55m
78

LSTM Gate Equations

Covers forget gate equations, input gate equations, cell state update, output gate equations, hidden state computation, LSTM parameter count, implementing LSTM from scratch.

40m
79

LSTM Gradient Flow

Covers constant error carousel, forget gate gradient highway, gradient flow analysis, LSTM vs vanilla RNN gradients, peephole connections, LSTM gradient clipping needs.

52m
80

GRU Architecture

Covers GRU vs LSTM comparison, reset gate function, update gate function, candidate hidden state, GRU equations, GRU parameter efficiency, when to choose GRU vs LSTM.

48m
81

Bidirectional RNNs

Covers forward and backward passes, hidden state concatenation, bidirectional architectures, bidirectionality for classification, limitations for generation, implementing bidirectional RNNs.

50m
82

Stacked RNNs

Covers multiple RNN layers, residual connections for depth, layer normalization in RNNs, depth vs width trade-offs, gradient flow in deep RNNs, practical depth limits.

41m

Part XII: Sequence-to-Sequence

7 chapters

Part XIII: Self-Attention

6 chapters

Part XIV: Positional Encoding

7 chapters

Part XV: Transformer Blocks

8 chapters

Part XVI: Transformer Architectures

6 chapters

Part XVII: Efficient Attention

9 chapters
117

Quadratic Attention Bottleneck

Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.

48m
118

Sparse Attention Patterns

Covers local attention windows, strided attention patterns, block-sparse attention, combining sparse patterns, sparse attention implementation.

61m
119

Sliding Window Attention

Covers sliding window formulation, window size selection, dilated sliding windows, sliding window for long sequences, Mistral-style windowed attention.

43m
120

Global Tokens

Covers CLS token global attention, learned global tokens, global-local attention mixing, global token count, implementation strategies.

51m
121

Longformer

Covers Longformer attention pattern, global attention configuration, Longformer complexity, Longformer for documents, Longformer implementation.

55m
122

BigBird

Covers BigBird attention pattern, random attention benefits, BigBird theoretical guarantees, BigBird vs Longformer, BigBird applications.

40m
123

Linear Attention

Covers softmax attention reformulation, kernel feature maps, linear complexity attention, linear attention limitations, Performer and variants.

42m
124

FlashAttention Algorithm

Covers GPU memory hierarchy, tiling for SRAM, online softmax computation, recomputation strategy, FlashAttention complexity, FlashAttention benefits.

45m
125

FlashAttention Implementation

Covers CUDA kernel basics, memory access patterns, FlashAttention-2 improvements, using FlashAttention in PyTorch, FlashAttention limitations.

51m

Part XVIII: Long Context

7 chapters

Part XIX: Alternative Sequence and Generative Architectures

6 chapters
133

State Space Models for Language

Soon

Structured state spaces, selective state updates, Mamba-style sequence modeling, and linear-time claims.

134

Hybrid Attention and State Space Models

Soon

Jamba-style hybrids, layer allocation, memory-throughput trade-offs, and architecture ablations.

135

Token-Free and Byte-Level Language Models

Soon

Bytes, learned patches, dynamic segmentation, and the efficiency and robustness trade-offs of removing fixed tokenizers.

136

Diffusion Language Models

Soon

Discrete diffusion, denoising objectives, parallel refinement, controllability, and likelihood evaluation.

137

Blockwise and Semi-Autoregressive Generation

Soon

Generating multiple tokens per step, speculative blocks, verification, and quality-latency trade-offs.

138

Comparing Architecture Families Fairly

Soon

Matched-compute experiments, hardware efficiency, memory scaling, long-context quality, and benchmark leakage.

Part XX: Data Curation

10 chapters

Part XXI: Data Governance, Provenance, and Synthetic Ecosystems

6 chapters
149

Dataset Lineage and Provenance

Soon

Tracing documents through collection, filtering, mixing, training, evaluation, and model release.

150

Licensing, Consent, and Data-Use Signals

Soon

Licenses, terms, robots directives, creator consent, jurisdictional uncertainty, and defensible data decisions.

151

Attribution and Training-Data Transparency

Soon

Data statements, source disclosure, influence estimation, attribution limits, and communicating uncertainty.

152

Synthetic Data Pipelines

Soon

Generation, filtering, diversity control, verification, curriculum design, and provenance for synthetic corpora.

153

Recursive Training and Model Collapse

Soon

Feedback loops from model-generated data, tail loss, distribution drift, and mitigations for synthetic ecosystems.

154

Data Audits and Release Governance

Soon

Pre-training audits, sensitive-content review, documentation, approval gates, and post-release traceability.

Part XXII: Pre-training Objectives

7 chapters

Part XXIII: Scaling Laws

7 chapters

Part XXIV: BERT and Variants

8 chapters

Part XXV: Encoder-Decoder Models

6 chapters

Part XXVI: Multilingual Language Models and Cross-Lingual Transfer

7 chapters
183

Language Diversity, Typology, and Scripts

Soon

Writing systems, morphology, word order, language families, and why English-centric assumptions fail.

184

Multilingual Tokenization and Vocabulary Allocation

Soon

Fertility, script coverage, shared vocabularies, byte models, and unequal token costs across languages.

185

Multilingual Pre-training and Data Balancing

Soon

Sampling temperatures, capacity allocation, transfer-interference trade-offs, and multilingual mixture design.

186

Cross-Lingual Transfer and Alignment

Soon

Shared representations, zero-shot transfer, alignment objectives, adapters, and transfer diagnostics.

187

Low-Resource and Endangered Languages

Soon

Data scarcity, community participation, transliteration, active learning, and responsible evaluation.

188

Code-Switching and Mixed-Language Text

Soon

Language identification, mixed scripts, code-switched generation, evaluation, and deployment failure modes.

189

Cultural and Multilingual Evaluation

Soon

Translationese, construct validity, cultural knowledge, local harms, and evaluation led by native speakers.

Part XXVII: Machine Translation and Speech Translation

7 chapters
190

Statistical Machine Translation Foundations

Soon

Word alignment, phrase tables, language models, log-linear decoding, and the ideas inherited by neural MT.

191

Neural Machine Translation

Soon

Encoder-decoder translation, attention, transformer MT, training objectives, and decoding.

192

Parallel Data Mining and Quality Control

Soon

Bitext mining, alignment, filtering, back-translation, synthetic parallel data, and contamination checks.

193

Many-to-Many and Low-Resource Translation

Soon

Multilingual transfer, mixture-of-experts translation, zero-shot directions, and capacity bottlenecks.

194

Translation Evaluation Beyond BLEU

Soon

Learned metrics, adequacy, fluency, terminology, human evaluation, and metric failure modes.

195

Document-Level and Context-Aware Translation

Soon

Terminology consistency, discourse context, pronouns, document memory, and long-form evaluation.

196

Speech-to-Speech Translation

Soon

Cascaded and end-to-end systems, latency, speaker preservation, prosody, and multilingual safety.

Part XXVIII: GPT Architecture

10 chapters

Part XXIX: Modern Decoder Models

7 chapters

Part XXX: Emergent Capabilities

6 chapters

Part XXXI: Training Infrastructure

11 chapters

Part XXXII: Training Optimization

8 chapters

Part XXXIII: Mixture of Experts

10 chapters

Part XXXIV: Fine-tuning Fundamentals

5 chapters

Part XXXV: Parameter-Efficient Fine-tuning

12 chapters

Part XXXVI: Instruction Tuning

6 chapters

Part XXXVII: Alignment and RLHF

16 chapters
272

Alignment Problem

Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.

70m
273

Human Preference Data

Covers preference collection UI, comparison design, annotator guidelines, preference data quality.

67m
274

Bradley-Terry Model

Covers pairwise comparison model, preference probability, Bradley-Terry likelihood, preference strength.

63m
275

Reward Modeling

Covers reward model architecture, preference loss function, reward model training, reward model evaluation.

68m
276

Reward Hacking

Covers reward hacking examples, distribution shift, over-optimization, reward hacking mitigation.

68m
277

Policy Gradient Methods

Covers policy definition, REINFORCE algorithm, policy gradient derivation, variance reduction.

59m
278

PPO Algorithm

Covers clipped objective, PPO derivation, trust region intuition, PPO implementation.

71m
279

PPO for Language Models

Covers LLM as policy, action space (tokens), reward assignment, KL penalty importance.

64m
280

RLHF Pipeline

Covers SFT stage, reward model training, PPO fine-tuning, RLHF hyperparameters, RLHF debugging.

61m
281

KL Divergence Penalty

Covers KL penalty motivation, KL coefficient selection, adaptive KL, KL effects on training.

69m
282

DPO Concept

Covers DPO motivation, removing reward model, DPO intuition, DPO benefits.

53m
283

DPO Derivation

Covers DPO from RLHF objective, optimal policy derivation, DPO loss function, DPO as classification.

45m
284

DPO Implementation

Covers DPO data format, DPO loss computation, DPO training procedure, DPO hyperparameters.

56m
285

DPO Variants

Covers IPO formulation, KTO for unpaired feedback, ORPO, cDPO, comparing alignment methods.

65m
286

RLAIF

Covers AI as annotator, constitutional AI principles, AI preference generation, RLAIF scalability.

51m
287

Iterative Alignment

Covers iterative DPO, online preference learning, self-improvement loops, alignment stability.

48m

Part XXXVIII: Scalable Oversight and Safety Training

5 chapters
288

Constitutional and Principle-Based Alignment

Soon

Written principles, critique and revision, AI feedback, evaluation, and failure modes.

289

Scalable Oversight

Soon

Decomposition, debate, recursive supervision, weak-to-strong generalization, and oversight bottlenecks.

290

Process Supervision and Behavioral Specifications

Soon

Supervising intermediate behavior, specifying policies, resolving conflicts, and testing adherence.

291

Alignment Data Quality and Rater Populations

Soon

Rater disagreement, cultural pluralism, annotator effects, preference aggregation, and auditability.

292

Alignment Robustness and Distribution Shift

Soon

Sycophancy, reward tampering, jailbreak pressure, deployment drift, and adversarial evaluation.

Part XXXIX: Reasoning

7 chapters

Part XL: Reasoning Post-Training and Verifiable Rewards

7 chapters
300

Reinforcement Learning from Verifiable Rewards

Soon

Rule-based and executable rewards, correctness verification, reward design, and domains where RLVR works.

301

Group-Based Policy Optimization

Soon

GRPO-style objectives, relative advantages, stability, sampling costs, and implementation choices.

302

Cold Starts, Rejection Sampling, and Distillation

Soon

Bootstrapping reasoning behavior, filtering traces, iterative training, and transferring reasoning to smaller models.

303

Outcome and Process Reward Models

Soon

Sparse outcomes, step-level supervision, verifier reliability, credit assignment, and reward hacking.

304

Search, Reflection, and Test-Time Compute

Soon

Sampling, verifier-guided search, self-correction, budget allocation, stopping, and compute-optimal reasoning.

305

Faithful, Hidden, and Latent Reasoning

Soon

When visible chains of thought are explanations, when they are not, and alternatives for latent deliberation.

306

Overthinking and Reasoning Efficiency

Soon

Unnecessary deliberation, error amplification, confidence, adaptive budgets, and concise reasoning.

Part XLI: Model Compression

6 chapters

Part XLII: Inference Optimization

14 chapters
313

KV Cache

Covers KV cache motivation, cache structure, cache memory requirements, cache management.

56m
314

KV Cache Memory

Covers cache size calculation, batch size effects, sequence length effects, memory bottleneck.

52m
315

Paged Attention

Covers memory fragmentation problem, page-based allocation, vLLM approach, paged attention benefits.

68m
316

KV Cache Compression

Covers cache eviction strategies, attention sink preservation, H2O algorithm, cache quantization.

56m
317

Weight Quantization Basics

Covers quantization fundamentals, per-tensor vs per-channel, symmetric vs asymmetric, calibration.

59m
318

INT8 Quantization

Covers INT8 range mapping, absmax quantization, smooth quantization, INT8 accuracy.

59m
319

INT4 Quantization

Covers 4-bit challenges, group-wise quantization, 4-bit accuracy trade-offs, 4-bit formats.

59m
320

GPTQ

Covers GPTQ algorithm, layer-wise quantization, Hessian approximation, GPTQ implementation.

49m
321

AWQ

Covers salient weight preservation, AWQ algorithm, AWQ vs GPTQ, AWQ benefits.

44m
322

GGUF Format

Covers GGML/GGUF history, quantization types, GGUF file format, llama.cpp integration.

55m
323

Speculative Decoding: Fast LLM Inference Without Quality Loss

Covers speculative decoding concept, draft model selection, verification procedure, acceptance rate.

54m
324

Speculative Decoding Math: Algorithms & Speedup Limits

Covers acceptance criterion, expected speedup, draft quality effects, optimal draft length.

52m
325

Continuous Batching: Optimizing LLM Inference Throughput

Covers static vs continuous batching, iteration-level scheduling, request completion handling, throughput gains.

58m
326

LLM Inference Serving: Architecture, Routing & Auto-Scaling

Master LLM inference serving architecture, token-aware load balancing, and auto-scaling. Optimize time-to-first-token and throughput for production systems.

73m

Part XLIII: Small and On-Device Language Models

5 chapters
327

Small Language Model Design

Soon

Capacity allocation, data quality, architecture choices, and capability boundaries below frontier scale.

328

Hardware-Aware Model Optimization

Soon

Memory bandwidth, kernels, quantization formats, sparsity, and co-design for phones, browsers, and edge devices.

329

On-Device Adaptation and Personalization

Soon

Private fine-tuning, adapters, local retrieval, user control, and update strategies under tight budgets.

330

Hybrid Edge-Cloud Inference

Soon

Routing, partitioning, privacy boundaries, offline fallbacks, latency, and cost-aware orchestration.

331

Evaluating Small Models in Context

Soon

Task fit, energy, latency, privacy, reliability, and comparing systems rather than parameter counts alone.

Part XLIV: Retrieval-Augmented Generation

14 chapters
332

RAG Motivation: Solving Hallucinations & Knowledge Gaps

Discover why LLMs need Retrieval-Augmented Generation. Learn how RAG bridges knowledge gaps, reduces hallucinations, and enables non-parametric memory.

59m
333

RAG Architecture: Components, Timing & Design Patterns

Master RAG system design by exploring retriever-generator interactions, timing strategies like iterative retrieval, and architectural variations like RETRO.

60m
334

Dense Retrieval: Semantic Search & Bi-Encoder Implementation

Master dense retrieval for semantic search. Explore bi-encoder architectures, embedding metrics, and contrastive learning to overcome keyword limitations.

62m
335

Contrastive Learning for Retrieval: InfoNCE & DPR Guide

Master contrastive learning for dense retrieval. Learn to train models using InfoNCE loss, in-batch negatives, and hard negative mining strategies effectively.

57m
336

Document Chunking: Optimizing RAG Retrieval Pipelines

Master document chunking for RAG systems. Explore fixed-size, recursive, and semantic strategies to balance retrieval precision with context window limits.

63m
337

Embedding Models: Architecture, Pooling & Selection

Learn how embedding models convert text to vectors for RAG. Covers bi-encoder architecture, pooling strategies, dimensionality trade-offs, and model selection.

65m
338

Vector Similarity Search: Metrics & Approximate Methods

Explore vector similarity search for RAG systems. Compare cosine, dot product, and Euclidean metrics, and implement exact vs. approximate search with FAISS.

65m
339

HNSW Index: Architecture for Fast Vector Search

Master Hierarchical Navigable Small World (HNSW) graphs for vector search. Learn graph architecture, construction, and tuning for high-speed retrieval.

67m
340

IVF Index: Clustering-Based Vector Search & Partitioning

Master IVF indexes for scalable vector search. Learn clustering-based partitioning, nprobe tuning, and IVF-PQ compression for billion-scale retrieval.

66m
341

Product Quantization: Vector Compression for ANN Search

Learn how Product Quantization compresses embeddings up to 100x using learned codebooks and asymmetric distance computation for scalable vector search.

62m
342

Hybrid Search: BM25 and Dense Retrieval Combined

Learn how hybrid search fuses BM25 keyword retrieval with dense vector retrieval using reciprocal rank fusion and weighted score combination to improve recall.

65m
343

Reranking: Cross-Encoders for Precise Information Retrieval

Learn how reranking with cross-encoders solves bi-encoder limitations. Master two-stage retrieval, training strategies, and latency optimization for production search systems.

61m
344

RAG Prompt Engineering: Context Placement & Citation Strategies

Master RAG prompt engineering with strategic context placement, citation formats, and truncation strategies to improve LLM accuracy and reduce hallucinations.

53m
345

RAG Evaluation: Metrics for Retrieval and Generation Quality

Master RAG evaluation with metrics for retrieval quality (Precision@K, NDCG, MRR) and generation faithfulness using the RAGAS framework for AI systems.

65m

Part XLV: Knowledge, Memory, Editing, and Unlearning

6 chapters
346

Parametric and External Knowledge

Soon

What models store, what retrieval stores, and how to choose among prompting, RAG, editing, and retraining.

347

Knowledge Editing

Soon

Localized weight updates, memory-based editors, specificity, generalization, and multi-hop consistency.

348

Knowledge Freshness and Temporal Updates

Soon

Time-sensitive facts, temporal benchmarks, update propagation, versioning, and rollback.

349

Personalization and Long-Term Memory

Soon

User models, episodic and semantic memory, consent, retention, conflict resolution, and forgetting.

350

Machine Unlearning

Soon

Forget sets, retraining baselines, approximate removal, privacy goals, and the limits of verification.

351

Evaluating Knowledge Interventions

Soon

Efficacy, locality, generalization, side effects, privacy leakage, and auditable change histories.

Part XLVI: Continual Learning

5 chapters

Part XLVII: Code Generation

6 chapters

Part XLVIII: Tool Use and Agents

10 chapters
363

Tool Use Motivation: Why LLMs Need External Tools for Accuracy

Discover why LLMs require external tools to overcome knowledge cutoffs, computational limits, and hallucinations. Learn about tool-augmented AI systems.

57m
364

Function Calling: Structured Tool Use for Large Language Models

Learn how function calling enables LLMs to invoke external tools and APIs through structured JSON schemas, bridging natural language and executable code.

56m
365

ReAct Pattern: Interleaving Reasoning and Action for LLM Agents

Learn how the ReAct pattern enables LLM agents to interleave reasoning with tool execution. Master thought-action-observation loops for autonomous AI systems.

50m
366

Tool Selection for LLM Agents: Routing Strategies and Implementation

Master LLM tool selection through embedding-based routing, hybrid strategies, and semantic interfaces. Learn to build scalable multi-tool agent systems.

58m
367

Agent Architectures: Control Loops, State & Planning

Master LLM agent architectures including control loops, state management strategies, planning mechanisms, and termination conditions for autonomous AI systems.

53m
368

Agent Memory Systems: From Context to Persistent Storage

Learn how AI agents manage short-term context and long-term memory. Explore vector databases, retrieval algorithms, and memory hierarchies that enable persistent, learning systems.

52m
369

Agent Evaluation: Metrics, Benchmarks and Safety Standards

Learn to evaluate AI agents with task completion metrics, trajectory analysis, and safety testing. Covers WebArena, SWE-bench, GAIA, and OSWorld benchmarks.

52m
370

Planning

Covers task decomposition, goal-directed planning, plan execution, plan revision and recovery.

61m
371

Multi-Agent Systems

Covers agent coordination, communication protocols, role assignment, multi-agent benchmarks.

55m
372

Agent Safety

Covers unsafe action prevention, agent alignment, sandboxing, monitoring and intervention.

55m

Part XLIX: Long-Horizon Agents and Interoperability

8 chapters
373

Computer-Use and Browser Agents

Soon

Screenshots, accessibility trees, action spaces, grounding, recovery, and evaluation on real interfaces.

374

Coding and Software-Engineering Agents

Soon

Repository navigation, test-driven changes, execution feedback, code review, and long-running tasks.

375

Deep-Research Agents

Soon

Search, source selection, evidence synthesis, citation integrity, uncertainty, and reproducible research traces.

376

Durable and Asynchronous Agent Work

Soon

Checkpoints, resumability, queues, idempotency, retries, deadlines, and state across long-running workflows.

377

Model Context Protocol

Soon

MCP clients, servers, tools, resources, prompts, capability negotiation, transport, and trust boundaries.

378

Agent-to-Agent Protocols

Soon

Discovery, task delegation, status, artifacts, interoperability, and failure semantics across agents.

379

Human Approval and Control Boundaries

Soon

Approval gates, reversibility, escalation, authority, interruption, and clear ownership of consequential actions.

380

Long-Horizon Agent Evaluation

Soon

Functional success, partial credit, recovery, cost, time, safety, and contamination-resistant task suites.

Part L: Multimodal Models

12 chapters

Part LI: Speech and Audio

5 chapters

Part LII: Omni-Modal and Embodied Language Systems

7 chapters
398

Unified Audio-Visual-Language Modeling

Soon

Modality encoders, shared token spaces, fusion, generation, and interference in omni models.

399

Video and Temporal Reasoning

Soon

Frame sampling, event localization, temporal memory, causal reasoning, and long-video evaluation.

400

Streaming Multimodal Interaction

Soon

Continuous perception, interruption, turn-taking, proactive responses, latency, and duplex generation.

401

Document, Diagram, and Interface Understanding

Soon

OCR, layout, charts, diagrams, screenshots, grounding, and visually situated tool use.

402

Speech-to-Speech and Expressive Voice

Soon

Direct speech interaction, prosody, speaker identity, emotion, latency, and voice-specific safety.

403

Vision-Language-Action Models

Soon

Grounded language for robotics, action tokenization, policy learning, embodiment, and simulation-to-reality gaps.

404

Multimodal Safety and Evaluation

Soon

Cross-modal injection, privacy, deepfakes, modality imbalance, interactive benchmarks, and human evaluation.

Part LIII: Evaluation Fundamentals

10 chapters

Part LIV: Benchmark Evaluation

8 chapters

Part LV: Human and Model Evaluation

6 chapters

Part LVI: Hallucination and Factuality

6 chapters

Part LVII: Dynamic, Interactive, and Agentic Evaluation

6 chapters
435

Live and Continuously Updated Benchmarks

Soon

Fresh questions, private test sets, rolling releases, objective scoring, and contamination controls.

436

Evaluation Harnesses and Reproducibility

Soon

Versioned prompts, model settings, graders, confidence intervals, artifacts, and comparable reports.

437

Interactive and Long-Horizon Task Evaluation

Soon

Stateful environments, functional outcomes, partial credit, recovery, efficiency, and trajectory analysis.

438

Judge Calibration and Meta-Evaluation

Soon

Bias, self-preference, verbosity effects, rubric validity, judge ensembles, and agreement with experts.

439

Online Experiments and Production Evaluation

Soon

Shadow traffic, interleaving, A/B tests, guardrail metrics, drift, and linking offline scores to user outcomes.

440

Evaluation Governance

Soon

Benchmark access, disclosure, privacy, gaming, deprecation, audit trails, and responsible leaderboard use.

Part LVIII: Bias and Fairness

5 chapters

Part LIX: Interpretability

11 chapters

Part LX: Safety and Security

8 chapters

Part LXI: Agentic AI and Protocol Security

7 chapters
465

Threat Modeling Compound AI Systems

Soon

Assets, actors, trust boundaries, control flow, data flow, and failure propagation across model-centered systems.

466

Indirect Prompt Injection

Soon

Malicious instructions in documents, web pages, messages, tool output, and retrieved context.

467

Tool, Retrieval, and Memory Poisoning

Soon

Schema attacks, compromised tools, poisoned indexes, persistent memory attacks, and integrity controls.

468

Identity, Authorization, and Least Privilege

Soon

User and agent identity, scoped credentials, delegation, expiry, consent, and policy enforcement outside the model.

469

Secrets, Sandboxes, and Safe Execution

Soon

Credential isolation, capability sandboxes, network and filesystem boundaries, output validation, and containment.

470

Agent Supply Chains and Protocol Security

Soon

Tool registries, dependency tampering, server discovery, trust-on-first-use, signing, and update risk.

471

Audit, Detection, and Incident Response

Soon

Tamper-evident traces, anomaly detection, kill switches, investigation, rollback, and post-incident learning.

Part LXII: Human-AI Interaction and Governance

6 chapters
472

Designing Human-AI Collaboration

Soon

Task allocation, mixed initiative, handoffs, feedback, shared context, and preserving human agency.

473

Trust, Reliance, and Automation Bias

Soon

Calibrated reliance, overtrust, undertrust, explanations, uncertainty, and behavior under time pressure.

474

Oversight, Appeals, and Escalation

Soon

Review queues, contestability, escalation paths, documentation, and accountability for consequential decisions.

475

Accessibility and Inclusive Language AI

Soon

Disability access, literacy, language inclusion, participatory design, and evaluating who benefits or is excluded.

476

System Cards, Audits, and Impact Assessments

Soon

Documenting complete systems, evaluating downstream use, independent audits, and monitoring material changes.

477

Risk Governance Across the Lifecycle

Soon

Ownership, risk tiers, review gates, incident reporting, change management, and retirement.

Part LXIII: LLM Applications

7 chapters

Part LXIV: Compound AI System Design

6 chapters
485

Models, Retrieval, Tools, and Control Flow

Soon

Decomposing systems into components, explicit orchestration, typed interfaces, and failure boundaries.

486

Routing, Fallbacks, and Graceful Degradation

Soon

Model cascades, confidence-based routing, deterministic fallbacks, abstention, and partial service.

487

Evaluation-Driven Development

Soon

Task suites, traces, regression gates, error taxonomies, and improving systems without benchmark theater.

488

Tracing and Observability for AI Systems

Soon

Prompt, retrieval, tool, latency, cost, and quality traces with privacy-aware retention.

489

Versioning Models, Prompts, Data, and Tools

Soon

Coordinated releases, compatibility, reproducibility, canaries, rollback, and artifact lineage.

490

Reliability and Cost Engineering

Soon

Service objectives, retries, idempotency, concurrency, capacity, caching, budgets, and failure testing.

Part LXV: Production Systems

9 chapters

Part LXVI: Frontier Methods of 2025

8 chapters

Part LXVII: Frontier Outlook from 2025

6 chapters

In Progress

This comprehensive handbook is currently in development. Each chapter will be published as it's completed, with practical examples, code implementations, and real-world applications.

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@book{brenndoerfer2025languageaihandbook, author = {Michael Brenndoerfer}, title = {Language AI Handbook}, year = {2025}, url = {https://mbrenndoerfer.com/books/language-ai-handbook}, publisher = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2025). Language AI Handbook. Retrieved from https://mbrenndoerfer.com/books/language-ai-handbook
MLAAcademic
Michael Brenndoerfer. "Language AI Handbook." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/books/language-ai-handbook>.
CHICAGOAcademic
Michael Brenndoerfer. "Language AI Handbook." Accessed September 27, 2026. https://mbrenndoerfer.com/books/language-ai-handbook.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Language AI Handbook'. Available at: https://mbrenndoerfer.com/books/language-ai-handbook (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Language AI Handbook. https://mbrenndoerfer.com/books/language-ai-handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.