Language AI Handbook

Read Language AI Handbook free online
The Language AI Handbook connects the pieces of modern language AI. We begin with classical NLP, build up to transformers and LLM training, then cover retrieval, evaluation, safety, and production deployment.
Begin with the fundamentals that never go out of style: tokenization, embeddings, and the statistical foundations that inform modern approaches. Then dive deep into the transformer architecture. Learn not just how to use it, but how it actually works. Understand self-attention mathematically, grasp why positional encodings matter, and see how architectural choices like layer normalization affect training dynamics.
Author and edition details
About the author and this edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- First edition
- Published
- Last reviewed
Track Your Progress
Sign in to mark chapters as complete, track quiz scores, and see your reading journey
Text as Data
Classical Text Representations
Distributional Semantics
Word Embeddings
Subword Tokenization
Sequence Labeling
Linguistic Form: Morphology and Syntax
Semantics and Information Extraction
Coreference, Discourse, Pragmatics, and Dialogue
Neural Network Foundations
Recurrent Neural Networks
Sequence-to-Sequence
Self-Attention
Positional Encoding
Transformer Blocks
Transformer Architectures
Efficient Attention
Long Context
Alternative Sequence and Generative Architectures
Data Curation
Data Governance, Provenance, and Synthetic Ecosystems
Pre-training Objectives
Scaling Laws
BERT and Variants
Encoder-Decoder Models
Multilingual Language Models and Cross-Lingual Transfer
Machine Translation and Speech Translation
GPT Architecture
Modern Decoder Models
Emergent Capabilities
Training Infrastructure
Training Optimization
Mixture of Experts
Fine-tuning Fundamentals
Parameter-Efficient Fine-tuning
Instruction Tuning
Alignment and RLHF
Scalable Oversight and Safety Training
Reasoning
Reasoning Post-Training and Verifiable Rewards
Model Compression
Inference Optimization
Small and On-Device Language Models
Retrieval-Augmented Generation
Knowledge, Memory, Editing, and Unlearning
Continual Learning
Code Generation
Tool Use and Agents
Long-Horizon Agents and Interoperability
Multimodal Models
Speech and Audio
Omni-Modal and Embodied Language Systems
Evaluation Fundamentals
Benchmark Evaluation
Human and Model Evaluation
Hallucination and Factuality
Dynamic, Interactive, and Agentic Evaluation
Bias and Fairness
Interpretability
Safety and Security
Agentic AI and Protocol Security
Human-AI Interaction and Governance
LLM Applications
Compound AI System Design
Production Systems
Frontier Methods of 2025
Frontier Outlook from 2025
The whole growing library, yours forever.
All 395 written chapters across 19 available volumes
19 written volume PDFs plus the Volume 3 placeholder. Files are delivered separately.
All future expansions, updates, and errata, free, forever
Every future volume included, including 104 planned chapters
Downloadable editions for durable offline reading
$199
instead of $456 for the 19 available volume PDFs bought separately

Secure checkout via Stripe · Individual PDF links delivered to your inbox
Reader verdict
5.0
out of 5
12 five-star reader reviews
Genuine feedback, with names abbreviated for privacy.
Meet the community“Really appreciate the way you have explained the concepts. Simple to understand, yet builds up knowledge exponentially, all while reinforcing it with examples repeatedly. I've read the last two chapters `74.Self-attention & 75.Q,K,V` and I've to admit, my concepts have never been clearer. You really are doing an amazing job breaking down maths and complex concepts like a story!”
“I found your book Language AI Handbook's explanations of Transformer architectures and production deployment strategies to be exceptionally clear and insightful”
“I came across one of your books, Language AI Handbook, and man, I'm loving it. I've read many books but have never found such a comprehensive book on large models that covers everything from start to finish with such depth.”
Read the other 9 reader reviews
“I would like to thank you for your work. Specifically, the chapter: 'LIME Explainability: Complete Guide to Local Interpretable Model-Agnostic Explanations' was extremely helpful for me in understanding LIME and communicating to colleagues. Your work is inspiring and valuable. Thank you for doing it.”
“Thanks so much for establishing such an amazing community. I appreciate the opportunity to be a learner and a contributor to the community.”
“I just came across your website where your books are available for free to read. Thank you so much for making this available to the public. This will help a lot of people.”
“I am enjoying reading your writings. I appreciate the effort you put into it. I want to drop a note of thanks. Thank you!”
“Thank you Michael. It's top!”
“Thanks for putting cool stuff out into the world.”
“Your writing feels like a treasure trove of great material. So, thank you for compiling everything on your website!”
“I came across your collection of online handbooks covering quantitative finance, AI, and machine learning. I wanted to reach out and thank you for making such high-quality material openly available. They are incredibly helpful for my CFA Level II review and for brushing up ahead of ML interviews.”
“I read a few of your blogs on NLP and oh WOW, they really are something, now I exactly know how all of the terms Entropy, Cross-entropy and how it is connected to Perplexity, will definitely read more of your blogs.”
Prefer to start with a single volume?
Buy 9 volumes and all-volume access unlocks free. Already own volumes? Upgrade to all-volume access and you only pay the difference.
Table of Contents
Part I: Text as Data
5 chapters
Part I: Text as Data
Character Encoding
Covers ASCII origins and 7-bit limitations, Unicode code points and planes, UTF-8 variable-width encoding scheme, byte order marks and endianness, encoding detection heuristics, common encoding errors and mojibake, practical encoding/decoding in Python.
Text Normalization
Covers Unicode normalization forms (NFC, NFD, NFKC, NFKD), case folding vs lowercasing, accent and diacritic handling, whitespace normalization, ligature expansion, full-width to half-width conversion, implementing a normalization pipeline.
Regular Expressions
Covers regex syntax and metacharacters, character classes and quantifiers, grouping and backreferences, lookahead and lookbehind assertions, greedy vs lazy matching, common NLP patterns (emails, URLs, dates), regex performance considerations.
Sentence Segmentation
Covers period disambiguation challenges, abbreviation handling, rule-based boundary detection, Punkt sentence tokenizer algorithm, evaluation metrics for segmentation, handling edge cases (quotes, parentheses, lists), multilingual segmentation issues.
Word Tokenization
Covers whitespace tokenization limitations, punctuation handling rules, contractions and clitics, language-specific challenges (Chinese, Japanese, German compounds), Penn Treebank tokenization standard, building a rule-based tokenizer, tokenization evaluation.
Part II: Classical Text Representations
9 chapters
Part II: Classical Text Representations
Bag of Words
Covers document-term matrix construction, vocabulary building from corpus, word counting and frequency vectors, sparse matrix representation (CSR/CSC formats), vocabulary pruning (min_df, max_df), binary vs count representations, limitations of word order loss.
N-grams
Covers bigram and trigram extraction, n-gram vocabulary explosion, n-gram frequency distributions, Zipf's law in n-grams, character n-grams for robustness, skip-grams and flexible windows, n-gram indexing for search.
N-gram Language Models
Covers Markov assumption and chain rule, maximum likelihood estimation, probability calculation for sequences, handling unseen n-grams, start and end tokens, generating text from n-gram models, model storage and lookup efficiency.
Smoothing Techniques
Covers add-one (Laplace) smoothing, add-k smoothing and tuning, Good-Turing smoothing derivation, Kneser-Ney smoothing intuition and formula, interpolation vs backoff, modified Kneser-Ney, comparing smoothing methods empirically.
Perplexity
Covers cross-entropy definition and derivation, perplexity as branching factor, relationship to bits-per-character, held-out evaluation methodology, perplexity vs downstream performance, comparing models with perplexity, perplexity limitations and caveats.
Term Frequency
Covers raw term frequency, log-scaled term frequency, boolean term frequency, augmented term frequency, L2-normalized frequency vectors, term frequency sparsity patterns, efficient term frequency computation.
Inverse Document Frequency
Covers document frequency calculation, IDF formula derivation, IDF intuition (rare words matter more), smoothed IDF variants, IDF across corpus splits, relationship to information theory, implementing IDF efficiently.
TF-IDF
Covers TF-IDF formula and variants, TF-IDF vector computation, TF-IDF normalization options, BM25 as TF-IDF extension, document similarity with TF-IDF, TF-IDF for feature extraction, sklearn TfidfVectorizer deep dive.
BM25
Covers BM25 derivation from probabilistic IR, saturation parameter k1, length normalization parameter b, BM25+ and BM25L variants, field-weighted BM25, implementing BM25 scoring, BM25 vs TF-IDF empirically.
Part III: Distributional Semantics
4 chapters
Part III: Distributional Semantics
The Distributional Hypothesis
Covers Firth's "you shall know a word by the company it keeps," distributional similarity intuition, context window definitions, paradigmatic vs syntagmatic relations, word similarity from distributions, limitations of distributional semantics.
Co-occurrence Matrices
Covers word-word co-occurrence matrices, word-document matrices, context window size effects, weighting by distance, symmetric vs directional contexts, matrix sparsity patterns, efficient construction algorithms.
Pointwise Mutual Information
Covers PMI formula derivation, PMI interpretation as association, positive PMI (PPMI), shifted PPMI variants, PMI matrix properties, PMI vs raw counts comparison, PMI for collocation extraction.
Singular Value Decomposition
Covers SVD mathematical formulation, truncated SVD for dimensionality reduction, LSA (Latent Semantic Analysis), choosing embedding dimensions, SVD computational complexity, randomized SVD for scale, interpreting SVD dimensions.
Part IV: Word Embeddings
9 chapters
Part IV: Word Embeddings
Skip-gram Model
Covers skip-gram architecture diagram, input/output representations, softmax over vocabulary, skip-gram objective function, training data generation, window size hyperparameter, skip-gram vs CBOW intuition.
CBOW Model
Covers CBOW architecture, context word averaging, CBOW objective function, CBOW vs skip-gram training speed, CBOW for frequent words, implementing CBOW forward pass, CBOW gradient derivation.
Negative Sampling
Covers softmax computational bottleneck, negative sampling objective derivation, sampling distribution (unigram^0.75), number of negatives hyperparameter, negative sampling gradient computation, NCE vs negative sampling, implementing efficient sampling.
Hierarchical Softmax
Covers binary tree construction (Huffman coding), path probability computation, hierarchical softmax objective, gradient computation along paths, tree structure impact on learning, hierarchical softmax vs negative sampling, when to use each approach.
Word2Vec Training
Covers data preprocessing pipeline, subsampling frequent words, learning rate scheduling, minibatch vs online training, convergence monitoring, gensim Word2Vec usage, training from scratch in PyTorch.
Word Analogy
Covers vector arithmetic for analogies, parallelogram model, analogy evaluation datasets, 3CosAdd vs 3CosMul methods, analogy accuracy metrics, limitations of analogy evaluation, what analogies reveal about embeddings.
GloVe
Covers GloVe objective function derivation, weighted least squares formulation, relationship to matrix factorization, weighting function design, bias terms in GloVe, GloVe vs Word2Vec comparison, training GloVe efficiently.
FastText
Covers character n-gram representation, word vector as n-gram sum, FastText architecture, handling OOV words, morphological awareness, FastText for morphologically rich languages, training FastText models.
Embedding Evaluation
Covers intrinsic vs extrinsic evaluation, word similarity datasets (SimLex, WordSim), analogy accuracy, embedding visualization (t-SNE, UMAP), downstream task evaluation, embedding bias detection, evaluation pitfalls.
Part V: Subword Tokenization
8 chapters
Part V: Subword Tokenization
The Vocabulary Problem
Covers OOV word problem, vocabulary size explosion, rare word representation, morphological productivity, compound words, code and technical text, the case for subword units.
Byte Pair Encoding
Covers BPE algorithm step-by-step, merge rules learning, vocabulary size control, BPE encoding procedure, BPE decoding procedure, BPE implementation from scratch, BPE hyperparameters.
WordPiece
Covers WordPiece vs BPE differences, likelihood objective for merges, greedy tokenization algorithm, ## prefix notation, WordPiece in BERT, training WordPiece tokenizers, handling unknown characters.
Unigram Language Model Tokenization
Covers unigram LM formulation, EM algorithm for training, Viterbi decoding for tokenization, sampling multiple segmentations, subword regularization, unigram vs BPE comparison, SentencePiece unigram mode.
SentencePiece
Covers treating text as raw bytes, whitespace handling (▁ prefix), BPE and unigram modes, training from raw text, pretokenization elimination, SentencePiece in production, multilingual tokenization.
Tokenizer Training
Covers corpus preparation, vocabulary size selection, special tokens configuration, training with HuggingFace tokenizers, saving and loading tokenizers, tokenizer versioning, domain-specific tokenizers.
Special Tokens
Covers [CLS], [SEP], [PAD], [MASK], [UNK] tokens, beginning/end of sequence tokens, custom special tokens, special token embeddings, token type IDs, handling special tokens in generation.
Tokenization Challenges
Covers number tokenization issues, code tokenization, multilingual text mixing, emoji and Unicode edge cases, tokenization artifacts, adversarial tokenization, measuring tokenization quality.
Part VI: Sequence Labeling
8 chapters
Part VI: Sequence Labeling
Part-of-Speech Tagging
Covers POS tag sets (Penn Treebank, Universal), POS tagging as classification, contextual disambiguation, POS tagging accuracy metrics, POS tagging for downstream tasks, rule-based vs statistical taggers.
Named Entity Recognition
Covers entity types (PER, ORG, LOC, etc.), NER as sequence labeling, nested entity challenges, entity boundary detection, NER evaluation (exact vs partial match), NER datasets and benchmarks.
BIO Tagging
Covers BIO scheme explanation, BIOES/BILOU variants, converting spans to BIO tags, BIO decoding to spans, handling tagging inconsistencies, BIO for multi-label scenarios, implementing BIO utilities.
Chunking
Covers noun phrase chunking, chunk types (NP, VP, PP), IOB tagging for chunks, chunking vs full parsing, chunking evaluation, chunking as preprocessing, regex chunking with NLTK.
Hidden Markov Models
Covers HMM components (states, observations, transitions), emission and transition probabilities, HMM assumptions (Markov, independence), HMM for POS tagging, HMM parameter estimation, HMM limitations for NLP.
Viterbi Algorithm
Covers optimal path problem formulation, Viterbi recursion derivation, backpointer tracking, Viterbi complexity analysis, log-space computation, implementing Viterbi efficiently, Viterbi for beam search foundation.
Conditional Random Fields
Covers CRF vs HMM comparison, CRF feature functions, log-linear formulation, partition function computation, CRF for NER, CRF inference complexity, neural CRF layers.
CRF Training
Covers CRF log-likelihood objective, forward-backward algorithm, gradient computation, L-BFGS optimization, feature template design, CRF regularization, CRF training convergence.
Part VII: Linguistic Form: Morphology and Syntax
6 chapters
Part VII: Linguistic Form: Morphology and Syntax
Morphology and Word Formation
Morphemes, inflection, derivation, compounding, and why word structure matters for language models.
Morphological Analysis and Generation
Finite-state and neural approaches to analyzing and generating morphologically rich language.
Grammars and Constituency Structure
SoonContext-free grammars, phrase structure, constituency trees, and the limits of symbolic grammars.
Dependency Syntax and Universal Dependencies
SoonHead-dependent relations, dependency trees, projectivity, and cross-lingual annotation with Universal Dependencies.
Constituency and Dependency Parsing
SoonTransition-based, graph-based, chart, and neural parsing methods with their decoding trade-offs.
Structured Prediction and Parsing Evaluation
SoonDynamic programming, constrained decoding, labeled attachment, span scores, and meaningful parser error analysis.
Part VIII: Semantics and Information Extraction
6 chapters
Part VIII: Semantics and Information Extraction
Lexical Semantics and Word Senses
SoonPolysemy, synonymy, lexical resources, contextual word senses, and word-sense disambiguation.
Sentence Meaning and Natural Language Inference
SoonCompositional meaning, entailment, contradiction, presupposition, and robust NLI evaluation.
Semantic Roles and Argument Structure
SoonPredicates, arguments, thematic roles, frame semantics, and semantic role labeling.
Relation, Event, and Temporal Extraction
SoonExtracting relations, events, arguments, and timelines from documents rather than isolated sentences.
Entity Linking and Knowledge Graph Construction
SoonLinking mentions to entities, resolving ambiguity, and building provenance-aware knowledge graphs.
Semantic Parsing and Meaning Representations
SoonLogical forms, AMR, text-to-SQL, executable meaning representations, and constrained generation.
Part IX: Coreference, Discourse, Pragmatics, and Dialogue
5 chapters
Part IX: Coreference, Discourse, Pragmatics, and Dialogue
Coreference Resolution
SoonMention detection, entity clusters, pronoun resolution, long-document coreference, and evaluation.
Discourse Relations and Coherence
SoonDiscourse structure, rhetorical relations, coherence modeling, and long-form generation failures.
Pragmatics, Implicature, and Presupposition
SoonSpeaker intent, common ground, deixis, implicature, and what literal next-token prediction misses.
Dialogue Acts and Conversation Structure
SoonTurn-taking, grounding, repair, dialogue acts, state tracking, and multi-party conversation.
Argumentation, Stance, and Persuasion
SoonClaims, evidence, stance, argument structure, persuasion, and responsible evaluation.
Part X: Neural Network Foundations
13 chapters
Part X: Neural Network Foundations
Linear Classifiers
Covers linear decision boundaries, weight vectors and bias, dot product interpretation, multiclass classification (softmax), linear classifier limitations, training with gradient descent.
Activation Functions
Covers sigmoid function and saturation, tanh properties, ReLU and dying ReLU, Leaky ReLU and PReLU, ELU and SELU, GELU derivation and properties, Swish and Mish, choosing activation functions.
Multilayer Perceptrons
Covers hidden layers and depth, weight matrices between layers, forward pass computation, representational capacity, MLP for classification, MLP for regression, MLP architecture design.
Loss Functions
Covers cross-entropy loss derivation, MSE for regression, binary vs multiclass cross-entropy, label smoothing, focal loss for imbalance, loss function numerical stability, custom loss functions.
Backpropagation
Covers computational graphs, chain rule review, forward and backward pass, gradient accumulation, backprop complexity analysis, automatic differentiation, implementing backprop from scratch.
Stochastic Gradient Descent
Covers batch vs stochastic gradient descent, minibatch gradient descent, learning rate selection, SGD convergence properties, SGD noise as regularization, learning rate schedules basics, SGD implementation.
Momentum
Covers momentum intuition (ball rolling), momentum update equations, momentum coefficient selection, dampening oscillations, momentum vs vanilla SGD, Nesterov momentum derivation, implementing momentum.
Adam Optimizer
Covers exponential moving averages, first moment (mean) estimation, second moment (variance) estimation, bias correction derivation, Adam update rule, Adam hyperparameters, Adam convergence properties.
AdamW
Covers L2 regularization vs weight decay, why they differ with Adam, AdamW formulation, weight decay coefficient selection, AdamW as default optimizer, AdamW vs Adam empirically.
Weight Initialization
Covers random initialization importance, Xavier/Glorot initialization derivation, He initialization for ReLU, initialization for different activations, layer-wise initialization, initialization debugging, modern initialization practices.
Batch Normalization
Covers internal covariate shift, batch statistics computation, learnable scale and shift, training vs inference mode, batch norm gradient flow, batch norm placement debates, batch norm limitations.
Dropout
Covers dropout as ensemble, dropout mask sampling, inverted dropout scaling, dropout rate selection, dropout at inference, spatial dropout for sequences, dropout in modern architectures.
Gradient Clipping
Covers gradient explosion detection, clip by value, clip by global norm, gradient clipping implementation, when to use gradient clipping, clipping threshold selection, monitoring gradient norms.
Part XI: Recurrent Neural Networks
9 chapters
Part XI: Recurrent Neural Networks
RNN Architecture
Covers recurrent connection intuition, hidden state as memory, unrolled computation graph, parameter sharing across time, RNN for sequence classification, RNN for sequence generation, RNN equations and dimensions.
Backpropagation Through Time
Covers BPTT derivation, gradient flow through time, truncated BPTT, BPTT memory requirements, BPTT implementation, gradient accumulation across timesteps.
Vanishing Gradients
Covers gradient product across timesteps, vanishing gradient analysis, long-range dependency failure, gradient visualization, vanishing vs exploding trade-off, architectural solutions overview.
LSTM Architecture
Covers cell state as information highway, gate mechanism intuition, LSTM diagram walkthrough, information flow in LSTMs, LSTM for long sequences, LSTM memory capacity.
LSTM Gate Equations
Covers forget gate equations, input gate equations, cell state update, output gate equations, hidden state computation, LSTM parameter count, implementing LSTM from scratch.
LSTM Gradient Flow
Covers constant error carousel, forget gate gradient highway, gradient flow analysis, LSTM vs vanilla RNN gradients, peephole connections, LSTM gradient clipping needs.
GRU Architecture
Covers GRU vs LSTM comparison, reset gate function, update gate function, candidate hidden state, GRU equations, GRU parameter efficiency, when to choose GRU vs LSTM.
Bidirectional RNNs
Covers forward and backward passes, hidden state concatenation, bidirectional architectures, bidirectionality for classification, limitations for generation, implementing bidirectional RNNs.
Stacked RNNs
Covers multiple RNN layers, residual connections for depth, layer normalization in RNNs, depth vs width trade-offs, gradient flow in deep RNNs, practical depth limits.
Part XII: Sequence-to-Sequence
7 chapters
Part XII: Sequence-to-Sequence
Encoder-Decoder Framework
Covers encoder role and design, decoder role and design, context vector as bottleneck, seq2seq for machine translation, seq2seq for summarization, seq2seq training setup.
Teacher Forcing
Covers teacher forcing procedure, exposure bias problem, teacher forcing efficiency, scheduled sampling, curriculum learning, teacher forcing vs autoregressive training.
Beam Search
Covers greedy decoding limitations, beam search algorithm, beam width selection, length normalization, diverse beam search, beam search implementation, beam search vs sampling.
Attention Intuition
Covers attention as soft lookup, attention weight interpretation, attention for variable-length inputs, attention visualization, attention vs pooling, attention computation overview.
Bahdanau Attention
Covers alignment model formulation, score function (additive), attention weight computation, context vector as weighted sum, attention in decoder, Bahdanau attention implementation.
Luong Attention
Covers dot product attention, general (bilinear) attention, concat attention variant, global vs local attention, Luong vs Bahdanau comparison, attention placement (input vs output).
Copy Mechanism
Covers pointer network motivation, copy probability computation, mixing generation and copying, pointer-generator networks, copy mechanism for summarization, OOV handling with copy.
Part XIII: Self-Attention
6 chapters
Part XIII: Self-Attention
Self-Attention Concept
Covers cross-attention vs self-attention, self-attention motivation, all-pairs interaction, self-attention for representation learning, self-attention computational pattern.
Query, Key, Value
Covers QKV intuition (database lookup), projection matrices Wq, Wk, Wv, query-key matching, value retrieval, QKV dimensions and shapes, QKV as learned transformations.
Scaled Dot-Product Attention
Covers dot product for similarity, softmax for normalization, scaling factor derivation (1/√dk), attention output computation, attention in matrix form, attention implementation.
Attention Masking
Covers padding masks, causal (look-ahead) masks, combining multiple masks, mask shapes and broadcasting, efficient masking implementation, custom attention patterns.
Multi-Head Attention
Covers multiple attention heads motivation, head dimension splitting, parallel attention computation, output concatenation and projection, head specialization, multi-head vs single head.
Attention Complexity
Covers O(n²) attention complexity, memory requirements, attention bottleneck in long sequences, FLOPs computation, attention vs RNN complexity, practical scaling limits.
Part XIV: Positional Encoding
7 chapters
Part XIV: Positional Encoding
Position Problem
Covers transformer position blindness, why position matters for language, position information requirements, position encoding vs position embedding, absolute vs relative position.
Sinusoidal Position Encoding
Covers sinusoidal formula derivation, wavelength intuition, position encoding visualization, extrapolation properties, sinusoidal encoding implementation, learned vs sinusoidal trade-offs.
Learned Position Embeddings
Covers position embedding table, position embedding training, maximum sequence length, learned embedding extrapolation, position embedding analysis, GPT-style position embeddings.
Relative Position Encoding
Covers relative position motivation, relative attention formulation, clipping relative positions, relative position in self-attention, Shaw et al. relative positions, relative bias implementation.
Rotary Position Embedding (RoPE)
Covers RoPE intuition, rotation matrix formulation, RoPE in complex numbers, relative position through rotation, RoPE implementation, RoPE frequency patterns.
ALiBi
Covers ALiBi motivation, linear bias by distance, head-specific slopes, ALiBi extrapolation properties, ALiBi simplicity advantages, ALiBi vs RoPE comparison.
Position Encoding Comparison
Covers extrapolation benchmarks, training efficiency comparison, implementation complexity, position encoding for long context, hybrid approaches, current best practices.
Part XV: Transformer Blocks
8 chapters
Part XV: Transformer Blocks
Residual Connections
Covers residual connection formulation, gradient highway interpretation, residual scaling, residual connections in transformers, pre-norm vs post-norm residuals.
Layer Normalization
Covers layer norm vs batch norm, layer norm formula, learnable affine parameters, layer norm placement, layer norm gradient flow, layer norm implementation.
RMSNorm
Covers RMSNorm derivation, removing mean centering, RMSNorm efficiency, RMSNorm vs LayerNorm performance, RMSNorm in modern architectures.
Pre-Norm vs Post-Norm
Covers original transformer (post-norm), pre-norm formulation, training stability comparison, gradient flow differences, when to use each, modern consensus.
Feed-Forward Networks
Covers FFN architecture, hidden dimension expansion, FFN as two linear layers, position independence, FFN parameter count, FFN computational cost.
FFN Activation Functions
Covers ReLU in original transformer, GELU adoption, GELU approximations, SiLU/Swish in modern models, activation function comparison.
Gated Linear Units
Covers GLU formulation, gating mechanism, SwiGLU derivation, GeGLU variant, GLU parameter efficiency, GLU in modern architectures.
Transformer Block Assembly
Covers standard block structure, component ordering, block implementation, block initialization, forward pass walkthrough, block hyperparameters.
Part XVI: Transformer Architectures
6 chapters
Part XVI: Transformer Architectures
Encoder Architecture
Covers encoder-only design, bidirectional self-attention, encoder for understanding tasks, encoder output usage, BERT-style encoder, encoder layer stacking.
Decoder Architecture
Covers decoder-only design, causal masking requirement, autoregressive generation, decoder for generation tasks, GPT-style decoder, decoder layer stacking.
Encoder-Decoder Architecture
Covers encoder-decoder interaction, cross-attention mechanism, encoder-decoder for seq2seq, T5-style architecture, information flow, when to use encoder-decoder.
Cross-Attention
Covers cross-attention formulation, KV from encoder, Q from decoder, cross-attention masking, cross-attention placement, cross-attention implementation.
Weight Tying
Covers input-output embedding tying, encoder-decoder tying, parameter reduction, weight tying effects on training, when to tie weights.
Architecture Hyperparameters
Covers depth vs width trade-offs, number of heads selection, hidden dimension ratios, FFN expansion ratio, total parameter calculation, architecture search.
Part XVII: Efficient Attention
9 chapters
Part XVII: Efficient Attention
Quadratic Attention Bottleneck
Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.
Sparse Attention Patterns
Covers local attention windows, strided attention patterns, block-sparse attention, combining sparse patterns, sparse attention implementation.
Sliding Window Attention
Covers sliding window formulation, window size selection, dilated sliding windows, sliding window for long sequences, Mistral-style windowed attention.
Global Tokens
Covers CLS token global attention, learned global tokens, global-local attention mixing, global token count, implementation strategies.
Longformer
Covers Longformer attention pattern, global attention configuration, Longformer complexity, Longformer for documents, Longformer implementation.
BigBird
Covers BigBird attention pattern, random attention benefits, BigBird theoretical guarantees, BigBird vs Longformer, BigBird applications.
Linear Attention
Covers softmax attention reformulation, kernel feature maps, linear complexity attention, linear attention limitations, Performer and variants.
FlashAttention Algorithm
Covers GPU memory hierarchy, tiling for SRAM, online softmax computation, recomputation strategy, FlashAttention complexity, FlashAttention benefits.
FlashAttention Implementation
Covers CUDA kernel basics, memory access patterns, FlashAttention-2 improvements, using FlashAttention in PyTorch, FlashAttention limitations.
Part XVIII: Long Context
7 chapters
Part XVIII: Long Context
Context Length Challenges
Covers training sequence length limits, attention memory scaling, position encoding extrapolation, long-range dependency learning, evaluation challenges.
Position Interpolation
Covers linear position scaling, interpolation vs extrapolation, position interpolation implementation, fine-tuning for longer context, interpolation limitations.
NTK-aware Scaling
Covers RoPE frequency analysis, high-frequency preservation, NTK-aware formula, dynamic NTK scaling, NTK vs linear interpolation.
YaRN
Covers YaRN motivation, attention scaling factor, YaRN formula, YaRN training requirements, YaRN vs alternatives.
Attention Sinks
Covers attention sink phenomenon, StreamingLLM approach, sink token design, streaming inference, infinite context generation.
Memory Augmentation
Covers memory network concepts, memory retrieval mechanisms, memory writing and updating, memory-augmented transformers, Memorizing Transformers.
Recurrent Memory
Covers Transformer-XL approach, segment-level processing, recurrent state passing, relative position in recurrence, recurrent memory limitations.
Part XIX: Alternative Sequence and Generative Architectures
6 chapters
Part XIX: Alternative Sequence and Generative Architectures
State Space Models for Language
SoonStructured state spaces, selective state updates, Mamba-style sequence modeling, and linear-time claims.
Hybrid Attention and State Space Models
SoonJamba-style hybrids, layer allocation, memory-throughput trade-offs, and architecture ablations.
Token-Free and Byte-Level Language Models
SoonBytes, learned patches, dynamic segmentation, and the efficiency and robustness trade-offs of removing fixed tokenizers.
Diffusion Language Models
SoonDiscrete diffusion, denoising objectives, parallel refinement, controllability, and likelihood evaluation.
Blockwise and Semi-Autoregressive Generation
SoonGenerating multiple tokens per step, speculative blocks, verification, and quality-latency trade-offs.
Comparing Architecture Families Fairly
SoonMatched-compute experiments, hardware efficiency, memory scaling, long-context quality, and benchmark leakage.
Part XX: Data Curation
10 chapters
Part XX: Data Curation
Web Crawling
Covers Common Crawl, crawling strategies, robots.txt respect, crawl freshness.
Document Extraction
Covers HTML parsing, boilerplate removal, content extraction, trafilatura and similar tools.
Language Identification
Covers language ID models, multilingual document handling, code-switching, language filtering.
Deduplication
Covers exact deduplication, near-duplicate detection, document vs substring dedup, dedup at scale.
MinHash
Covers MinHash algorithm, Jaccard similarity estimation, MinHash LSH, MinHash implementation.
Quality Filtering
Covers heuristic filters, perplexity filtering, classifier-based filtering, filter thresholds.
Toxicity Filtering
Covers toxicity classifiers, toxicity thresholds, over-filtering risks, toxicity filter evaluation.
PII Removal
Covers PII detection methods, PII removal strategies, PII removal evaluation, privacy preservation.
Data Mixing
Covers domain proportions, quality weighting, data mixing experiments, optimal mixing.
Synthetic Data
Covers synthetic data generation, quality verification, synthetic data diversity, distillation.
Part XXI: Data Governance, Provenance, and Synthetic Ecosystems
6 chapters
Part XXI: Data Governance, Provenance, and Synthetic Ecosystems
Dataset Lineage and Provenance
SoonTracing documents through collection, filtering, mixing, training, evaluation, and model release.
Licensing, Consent, and Data-Use Signals
SoonLicenses, terms, robots directives, creator consent, jurisdictional uncertainty, and defensible data decisions.
Attribution and Training-Data Transparency
SoonData statements, source disclosure, influence estimation, attribution limits, and communicating uncertainty.
Synthetic Data Pipelines
SoonGeneration, filtering, diversity control, verification, curriculum design, and provenance for synthetic corpora.
Recursive Training and Model Collapse
SoonFeedback loops from model-generated data, tail loss, distribution drift, and mitigations for synthetic ecosystems.
Data Audits and Release Governance
SoonPre-training audits, sensitive-content review, documentation, approval gates, and post-release traceability.
Part XXII: Pre-training Objectives
7 chapters
Part XXII: Pre-training Objectives
Causal Language Modeling
Covers CLM objective formulation, autoregressive factorization, CLM loss computation, CLM for generation, CLM training data, CLM scaling properties.
Masked Language Modeling
Covers MLM objective formulation, masking strategies (15% rule), [MASK] token usage, MLM for understanding, MLM training dynamics.
Whole Word Masking
Covers subword masking problems, whole word masking procedure, WWM implementation, WWM vs random masking, WWM for different tokenizers.
Span Corruption
Covers span selection strategies, span length distribution, sentinel tokens, T5-style corruption, span corruption benefits.
Prefix Language Modeling
Covers prefix LM formulation, prefix LM attention pattern, prefix LM for generation, prefix LM training, UniLM-style objectives.
Replaced Token Detection
Covers generator-discriminator setup, replaced vs original detection, RTD efficiency advantages, ELECTRA training procedure, RTD vs MLM comparison.
Denoising Objectives
Covers token deletion, token shuffling, sentence permutation, document rotation, BART-style denoising, combining denoising tasks.
Part XXIII: Scaling Laws
7 chapters
Part XXIII: Scaling Laws
Power Laws in Deep Learning
Covers power law definition, log-log linear relationships, power law fitting, power law universality, power law intuition.
Kaplan Scaling Laws
Covers loss vs parameters, loss vs data, loss vs compute, Kaplan optimal allocation, Kaplan predictions.
Chinchilla Scaling Laws
Covers Chinchilla experiments, revised scaling coefficients, optimal tokens per parameter, Chinchilla vs Kaplan, Chinchilla implications.
Compute-Optimal Training
Covers compute budget allocation, tokens vs parameters ratio, training efficiency, compute-optimal recipes, practical guidelines.
Data-Constrained Scaling
Covers data repetition effects, optimal repetition strategies, data augmentation scaling, synthetic data scaling.
Inference Scaling
Covers training vs inference compute, inference-optimal models, over-training for efficiency, deployment cost modeling.
Predicting Model Performance
Covers loss extrapolation, capability prediction, scaling law uncertainty, prediction reliability, practical forecasting.
Part XXIV: BERT and Variants
8 chapters
Part XXIV: BERT and Variants
BERT Architecture
Covers BERT model sizes, BERT layer configuration, BERT embedding layers, BERT attention patterns, BERT output representations.
BERT Pre-training
Covers pre-training data preparation, MLM implementation, NSP task design, pre-training hyperparameters, pre-training duration.
BERT Fine-tuning
Covers classification fine-tuning, sequence labeling fine-tuning, question answering fine-tuning, fine-tuning hyperparameters, catastrophic forgetting.
BERT Representations
Covers [CLS] token usage, layer selection strategies, pooling strategies, BERT as feature extractor, frozen vs fine-tuned representations.
RoBERTa
Covers dynamic masking, NSP removal, larger batches, more data, RoBERTa training recipe, RoBERTa vs BERT performance.
ALBERT
Covers factorized embeddings, cross-layer parameter sharing, sentence order prediction, ALBERT efficiency, ALBERT performance trade-offs.
ELECTRA
Covers generator training, discriminator training, RTD objective, ELECTRA sample efficiency, ELECTRA scaling, ELECTRA fine-tuning.
DeBERTa
Covers disentangled attention formulation, enhanced mask decoder, DeBERTa position encoding, DeBERTa improvements, DeBERTa-v3 advances.
Part XXV: Encoder-Decoder Models
6 chapters
Part XXV: Encoder-Decoder Models
T5 Architecture
Covers T5 encoder-decoder design, T5 attention patterns, T5 model sizes, T5 relative positions, T5 implementation.
T5 Pre-training
Covers span corruption procedure, sentinel tokens, corruption rate, T5 pre-training data, T5 training scale.
T5 Task Formatting
Covers task prefixes, classification as generation, NER as generation, QA as generation, task formatting examples.
BART Architecture
Covers BART encoder-decoder, BART attention configuration, BART vs T5 comparison, BART model sizes.
BART Pre-training
Covers token masking, token deletion, text infilling, sentence permutation, document rotation, objective combinations.
mT5
Covers mT5 training data, language sampling, cross-lingual transfer, mT5 vs T5 performance, multilingual tokenization.
Part XXVI: Multilingual Language Models and Cross-Lingual Transfer
7 chapters
Part XXVI: Multilingual Language Models and Cross-Lingual Transfer
Language Diversity, Typology, and Scripts
SoonWriting systems, morphology, word order, language families, and why English-centric assumptions fail.
Multilingual Tokenization and Vocabulary Allocation
SoonFertility, script coverage, shared vocabularies, byte models, and unequal token costs across languages.
Multilingual Pre-training and Data Balancing
SoonSampling temperatures, capacity allocation, transfer-interference trade-offs, and multilingual mixture design.
Cross-Lingual Transfer and Alignment
SoonShared representations, zero-shot transfer, alignment objectives, adapters, and transfer diagnostics.
Low-Resource and Endangered Languages
SoonData scarcity, community participation, transliteration, active learning, and responsible evaluation.
Code-Switching and Mixed-Language Text
SoonLanguage identification, mixed scripts, code-switched generation, evaluation, and deployment failure modes.
Cultural and Multilingual Evaluation
SoonTranslationese, construct validity, cultural knowledge, local harms, and evaluation led by native speakers.
Part XXVII: Machine Translation and Speech Translation
7 chapters
Part XXVII: Machine Translation and Speech Translation
Statistical Machine Translation Foundations
SoonWord alignment, phrase tables, language models, log-linear decoding, and the ideas inherited by neural MT.
Neural Machine Translation
SoonEncoder-decoder translation, attention, transformer MT, training objectives, and decoding.
Parallel Data Mining and Quality Control
SoonBitext mining, alignment, filtering, back-translation, synthetic parallel data, and contamination checks.
Many-to-Many and Low-Resource Translation
SoonMultilingual transfer, mixture-of-experts translation, zero-shot directions, and capacity bottlenecks.
Translation Evaluation Beyond BLEU
SoonLearned metrics, adequacy, fluency, terminology, human evaluation, and metric failure modes.
Document-Level and Context-Aware Translation
SoonTerminology consistency, discourse context, pronouns, document memory, and long-form evaluation.
Speech-to-Speech Translation
SoonCascaded and end-to-end systems, latency, speaker preservation, prosody, and multilingual safety.
Part XXVIII: GPT Architecture
10 chapters
Part XXVIII: GPT Architecture
GPT-1
Covers GPT-1 architecture, GPT-1 pre-training, GPT-1 fine-tuning approach, GPT-1 transfer learning, GPT-1 historical significance.
GPT-2
Covers GPT-2 model sizes, GPT-2 architectural changes, zero-shot task performance, GPT-2 training data (WebText), GPT-2 generation quality.
GPT-3
Covers GPT-3 scale (175B), few-shot prompting discovery, in-context learning analysis, GPT-3 capabilities, GPT-3 limitations.
In-Context Learning
Covers ICL phenomenon, ICL vs fine-tuning, example selection strategies, ICL scaling behavior, ICL theoretical understanding.
Autoregressive Generation
Covers generation procedure, KV caching for efficiency, generation stopping criteria, generation speed optimization, generation code implementation.
Decoding Temperature
Covers temperature scaling, temperature effects on distribution, temperature selection guidelines, temperature vs quality trade-off.
Top-k Sampling
Covers top-k truncation, k selection strategies, top-k limitations, top-k implementation, combining with temperature.
Nucleus Sampling
Covers top-p formulation, cumulative probability threshold, nucleus sampling benefits, p selection guidelines, nucleus vs top-k.
Repetition Penalties
Covers repetition in generation, repetition penalty formulation, frequency penalty, presence penalty, n-gram blocking.
Constrained Decoding
Covers grammar-guided generation, JSON schema constraints, regex constraints, constrained beam search, constrained sampling.
Part XXIX: Modern Decoder Models
7 chapters
Part XXIX: Modern Decoder Models
LLaMA Architecture
Covers LLaMA design philosophy, LLaMA architectural choices, LLaMA training data, LLaMA efficiency, LLaMA significance.
LLaMA Components
Covers pre-norm with RMSNorm, SwiGLU FFN, RoPE implementation, component interactions, implementation details.
Grouped Query Attention
Covers GQA motivation, GQA formulation, KV head grouping, GQA memory savings, GQA vs MHA performance, GQA implementation.
Multi-Query Attention
Covers MQA extreme sharing, MQA memory benefits, MQA quality trade-offs, MQA for inference, MQA vs GQA.
Mistral Architecture
Covers Mistral design choices, sliding window attention, Mistral efficiency, Mistral performance, Mistral vs LLaMA.
Qwen Architecture
Covers Qwen architectural choices, Qwen training approach, Qwen multilingual capabilities, Qwen variants.
Phi Models
Covers Phi design philosophy, textbook-quality data, Phi training approach, Phi efficiency, small model capabilities.
Part XXX: Emergent Capabilities
6 chapters
Part XXX: Emergent Capabilities
Emergence in Neural Networks
Covers emergence definition, phase transitions, emergence examples, emergence mechanisms, emergence debate.
In-Context Learning Emergence
Covers ICL emergence curves, ICL vs fine-tuning scaling, ICL mechanism hypotheses, ICL as meta-learning.
Chain-of-Thought Emergence
Covers CoT emergence observations, CoT elicitation, CoT scaling behavior, CoT mechanism theories.
Emergence vs Metrics
Covers discontinuous metrics, accuracy threshold effects, smooth underlying capabilities, re-examining emergence claims.
Inverse Scaling
Covers inverse scaling phenomena, distractor tasks, sycophancy scaling, inverse scaling prize findings.
Grokking
Covers grokking phenomenon, grokking in arithmetic, grokking mechanism theories, grokking phase transitions, practical implications.
Part XXXI: Training Infrastructure
11 chapters
Part XXXI: Training Infrastructure
GPU Architecture
Covers GPU memory hierarchy, CUDA cores, tensor cores, GPU specifications.
Memory Management
Covers memory breakdown (activations, parameters, gradients, optimizer states), memory estimation, OOM debugging.
Data Parallelism
Covers DDP algorithm, gradient synchronization, all-reduce operations, DDP scaling.
Tensor Parallelism
Covers column parallelism, row parallelism, communication patterns, Megatron-style parallelism.
Pipeline Parallelism
Covers pipeline stages, micro-batching, pipeline bubbles, pipeline schedules (GPipe, 1F1B).
ZeRO Optimization
Covers ZeRO stage 1 (optimizer state partitioning), ZeRO stage 2 (gradient partitioning), ZeRO stage 3 (parameter partitioning), ZeRO memory savings.
FSDP
Covers FSDP concepts, FSDP vs ZeRO, FSDP sharding strategies, FSDP usage.
Activation Checkpointing
Covers checkpointing concept, checkpoint selection, checkpointing overhead, selective checkpointing.
Mixed Precision Training
Covers floating point formats, loss scaling, BF16 advantages, mixed precision implementation.
Communication Optimization
Covers gradient compression, communication overlap, topology-aware communication, NCCL optimization.
Checkpointing and Recovery
Covers checkpoint contents, checkpoint frequency, async checkpointing, fault recovery.
Part XXXII: Training Optimization
8 chapters
Part XXXII: Training Optimization
Learning Rate Warmup
Covers warmup motivation, linear warmup, warmup duration, warmup for large batches.
Learning Rate Decay
Covers step decay, exponential decay, inverse square root decay, decay scheduling.
Cosine Learning Rate Schedule
Covers cosine decay formula, cosine with restarts, cosine schedule parameters, cosine vs linear.
Large Batch Training
Covers batch size effects, learning rate scaling, batch size limits, LAMB optimizer.
Weight Decay
Covers weight decay formula, decoupled weight decay, weight decay selection, weight decay interaction with Adam.
Gradient Accumulation
Covers accumulation procedure, accumulation steps, accumulation for memory, accumulation correctness.
Training Stability
Covers loss spikes, gradient norm monitoring, stability techniques, training stability debugging.
Hyperparameter Selection
Covers hyperparameter search, hyperparameter transfer, critical vs robust hyperparameters, default recipes.
Part XXXIII: Mixture of Experts
10 chapters
Part XXXIII: Mixture of Experts
Sparse Models
Covers dense vs sparse trade-offs, conditional computation motivation, sparse model efficiency, sparse model challenges.
Expert Networks
Covers expert architecture, expert as FFN, expert capacity, expert count selection, expert placement in transformer.
Gating Networks
Covers router architecture, routing score computation, router training, router learned behavior.
Top-K Routing
Covers top-1 routing, top-2 routing, k selection trade-offs, routing implementation, combining expert outputs.
Load Balancing
Covers expert utilization imbalance, collapse failure mode, load metrics, balanced routing importance.
Auxiliary Balancing Loss
Covers load balancing loss formulation, loss coefficient tuning, balancing vs task loss, auxiliary loss implementation.
Router Z-Loss
Covers router instability, z-loss formulation, z-loss benefits, z-loss coefficient, combined auxiliary losses.
Expert Parallelism
Covers expert placement strategies, all-to-all communication, communication overhead, expert parallelism implementation.
Switch Transformer
Covers Switch Transformer design, top-1 routing choice, capacity factor, Switch scaling results.
Mixtral
Covers Mixtral architecture, Mixtral expert design, Mixtral performance, Mixtral efficiency, Mixtral vs dense models.
Part XXXIV: Fine-tuning Fundamentals
5 chapters
Part XXXIV: Fine-tuning Fundamentals
Transfer Learning
Covers transfer learning paradigm, pre-training/fine-tuning split, what transfers, transfer learning efficiency.
Full Fine-tuning
Covers full fine-tuning procedure, fine-tuning hyperparameters, learning rate selection, batch size effects.
Catastrophic Forgetting
Covers forgetting phenomenon, forgetting measurement, forgetting mitigation, pre-trained capability preservation.
Fine-tuning Learning Rates
Covers discriminative fine-tuning, layer-wise learning rates, warmup for fine-tuning, learning rate decay.
Fine-tuning Data Efficiency
Covers few-shot fine-tuning, data augmentation, sample efficiency patterns, small data strategies.
Part XXXV: Parameter-Efficient Fine-tuning
12 chapters
Part XXXV: Parameter-Efficient Fine-tuning
PEFT Motivation
Covers parameter storage costs, multi-task deployment, PEFT efficiency, PEFT quality trade-offs.
LoRA Concept
Covers weight update decomposition, low-rank assumption, LoRA efficiency gains, LoRA flexibility.
LoRA Mathematics
Covers LoRA formulation W + BA, rank selection, initialization scheme, LoRA gradient computation.
LoRA Implementation
Covers LoRA module design, merging weights, LoRA training loop, LoRA in PyTorch, HuggingFace PEFT usage.
LoRA Hyperparameters
Covers rank selection guidelines, alpha/rank ratio, which layers to adapt, LoRA dropout.
QLoRA
Covers 4-bit quantization for base model, NF4 data type, double quantization, QLoRA memory savings.
AdaLoRA
Covers importance-based pruning, SVD-based adaptation, dynamic rank, AdaLoRA training procedure.
IA3
Covers IA3 formulation, learned rescaling vectors, IA3 parameter efficiency, IA3 vs LoRA.
Prefix Tuning
Covers prefix tuning formulation, prefix length selection, prefix tuning for generation, prefix vs LoRA.
Prompt Tuning
Covers prompt tuning formulation, prompt initialization, prompt tuning scaling, prompt length effects.
Adapter Layers
Covers adapter architecture, adapter placement, adapter dimensionality, adapter fusion.
PEFT Comparison
Covers performance comparison, parameter efficiency comparison, task suitability, practical recommendations.
Part XXXVI: Instruction Tuning
6 chapters
Part XXXVI: Instruction Tuning
Instruction Following
Covers instruction tuning motivation, instruction format design, instruction diversity, instruction quality.
Instruction Data Creation
Covers human annotation, template-based generation, seed task expansion, quality filtering.
Self-Instruct
Covers self-instruct procedure, instruction generation, response generation, filtering strategies.
Instruction Format
Covers prompt templates, system messages, multi-turn format, chat templates, role definitions.
Instruction Tuning Training
Covers instruction tuning data mixing, training hyperparameters, loss masking, multi-task learning.
Instruction Following Evaluation
Covers instruction following benchmarks, human evaluation, automatic evaluation, instruction difficulty.
Part XXXVII: Alignment and RLHF
16 chapters
Part XXXVII: Alignment and RLHF
Alignment Problem
Covers alignment definition, helpfulness vs harmlessness, alignment challenges, alignment approaches overview.
Human Preference Data
Covers preference collection UI, comparison design, annotator guidelines, preference data quality.
Bradley-Terry Model
Covers pairwise comparison model, preference probability, Bradley-Terry likelihood, preference strength.
Reward Modeling
Covers reward model architecture, preference loss function, reward model training, reward model evaluation.
Reward Hacking
Covers reward hacking examples, distribution shift, over-optimization, reward hacking mitigation.
Policy Gradient Methods
Covers policy definition, REINFORCE algorithm, policy gradient derivation, variance reduction.
PPO Algorithm
Covers clipped objective, PPO derivation, trust region intuition, PPO implementation.
PPO for Language Models
Covers LLM as policy, action space (tokens), reward assignment, KL penalty importance.
RLHF Pipeline
Covers SFT stage, reward model training, PPO fine-tuning, RLHF hyperparameters, RLHF debugging.
KL Divergence Penalty
Covers KL penalty motivation, KL coefficient selection, adaptive KL, KL effects on training.
DPO Concept
Covers DPO motivation, removing reward model, DPO intuition, DPO benefits.
DPO Derivation
Covers DPO from RLHF objective, optimal policy derivation, DPO loss function, DPO as classification.
DPO Implementation
Covers DPO data format, DPO loss computation, DPO training procedure, DPO hyperparameters.
DPO Variants
Covers IPO formulation, KTO for unpaired feedback, ORPO, cDPO, comparing alignment methods.
RLAIF
Covers AI as annotator, constitutional AI principles, AI preference generation, RLAIF scalability.
Iterative Alignment
Covers iterative DPO, online preference learning, self-improvement loops, alignment stability.
Part XXXVIII: Scalable Oversight and Safety Training
5 chapters
Part XXXVIII: Scalable Oversight and Safety Training
Constitutional and Principle-Based Alignment
SoonWritten principles, critique and revision, AI feedback, evaluation, and failure modes.
Scalable Oversight
SoonDecomposition, debate, recursive supervision, weak-to-strong generalization, and oversight bottlenecks.
Process Supervision and Behavioral Specifications
SoonSupervising intermediate behavior, specifying policies, resolving conflicts, and testing adherence.
Alignment Data Quality and Rater Populations
SoonRater disagreement, cultural pluralism, annotator effects, preference aggregation, and auditability.
Alignment Robustness and Distribution Shift
SoonSycophancy, reward tampering, jailbreak pressure, deployment drift, and adversarial evaluation.
Part XXXIX: Reasoning
7 chapters
Part XXXIX: Reasoning
Reasoning Foundations
Covers reasoning types, reasoning in LLMs, reasoning failure modes, reasoning evaluation.
Chain-of-Thought
Covers CoT prompting, zero-shot CoT, CoT fine-tuning, CoT limitations.
Reasoning Strategies
Covers self-consistency, tree of thought, least-to-most prompting, decomposition strategies.
Reasoning Verification
Covers step verification, process reward models, verification-guided search, self-correction.
Mathematical Reasoning
Covers math problem solving, symbolic integration, math benchmarks, math reasoning training.
Reasoning Limitations
Covers systematic failures, spurious correlations, reasoning shortcuts, robustness challenges.
Reasoning Frontiers
Covers o1-style reasoning, test-time compute scaling, reasoning-capable models, open research questions.
Part XL: Reasoning Post-Training and Verifiable Rewards
7 chapters
Part XL: Reasoning Post-Training and Verifiable Rewards
Reinforcement Learning from Verifiable Rewards
SoonRule-based and executable rewards, correctness verification, reward design, and domains where RLVR works.
Group-Based Policy Optimization
SoonGRPO-style objectives, relative advantages, stability, sampling costs, and implementation choices.
Cold Starts, Rejection Sampling, and Distillation
SoonBootstrapping reasoning behavior, filtering traces, iterative training, and transferring reasoning to smaller models.
Outcome and Process Reward Models
SoonSparse outcomes, step-level supervision, verifier reliability, credit assignment, and reward hacking.
Search, Reflection, and Test-Time Compute
SoonSampling, verifier-guided search, self-correction, budget allocation, stopping, and compute-optimal reasoning.
Faithful, Hidden, and Latent Reasoning
SoonWhen visible chains of thought are explanations, when they are not, and alternatives for latent deliberation.
Overthinking and Reasoning Efficiency
SoonUnnecessary deliberation, error amplification, confidence, adaptive budgets, and concise reasoning.
Part XLI: Model Compression
6 chapters
Part XLI: Model Compression
Knowledge Distillation
Covers distillation objective, temperature in distillation, teacher selection, distillation for LLMs.
Distillation Variants
Covers feature distillation, attention transfer, progressive distillation, on-policy distillation.
Pruning Basics
Covers weight pruning, structured vs unstructured, pruning criteria, pruning schedule.
Structured Pruning
Covers head pruning, layer pruning, width pruning, structured pruning implementation.
Model Merging
Covers weight averaging, task arithmetic, TIES merging, DARE merging.
Model Merging Applications
Covers multi-task merging, style merging, capability composition, merging evaluation.
Part XLII: Inference Optimization
14 chapters
Part XLII: Inference Optimization
KV Cache
Covers KV cache motivation, cache structure, cache memory requirements, cache management.
KV Cache Memory
Covers cache size calculation, batch size effects, sequence length effects, memory bottleneck.
Paged Attention
Covers memory fragmentation problem, page-based allocation, vLLM approach, paged attention benefits.
KV Cache Compression
Covers cache eviction strategies, attention sink preservation, H2O algorithm, cache quantization.
Weight Quantization Basics
Covers quantization fundamentals, per-tensor vs per-channel, symmetric vs asymmetric, calibration.
INT8 Quantization
Covers INT8 range mapping, absmax quantization, smooth quantization, INT8 accuracy.
INT4 Quantization
Covers 4-bit challenges, group-wise quantization, 4-bit accuracy trade-offs, 4-bit formats.
GPTQ
Covers GPTQ algorithm, layer-wise quantization, Hessian approximation, GPTQ implementation.
AWQ
Covers salient weight preservation, AWQ algorithm, AWQ vs GPTQ, AWQ benefits.
GGUF Format
Covers GGML/GGUF history, quantization types, GGUF file format, llama.cpp integration.
Speculative Decoding: Fast LLM Inference Without Quality Loss
Covers speculative decoding concept, draft model selection, verification procedure, acceptance rate.
Speculative Decoding Math: Algorithms & Speedup Limits
Covers acceptance criterion, expected speedup, draft quality effects, optimal draft length.
Continuous Batching: Optimizing LLM Inference Throughput
Covers static vs continuous batching, iteration-level scheduling, request completion handling, throughput gains.
LLM Inference Serving: Architecture, Routing & Auto-Scaling
Master LLM inference serving architecture, token-aware load balancing, and auto-scaling. Optimize time-to-first-token and throughput for production systems.
Part XLIII: Small and On-Device Language Models
5 chapters
Part XLIII: Small and On-Device Language Models
Small Language Model Design
SoonCapacity allocation, data quality, architecture choices, and capability boundaries below frontier scale.
Hardware-Aware Model Optimization
SoonMemory bandwidth, kernels, quantization formats, sparsity, and co-design for phones, browsers, and edge devices.
On-Device Adaptation and Personalization
SoonPrivate fine-tuning, adapters, local retrieval, user control, and update strategies under tight budgets.
Hybrid Edge-Cloud Inference
SoonRouting, partitioning, privacy boundaries, offline fallbacks, latency, and cost-aware orchestration.
Evaluating Small Models in Context
SoonTask fit, energy, latency, privacy, reliability, and comparing systems rather than parameter counts alone.
Part XLIV: Retrieval-Augmented Generation
14 chapters
Part XLIV: Retrieval-Augmented Generation
RAG Motivation: Solving Hallucinations & Knowledge Gaps
Discover why LLMs need Retrieval-Augmented Generation. Learn how RAG bridges knowledge gaps, reduces hallucinations, and enables non-parametric memory.
RAG Architecture: Components, Timing & Design Patterns
Master RAG system design by exploring retriever-generator interactions, timing strategies like iterative retrieval, and architectural variations like RETRO.
Dense Retrieval: Semantic Search & Bi-Encoder Implementation
Master dense retrieval for semantic search. Explore bi-encoder architectures, embedding metrics, and contrastive learning to overcome keyword limitations.
Contrastive Learning for Retrieval: InfoNCE & DPR Guide
Master contrastive learning for dense retrieval. Learn to train models using InfoNCE loss, in-batch negatives, and hard negative mining strategies effectively.
Document Chunking: Optimizing RAG Retrieval Pipelines
Master document chunking for RAG systems. Explore fixed-size, recursive, and semantic strategies to balance retrieval precision with context window limits.
Embedding Models: Architecture, Pooling & Selection
Learn how embedding models convert text to vectors for RAG. Covers bi-encoder architecture, pooling strategies, dimensionality trade-offs, and model selection.
Vector Similarity Search: Metrics & Approximate Methods
Explore vector similarity search for RAG systems. Compare cosine, dot product, and Euclidean metrics, and implement exact vs. approximate search with FAISS.
HNSW Index: Architecture for Fast Vector Search
Master Hierarchical Navigable Small World (HNSW) graphs for vector search. Learn graph architecture, construction, and tuning for high-speed retrieval.
IVF Index: Clustering-Based Vector Search & Partitioning
Master IVF indexes for scalable vector search. Learn clustering-based partitioning, nprobe tuning, and IVF-PQ compression for billion-scale retrieval.
Product Quantization: Vector Compression for ANN Search
Learn how Product Quantization compresses embeddings up to 100x using learned codebooks and asymmetric distance computation for scalable vector search.
Hybrid Search: BM25 and Dense Retrieval Combined
Learn how hybrid search fuses BM25 keyword retrieval with dense vector retrieval using reciprocal rank fusion and weighted score combination to improve recall.
Reranking: Cross-Encoders for Precise Information Retrieval
Learn how reranking with cross-encoders solves bi-encoder limitations. Master two-stage retrieval, training strategies, and latency optimization for production search systems.
RAG Prompt Engineering: Context Placement & Citation Strategies
Master RAG prompt engineering with strategic context placement, citation formats, and truncation strategies to improve LLM accuracy and reduce hallucinations.
RAG Evaluation: Metrics for Retrieval and Generation Quality
Master RAG evaluation with metrics for retrieval quality (Precision@K, NDCG, MRR) and generation faithfulness using the RAGAS framework for AI systems.
Part XLV: Knowledge, Memory, Editing, and Unlearning
6 chapters
Part XLV: Knowledge, Memory, Editing, and Unlearning
Parametric and External Knowledge
SoonWhat models store, what retrieval stores, and how to choose among prompting, RAG, editing, and retraining.
Knowledge Editing
SoonLocalized weight updates, memory-based editors, specificity, generalization, and multi-hop consistency.
Knowledge Freshness and Temporal Updates
SoonTime-sensitive facts, temporal benchmarks, update propagation, versioning, and rollback.
Personalization and Long-Term Memory
SoonUser models, episodic and semantic memory, consent, retention, conflict resolution, and forgetting.
Machine Unlearning
SoonForget sets, retraining baselines, approximate removal, privacy goals, and the limits of verification.
Evaluating Knowledge Interventions
SoonEfficacy, locality, generalization, side effects, privacy leakage, and auditable change histories.
Part XLVI: Continual Learning
5 chapters
Part XLVI: Continual Learning
Continual Learning Problem
Covers continual learning definition, catastrophic forgetting, continual learning scenarios.
Regularization Methods
Covers elastic weight consolidation, synaptic intelligence, parameter importance, regularization trade-offs.
Replay Methods
Covers replay buffer design, pseudo-rehearsal, generative replay, replay selection.
Architecture Methods
Covers progressive networks, expert expansion, architecture search, modular approaches.
Continual Learning Evaluation
Covers forward transfer, backward transfer, evaluation protocols, continual benchmarks.
Part XLVII: Code Generation
6 chapters
Part XLVII: Code Generation
Code LLM Training
Covers code training data, code tokenization, fill-in-the-middle training, code pre-training objectives.
Code Understanding
Covers code explanation, bug detection, code review, code search.
Code Completion
Covers completion context, completion ranking, completion latency, completion UX.
Code Generation
Covers docstring-to-code, test-to-code, code generation strategies, generation quality.
Code Execution
Covers sandboxed execution, execution feedback, iterative refinement, execution safety.
Code Evaluation
Covers functional correctness, pass@k metric, code benchmarks, beyond correctness.
Part XLVIII: Tool Use and Agents
10 chapters
Part XLVIII: Tool Use and Agents
Tool Use Motivation: Why LLMs Need External Tools for Accuracy
Discover why LLMs require external tools to overcome knowledge cutoffs, computational limits, and hallucinations. Learn about tool-augmented AI systems.
Function Calling: Structured Tool Use for Large Language Models
Learn how function calling enables LLMs to invoke external tools and APIs through structured JSON schemas, bridging natural language and executable code.
ReAct Pattern: Interleaving Reasoning and Action for LLM Agents
Learn how the ReAct pattern enables LLM agents to interleave reasoning with tool execution. Master thought-action-observation loops for autonomous AI systems.
Tool Selection for LLM Agents: Routing Strategies and Implementation
Master LLM tool selection through embedding-based routing, hybrid strategies, and semantic interfaces. Learn to build scalable multi-tool agent systems.
Agent Architectures: Control Loops, State & Planning
Master LLM agent architectures including control loops, state management strategies, planning mechanisms, and termination conditions for autonomous AI systems.
Agent Memory Systems: From Context to Persistent Storage
Learn how AI agents manage short-term context and long-term memory. Explore vector databases, retrieval algorithms, and memory hierarchies that enable persistent, learning systems.
Agent Evaluation: Metrics, Benchmarks and Safety Standards
Learn to evaluate AI agents with task completion metrics, trajectory analysis, and safety testing. Covers WebArena, SWE-bench, GAIA, and OSWorld benchmarks.
Planning
Covers task decomposition, goal-directed planning, plan execution, plan revision and recovery.
Multi-Agent Systems
Covers agent coordination, communication protocols, role assignment, multi-agent benchmarks.
Agent Safety
Covers unsafe action prevention, agent alignment, sandboxing, monitoring and intervention.
Part XLIX: Long-Horizon Agents and Interoperability
8 chapters
Part XLIX: Long-Horizon Agents and Interoperability
Computer-Use and Browser Agents
SoonScreenshots, accessibility trees, action spaces, grounding, recovery, and evaluation on real interfaces.
Coding and Software-Engineering Agents
SoonRepository navigation, test-driven changes, execution feedback, code review, and long-running tasks.
Deep-Research Agents
SoonSearch, source selection, evidence synthesis, citation integrity, uncertainty, and reproducible research traces.
Durable and Asynchronous Agent Work
SoonCheckpoints, resumability, queues, idempotency, retries, deadlines, and state across long-running workflows.
Model Context Protocol
SoonMCP clients, servers, tools, resources, prompts, capability negotiation, transport, and trust boundaries.
Agent-to-Agent Protocols
SoonDiscovery, task delegation, status, artifacts, interoperability, and failure semantics across agents.
Human Approval and Control Boundaries
SoonApproval gates, reversibility, escalation, authority, interruption, and clear ownership of consequential actions.
Long-Horizon Agent Evaluation
SoonFunctional success, partial credit, recovery, cost, time, safety, and contamination-resistant task suites.
Part L: Multimodal Models
12 chapters
Part L: Multimodal Models
Vision Transformer
Covers image patching, patch embeddings, ViT architecture, ViT pre-training.
CLIP
Covers CLIP architecture, CLIP training objective, CLIP zero-shot classification, CLIP embeddings.
Vision Encoders for VLMs
Covers ViT variants for VLMs, SigLIP improvements, image resolution handling, encoder selection.
Vision-Language Projection
Covers linear projection, MLP projection, Q-Former approach, projection training.
LLaVA Architecture
Covers LLaVA design, two-stage training, visual conversation, LLaVA variants.
Flamingo Architecture
Covers cross-attention to images, gated cross-attention, few-shot visual learning, Flamingo training.
Multimodal Training Data
Covers image-text pairs, interleaved documents, visual instruction data, data quality.
Multimodal Evaluation
Covers VQA benchmarks, multimodal understanding benchmarks, multimodal generation evaluation.
Multimodal Foundations
Covers multimodal learning principles, cross-modal alignment, joint embedding spaces, multimodal challenges.
Image Understanding
Covers visual question answering, image captioning, visual grounding, scene understanding.
Image Generation
Covers diffusion models, text-to-image generation, image editing, generation evaluation.
Multimodal Applications
Covers document understanding, medical imaging, video understanding, multimodal reasoning tasks.
Part LI: Speech and Audio
5 chapters
Part LI: Speech and Audio
Speech Representations
Covers mel spectrograms, mel filterbanks, feature normalization, audio preprocessing.
Whisper Architecture
Covers Whisper encoder-decoder, multitask training, language tokens, timestamp prediction.
Whisper Training
Covers Whisper training data, weak supervision, multilingual training, Whisper capabilities.
Speech-Language Integration
Covers speech encoder + LLM, audio tokens, speech-to-text-to-LLM vs end-to-end, speech LLM architectures.
Text-to-Speech
Covers TTS architecture overview, vocoder role, TTS quality metrics, neural TTS approaches.
Part LII: Omni-Modal and Embodied Language Systems
7 chapters
Part LII: Omni-Modal and Embodied Language Systems
Unified Audio-Visual-Language Modeling
SoonModality encoders, shared token spaces, fusion, generation, and interference in omni models.
Video and Temporal Reasoning
SoonFrame sampling, event localization, temporal memory, causal reasoning, and long-video evaluation.
Streaming Multimodal Interaction
SoonContinuous perception, interruption, turn-taking, proactive responses, latency, and duplex generation.
Document, Diagram, and Interface Understanding
SoonOCR, layout, charts, diagrams, screenshots, grounding, and visually situated tool use.
Speech-to-Speech and Expressive Voice
SoonDirect speech interaction, prosody, speaker identity, emotion, latency, and voice-specific safety.
Vision-Language-Action Models
SoonGrounded language for robotics, action tokenization, policy learning, embodiment, and simulation-to-reality gaps.
Multimodal Safety and Evaluation
SoonCross-modal injection, privacy, deepfakes, modality imbalance, interactive benchmarks, and human evaluation.
Part LIII: Evaluation Fundamentals
10 chapters
Part LIII: Evaluation Fundamentals
Perplexity Evaluation
Covers perplexity calculation, perplexity interpretation, perplexity limitations, comparing perplexities.
Cross-Entropy Loss
Covers cross-entropy definition, bits-per-character, cross-entropy vs perplexity, loss curves.
BLEU Score
Covers n-gram precision, brevity penalty, BLEU formula, BLEU limitations, corpus vs sentence BLEU.
ROUGE Scores
Covers ROUGE-N, ROUGE-L, ROUGE-W, ROUGE interpretation, ROUGE limitations.
BERTScore
Covers BERTScore computation, token alignment, BERTScore variants, BERTScore vs BLEU.
Exact Match and F1
Covers exact match scoring, token-level F1, normalization for matching, metric selection.
Calibration
Covers calibration definition, expected calibration error, calibration plots, calibration methods.
Evaluation Fundamentals
Covers evaluation design principles, metric selection, evaluation pitfalls, evaluation frameworks.
Benchmark Design
Covers benchmark construction, dataset collection, annotation guidelines, benchmark validity.
Evaluation Challenges
Covers benchmark contamination, evaluation brittleness, gaming metrics, evaluation best practices.
Part LIV: Benchmark Evaluation
8 chapters
Part LIV: Benchmark Evaluation
MMLU
Covers MMLU structure, subject coverage, MMLU evaluation protocol, MMLU limitations.
HellaSwag
Covers HellaSwag task design, adversarial filtering, HellaSwag evaluation, HellaSwag saturation.
GSM8K
Covers GSM8K problem types, chain-of-thought evaluation, GSM8K accuracy metrics, math reasoning assessment.
HumanEval
Covers HumanEval structure, functional correctness, pass@k metric, HumanEval limitations.
MBPP
Covers MBPP dataset, MBPP vs HumanEval, code evaluation challenges.
TruthfulQA
Covers TruthfulQA design, truthfulness vs informativeness, TruthfulQA evaluation methods.
Benchmark Contamination
Covers contamination problem, contamination detection methods, n-gram overlap analysis, contamination mitigation.
Benchmark Saturation
Covers ceiling effects, benchmark retirement, dynamic benchmarks, benchmark evolution.
Part LV: Human and Model Evaluation
6 chapters
Part LV: Human and Model Evaluation
Human Evaluation Design
Covers evaluation interface design, task instructions, annotator selection, evaluation cost.
Inter-Annotator Agreement
Covers Cohen's kappa, Fleiss' kappa, Krippendorff's alpha, handling disagreement.
Preference Evaluation
Covers A/B comparison design, Elo rating systems, preference aggregation, statistical significance.
LLM-as-Judge: Scalable AI Evaluation with Language Models
Learn how to build LLM-as-Judge evaluation pipelines: prompt design, judge model selection, calibration against human annotations, and bias mitigation.
Position Bias in LLM Judges
Covers position bias measurement, bias mitigation (swapping), verbosity bias, sycophancy.
Evaluation Prompt Engineering: Designing Reliable LLM Judges
Learn how to design reliable LLM judge prompts using explicit criteria, few-shot examples, and chain-of-thought formatting to maximize evaluation accuracy.
Part LVI: Hallucination and Factuality
6 chapters
Part LVI: Hallucination and Factuality
Hallucination Types in Language Models: A Complete Guide
Learn how language models hallucinate: intrinsic and extrinsic hallucination, factual errors, fabrication, and inconsistency with NLI-based detection.
Hallucination Detection: NLI, Self-Consistency & Learned Models
Learn four methods for detecting LLM hallucinations: entailment-based scoring, knowledge base verification, self-consistency checks, and learned detection models.
Hallucination Causes
Covers training data issues, exposure bias, knowledge gaps, generation pressure.
Hallucination Mitigation: RAG, Decoding, and Training
Learn how to reduce LLM hallucination using retrieval augmentation, self-consistency decoding, DPO training, and calibrated uncertainty expression.
Attribution and Citation: Sourcing LLM Outputs
Learn how language models link generated claims to source documents, evaluate citation accuracy with NLI, and measure attribution precision and recall.
Uncertainty Quantification
Covers confidence calibration, verbalized uncertainty, sampling-based uncertainty, uncertainty communication.
Part LVII: Dynamic, Interactive, and Agentic Evaluation
6 chapters
Part LVII: Dynamic, Interactive, and Agentic Evaluation
Live and Continuously Updated Benchmarks
SoonFresh questions, private test sets, rolling releases, objective scoring, and contamination controls.
Evaluation Harnesses and Reproducibility
SoonVersioned prompts, model settings, graders, confidence intervals, artifacts, and comparable reports.
Interactive and Long-Horizon Task Evaluation
SoonStateful environments, functional outcomes, partial credit, recovery, efficiency, and trajectory analysis.
Judge Calibration and Meta-Evaluation
SoonBias, self-preference, verbosity effects, rubric validity, judge ensembles, and agreement with experts.
Online Experiments and Production Evaluation
SoonShadow traffic, interleaving, A/B tests, guardrail metrics, drift, and linking offline scores to user outcomes.
Evaluation Governance
SoonBenchmark access, disclosure, privacy, gaming, deprecation, audit trails, and responsible leaderboard use.
Part LVIII: Bias and Fairness
5 chapters
Part LVIII: Bias and Fairness
Bias in Language Models
Covers bias sources, bias types (demographic, cultural), bias in training data, bias amplification.
Bias Measurement
Covers embedding association tests, generation bias metrics, classification bias metrics, bias benchmarks.
Bias Mitigation: Debiasing, CDA, and Fair Fine-tuning
Practical techniques for reducing demographic bias in language models: data balancing, embedding debiasing, adversarial training, and prompt-based interventions.
Fairness Metrics: Demographic Parity, Equalized Odds, and Trade-offs
Learn the key mathematical definitions of algorithmic fairness, from demographic parity to equalized odds, and why satisfying all metrics simultaneously is impossible.
Representation Harms
Covers stereotyping, erasure, demeaning associations, measuring representation harms.
Part LIX: Interpretability
11 chapters
Part LIX: Interpretability
Interpretability Goals
Covers debugging, trust, safety, scientific understanding, interpretability approaches overview.
Attention Visualization
Covers attention weight extraction, attention head visualization, attention interpretation caveats, attention tools.
Attention Analysis Limitations
Covers attention vs importance, attention manipulation studies, gradient-based alternatives.
Probing Classifiers
Covers linear probing methodology, probing task design, probing interpretation, control tasks.
Probing Layers
Covers layer selection, representation evolution, task localization, layer probing patterns.
Activation Patching
Covers patching methodology, locating information, patching experiments, causal tracing.
Logit Lens
Covers logit lens concept, intermediate vocabulary projection, tuned lens, lens interpretation.
Sparse Autoencoders
Covers SAE architecture, sparsity constraints, dictionary learning, SAE for LLMs.
Feature Interpretation
Covers feature activation patterns, feature naming, automated interpretation, feature circuits.
Mechanistic Interpretability
Reverse-engineer transformer networks into human-understandable algorithms by identifying circuits, induction heads, and mechanistic discoveries.
Activation Steering: Steering Vectors and Representation Engineering
Learn how steering vectors and activation addition let you modify language model behavior at inference time by injecting directional nudges into the residual stream.
Part LX: Safety and Security
8 chapters
Part LX: Safety and Security
Safety Risks
Covers harmful content generation, misuse scenarios, unintended harms, safety threat models.
Red Teaming
Learn how red teams systematically probe language models for safety failures, covering attack taxonomies, ASR metrics, RL-based automated attack generation, and real-world findings.
Jailbreaking
Covers jailbreak techniques, prompt injection, adversarial suffixes, jailbreak defenses.
Prompt Injection
Covers direct prompt injection, indirect prompt injection, injection in RAG, injection defenses.
Content Filtering
Covers classification-based filtering, rule-based filtering, filter placement, filter evaluation.
Guardrails
Covers input guardrails, output guardrails, guardrail frameworks, guardrail design.
Memorization and Privacy in Language Models
How language models memorize training data, methods for measuring extractable memorization, PII risks in web-scale corpora, and practical privacy mitigations.
Differential Privacy
Learn how differential privacy protects training data in language models, from the mathematical guarantee to DP-SGD, privacy budgets, and practical LLM fine-tuning.
Part LXI: Agentic AI and Protocol Security
7 chapters
Part LXI: Agentic AI and Protocol Security
Threat Modeling Compound AI Systems
SoonAssets, actors, trust boundaries, control flow, data flow, and failure propagation across model-centered systems.
Indirect Prompt Injection
SoonMalicious instructions in documents, web pages, messages, tool output, and retrieved context.
Tool, Retrieval, and Memory Poisoning
SoonSchema attacks, compromised tools, poisoned indexes, persistent memory attacks, and integrity controls.
Identity, Authorization, and Least Privilege
SoonUser and agent identity, scoped credentials, delegation, expiry, consent, and policy enforcement outside the model.
Secrets, Sandboxes, and Safe Execution
SoonCredential isolation, capability sandboxes, network and filesystem boundaries, output validation, and containment.
Agent Supply Chains and Protocol Security
SoonTool registries, dependency tampering, server discovery, trust-on-first-use, signing, and update risk.
Audit, Detection, and Incident Response
SoonTamper-evident traces, anomaly detection, kill switches, investigation, rollback, and post-incident learning.
Part LXII: Human-AI Interaction and Governance
6 chapters
Part LXII: Human-AI Interaction and Governance
Designing Human-AI Collaboration
SoonTask allocation, mixed initiative, handoffs, feedback, shared context, and preserving human agency.
Trust, Reliance, and Automation Bias
SoonCalibrated reliance, overtrust, undertrust, explanations, uncertainty, and behavior under time pressure.
Oversight, Appeals, and Escalation
SoonReview queues, contestability, escalation paths, documentation, and accountability for consequential decisions.
Accessibility and Inclusive Language AI
SoonDisability access, literacy, language inclusion, participatory design, and evaluating who benefits or is excluded.
System Cards, Audits, and Impact Assessments
SoonDocumenting complete systems, evaluating downstream use, independent audits, and monitoring material changes.
Risk Governance Across the Lifecycle
SoonOwnership, risk tiers, review gates, incident reporting, change management, and retirement.
Part LXIII: LLM Applications
7 chapters
Part LXIII: LLM Applications
Text Generation Applications
Covers creative writing, content generation, text transformation, generation quality control.
Summarization
Covers extractive vs abstractive, length control, faithfulness, multi-document summarization.
Question Answering
Covers open-domain QA, reading comprehension, knowledge-intensive QA, QA evaluation.
Information Extraction
Covers named entity recognition, relation extraction, event extraction, structured output generation.
Classification Applications
Covers sentiment analysis, intent detection, topic classification, zero-shot classification.
Conversational AI
Covers dialogue management, context tracking, persona consistency, conversation evaluation.
Creative Applications
Covers story generation, poetry, code creativity, creative constraints and control.
Part LXIV: Compound AI System Design
6 chapters
Part LXIV: Compound AI System Design
Models, Retrieval, Tools, and Control Flow
SoonDecomposing systems into components, explicit orchestration, typed interfaces, and failure boundaries.
Routing, Fallbacks, and Graceful Degradation
SoonModel cascades, confidence-based routing, deterministic fallbacks, abstention, and partial service.
Evaluation-Driven Development
SoonTask suites, traces, regression gates, error taxonomies, and improving systems without benchmark theater.
Tracing and Observability for AI Systems
SoonPrompt, retrieval, tool, latency, cost, and quality traces with privacy-aware retention.
Versioning Models, Prompts, Data, and Tools
SoonCoordinated releases, compatibility, reproducibility, canaries, rollback, and artifact lineage.
Reliability and Cost Engineering
SoonService objectives, retries, idempotency, concurrency, capacity, caching, budgets, and failure testing.
Part LXV: Production Systems
9 chapters
Part LXV: Production Systems
Model Serving
Covers serving frameworks, model loading, request handling, serving configuration.
Latency Optimization
Covers latency breakdown, batching latency, streaming responses, latency monitoring.
Throughput Optimization
Covers batch size tuning, GPU utilization, concurrent requests, throughput measurement.
Auto-scaling
Covers scaling metrics, horizontal scaling, scale-up vs scale-out, scaling policies.
Model Routing
Covers model selection, A/B testing, model cascades, routing strategies.
Caching
Covers prompt caching, semantic caching, cache invalidation, cache hit rates.
Monitoring
Covers metrics collection, alerting, logging, dashboards.
Quality Monitoring
Covers output quality metrics, drift detection, regression detection, quality alerts.
Cost Management
Covers cost modeling, cost optimization, cost allocation, cost monitoring.
Part LXVI: Frontier Methods of 2025
8 chapters
Part LXVI: Frontier Methods of 2025
Constitutional AI
Covers constitutional principles, critique and revision, CAI training, CAI effectiveness.
Process Reward Models
Covers outcome vs process reward, PRM training, PRM for math, PRM limitations.
Test-Time Compute
Covers multiple sampling, iterative refinement, compute-optimal inference, scaling test-time compute.
Retrieval-Augmented Training
Covers RETRO architecture, retrieval during training, retrieved context integration.
Long-Form Generation
Covers outline-based generation, hierarchical generation, coherence maintenance, long-form evaluation.
Watermarking
Covers watermarking schemes, statistical detection, watermark robustness, watermark evaluation.
Model Cards
Covers model card contents, intended use documentation, limitation documentation, model card best practices.
Responsible Deployment
Covers release decisions, staged release, access control, deployment monitoring.
Part LXVII: Frontier Outlook from 2025
6 chapters
Part LXVII: Frontier Outlook from 2025
Scaling Frontiers
Covers scaling limits, beyond power laws, data wall, architectural innovations for scale.
Efficiency Frontiers
Covers efficient architectures, hardware co-design, inference efficiency, training efficiency advances.
Capability Frontiers
Covers emerging capabilities, world models, planning and agency, multimodal reasoning advances.
Alignment Challenges
Covers scalable oversight, alignment tax, goal mis-specification, long-term alignment research.
Societal Implications
Covers labor market effects, access and equity, regulation landscape, governance frameworks.
Research Directions
Covers open problems, promising research areas, benchmark gaps, community priorities.
In Progress
This comprehensive handbook is currently in development. Each chapter will be published as it's completed, with practical examples, code implementations, and real-world applications.
Reference
Citation details
Cite or share this article.
Stay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.



















