Volume 4 of 20 · PDF edition
Available nowNeural NLP and Sequence Models
For readers who want the neural and encoder-decoder foundations behind modern language systems, including what transformers replaced and retained.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 29 written chapters
- Approximate pages
- ~920 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 4 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 4 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Neural Network Foundations
- Recurrent Neural Networks
- Sequence-to-Sequence
Audience and prerequisites
Where this volume fits
For readers who want the neural and encoder-decoder foundations behind modern language systems, including what transformers replaced and retained.
Prerequisites: Assumes basic linear algebra, calculus, probability, and Volume 2.
Free online preview
Start with “Linear Classifiers”
Covers linear decision boundaries, weight vectors and bias, dot product interpretation, multiclass classification (softmax), linear classifier limitations, training with gradient descent.
Exact contents
29 chapters available now
Part X: Neural Network Foundations
- 01Linear Classifiers
Covers linear decision boundaries, weight vectors and bias, dot product interpretation, multiclass classification (softmax), linear classifier limitations, training with gradient descent.
- 02Activation Functions
Covers sigmoid function and saturation, tanh properties, ReLU and dying ReLU, Leaky ReLU and PReLU, ELU and SELU, GELU derivation and properties, Swish and Mish, choosing activation functions.
- 03Multilayer Perceptrons
Covers hidden layers and depth, weight matrices between layers, forward pass computation, representational capacity, MLP for classification, MLP for regression, MLP architecture design.
- 04Loss Functions
Covers cross-entropy loss derivation, MSE for regression, binary vs multiclass cross-entropy, label smoothing, focal loss for imbalance, loss function numerical stability, custom loss functions.
- 05Backpropagation
Covers computational graphs, chain rule review, forward and backward pass, gradient accumulation, backprop complexity analysis, automatic differentiation, implementing backprop from scratch.
- 06Stochastic Gradient Descent
Covers batch vs stochastic gradient descent, minibatch gradient descent, learning rate selection, SGD convergence properties, SGD noise as regularization, learning rate schedules basics, SGD implementation.
- 07Momentum
Covers momentum intuition (ball rolling), momentum update equations, momentum coefficient selection, dampening oscillations, momentum vs vanilla SGD, Nesterov momentum derivation, implementing momentum.
- 08Adam Optimizer
Covers exponential moving averages, first moment (mean) estimation, second moment (variance) estimation, bias correction derivation, Adam update rule, Adam hyperparameters, Adam convergence properties.
- 09AdamW
Covers L2 regularization vs weight decay, why they differ with Adam, AdamW formulation, weight decay coefficient selection, AdamW as default optimizer, AdamW vs Adam empirically.
- 10Weight Initialization
Covers random initialization importance, Xavier/Glorot initialization derivation, He initialization for ReLU, initialization for different activations, layer-wise initialization, initialization debugging, modern initialization practices.
- 11Batch Normalization
Covers internal covariate shift, batch statistics computation, learnable scale and shift, training vs inference mode, batch norm gradient flow, batch norm placement debates, batch norm limitations.
- 12Dropout
Covers dropout as ensemble, dropout mask sampling, inverted dropout scaling, dropout rate selection, dropout at inference, spatial dropout for sequences, dropout in modern architectures.
- 13Gradient Clipping
Covers gradient explosion detection, clip by value, clip by global norm, gradient clipping implementation, when to use gradient clipping, clipping threshold selection, monitoring gradient norms.
Part XI: Recurrent Neural Networks
- 01RNN Architecture
Covers recurrent connection intuition, hidden state as memory, unrolled computation graph, parameter sharing across time, RNN for sequence classification, RNN for sequence generation, RNN equations and dimensions.
- 02Backpropagation Through Time
Covers BPTT derivation, gradient flow through time, truncated BPTT, BPTT memory requirements, BPTT implementation, gradient accumulation across timesteps.
- 03Vanishing Gradients
Covers gradient product across timesteps, vanishing gradient analysis, long-range dependency failure, gradient visualization, vanishing vs exploding trade-off, architectural solutions overview.
- 04LSTM Architecture
Covers cell state as information highway, gate mechanism intuition, LSTM diagram walkthrough, information flow in LSTMs, LSTM for long sequences, LSTM memory capacity.
- 05LSTM Gate Equations
Covers forget gate equations, input gate equations, cell state update, output gate equations, hidden state computation, LSTM parameter count, implementing LSTM from scratch.
- 06LSTM Gradient Flow
Covers constant error carousel, forget gate gradient highway, gradient flow analysis, LSTM vs vanilla RNN gradients, peephole connections, LSTM gradient clipping needs.
- 07GRU Architecture
Covers GRU vs LSTM comparison, reset gate function, update gate function, candidate hidden state, GRU equations, GRU parameter efficiency, when to choose GRU vs LSTM.
- 08Bidirectional RNNs
Covers forward and backward passes, hidden state concatenation, bidirectional architectures, bidirectionality for classification, limitations for generation, implementing bidirectional RNNs.
- 09Stacked RNNs
Covers multiple RNN layers, residual connections for depth, layer normalization in RNNs, depth vs width trade-offs, gradient flow in deep RNNs, practical depth limits.
Part XII: Sequence-to-Sequence
- 01Encoder-Decoder Framework
Covers encoder role and design, decoder role and design, context vector as bottleneck, seq2seq for machine translation, seq2seq for summarization, seq2seq training setup.
- 02Teacher Forcing
Covers teacher forcing procedure, exposure bias problem, teacher forcing efficiency, scheduled sampling, curriculum learning, teacher forcing vs autoregressive training.
- 03Beam Search
Covers greedy decoding limitations, beam search algorithm, beam width selection, length normalization, diverse beam search, beam search implementation, beam search vs sampling.
- 04Attention Intuition
Covers attention as soft lookup, attention weight interpretation, attention for variable-length inputs, attention visualization, attention vs pooling, attention computation overview.
- 05Bahdanau Attention
Covers alignment model formulation, score function (additive), attention weight computation, context vector as weighted sum, attention in decoder, Bahdanau attention implementation.
- 06Luong Attention
Covers dot product attention, general (bilinear) attention, concat attention variant, global vs local attention, Luong vs Bahdanau comparison, attention placement (input vs output).
- 07Copy Mechanism
Covers pointer network motivation, copy probability computation, mixing generation and copying, pointer-generator networks, copy mechanism for summarization, OOV handling with copy.
Volume 4 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 4 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.