Volume 5 of 20 · PDF edition
Available nowTransformers from First Principles
For engineers and researchers who want to derive, implement, and reason about every component of a transformer.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 27 written chapters
- Approximate pages
- ~860 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 5 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 5 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Self-Attention
- Positional Encoding
- Transformer Blocks
- Transformer Architectures
Audience and prerequisites
Where this volume fits
For engineers and researchers who want to derive, implement, and reason about every component of a transformer.
Prerequisites: Assumes the neural foundations in Volume 4.
Free online preview
Start with “Self-Attention Concept”
Covers cross-attention vs self-attention, self-attention motivation, all-pairs interaction, self-attention for representation learning, self-attention computational pattern.
Exact contents
27 chapters available now
Part XIII: Self-Attention
- 01Self-Attention Concept
Covers cross-attention vs self-attention, self-attention motivation, all-pairs interaction, self-attention for representation learning, self-attention computational pattern.
- 02Query, Key, Value
Covers QKV intuition (database lookup), projection matrices Wq, Wk, Wv, query-key matching, value retrieval, QKV dimensions and shapes, QKV as learned transformations.
- 03Scaled Dot-Product Attention
Covers dot product for similarity, softmax for normalization, scaling factor derivation (1/√dk), attention output computation, attention in matrix form, attention implementation.
- 04Attention Masking
Covers padding masks, causal (look-ahead) masks, combining multiple masks, mask shapes and broadcasting, efficient masking implementation, custom attention patterns.
- 05Multi-Head Attention
Covers multiple attention heads motivation, head dimension splitting, parallel attention computation, output concatenation and projection, head specialization, multi-head vs single head.
- 06Attention Complexity
Covers O(n²) attention complexity, memory requirements, attention bottleneck in long sequences, FLOPs computation, attention vs RNN complexity, practical scaling limits.
Part XIV: Positional Encoding
- 01Position Problem
Covers transformer position blindness, why position matters for language, position information requirements, position encoding vs position embedding, absolute vs relative position.
- 02Sinusoidal Position Encoding
Covers sinusoidal formula derivation, wavelength intuition, position encoding visualization, extrapolation properties, sinusoidal encoding implementation, learned vs sinusoidal trade-offs.
- 03Learned Position Embeddings
Covers position embedding table, position embedding training, maximum sequence length, learned embedding extrapolation, position embedding analysis, GPT-style position embeddings.
- 04Relative Position Encoding
Covers relative position motivation, relative attention formulation, clipping relative positions, relative position in self-attention, Shaw et al. relative positions, relative bias implementation.
- 05Rotary Position Embedding (RoPE)
Covers RoPE intuition, rotation matrix formulation, RoPE in complex numbers, relative position through rotation, RoPE implementation, RoPE frequency patterns.
- 06ALiBi
Covers ALiBi motivation, linear bias by distance, head-specific slopes, ALiBi extrapolation properties, ALiBi simplicity advantages, ALiBi vs RoPE comparison.
- 07Position Encoding Comparison
Covers extrapolation benchmarks, training efficiency comparison, implementation complexity, position encoding for long context, hybrid approaches, current best practices.
Part XV: Transformer Blocks
- 01Residual Connections
Covers residual connection formulation, gradient highway interpretation, residual scaling, residual connections in transformers, pre-norm vs post-norm residuals.
- 02Layer Normalization
Covers layer norm vs batch norm, layer norm formula, learnable affine parameters, layer norm placement, layer norm gradient flow, layer norm implementation.
- 03RMSNorm
Covers RMSNorm derivation, removing mean centering, RMSNorm efficiency, RMSNorm vs LayerNorm performance, RMSNorm in modern architectures.
- 04Pre-Norm vs Post-Norm
Covers original transformer (post-norm), pre-norm formulation, training stability comparison, gradient flow differences, when to use each, modern consensus.
- 05Feed-Forward Networks
Covers FFN architecture, hidden dimension expansion, FFN as two linear layers, position independence, FFN parameter count, FFN computational cost.
- 06FFN Activation Functions
Covers ReLU in original transformer, GELU adoption, GELU approximations, SiLU/Swish in modern models, activation function comparison.
- 07Gated Linear Units
Covers GLU formulation, gating mechanism, SwiGLU derivation, GeGLU variant, GLU parameter efficiency, GLU in modern architectures.
- 08Transformer Block Assembly
Covers standard block structure, component ordering, block implementation, block initialization, forward pass walkthrough, block hyperparameters.
Part XVI: Transformer Architectures
- 01Encoder Architecture
Covers encoder-only design, bidirectional self-attention, encoder for understanding tasks, encoder output usage, BERT-style encoder, encoder layer stacking.
- 02Decoder Architecture
Covers decoder-only design, causal masking requirement, autoregressive generation, decoder for generation tasks, GPT-style decoder, decoder layer stacking.
- 03Encoder-Decoder Architecture
Covers encoder-decoder interaction, cross-attention mechanism, encoder-decoder for seq2seq, T5-style architecture, information flow, when to use encoder-decoder.
- 04Cross-Attention
Covers cross-attention formulation, KV from encoder, Q from decoder, cross-attention masking, cross-attention placement, cross-attention implementation.
- 05Weight Tying
Covers input-output embedding tying, encoder-decoder tying, parameter reduction, weight tying effects on training, when to tie weights.
- 06Architecture Hyperparameters
Covers depth vs width trade-offs, number of heads selection, hidden dimension ratios, FFN expansion ratio, total parameter calculation, architecture search.
Volume 5 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 5 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.