Volume 5 of 20 · PDF edition

Available now

Transformers from First Principles

For engineers and researchers who want to derive, implement, and reason about every component of a transformer.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
27 written chapters
Approximate pages
~860 pages
Price
$24 one-time
Language AI Handbook, Volume 5: Transformers from First Principles cover

Author and edition details

About the author and Volume 5 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 5 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Self-Attention
  • Positional Encoding
  • Transformer Blocks
  • Transformer Architectures

Audience and prerequisites

Where this volume fits

For engineers and researchers who want to derive, implement, and reason about every component of a transformer.

Prerequisites: Assumes the neural foundations in Volume 4.

Free online preview

Start with “Self-Attention Concept

Covers cross-attention vs self-attention, self-attention motivation, all-pairs interaction, self-attention for representation learning, self-attention computational pattern.

Read the chapter

Exact contents

27 chapters available now

Part XIII: Self-Attention

  1. 01
    Self-Attention Concept

    Covers cross-attention vs self-attention, self-attention motivation, all-pairs interaction, self-attention for representation learning, self-attention computational pattern.

  2. 02
    Query, Key, Value

    Covers QKV intuition (database lookup), projection matrices Wq, Wk, Wv, query-key matching, value retrieval, QKV dimensions and shapes, QKV as learned transformations.

  3. 03
    Scaled Dot-Product Attention

    Covers dot product for similarity, softmax for normalization, scaling factor derivation (1/√dk), attention output computation, attention in matrix form, attention implementation.

  4. 04
    Attention Masking

    Covers padding masks, causal (look-ahead) masks, combining multiple masks, mask shapes and broadcasting, efficient masking implementation, custom attention patterns.

  5. 05
    Multi-Head Attention

    Covers multiple attention heads motivation, head dimension splitting, parallel attention computation, output concatenation and projection, head specialization, multi-head vs single head.

  6. 06
    Attention Complexity

    Covers O(n²) attention complexity, memory requirements, attention bottleneck in long sequences, FLOPs computation, attention vs RNN complexity, practical scaling limits.

Part XIV: Positional Encoding

  1. 01
    Position Problem

    Covers transformer position blindness, why position matters for language, position information requirements, position encoding vs position embedding, absolute vs relative position.

  2. 02
    Sinusoidal Position Encoding

    Covers sinusoidal formula derivation, wavelength intuition, position encoding visualization, extrapolation properties, sinusoidal encoding implementation, learned vs sinusoidal trade-offs.

  3. 03
    Learned Position Embeddings

    Covers position embedding table, position embedding training, maximum sequence length, learned embedding extrapolation, position embedding analysis, GPT-style position embeddings.

  4. 04
    Relative Position Encoding

    Covers relative position motivation, relative attention formulation, clipping relative positions, relative position in self-attention, Shaw et al. relative positions, relative bias implementation.

  5. 05
    Rotary Position Embedding (RoPE)

    Covers RoPE intuition, rotation matrix formulation, RoPE in complex numbers, relative position through rotation, RoPE implementation, RoPE frequency patterns.

  6. 06
    ALiBi

    Covers ALiBi motivation, linear bias by distance, head-specific slopes, ALiBi extrapolation properties, ALiBi simplicity advantages, ALiBi vs RoPE comparison.

  7. 07
    Position Encoding Comparison

    Covers extrapolation benchmarks, training efficiency comparison, implementation complexity, position encoding for long context, hybrid approaches, current best practices.

Part XV: Transformer Blocks

  1. 01
    Residual Connections

    Covers residual connection formulation, gradient highway interpretation, residual scaling, residual connections in transformers, pre-norm vs post-norm residuals.

  2. 02
    Layer Normalization

    Covers layer norm vs batch norm, layer norm formula, learnable affine parameters, layer norm placement, layer norm gradient flow, layer norm implementation.

  3. 03
    RMSNorm

    Covers RMSNorm derivation, removing mean centering, RMSNorm efficiency, RMSNorm vs LayerNorm performance, RMSNorm in modern architectures.

  4. 04
    Pre-Norm vs Post-Norm

    Covers original transformer (post-norm), pre-norm formulation, training stability comparison, gradient flow differences, when to use each, modern consensus.

  5. 05
    Feed-Forward Networks

    Covers FFN architecture, hidden dimension expansion, FFN as two linear layers, position independence, FFN parameter count, FFN computational cost.

  6. 06
    FFN Activation Functions

    Covers ReLU in original transformer, GELU adoption, GELU approximations, SiLU/Swish in modern models, activation function comparison.

  7. 07
    Gated Linear Units

    Covers GLU formulation, gating mechanism, SwiGLU derivation, GeGLU variant, GLU parameter efficiency, GLU in modern architectures.

  8. 08
    Transformer Block Assembly

    Covers standard block structure, component ordering, block implementation, block initialization, forward pass walkthrough, block hyperparameters.

Part XVI: Transformer Architectures

  1. 01
    Encoder Architecture

    Covers encoder-only design, bidirectional self-attention, encoder for understanding tasks, encoder output usage, BERT-style encoder, encoder layer stacking.

  2. 02
    Decoder Architecture

    Covers decoder-only design, causal masking requirement, autoregressive generation, decoder for generation tasks, GPT-style decoder, decoder layer stacking.

  3. 03
    Encoder-Decoder Architecture

    Covers encoder-decoder interaction, cross-attention mechanism, encoder-decoder for seq2seq, T5-style architecture, information flow, when to use encoder-decoder.

  4. 04
    Cross-Attention

    Covers cross-attention formulation, KV from encoder, Q from decoder, cross-attention masking, cross-attention placement, cross-attention implementation.

  5. 05
    Weight Tying

    Covers input-output embedding tying, encoder-decoder tying, parameter reduction, weight tying effects on training, when to tie weights.

  6. 06
    Architecture Hyperparameters

    Covers depth vs width trade-offs, number of heads selection, hidden dimension ratios, FFN expansion ratio, total parameter calculation, architecture search.

Volume 5 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 5 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.