Volume 6 of 20 · PDF edition
In progressEfficient, Long-Context, and Alternative Architectures
For readers choosing or designing architectures under sequence-length, memory, latency, and hardware constraints.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 16 written chapters
- Planned chapters
- 6 planned chapters
- Approximate pages
- ~510 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 6 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 6 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Efficient Attention
- Long Context
- Alternative Sequence and Generative Architectures
Audience and prerequisites
Where this volume fits
For readers choosing or designing architectures under sequence-length, memory, latency, and hardware constraints.
Prerequisites: Assumes Volume 5 or equivalent transformer knowledge.
Free online preview
Start with “Quadratic Attention Bottleneck”
Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.
Exact contents
16 chapters available now
The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.
Part XVII: Efficient Attention
- 01Quadratic Attention Bottleneck
Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.
- 02Sparse Attention Patterns
Covers local attention windows, strided attention patterns, block-sparse attention, combining sparse patterns, sparse attention implementation.
- 03Sliding Window Attention
Covers sliding window formulation, window size selection, dilated sliding windows, sliding window for long sequences, Mistral-style windowed attention.
- 04Global Tokens
Covers CLS token global attention, learned global tokens, global-local attention mixing, global token count, implementation strategies.
- 05Longformer
Covers Longformer attention pattern, global attention configuration, Longformer complexity, Longformer for documents, Longformer implementation.
- 06BigBird
Covers BigBird attention pattern, random attention benefits, BigBird theoretical guarantees, BigBird vs Longformer, BigBird applications.
- 07Linear Attention
Covers softmax attention reformulation, kernel feature maps, linear complexity attention, linear attention limitations, Performer and variants.
- 08FlashAttention Algorithm
Covers GPU memory hierarchy, tiling for SRAM, online softmax computation, recomputation strategy, FlashAttention complexity, FlashAttention benefits.
- 09FlashAttention Implementation
Covers CUDA kernel basics, memory access patterns, FlashAttention-2 improvements, using FlashAttention in PyTorch, FlashAttention limitations.
Part XVIII: Long Context
- 01Context Length Challenges
Covers training sequence length limits, attention memory scaling, position encoding extrapolation, long-range dependency learning, evaluation challenges.
- 02Position Interpolation
Covers linear position scaling, interpolation vs extrapolation, position interpolation implementation, fine-tuning for longer context, interpolation limitations.
- 03NTK-aware Scaling
Covers RoPE frequency analysis, high-frequency preservation, NTK-aware formula, dynamic NTK scaling, NTK vs linear interpolation.
- 04YaRN
Covers YaRN motivation, attention scaling factor, YaRN formula, YaRN training requirements, YaRN vs alternatives.
- 05Attention Sinks
Covers attention sink phenomenon, StreamingLLM approach, sink token design, streaming inference, infinite context generation.
- 06Memory Augmentation
Covers memory network concepts, memory retrieval mechanisms, memory writing and updating, memory-augmented transformers, Memorizing Transformers.
- 07Recurrent Memory
Covers Transformer-XL approach, segment-level processing, recurrent state passing, relative position in recurrence, recurrent memory limitations.
Part XIX: Alternative Sequence and Generative Architectures
- 01State Space Models for LanguagePlanned
Structured state spaces, selective state updates, Mamba-style sequence modeling, and linear-time claims.
- 02Hybrid Attention and State Space ModelsPlanned
Jamba-style hybrids, layer allocation, memory-throughput trade-offs, and architecture ablations.
- 03Token-Free and Byte-Level Language ModelsPlanned
Bytes, learned patches, dynamic segmentation, and the efficiency and robustness trade-offs of removing fixed tokenizers.
- 04Diffusion Language ModelsPlanned
Discrete diffusion, denoising objectives, parallel refinement, controllability, and likelihood evaluation.
- 05Blockwise and Semi-Autoregressive GenerationPlanned
Generating multiple tokens per step, speculative blocks, verification, and quality-latency trade-offs.
- 06Comparing Architecture Families FairlyPlanned
Matched-compute experiments, hardware efficiency, memory scaling, long-context quality, and benchmark leakage.
Volume 6 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 6 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.