Volume 6 of 20 · PDF edition

In progress

Efficient, Long-Context, and Alternative Architectures

For readers choosing or designing architectures under sequence-length, memory, latency, and hardware constraints.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
16 written chapters
Planned chapters
6 planned chapters
Approximate pages
~510 pages
Price
$24 one-time
Language AI Handbook, Volume 6: Efficient, Long-Context, and Alternative Architectures cover

Author and edition details

About the author and Volume 6 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 6 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Efficient Attention
  • Long Context
  • Alternative Sequence and Generative Architectures

Audience and prerequisites

Where this volume fits

For readers choosing or designing architectures under sequence-length, memory, latency, and hardware constraints.

Prerequisites: Assumes Volume 5 or equivalent transformer knowledge.

Free online preview

Start with “Quadratic Attention Bottleneck

Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.

Read the chapter

Exact contents

16 chapters available now

The current PDF contains every linked chapter below. The remaining 6 planned chapters will be added through free volume updates.

Part XVII: Efficient Attention

  1. 01
    Quadratic Attention Bottleneck

    Covers O(n²) memory analysis, O(n²) compute analysis, attention matrix size, practical sequence limits, bottleneck visualization, motivation for efficiency.

  2. 02
    Sparse Attention Patterns

    Covers local attention windows, strided attention patterns, block-sparse attention, combining sparse patterns, sparse attention implementation.

  3. 03
    Sliding Window Attention

    Covers sliding window formulation, window size selection, dilated sliding windows, sliding window for long sequences, Mistral-style windowed attention.

  4. 04
    Global Tokens

    Covers CLS token global attention, learned global tokens, global-local attention mixing, global token count, implementation strategies.

  5. 05
    Longformer

    Covers Longformer attention pattern, global attention configuration, Longformer complexity, Longformer for documents, Longformer implementation.

  6. 06
    BigBird

    Covers BigBird attention pattern, random attention benefits, BigBird theoretical guarantees, BigBird vs Longformer, BigBird applications.

  7. 07
    Linear Attention

    Covers softmax attention reformulation, kernel feature maps, linear complexity attention, linear attention limitations, Performer and variants.

  8. 08
    FlashAttention Algorithm

    Covers GPU memory hierarchy, tiling for SRAM, online softmax computation, recomputation strategy, FlashAttention complexity, FlashAttention benefits.

  9. 09
    FlashAttention Implementation

    Covers CUDA kernel basics, memory access patterns, FlashAttention-2 improvements, using FlashAttention in PyTorch, FlashAttention limitations.

Part XVIII: Long Context

  1. 01
    Context Length Challenges

    Covers training sequence length limits, attention memory scaling, position encoding extrapolation, long-range dependency learning, evaluation challenges.

  2. 02
    Position Interpolation

    Covers linear position scaling, interpolation vs extrapolation, position interpolation implementation, fine-tuning for longer context, interpolation limitations.

  3. 03
    NTK-aware Scaling

    Covers RoPE frequency analysis, high-frequency preservation, NTK-aware formula, dynamic NTK scaling, NTK vs linear interpolation.

  4. 04
    YaRN

    Covers YaRN motivation, attention scaling factor, YaRN formula, YaRN training requirements, YaRN vs alternatives.

  5. 05
    Attention Sinks

    Covers attention sink phenomenon, StreamingLLM approach, sink token design, streaming inference, infinite context generation.

  6. 06
    Memory Augmentation

    Covers memory network concepts, memory retrieval mechanisms, memory writing and updating, memory-augmented transformers, Memorizing Transformers.

  7. 07
    Recurrent Memory

    Covers Transformer-XL approach, segment-level processing, recurrent state passing, relative position in recurrence, recurrent memory limitations.

Part XIX: Alternative Sequence and Generative Architectures

  1. 01
    State Space Models for LanguagePlanned

    Structured state spaces, selective state updates, Mamba-style sequence modeling, and linear-time claims.

  2. 02
    Hybrid Attention and State Space ModelsPlanned

    Jamba-style hybrids, layer allocation, memory-throughput trade-offs, and architecture ablations.

  3. 03
    Token-Free and Byte-Level Language ModelsPlanned

    Bytes, learned patches, dynamic segmentation, and the efficiency and robustness trade-offs of removing fixed tokenizers.

  4. 04
    Diffusion Language ModelsPlanned

    Discrete diffusion, denoising objectives, parallel refinement, controllability, and likelihood evaluation.

  5. 05
    Blockwise and Semi-Autoregressive GenerationPlanned

    Generating multiple tokens per step, speculative blocks, verification, and quality-latency trade-offs.

  6. 06
    Comparing Architecture Families FairlyPlanned

    Matched-compute experiments, hardware efficiency, memory scaling, long-context quality, and benchmark leakage.

Volume 6 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 6 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.