Volume 2 of 3 · PDF edition
Volume 2: Neural Networks, Attention, and Transformers
The transition from structured prediction and benchmarks to embeddings, sequence-to-sequence learning, attention, transformers, BERT, and GPT.
The complete book remains free to read online. This paid edition is the carefully typeset PDF you keep, with every future update included.
- Pages
- 319 pages
- Chapters
- 34 chapters
- Edition
- Version 2026.08.2
- Price
- $14 one-time

Author and edition details
About the author and Volume 2 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 2 PDF 2026.08.2
- Published
- Last reviewed
Historical through-line
What this volume explains
- See how shared benchmarks and structured learning made progress measurable.
- Follow the sequence from word embeddings and neural translation to attention and transformer architectures.
- Understand why pretraining changed the economics and capabilities of language AI.
Audience and prerequisites
Who this volume is for
Practitioners and students who want the historical path into neural language models and transformers.
Prerequisites: Readable on its own; Volume 1 supplies useful background for the statistical methods referenced here.
Free online preview
Start with “Transformer Architecture (2017)”
Vaswani et al.'s 'Attention Is All You Need' introduced the transformer, replacing recurrence with self-attention and establishing the architecture that would dominate all of NLP
Read the chapterExact contents
34 chapters across 3 parts
Structured Learning & Benchmarks
Deep Learning Arrives
- 01IBM Watson on Jeopardy! (2011)
- 02Deep Learning for Speech Recognition (2012)
- 03Wikidata
- 04Word2Vec (2013)
- 05GloVe & Adam Optimizer (2014)
- 06Seq2Seq for MT (2014)
- 07Memory Networks (2014)
- 08Attention Mechanism (2015)
- 09Residual Connections (2015)
- 10Layer Normalization (2016)
- 11Subword Tokenization & FastText (2016)
- 12SQuAD (2016)
- 13Neural Information Retrieval
- 14Google Neural Machine Translation (2016)
- 15WaveNet (2016)
Transformers & Pretraining
Volume 2 PDF
Own this 319-page edition
Get the PDF and every future update and erratum for this volume. Your purchase is fully credited toward the complete edition.
- Carefully typeset PDF
- Future updates and errata included
- Portable offline reading
- Delivered by email and added to My books
Volume 2 of 3
$14
one-timeVersion 2026.08.2 · 319 pages · secure checkout via Stripe
Delivered by email · Free updates included
Compare editions
This volume or the complete book?
Volume 2: Neural Networks, Attention, and Transformers
319 pages focused on Structured Learning & Benchmarks and Deep Learning Arrives and Transformers & Pretraining.
$14
The complete History of Language AI
All three volumes in one 968-page PDF. Three distinct volume purchases unlock it automatically; two volumes leave a $6 complete-edition upgrade.
$34
Compare all PDF editionsAlso in the series
Explore the other volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.2
Editorial improvements.
Earlier releases2
2026.08.1
Editorial improvements throughout the book, with clearer explanations, refined presentation, and improved plots and visualizations.
2026.08.0
Initial complete-book and three-volume PDF release with perpetual updates.