Volume 17 of 20 · PDF edition
In progressSpeech, Multimodal, and Embodied Language
For readers building systems that see, hear, speak, understand documents and video, or connect language to action.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 17 written chapters
- Planned chapters
- 7 planned chapters
- Approximate pages
- ~540 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 17 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 17 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Multimodal Models
- Speech and Audio
- Omni-Modal and Embodied Language Systems
Audience and prerequisites
Where this volume fits
For readers building systems that see, hear, speak, understand documents and video, or connect language to action.
Prerequisites: Assumes transformer and model-family foundations from Volumes 5, 8, and 9.
Free online preview
Start with “Vision Transformer”
Covers image patching, patch embeddings, ViT architecture, ViT pre-training.
Exact contents
17 chapters available now
The current PDF contains every linked chapter below. The remaining 7 planned chapters will be added through free volume updates.
Part L: Multimodal Models
- 01Vision Transformer
Covers image patching, patch embeddings, ViT architecture, ViT pre-training.
- 02CLIP
Covers CLIP architecture, CLIP training objective, CLIP zero-shot classification, CLIP embeddings.
- 03Vision Encoders for VLMs
Covers ViT variants for VLMs, SigLIP improvements, image resolution handling, encoder selection.
- 04Vision-Language Projection
Covers linear projection, MLP projection, Q-Former approach, projection training.
- 05LLaVA Architecture
Covers LLaVA design, two-stage training, visual conversation, LLaVA variants.
- 06Flamingo Architecture
Covers cross-attention to images, gated cross-attention, few-shot visual learning, Flamingo training.
- 07Multimodal Training Data
Covers image-text pairs, interleaved documents, visual instruction data, data quality.
- 08Multimodal Evaluation
Covers VQA benchmarks, multimodal understanding benchmarks, multimodal generation evaluation.
- 09Multimodal Foundations
Covers multimodal learning principles, cross-modal alignment, joint embedding spaces, multimodal challenges.
- 10Image Understanding
Covers visual question answering, image captioning, visual grounding, scene understanding.
- 11Image Generation
Covers diffusion models, text-to-image generation, image editing, generation evaluation.
- 12Multimodal Applications
Covers document understanding, medical imaging, video understanding, multimodal reasoning tasks.
Part LI: Speech and Audio
- 01Speech Representations
Covers mel spectrograms, mel filterbanks, feature normalization, audio preprocessing.
- 02Whisper Architecture
Covers Whisper encoder-decoder, multitask training, language tokens, timestamp prediction.
- 03Whisper Training
Covers Whisper training data, weak supervision, multilingual training, Whisper capabilities.
- 04Speech-Language Integration
Covers speech encoder + LLM, audio tokens, speech-to-text-to-LLM vs end-to-end, speech LLM architectures.
- 05Text-to-Speech
Covers TTS architecture overview, vocoder role, TTS quality metrics, neural TTS approaches.
Part LII: Omni-Modal and Embodied Language Systems
- 01Unified Audio-Visual-Language ModelingPlanned
Modality encoders, shared token spaces, fusion, generation, and interference in omni models.
- 02Video and Temporal ReasoningPlanned
Frame sampling, event localization, temporal memory, causal reasoning, and long-video evaluation.
- 03Streaming Multimodal InteractionPlanned
Continuous perception, interruption, turn-taking, proactive responses, latency, and duplex generation.
- 04Document, Diagram, and Interface UnderstandingPlanned
OCR, layout, charts, diagrams, screenshots, grounding, and visually situated tool use.
- 05Speech-to-Speech and Expressive VoicePlanned
Direct speech interaction, prosody, speaker identity, emotion, latency, and voice-specific safety.
- 06Vision-Language-Action ModelsPlanned
Grounded language for robotics, action tokenization, policy learning, embodiment, and simulation-to-reality gaps.
- 07Multimodal Safety and EvaluationPlanned
Cross-modal injection, privacy, deepfakes, modality imbalance, interactive benchmarks, and human evaluation.
Volume 17 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 17 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.