Volume 17 of 20 · PDF edition

In progress

Speech, Multimodal, and Embodied Language

For readers building systems that see, hear, speak, understand documents and video, or connect language to action.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
17 written chapters
Planned chapters
7 planned chapters
Approximate pages
~540 pages
Price
$24 one-time
Language AI Handbook, Volume 17: Speech, Multimodal, and Embodied Language cover

Author and edition details

About the author and Volume 17 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 17 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Multimodal Models
  • Speech and Audio
  • Omni-Modal and Embodied Language Systems

Audience and prerequisites

Where this volume fits

For readers building systems that see, hear, speak, understand documents and video, or connect language to action.

Prerequisites: Assumes transformer and model-family foundations from Volumes 5, 8, and 9.

Free online preview

Start with “Vision Transformer

Covers image patching, patch embeddings, ViT architecture, ViT pre-training.

Read the chapter

Exact contents

17 chapters available now

The current PDF contains every linked chapter below. The remaining 7 planned chapters will be added through free volume updates.

Part L: Multimodal Models

  1. 01
    Vision Transformer

    Covers image patching, patch embeddings, ViT architecture, ViT pre-training.

  2. 02
    CLIP

    Covers CLIP architecture, CLIP training objective, CLIP zero-shot classification, CLIP embeddings.

  3. 03
    Vision Encoders for VLMs

    Covers ViT variants for VLMs, SigLIP improvements, image resolution handling, encoder selection.

  4. 04
    Vision-Language Projection

    Covers linear projection, MLP projection, Q-Former approach, projection training.

  5. 05
    LLaVA Architecture

    Covers LLaVA design, two-stage training, visual conversation, LLaVA variants.

  6. 06
    Flamingo Architecture

    Covers cross-attention to images, gated cross-attention, few-shot visual learning, Flamingo training.

  7. 07
    Multimodal Training Data

    Covers image-text pairs, interleaved documents, visual instruction data, data quality.

  8. 08
    Multimodal Evaluation

    Covers VQA benchmarks, multimodal understanding benchmarks, multimodal generation evaluation.

  9. 09
    Multimodal Foundations

    Covers multimodal learning principles, cross-modal alignment, joint embedding spaces, multimodal challenges.

  10. 10
    Image Understanding

    Covers visual question answering, image captioning, visual grounding, scene understanding.

  11. 11
    Image Generation

    Covers diffusion models, text-to-image generation, image editing, generation evaluation.

  12. 12
    Multimodal Applications

    Covers document understanding, medical imaging, video understanding, multimodal reasoning tasks.

Part LI: Speech and Audio

  1. 01
    Speech Representations

    Covers mel spectrograms, mel filterbanks, feature normalization, audio preprocessing.

  2. 02
    Whisper Architecture

    Covers Whisper encoder-decoder, multitask training, language tokens, timestamp prediction.

  3. 03
    Whisper Training

    Covers Whisper training data, weak supervision, multilingual training, Whisper capabilities.

  4. 04
    Speech-Language Integration

    Covers speech encoder + LLM, audio tokens, speech-to-text-to-LLM vs end-to-end, speech LLM architectures.

  5. 05
    Text-to-Speech

    Covers TTS architecture overview, vocoder role, TTS quality metrics, neural TTS approaches.

Part LII: Omni-Modal and Embodied Language Systems

  1. 01
    Unified Audio-Visual-Language ModelingPlanned

    Modality encoders, shared token spaces, fusion, generation, and interference in omni models.

  2. 02
    Video and Temporal ReasoningPlanned

    Frame sampling, event localization, temporal memory, causal reasoning, and long-video evaluation.

  3. 03
    Streaming Multimodal InteractionPlanned

    Continuous perception, interruption, turn-taking, proactive responses, latency, and duplex generation.

  4. 04
    Document, Diagram, and Interface UnderstandingPlanned

    OCR, layout, charts, diagrams, screenshots, grounding, and visually situated tool use.

  5. 05
    Speech-to-Speech and Expressive VoicePlanned

    Direct speech interaction, prosody, speaker identity, emotion, latency, and voice-specific safety.

  6. 06
    Vision-Language-Action ModelsPlanned

    Grounded language for robotics, action tokenization, policy learning, embodiment, and simulation-to-reality gaps.

  7. 07
    Multimodal Safety and EvaluationPlanned

    Cross-modal injection, privacy, deepfakes, modality imbalance, interactive benchmarks, and human evaluation.

Volume 17 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 17 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.