Volume 14 of 20 · PDF edition

In progress

Efficient Inference and On-Device Models

For engineers reducing latency, memory, energy, and cost from datacenter serving to private edge deployment.

Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.

Written chapters
20 written chapters
Planned chapters
5 planned chapters
Approximate pages
~640 pages
Price
$24 one-time
Language AI Handbook, Volume 14: Efficient Inference and On-Device Models cover

Author and edition details

About the author and Volume 14 PDF edition

Michael Brenndoerfer, author of Language AI Handbook

Michael Brenndoerfer

Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.

Edition
Volume 14 PDF 2026.08.0
Published
Last reviewed

Focused learning path

What this volume covers

  • Model Compression
  • Inference Optimization
  • Small and On-Device Language Models

Audience and prerequisites

Where this volume fits

For engineers reducing latency, memory, energy, and cost from datacenter serving to private edge deployment.

Prerequisites: Assumes decoder architecture basics from Volume 9.

Free online preview

Start with “Knowledge Distillation

Covers distillation objective, temperature in distillation, teacher selection, distillation for LLMs.

Read the chapter

Exact contents

20 chapters available now

The current PDF contains every linked chapter below. The remaining 5 planned chapters will be added through free volume updates.

Part XLI: Model Compression

  1. 01
    Knowledge Distillation

    Covers distillation objective, temperature in distillation, teacher selection, distillation for LLMs.

  2. 02
    Distillation Variants

    Covers feature distillation, attention transfer, progressive distillation, on-policy distillation.

  3. 03
    Pruning Basics

    Covers weight pruning, structured vs unstructured, pruning criteria, pruning schedule.

  4. 04
    Structured Pruning

    Covers head pruning, layer pruning, width pruning, structured pruning implementation.

  5. 05
    Model Merging

    Covers weight averaging, task arithmetic, TIES merging, DARE merging.

  6. 06
    Model Merging Applications

    Covers multi-task merging, style merging, capability composition, merging evaluation.

Part XLII: Inference Optimization

  1. 01
    KV Cache

    Covers KV cache motivation, cache structure, cache memory requirements, cache management.

  2. 02
    KV Cache Memory

    Covers cache size calculation, batch size effects, sequence length effects, memory bottleneck.

  3. 03
    Paged Attention

    Covers memory fragmentation problem, page-based allocation, vLLM approach, paged attention benefits.

  4. 04
    KV Cache Compression

    Covers cache eviction strategies, attention sink preservation, H2O algorithm, cache quantization.

  5. 05
    Weight Quantization Basics

    Covers quantization fundamentals, per-tensor vs per-channel, symmetric vs asymmetric, calibration.

  6. 06
    INT8 Quantization

    Covers INT8 range mapping, absmax quantization, smooth quantization, INT8 accuracy.

  7. 07
    INT4 Quantization

    Covers 4-bit challenges, group-wise quantization, 4-bit accuracy trade-offs, 4-bit formats.

  8. 08
    GPTQ

    Covers GPTQ algorithm, layer-wise quantization, Hessian approximation, GPTQ implementation.

  9. 09
    AWQ

    Covers salient weight preservation, AWQ algorithm, AWQ vs GPTQ, AWQ benefits.

  10. 10
    GGUF Format

    Covers GGML/GGUF history, quantization types, GGUF file format, llama.cpp integration.

  11. 11
    Speculative Decoding: Fast LLM Inference Without Quality Loss

    Covers speculative decoding concept, draft model selection, verification procedure, acceptance rate.

  12. 12
    Speculative Decoding Math: Algorithms & Speedup Limits

    Covers acceptance criterion, expected speedup, draft quality effects, optimal draft length.

  13. 13
    Continuous Batching: Optimizing LLM Inference Throughput

    Covers static vs continuous batching, iteration-level scheduling, request completion handling, throughput gains.

  14. 14
    LLM Inference Serving: Architecture, Routing & Auto-Scaling

    Master LLM inference serving architecture, token-aware load balancing, and auto-scaling. Optimize time-to-first-token and throughput for production systems.

Part XLIII: Small and On-Device Language Models

  1. 01
    Small Language Model DesignPlanned

    Capacity allocation, data quality, architecture choices, and capability boundaries below frontier scale.

  2. 02
    Hardware-Aware Model OptimizationPlanned

    Memory bandwidth, kernels, quantization formats, sparsity, and co-design for phones, browsers, and edge devices.

  3. 03
    On-Device Adaptation and PersonalizationPlanned

    Private fine-tuning, adapters, local retrieval, user control, and update strategies under tight budgets.

  4. 04
    Hybrid Edge-Cloud InferencePlanned

    Routing, partitioning, privacy boundaries, offline fallbacks, latency, and cost-aware orchestration.

  5. 05
    Evaluating Small Models in ContextPlanned

    Task fit, energy, latency, privacy, reliability, and comparing systems rather than parameter counts alone.

Volume 14 PDF

Own this focused edition

Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.

  • Carefully typeset standalone volume PDF
  • Future chapters, updates, and errata included
  • No oversized combined edition to render or download

Volume 14 of 20

$24

one-time

Version 2026.08.0 · secure checkout via Stripe

Delivered by email · Free volume updates included

All-volume access

Every current and future volume

Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.

Continue through the library

Explore adjacent volumes

Version history

Kept current, not frozen in time

Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.

Current release

Edition 2026.08.0

Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.