Volume 14 of 20 · PDF edition
In progressEfficient Inference and On-Device Models
For engineers reducing latency, memory, energy, and cost from datacenter serving to private edge deployment.
Read every available chapter online for free. The paid edition is a focused, carefully typeset volume PDF with future chapters, updates, and errata included.
- Written chapters
- 20 written chapters
- Planned chapters
- 5 planned chapters
- Approximate pages
- ~640 pages
- Edition
- Version 2026.08.0
- Price
- $24 one-time

Author and edition details
About the author and Volume 14 PDF edition

Michael Brenndoerfer
Michael has spent more than a decade working across software engineering, data, AI, and business. He writes to understand difficult ideas more deeply and to share what he learns in a clear, practical way.
- Edition
- Volume 14 PDF 2026.08.0
- Published
- Last reviewed
Focused learning path
What this volume covers
- Model Compression
- Inference Optimization
- Small and On-Device Language Models
Audience and prerequisites
Where this volume fits
For engineers reducing latency, memory, energy, and cost from datacenter serving to private edge deployment.
Prerequisites: Assumes decoder architecture basics from Volume 9.
Free online preview
Start with “Knowledge Distillation”
Covers distillation objective, temperature in distillation, teacher selection, distillation for LLMs.
Exact contents
20 chapters available now
The current PDF contains every linked chapter below. The remaining 5 planned chapters will be added through free volume updates.
Part XLI: Model Compression
- 01Knowledge Distillation
Covers distillation objective, temperature in distillation, teacher selection, distillation for LLMs.
- 02Distillation Variants
Covers feature distillation, attention transfer, progressive distillation, on-policy distillation.
- 03Pruning Basics
Covers weight pruning, structured vs unstructured, pruning criteria, pruning schedule.
- 04Structured Pruning
Covers head pruning, layer pruning, width pruning, structured pruning implementation.
- 05Model Merging
Covers weight averaging, task arithmetic, TIES merging, DARE merging.
- 06Model Merging Applications
Covers multi-task merging, style merging, capability composition, merging evaluation.
Part XLII: Inference Optimization
- 01KV Cache
Covers KV cache motivation, cache structure, cache memory requirements, cache management.
- 02KV Cache Memory
Covers cache size calculation, batch size effects, sequence length effects, memory bottleneck.
- 03Paged Attention
Covers memory fragmentation problem, page-based allocation, vLLM approach, paged attention benefits.
- 04KV Cache Compression
Covers cache eviction strategies, attention sink preservation, H2O algorithm, cache quantization.
- 05Weight Quantization Basics
Covers quantization fundamentals, per-tensor vs per-channel, symmetric vs asymmetric, calibration.
- 06INT8 Quantization
Covers INT8 range mapping, absmax quantization, smooth quantization, INT8 accuracy.
- 07INT4 Quantization
Covers 4-bit challenges, group-wise quantization, 4-bit accuracy trade-offs, 4-bit formats.
- 08GPTQ
Covers GPTQ algorithm, layer-wise quantization, Hessian approximation, GPTQ implementation.
- 09AWQ
Covers salient weight preservation, AWQ algorithm, AWQ vs GPTQ, AWQ benefits.
- 10GGUF Format
Covers GGML/GGUF history, quantization types, GGUF file format, llama.cpp integration.
- 11Speculative Decoding: Fast LLM Inference Without Quality Loss
Covers speculative decoding concept, draft model selection, verification procedure, acceptance rate.
- 12Speculative Decoding Math: Algorithms & Speedup Limits
Covers acceptance criterion, expected speedup, draft quality effects, optimal draft length.
- 13Continuous Batching: Optimizing LLM Inference Throughput
Covers static vs continuous batching, iteration-level scheduling, request completion handling, throughput gains.
- 14LLM Inference Serving: Architecture, Routing & Auto-Scaling
Master LLM inference serving architecture, token-aware load balancing, and auto-scaling. Optimize time-to-first-token and throughput for production systems.
Part XLIII: Small and On-Device Language Models
- 01Small Language Model DesignPlanned
Capacity allocation, data quality, architecture choices, and capability boundaries below frontier scale.
- 02Hardware-Aware Model OptimizationPlanned
Memory bandwidth, kernels, quantization formats, sparsity, and co-design for phones, browsers, and edge devices.
- 03On-Device Adaptation and PersonalizationPlanned
Private fine-tuning, adapters, local retrieval, user control, and update strategies under tight budgets.
- 04Hybrid Edge-Cloud InferencePlanned
Routing, partitioning, privacy boundaries, offline fallbacks, latency, and cost-aware orchestration.
- 05Evaluating Small Models in ContextPlanned
Task fit, energy, latency, privacy, reliability, and comparing systems rather than parameter counts alone.
Volume 14 PDF
Own this focused edition
Get the carefully typeset PDF to keep, plus every future chapter, revision, and erratum for this volume.
- Carefully typeset standalone volume PDF
- Future chapters, updates, and errata included
- No oversized combined edition to render or download
Volume 14 of 20
$24
one-timeVersion 2026.08.0 · secure checkout via Stripe
Delivered by email · Free volume updates included
All-volume access
Every current and future volume
Get all 20 current volume slots, every future volume, and every update for $199. PDFs are always delivered as manageable individual volumes, never as one combined 12,500-page file.
$199 one-time
Get all-volume accessContinue through the library
Explore adjacent volumes
Version history
Kept current, not frozen in time
Each PDF purchase includes future editions. When the book changes, the updated copy appears in My books at no extra cost.
Current release
Edition 2026.08.0
Initial volume-library release with 19 written PDFs, a Volume 3 placeholder, and no combined edition.