Model Serving: Frameworks, Loading

Michael BrenndoerferFebruary 6, 202657 min read

Part of Language AI Handbook

Serve language models in production: weight loading strategies, tensor parallelism, request batching, rate limiting.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Model Serving

Training a language model is only half the work. The other half is getting it into the hands of users reliably, at scale, and within latency budgets that real applications demand. Model serving is the discipline of taking a trained model and turning it into a running service that handles requests from clients, manages computational resources, and behaves predictably under load.

The challenges of serving language models differ substantially from serving traditional machine learning models. A sentiment classifier produces one prediction per input in a fixed amount of compute. A language model generates tokens one at a time in a loop, making each request variable in length, variable in cost, and fundamentally sequential in its output. A request for a two-sentence summary consumes far less GPU time and memory than a request for a detailed report, yet both enter the same request queue and compete for the same hardware. Serving infrastructure must account for this variability at every layer.

This variability drives the design of every major serving technique. When requests have predictable, uniform cost, you can plan capacity with simple arithmetic. When each request has an unknown compute cost, you need adaptive queuing, batching, and memory management to keep the GPU busy without exhausting memory or exceeding latency budgets.

This chapter covers the full stack of model serving: loading and initializing models in memory, accepting and routing requests, configuring serving parameters, and the frameworks that tie it all together. As we saw in Part XXXI on training infrastructure, getting software and hardware to work together efficiently requires careful configuration at every layer. Serving is no different, but the optimization objectives shift from training throughput to inference latency and concurrency.

Model Loading and Initialization

Before any request can be served, the model must be loaded into memory and prepared for inference. This sounds straightforward but involves several design decisions that affect startup time, memory usage, and steady-state serving performance.

Weight Loading Strategies

The most basic approach is loading all model weights into GPU memory at startup. For a 7-billion-parameter model in 16-bit floating point, this requires approximately 14 GB of GPU memory just for weights, before accounting for activations and KV caches during inference. Larger models like 70B or 405B parameters require distributed loading across multiple GPUs.

The weight loading process involves several stages. First, the checkpoint files are read from disk. PyTorch checkpoints are typically stored as sharded files that are memory-mapped to avoid loading everything into CPU RAM simultaneously. Memory mapping lets the operating system handle the actual data movement on demand, paging in portions of the weight file only when the corresponding tensor is accessed. Then weights are transferred to GPU memory, which involves PCIe bandwidth if the GPU is discrete. For large models, this transfer alone can take tens of seconds.

Lazy loading defers this work, loading tensor shards only when they are first accessed. This reduces startup time but introduces latency spikes on first access, which is unacceptable in production where the first user request should not bear the cost of initializing the system. In practice, inference servers perform eager loading with warmup requests that pre-populate caches and ensure all code paths have been compiled before serving real traffic.

Why warmup requests? Many deep learning frameworks use just-in-time compilation: the first time a particular sequence length or batch size is encountered, the framework compiles a specialized CUDA kernel for that shape. Subsequent requests with the same shape reuse the compiled kernel and are much faster. Running warmup requests before opening the server to production traffic forces this compilation to happen in advance, preventing the first real users from experiencing artificially slow responses.

In[3]:
Code
def simulate_model_loading(
    num_params_billions: float, dtype_bytes: int = 2
) -> dict:
    """
    Simulate the memory and time cost of loading model weights.
    dtype_bytes=2 for float16, 4 for float32.
    """
    params = num_params_billions * 1e9
    weight_bytes = params * dtype_bytes
    weight_gb = weight_bytes / (1024**3)

    # Rough estimates based on NVMe SSD (~3 GB/s) and PCIe bandwidth (~25 GB/s)
    disk_read_time_s = weight_gb / 3.0
    pcie_transfer_time_s = weight_gb / 25.0
    total_load_time_s = disk_read_time_s + pcie_transfer_time_s

    return {
        "model_size_gb": weight_gb,
        "disk_read_time_s": disk_read_time_s,
        "pcie_transfer_time_s": pcie_transfer_time_s,
        "total_load_time_s": total_load_time_s,
    }


model_sizes = [7, 13, 34, 70]
loading_stats = {size: simulate_model_loading(size) for size in model_sizes}
Out[4]:
Console
   Model | Weights (GB) |  Disk Read | PCIe Transfer | Total Load
-----------------------------------------------------------------
7B params |         13.0 |      4.3s |         0.5s |      4.9s
13B params |         24.2 |      8.1s |         1.0s |      9.0s
34B params |         63.3 |     21.1s |         2.5s |     23.6s
70B params |        130.4 |     43.5s |         5.2s |     48.7s

These estimates reveal why model serving infrastructure is designed to avoid frequent restarts. A 70B model takes over a minute just to load from disk and transfer to GPU memory, not counting model initialization and warmup. Production serving systems aim for zero-downtime deployments that gradually shift traffic to new instances rather than restarting existing ones. A cold start causes seconds of unavailability and delays the point when a new instance can absorb its share of production traffic. Other instances must carry additional load during that window.

Safetensors is an increasingly common checkpoint format that addresses some of these loading bottlenecks. Unlike PyTorch's native format, safetensors files store tensor data with a simple binary layout and a JSON header describing tensor shapes and offsets. This allows memory-mapped loading without any deserialization overhead, and different tensors can be loaded by different threads simultaneously because their file offsets are independent. For large models, safetensors reduces load time by up to 40% compared to traditional PyTorch checkpoints.

Tensor Parallelism for Large Models

When a model is too large for a single GPU, tensor parallelism splits individual weight matrices across multiple GPUs. Each GPU holds a shard of every layer's weights, and all-reduce operations synchronize intermediate results between GPUs during each forward pass.

Consider a linear layer with weight matrix W∈Rd×h\mathbf{W} \in \mathbb{R}^{d \times h}. With tensor parallelism across NN GPUs, each GPU holds a column shard Wi∈Rd×h/N\mathbf{W}_i \in \mathbb{R}^{d \times h/N}. For an input x\mathbf{x}, each GPU computes its partial output yi=xWi\mathbf{y}_i = \mathbf{x} \mathbf{W}_i, then the results are concatenated or summed across GPUs using collective operations.

where:

  • dd: the input dimension of the linear layer
  • hh: the output dimension of the linear layer
  • NN: the number of GPUs participating in tensor parallelism
  • Wi\mathbf{W}_i: the weight shard held on GPU ii, containing columns [(i−1)h/N,ih/N)[(i-1)h/N, ih/N)
  • yi\mathbf{y}_i: the partial output computed on GPU ii

The trade-off is communication overhead. Every layer requires GPU-to-GPU communication, which adds latency proportional to the number of layers and the bandwidth between GPUs. NVLink (on server-class hardware) provides much higher bandwidth than PCIe, making tensor parallelism more viable on GPU clusters connected with high-speed interconnects. On PCIe-only systems, the communication overhead for 8-way tensor parallelism can add tens of milliseconds per forward pass, enough to substantially increase response latency for interactive applications.

Pipeline parallelism is an alternative that assigns different layers to different GPUs rather than splitting individual layers. This reduces inter-GPU communication to just the activations passed between stages, but introduces pipeline bubbles where GPUs wait idle while upstream stages complete. Practical systems often combine tensor and pipeline parallelism for very large models: tensor parallelism within a node (where GPUs share NVLink) and pipeline parallelism across nodes (where communication is limited to Ethernet or InfiniBand).

In[5]:
Code
def estimate_tensor_parallel_overhead(
    num_gpus: int, layer_hidden_dim: int, nvlink: bool = True
) -> dict:
    """
    Estimate communication overhead per layer with tensor parallelism.
    All-reduce communicates 2 * (N-1)/N * message_size bytes total.
    """
    # Activation size in bytes (float16, batch=1, seq=512)
    activation_bytes = 512 * layer_hidden_dim * 2

    # All-reduce volume: 2 * (N-1)/N * size
    allreduce_bytes = 2 * ((num_gpus - 1) / num_gpus) * activation_bytes

    # Bandwidth: NVLink ~600 GB/s bidirectional, PCIe ~25 GB/s
    bandwidth_gbs = 600 if nvlink else 25
    allreduce_time_us = (allreduce_bytes / (bandwidth_gbs * 1e9)) * 1e6

    return {
        "num_gpus": num_gpus,
        "activation_mb": activation_bytes / (1024**2),
        "allreduce_bytes": allreduce_bytes,
        "allreduce_time_us": allreduce_time_us,
        "interconnect": "NVLink" if nvlink else "PCIe",
    }


gpu_configs = [2, 4, 8]
dim = 8192  # typical hidden dim for a 70B model
results_nvlink = [
    estimate_tensor_parallel_overhead(n, dim, nvlink=True) for n in gpu_configs
]
results_pcie = [
    estimate_tensor_parallel_overhead(n, dim, nvlink=False) for n in gpu_configs
]
Out[6]:
Console
Tensor Parallelism Communication Overhead per Layer
  Hidden dimension: 8192, Batch=1, Seq=512 tokens

GPUs |  NVLink (us) |    PCIe (us)
-----------------------------------
   2 |         14.0 |        335.5
   4 |         21.0 |        503.3
   8 |         24.5 |        587.2

For a 80-layer model, multiply per-layer overhead by 80
  8-GPU NVLink total comm: 1957 us = 2.0 ms
  8-GPU PCIe total comm:   46976 us = 47.0 ms

The difference between NVLink and PCIe interconnects becomes critical for large-scale tensor parallelism. On PCIe, the communication overhead for an 8-GPU configuration with 80 layers adds tens of milliseconds per forward pass, substantially increasing latency. NVLink reduces this to a few milliseconds, making 8-way tensor parallelism practical for serving. This is one reason that server-class GPU machines with NVLink (such as DGX systems) command a substantial price premium over workstation-class hardware with the same number of GPUs but only PCIe connectivity.

Memory Budgeting

Planning memory allocation before deployment prevents out-of-memory failures under load. GPU memory serves three main consumers during LLM inference.

Model weights occupy a fixed amount of memory determined by parameter count and data type. A 7B-parameter model in float16 uses 14 GB. Using int8 quantization (as discussed in Part XXXI) halves this to 7 GB, and int4 quantization halves it again to 3.5 GB, at the cost of some quality degradation.

Activations are the intermediate tensors computed during the forward pass. Their size depends on batch size and sequence length. For autoregressive decoding, activations are relatively small because only one token is processed per step. During prefill, however, the entire prompt is processed in parallel, and activation memory scales with sequence length. A system that accepts very long prompts must plan for the activation spike that occurs when a large prefill batch is processed.

The KV cache stores the key and value tensors for all tokens in currently active sequences. This is typically the most dynamic component: more concurrent requests means more KV cache consumption. Understanding the KV cache formula is essential for capacity planning. The KV cache size per request is:

KV cache bytes=2×n_layers×n_heads×d_head×seq_len×bytes_per_element\text{KV cache bytes} = 2 \times n\_layers \times n\_heads \times d\_head \times \text{seq\_len} \times \text{bytes\_per\_element}

where:

  • 22: one tensor for keys, one for values
  • n_layersn\_layers: number of transformer layers in the model
  • n_headsn\_heads: number of attention heads per layer
  • d_headd\_head: dimensionality of each attention head
  • seq_len\text{seq\_len}: current sequence length for the request
  • bytes_per_element\text{bytes\_per\_element}: 2 for float16, 1 for int8

Notice that n_heads×d_head=d_modeln\_heads \times d\_head = d\_model, the model's hidden dimension. So the formula simplifies to 2×n_layers×d_model×seq_len×bytes2 \times n\_layers \times d\_model \times \text{seq\_len} \times \text{bytes}, which makes intuitive sense: each of the n_layersn\_layers layers stores a full-hidden-dimension key and value vector for every token position. As sequences grow longer or as more requests are served concurrently, KV cache consumption grows proportionally.

In[7]:
Code
def compute_kv_cache_bytes(
    n_layers: int,
    n_heads: int,
    d_head: int,
    seq_len: int,
    bytes_per_element: int = 2,
) -> int:
    """
    Compute KV cache memory for a single sequence.
    Factor of 2 accounts for both K and V tensors.
    """
    return 2 * n_layers * n_heads * d_head * seq_len * bytes_per_element


# Model configs: (layers, heads, d_head)
model_configs = {
    "Llama-2-7B": (32, 32, 128),
    "Llama-2-13B": (40, 40, 128),
    "Llama-2-70B": (80, 64, 128),
}

seq_lengths = [512, 2048, 4096]
kv_results = {}
for model_name, (layers, heads, d_head) in model_configs.items():
    kv_results[model_name] = {}
    for seq_len in seq_lengths:
        kv_gb = compute_kv_cache_bytes(layers, heads, d_head, seq_len) / (
            1024**3
        )
        kv_results[model_name][seq_len] = kv_gb
Out[8]:
Console
KV Cache Memory per Request (GB, float16)

Model           |      512 tokens |     2048 tokens |     4096 tokens
---------------------------------------------------------------------
Llama-2-7B      |         0.250 |         1.000 |         2.000
Llama-2-13B     |         0.391 |         1.562 |         3.125
Llama-2-70B     |         1.250 |         5.000 |        10.000

With 100 concurrent requests at 2048 tokens:
  Llama-2-7B: 100.0 GB KV cache
  Llama-2-13B: 156.2 GB KV cache
  Llama-2-70B: 500.0 GB KV cache

These numbers show why KV cache management is so critical for serving. A system handling 100 concurrent requests at 2048 tokens with Llama-2-7B needs over 50 GB of KV cache alone, well exceeding the memory of a single A100 GPU. Effective serving requires careful memory budgeting and dynamic allocation strategies like PagedAttention to maximize the number of concurrent requests per GPU.

The interplay between these three memory consumers creates a planning problem. Weights are fixed, so that term is determined by the model choice. Activations are bounded by batch size and sequence length limits you configure. KV cache is the flexible term that fills remaining GPU memory and determines how many concurrent requests you can handle. The formula for maximum concurrent requests, given MtotalM_{total} GPU memory, is approximately:

max_concurrent≈Mtotal−Mweights−MactivationsMKV_per_req\text{max\_concurrent} \approx \frac{M_{total} - M_{weights} - M_{activations}}{M_{KV\_per\_req}}

This directly motivates quantization: halving weight memory with int8 quantization doubles the headroom available for KV cache, allowing roughly twice as many concurrent requests. The trade-off is inference quality, which must be validated for each specific model and task.

Out[9]:
Visualization
Stacked bar chart showing GPU memory consumed by model weights and KV cache for 20 to 120 concurrent requests.
GPU memory budget for a 7B-parameter model in float16 as concurrency rises. Weights consume a fixed 14 GB and each 2048-token request adds 1 GB of KV cache. The combined requirement crosses the A100-80GB limit above 66 requests; among the plotted bars, 70 requests is the first over-capacity case.

KV Cache Management and PagedAttention

Traditional serving systems allocated a contiguous memory block for each request's KV cache at the start of the request, sized to the maximum allowed sequence length. This caused two problems. First, most sequences end long before reaching the maximum length, leaving most of the pre-allocated memory wasted. Second, memory fragmentation gradually reduced the effective memory available as many small gaps between allocations accumulated, making it impossible to serve new requests even when there seemed to be enough free memory in aggregate.

PagedAttention, introduced by the vLLM team in 2023, solved both problems by borrowing the virtual memory paging model from operating systems. Instead of allocating one contiguous block per request, PagedAttention divides KV cache memory into fixed-size blocks (pages), typically holding 16 or 32 token positions each. Each request is assigned pages on demand as new tokens are generated. When a request finishes, its pages are returned to a free pool and immediately available for new requests.

This design has three practical consequences. First, memory waste from over-allocation nearly disappears because pages are allocated incrementally. Second, fragmentation is eliminated because pages are interchangeable units of fixed size. Third, pages from different requests can share physical memory blocks if they have identical prefixes (prompt caching), because two requests starting with the same system prompt can point their early KV cache pages at the same physical memory rather than each maintaining their own copy. This prefix caching trick is particularly valuable for chat applications where most requests begin with the same assistant persona and instructions.

PagedAttention

PagedAttention manages KV cache memory as fixed-size pages, similar to how an operating system manages virtual memory. Each request's key and value tensors are stored in non-contiguous pages that may be scattered across GPU memory. An attention lookup translates logical page numbers to physical page addresses at runtime. The resulting memory efficiency allows 2-4 times more concurrent requests per GPU compared to traditional contiguous allocation.

Request Handling

Once a model is loaded, the serving system must accept requests from clients, schedule them for execution, and return results. This layer handles the coordination between user-facing APIs and the GPU-side computation.

The Prefill-Decode Asymmetry

Before discussing scheduling and batching, it helps to understand the two-phase structure of LLM inference, because virtually every serving optimization exploits the difference between these phases.

The prefill phase processes the entire prompt in parallel. All input tokens are fed through the model simultaneously, producing the initial KV cache and the logits for the first output token. Prefill is compute-bound: it saturates the GPU's matrix multiplication units because the batch of token representations is large. A 2048-token prompt produces a matrix multiply involving a [2048,d_model][2048, d\_model] activation tensor, which fully occupies tensor cores.

The decode phase generates output tokens one by one. At each step, the model processes only the single most recently generated token, using the stored KV cache to attend over the full context. Decode is memory-bandwidth-bound: the computational work per step is tiny (just one token's worth of matrix multiplies), but each step must read the entire KV cache from GPU memory. The GPU's compute units sit mostly idle while waiting for memory fetches.

This asymmetry has far-reaching consequences for serving design. Prefill and decode compete for the same GPU resources but have different bottlenecks. A long prefill request arriving while many decode steps are in progress will delay those decode steps, spiking their time-to-first-token. Serving architectures that ignore this asymmetry end up with poor interactive latency even when overall throughput looks fine.

Request Queuing and Scheduling

Requests arrive asynchronously and at variable rates. When demand exceeds current processing capacity, requests must wait. How they wait, and how the system decides which request to process next, has significant effects on both latency and throughput.

First-come, first-served (FCFS) scheduling is the simplest policy: process requests in the order they arrive. It is fair and easy to implement, but ignores request characteristics that could improve overall efficiency. A single very long request can block many short requests that would complete quickly, wasting GPU capacity. FCFS is appropriate when fairness across all users has priority and all requests have roughly similar cost.

Shortest-job-first (SJF) scheduling prioritizes requests expected to complete quickly, minimizing average waiting time across all requests. The challenge is estimating job length without executing it. For language models, the prompt length provides a rough estimate of prefill time, but the decode length is typically unknown at scheduling time (the model generates until it decides to stop). Heuristics like maximum generation length or user-supplied length hints help, but length estimates are inherently uncertain. In pathological cases, SJF starves long requests indefinitely: if short requests keep arriving, a long request may never run.

Priority-based scheduling introduces service tiers where premium users or critical applications receive preferential queue position. A production system might define priority levels where API customers always receive faster service than background batch jobs running on the same infrastructure. Priority scheduling requires carefully calibrating the scheduling function to prevent low-priority work from perpetual starvation while still guaranteeing good service for high-priority requests.

Preemptive scheduling goes one step further: if a high-priority request arrives while a lower-priority request is being processed, the server can pause the lower-priority request (swapping its KV cache to CPU memory or discarding and recomputing it later) to immediately serve the high-priority one. This adds complexity but provides much stronger latency guarantees for critical traffic.

In[10]:
Code
import heapq
import time
from dataclasses import dataclass, field
from typing import List, Optional


@dataclass(order=True)
class ScheduledRequest:
    """A request in the priority queue. Lower priority_score = processed first."""

    priority_score: float
    arrival_time: float = field(compare=False)
    request_id: str = field(compare=False)
    prompt_tokens: int = field(compare=False)
    max_new_tokens: int = field(compare=False)
    user_priority: int = field(compare=False)  # 0=low, 10=high


class RequestQueue:
    """Priority queue supporting FCFS, SJF, and priority-based scheduling."""

    def __init__(self, policy: str = "fcfs"):
        self._heap: List[ScheduledRequest] = []
        self.policy = policy
        self.enqueued = 0

    def enqueue(
        self,
        request_id: str,
        prompt_tokens: int,
        max_new_tokens: int,
        user_priority: int = 5,
    ) -> None:
        arrival = time.time()
        if self.policy == "fcfs":
            score = arrival
        elif self.policy == "sjf":
            # Estimate total tokens as proxy for job length
            score = prompt_tokens + max_new_tokens
        elif self.policy == "priority":
            # Invert priority so higher priority = lower score = processed first
            score = -user_priority + arrival * 1e-9
        else:
            score = arrival

        req = ScheduledRequest(
            priority_score=score,
            arrival_time=arrival,
            request_id=request_id,
            prompt_tokens=prompt_tokens,
            max_new_tokens=max_new_tokens,
            user_priority=user_priority,
        )
        heapq.heappush(self._heap, req)
        self.enqueued += 1

    def dequeue(self) -> Optional[ScheduledRequest]:
        if self._heap:
            return heapq.heappop(self._heap)
        return None

    def __len__(self):
        return len(self._heap)
In[11]:
Code
import random

random.seed(42)

# Simulate a batch of requests with different characteristics
requests_data = [
    ("req-001", 50, 200, 9),  # short prompt, short output, high priority
    ("req-002", 800, 50, 3),  # long prompt, short output, low priority
    ("req-003", 100, 500, 5),  # medium prompt, long output, medium priority
    ("req-004", 20, 30, 7),  # very short, high priority
    ("req-005", 400, 400, 2),  # large job, low priority
]

# Compare scheduling orders
schedule_comparisons = {}
for policy in ["fcfs", "sjf", "priority"]:
    q = RequestQueue(policy=policy)
    for req_id, p_tokens, n_tokens, upriority in requests_data:
        q.enqueue(req_id, p_tokens, n_tokens, upriority)
    order = []
    while len(q) > 0:
        r = q.dequeue()
        order.append(
            (r.request_id, r.prompt_tokens, r.max_new_tokens, r.user_priority)
        )
    schedule_comparisons[policy] = order
Out[12]:
Console
Scheduling Order by Policy

Request     Prompt Tokens  Max New  Priority
Original order:
  req-001                50      200         9
  req-002               800       50         3
  req-003               100      500         5
  req-004                20       30         7
  req-005               400      400         2

FCFS order:
  req-001                50      200         9
  req-002               800       50         3
  req-003               100      500         5
  req-004                20       30         7
  req-005               400      400         2

SJF order:
  req-004                20       30         7
  req-001                50      200         9
  req-003               100      500         5
  req-005               400      400         2
  req-002               800       50         3

PRIORITY order:
  req-001                50      200         9
  req-004                20       30         7
  req-003               100      500         5
  req-002               800       50         3
  req-005               400      400         2

The three policies produce different orderings with distinct trade-offs. FCFS preserves arrival order and is maximally fair, but req-002 (a bulky request) holds up shorter jobs behind it. SJF reorders so small jobs run first, potentially improving average latency, but risks starving large requests indefinitely if small requests keep arriving. Priority-based ordering ensures high-priority users receive fast service regardless of request size, but requires careful calibration to prevent low-priority work from never executing.

In practice, most production systems use FCFS as their default policy because fairness and simplicity are easier to reason about than the subtle bugs that poorly tuned SJF or priority schedulers can introduce. Priority-based scheduling is layered on top of FCFS for systems with clearly differentiated service tiers, such as when a team needs to guarantee latency for interactive users while also running background batch jobs on the same infrastructure.

Batching Strategies

Processing multiple requests together in a single forward pass improves GPU utilization by amortizing the fixed overhead of each forward pass across many requests. This is essential for cost-efficient serving.

Static batching groups requests into fixed-size batches that are submitted together. All requests in a batch are processed in parallel during prefill, and all continue decoding together until every request in the batch completes. The problem is the "straggler" effect: if one request in a batch requires 500 tokens while others need only 50, the GPU sits partially idle during the extra 450 decode steps as the short requests have already finished. The partially complete batch cannot accept new requests to fill the vacated slots. Static batching is simple but wasteful.

Consider a concrete example. A batch of 8 requests is submitted together. Six requests finish after 60 tokens each. Two requests continue generating for 300 more tokens. During those 300 extra decode steps, the batch is operating at 25% capacity rather than 100%. The GPU is doing real work for 2 sequences but could be doing work for 8. The missed throughput is unrecoverable: those cycles are gone.

Dynamic batching relaxes the batch-size constraint and collects requests over a short time window before submitting them as a batch. This improves batch fullness compared to processing requests individually but introduces a small amount of additional latency as requests wait for the window to close. A 20-millisecond collection window, for instance, allows a burst of requests to accumulate into a larger batch, improving amortization. The window length creates a direct latency floor: the fastest a request can be processed is at least as long as the window.

Continuous batching (also called in-flight batching) is the most sophisticated approach. Rather than waiting for all requests in a batch to finish before accepting new ones, the server continuously adds new requests to the active batch as completed requests free up their slots. Each decode step potentially includes different requests, with arrivals inserting after their prefill completes and completions freeing space for waiting requests. As we covered in Part XXXI, continuous batching dramatically improves GPU utilization for workloads with variable-length outputs.

The key insight in continuous batching is that the GPU does not care whether the 8 sequences in the current decode step are the same 8 sequences that were in the previous step. As long as the batch is full of valid requests, the GPU does productive work. Continuous batching keeps the batch full by treating request boundaries as opportunities to swap in waiting requests rather than as synchronization points for the entire batch.

Out[13]:
Visualization
Line chart comparing GPU utilization over time for static batching versus continuous batching strategies.
Simulated GPU utilization over 10 seconds. Static batching repeatedly drops into idle gaps, while continuous batching stays near 92% by refilling vacated slots. The dashed mean lines show about 64% average utilization for static batching versus 92% for continuous batching, a gap of roughly 28 percentage points.

Prefill-Decode Disaggregation

A more recent architectural pattern separates the prefill and decode phases of inference onto different hardware. This approach, sometimes called disaggregated serving or chunked prefill, recognizes that prefill and decode have different computational profiles and can benefit from different hardware or scheduling priorities.

Prefill is compute-bound: it processes all prompt tokens in parallel using dense matrix operations, fully utilizing GPU compute units. Decode is memory-bandwidth-bound: each step reads the entire KV cache from GPU memory to generate a single token. Mixing prefill and decode in the same batch means that long prefill phases for new requests delay the ongoing decoding for existing requests, spiking their latency.

In a fully disaggregated architecture, a pool of prefill instances handles incoming requests and computes their initial KV states, then transfers those KV tensors over a high-bandwidth network to a pool of decode instances that continue token generation. This separation allows each pool to be optimized for its task: prefill instances might use smaller GPUs with more compute throughput per dollar, while decode instances benefit from GPUs with high memory bandwidth. The trade-off is the KV cache transfer overhead: moving potentially gigabytes of KV state between machines requires high-bandwidth interconnects and adds a fixed delay between prefill completion and decode start.

Chunked prefill is a lighter-weight middle ground that keeps prefill and decode on the same GPU but limits how many prefill tokens are processed per step. Rather than completing an entire 2048-token prefill before resuming decode, the server processes 512 tokens of prefill per step, interleaving with ongoing decode steps. This prevents any single prefill request from monopolizing the GPU for an extended period, keeping time-to-first-token stable for concurrent decode requests.

Serving Frameworks

Several mature frameworks handle LLM serving at production scale. Each makes different trade-offs among deployment flexibility, inference performance, and operational complexity.

vLLM

vLLM is one of the most widely used open-source LLM serving frameworks, built around the PagedAttention memory management technique. Instead of allocating contiguous memory blocks for KV caches, vLLM manages KV cache memory as fixed-size pages, similar to how an operating system manages virtual memory. This eliminates fragmentation and allows many more concurrent requests to share GPU memory efficiently.

vLLM exposes an OpenAI-compatible API endpoint, making it a drop-in replacement for applications built against the OpenAI API. It supports tensor parallelism across multiple GPUs, quantization for reducing memory requirements, and continuous batching for high throughput. The framework handles most low-level serving concerns automatically, letting engineers focus on deployment configuration rather than inference implementation.

vLLM also supports prefix caching: if multiple requests share a common prefix (such as a long system prompt), the KV cache for that prefix is computed once and shared across all requests. For applications where 80% of every request is identical system instructions, this can dramatically reduce both compute and memory usage. A 1000-token system prompt that appears in every request needs to be prefilled only once per hour (when the cached KV blocks expire) rather than once per request.

In[14]:
Code
# Illustrate key vLLM configuration parameters conceptually
# (Actual vLLM requires GPU runtime; this demonstrates the config structure)

vllm_config = {
    # Model and precision
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "dtype": "bfloat16",  # Weight dtype: float32, float16, bfloat16
    "quantization": None,  # Options: "awq", "gptq", "squeezellm", None
    # Parallelism
    "tensor_parallel_size": 1,  # Number of GPUs for tensor parallelism
    "pipeline_parallel_size": 1,  # Number of pipeline stages
    # Memory
    "gpu_memory_utilization": 0.90,  # Fraction of GPU memory reserved
    "max_model_len": 8192,  # Maximum total sequence length (prompt + output)
    "max_num_seqs": 256,  # Maximum concurrent sequences
    # Batching
    "max_num_batched_tokens": 16384,  # Max total tokens in a single forward pass
    "enable_chunked_prefill": True,  # Split long prefills to reduce decode latency spikes
    "max_num_partial_prefills": 1,  # Max chunks for chunked prefill
    # Scheduling
    "scheduling_policy": "fcfs",  # "fcfs" or "priority"
    "preemption_mode": "recompute",  # "recompute" or "swap" on memory pressure
}
Out[15]:
Console
vLLM Server Configuration

  [Model & Precision]
    model: meta-llama/Meta-Llama-3-8B-Instruct
    dtype: bfloat16
    quantization: None

  [Parallelism]
    tensor_parallel_size: 1
    pipeline_parallel_size: 1

  [Memory]
    gpu_memory_utilization: 0.9
    max_model_len: 8192
    max_num_seqs: 256

  [Batching]
    max_num_batched_tokens: 16384
    enable_chunked_prefill: True
    max_num_partial_prefills: 1

  [Scheduling]
    scheduling_policy: fcfs
    preemption_mode: recompute

The preemption_mode parameter is worth a closer look. When the server runs out of KV cache memory while active requests are in progress, it must free memory from some requests so others can continue. Two strategies exist. With "recompute" mode, the server discards the KV cache of the preempted request entirely, forcing it to rerun its prefill phase from scratch when memory becomes available. With "swap" mode, the KV cache is copied to CPU memory (a slower, larger storage tier), freeing GPU memory, and later copied back when the request resumes. Recompute is simpler and requires no CPU memory budget but wastes compute. Swap preserves work but requires careful bandwidth management to avoid the swap itself becoming a bottleneck.

Text Generation Inference (TGI)

Text Generation Inference from Hugging Face offers tight integration with the Hugging Face model hub and the Transformers library. TGI provides production-grade serving with features like continuous batching, tensor parallelism, quantization, and token streaming. Its primary advantage is direct ecosystem integration: any model available on the Hugging Face hub can typically be deployed with TGI without custom loading code.

TGI exposes both a REST API and a gRPC interface. The REST API is compatible with many existing LLM client libraries. TGI also supports speculative decoding out of the box, pairing a small draft model with the main model to accelerate generation for common token sequences. The Hugging Face team maintains a Docker image that bundles TGI with optimized kernels, making deployment as straightforward as running a single docker run command with the model name as an argument.

TGI's tight integration with the Hugging Face hub also means that model updates are simple: changing the model name in the deployment configuration and restarting the container loads the new model automatically, without any changes to serving infrastructure code.

TensorRT-LLM

NVIDIA's TensorRT-LLM takes a compilation-based approach. Models are compiled into highly optimized CUDA inference engines that use custom kernels for attention, layer normalization, and other operations. The compilation step is slow and hardware-specific, but the resulting engines often achieve the highest raw inference throughput on NVIDIA GPUs.

TensorRT-LLM supports in-flight batching, INT8 and FP8 quantization, tensor parallelism, and speculative decoding. It integrates with NVIDIA Triton Inference Server for production deployment, giving load balancing, model versioning, and monitoring capabilities. The main trade-off is the setup complexity: compiling a new model or changing hardware requires rerunning the optimization pipeline, which can take tens of minutes for large models.

For organizations that standardize on NVIDIA hardware and prioritize maximum throughput over flexibility, TensorRT-LLM can deliver higher tokens-per-second than framework-agnostic alternatives. The compilation process applies hardware-specific optimizations that generic runtimes cannot perform at execution time: custom fused kernels that merge multiple operations into a single GPU pass, memory layouts that minimize data movement, and quantization-aware kernel selection that maps FP8 operations to the hardware's native accelerators. These optimizations compound, making TensorRT-LLM a strong choice for large-scale serving where hardware is homogeneous and setup cost is a one-time investment.

Ollama

Ollama is designed for local and single-machine deployment rather than distributed serving at scale. It provides an easy-to-use interface for running open-source models locally, handling model downloads, quantization, and serving configuration automatically. Ollama uses llama.cpp as its backend, which supports CPU inference in addition to GPU acceleration and implements highly optimized quantization formats like GGUF.

For development and small-scale deployment, Ollama's simplicity is its primary advantage. A developer can run ollama run llama3 to download and serve a model in a single command. There is no YAML configuration file, no GPU driver management, and no serving framework to configure. For production serving at scale, dedicated serving frameworks like vLLM or TGI offer better throughput and configuration options.

Ollama's CPU support suits certain workloads. A quantized 7B model in 4-bit GGUF format fits in roughly 4 GB of RAM and can generate tokens at a few tokens per second on a modern CPU. This is too slow for interactive use but sufficient for batch processing tasks where cost matters more than latency and GPU hardware is unavailable.

Out[16]:
Visualization
Radar chart comparing four LLM serving frameworks across five attributes: throughput, ease of setup, ecosystem integration, multi-GPU support, and quantization breadth.
Qualitative comparison of LLM serving frameworks across five dimensions: throughput, ease of setup, ecosystem integration, multi-GPU support, and quantization breadth. Scores are approximate and reflect general community assessments as of early 2026. vLLM and TGI balance high throughput with reasonable ease of use, while TensorRT-LLM offers peak performance at the cost of setup complexity. Ollama prioritizes simplicity for local deployments over throughput and multi-GPU scaling.

Choosing a Framework

No single framework is best for every situation. The right choice depends on your hardware, team capabilities, traffic patterns, and latency requirements.

vLLM is the default starting point for most teams because it combines strong throughput, straightforward configuration, and an OpenAI-compatible API. It handles the widest range of use cases adequately, even if it is not optimal for any single specific workload.

TGI is the natural choice for teams already using the Hugging Face ecosystem. If your model is on the Hugging Face hub and your inference pipeline uses Transformers, TGI adds the minimum possible friction.

TensorRT-LLM is worth the setup cost for teams running at very large scale on standardized NVIDIA hardware. The throughput advantage compounds at scale: if you serve billions of tokens per day, a 20% throughput improvement translates directly into 20% lower infrastructure costs.

Ollama serves development and small-scale deployment where ease of use is more important than performance. It is an excellent tool for testing models locally before deciding which production framework to use.

Serving Configuration

The performance of a serving system depends heavily on how it is configured. The key parameters control the trade-off between throughput, latency, memory usage, and resource cost.

Throughput vs. Latency Trade-offs

Serving configuration almost always involves trading throughput for latency or vice versa. Understanding which direction to optimize requires understanding your workload and user requirements.

Batch size is the most direct lever. Larger batches amortize GPU overhead across more requests, improving tokens-per-second throughput but increasing the time each individual request waits for other requests in the batch to complete. If your application cares about response time (interactive chatbots, real-time translation), smaller batch sizes with lower latency are preferable. If you are running overnight batch processing where per-request latency matters less than total throughput, larger batches maximize efficiency.

Queue depth controls how many requests accumulate before the system starts applying backpressure to callers. A deep queue absorbs bursty traffic by buffering requests, but at the cost of increased tail latency when the queue fills. A shallow queue forces callers to retry quickly, which works well when average load is below capacity but provides less smoothing for traffic spikes.

Max sequence length limits how long prompts and outputs can be. Shorter limits conserve KV cache memory and allow more concurrent requests, but reject or truncate long inputs. Setting this appropriately for your use case avoids wasting memory on pathologically long sequences while not frustrating users with legitimate long-context needs.

It is worth understanding why these three knobs interact. Increasing max sequence length increases KV cache consumption per request. This forces a reduction in max concurrent requests to stay within memory limits, which reduces effective batch size, which reduces throughput. Everything connects. A change to one parameter ripples through the others, which is why capacity planning requires thinking about the full system rather than individual knobs in isolation.

In[17]:
Code
def estimate_throughput_latency(
    batch_size: int,
    tokens_per_step_ms: float = 2.0,
    overhead_ms: float = 5.0,
    avg_output_tokens: int = 200,
) -> dict:
    """
    Simplified throughput and latency model for a serving configuration.

    tokens_per_step_ms: Time to generate one token for a single sequence (ms)
    overhead_ms: Fixed overhead per batch (scheduling, data movement)
    """
    # Total decode time per batch (serialized token steps)
    decode_time_ms = avg_output_tokens * tokens_per_step_ms
    total_batch_time_ms = overhead_ms + decode_time_ms

    # Average latency each request experiences (FCFS: waits for entire batch)
    avg_latency_ms = total_batch_time_ms

    # Throughput: requests completed per second
    throughput_rps = (batch_size / total_batch_time_ms) * 1000

    return {
        "batch_size": batch_size,
        "avg_latency_ms": avg_latency_ms,
        "throughput_rps": throughput_rps,
    }


batch_sizes = [1, 4, 8, 16, 32, 64]
configs = [estimate_throughput_latency(bs) for bs in batch_sizes]
Out[18]:
Console
Batch Size |   Avg Latency (ms) |   Throughput (req/s)
-------------------------------------------------------
         1 |                405 |                  2.5
         4 |                405 |                  9.9
         8 |                405 |                 19.8
        16 |                405 |                 39.5
        32 |                405 |                 79.0
        64 |                405 |                158.0

The table isolates one part of the serving trade-off. Its simplified calculation holds per-batch service time at 405 ms, so throughput rises with batch size while service latency stays constant. It does not model the time spent waiting for enough requests to form a batch or waiting in a queue. Those omitted costs grow with batching pressure in a real deployment, which is why practical systems tune batch size against measured end-to-end latency rather than relying on service-time arithmetic alone.

Out[19]:
Visualization
Dual-axis line chart showing throughput rising with batch size while the simplified model's latency stays flat at 405 milliseconds.
Output from the simplified serving model as batch size increases from 1 to 64. Throughput rises linearly from about 2.5 to 158 requests per second, while modeled service latency remains fixed at 405 ms because the calculation excludes batch-formation and queue waiting. The shaded band marks the sub-second target; real systems must add those omitted waiting costs when choosing a batch size.

Serving Parameters in Detail

The parameters exposed by modern LLM serving frameworks cluster into several categories, each controlling a different aspect of serving behavior.

Memory configuration parameters control how much GPU memory the server reserves and how it is divided between weights, activations, and KV cache:

  • gpu_memory_utilization: The fraction of total GPU memory the server claims. Setting this to 0.90 reserves 90% for the model and KV cache, leaving 10% as headroom for PyTorch internal allocations and safety margin. Values above 0.95 risk out-of-memory errors under peak load.
  • max_model_len: The maximum total sequence length (prompt plus generated tokens) the server will accept. Reducing this limit frees KV cache memory and allows more concurrent requests.
  • max_num_seqs: The maximum number of requests the server will process concurrently. This caps KV cache usage and prevents memory exhaustion when many long requests arrive simultaneously.

Batching parameters control how requests are grouped for GPU execution:

  • max_num_batched_tokens: The total number of tokens across all requests processed in a single forward pass. This controls the trade-off between batch efficiency and the latency impact of processing large prefill requests.
  • enable_chunked_prefill: When enabled, long prefill phases are split into smaller chunks. Each chunk interleaves with ongoing decode steps for other requests, preventing long prefills from blocking decode latency.

Scheduling parameters determine how the server handles resource contention:

  • scheduling_policy: Whether to process requests strictly in arrival order (FCFS) or by priority score.
  • preemption_mode: When memory pressure requires freeing a request's KV cache, "recompute" discards and later regenerates it, while "swap" moves it to CPU memory. Recompute is simpler but wastes compute; swap preserves work but requires CPU memory bandwidth.

Understanding when to adjust these parameters requires thinking about your workload characteristics. For a customer service chatbot handling short conversations from many simultaneous users, you want high max_num_seqs and modest max_model_len. For a document summarization pipeline processing long articles one at a time, you want high max_model_len and accept lower concurrency. For a mixed workload, chunked prefill becomes essential to prevent long document requests from disrupting the latency of concurrent short conversational requests.

Out[20]:
Visualization
Bar chart comparing time-to-first-token with and without chunked prefill for concurrent decode requests blocked by a large prefill.
Effect of chunked prefill on time-to-first-token (TTFT) for decode requests when a large prefill request arrives. Without chunked prefill, a 2048-token prefill blocks all ongoing decode steps for the duration of the prefill phase, spiking TTFT for affected requests to around 210ms. With chunked prefill enabled, the long prefill is interleaved with decode steps in 512-token chunks, keeping TTFT near 55ms for concurrent requests. The improvement is approximately 75%, at the cost of slightly longer total prefill duration for the large request.

Latency Metrics and SLOs

Production serving systems track multiple latency metrics because different aspects of latency matter for different use cases.

Time-to-first-token (TTFT) measures how long a user waits before seeing any output. For streaming applications where tokens appear progressively in the UI, TTFT determines how long the screen sits blank after a user submits a request. High TTFT makes applications feel unresponsive even if tokens arrive quickly after the first one. TTFT is dominated by prefill time for long prompts and by queue wait time when the server is busy.

Time-per-output-token (TPOT) measures the average delay between consecutive tokens during streaming generation. This determines the speed at which text appears on screen. Users perceive fast token generation (low TPOT) as a smooth, fluid experience. TPOT is dominated by the KV cache read bandwidth during the decode phase.

End-to-end latency is the total time from request submission to final token received. This matters for non-streaming applications that wait for the complete response before displaying anything. It equals TTFT plus (output tokens ×\times TPOT).

Throughput is tokens generated per second across all concurrent requests. This is the primary efficiency metric for batch workloads where latency is secondary to cost.

Service level objectives (SLOs) define the acceptable range for each metric. A typical interactive application might require p50 TTFT under 300ms, p99 TTFT under 1000ms, and p50 TPOT under 50ms. Setting SLOs requires understanding user expectations: research suggests users begin to perceive applications as "slow" when TTFT exceeds 500ms, but smooth streaming at 20+ tokens per second feels fast even for long responses.

Tracking the full distribution of latency (p50, p90, p95, p99) rather than just averages is essential. The tail of the latency distribution represents your worst-performing requests, which are often the most memorable for users. A system with a good average latency but a long tail leaves a fraction of users with a terrible experience on every visit.

Rate Limiting and Backpressure

A production serving system must protect itself from being overwhelmed by more requests than it can handle. Without rate limiting, a traffic spike can fill the request queue, causing all in-flight requests to experience extreme latency while new requests continue to pile up, making the system slow for everyone.

Rate limiting enforces an upper bound on how many requests a client or tenant can submit per unit time. Token-bucket algorithms are commonly used: each client has a token bucket that refills at a fixed rate. Submitting a request costs one token. When the bucket is empty, further requests are rejected with a 429 (Too Many Requests) HTTP response, signaling to the caller that it should back off and retry later.

Backpressure propagates this signal upstream through the entire call chain. Instead of queueing indefinitely, a saturated server rejects new requests immediately rather than accepting them into a long queue where they would wait and eventually time out. This fail-fast behavior gives clients accurate information about server state and enables intelligent retry logic, such as exponential backoff.

In[21]:
Code
import time
from dataclasses import dataclass


@dataclass
class TokenBucket:
    """
    Token bucket rate limiter.
    capacity: max burst size in tokens
    refill_rate: tokens replenished per second
    """

    capacity: float
    refill_rate: float
    _tokens: float = 0.0
    _last_refill: float = 0.0

    def __post_init__(self):
        self._tokens = self.capacity
        self._last_refill = time.time()

    def _refill(self) -> None:
        now = time.time()
        elapsed = now - self._last_refill
        self._tokens = min(
            self.capacity, self._tokens + elapsed * self.refill_rate
        )
        self._last_refill = now

    def try_consume(self, tokens: float = 1.0) -> bool:
        """Returns True if the request is allowed, False if rate limited."""
        self._refill()
        if self._tokens >= tokens:
            self._tokens -= tokens
            return True
        return False

    @property
    def available_tokens(self) -> float:
        self._refill()
        return self._tokens
In[22]:
Code
# Simulate a burst of 20 requests against a bucket with capacity=10, rate=5/sec
bucket = TokenBucket(capacity=10.0, refill_rate=5.0)
results = []
for i in range(20):
    allowed = bucket.try_consume()
    results.append((i + 1, allowed, round(bucket.available_tokens, 2)))

allowed_count = sum(1 for _, allowed, _ in results if allowed)
rejected_count = sum(1 for _, allowed, _ in results if not allowed)
Out[23]:
Console
Token Bucket Rate Limiting (capacity=10, rate=5/sec)
Burst of 20 simultaneous requests:

 Request |  Allowed |  Tokens Left
-----------------------------------
       1 |      YES |         9.00
       2 |      YES |         8.00
       3 |      YES |         7.00
       4 |      YES |         6.00
       5 |      YES |         5.00
       6 |      YES |         4.00
       7 |      YES |         3.00
       8 |      YES |         2.00
       9 |      YES |         1.00
      10 |      YES |         0.00
      11 |      NO  |         0.00
      12 |      NO  |         0.00
      13 |      NO  |         0.00
      14 |      NO  |         0.00
      15 |      NO  |         0.00
      16 |      NO  |         0.00
      17 |      NO  |         0.00
      18 |      NO  |         0.00
      19 |      NO  |         0.00
      20 |      NO  |         0.00

Summary: 10 allowed, 10 rejected

The first 10 requests drain the bucket immediately. Requests 11 through 20 arrive before the bucket can replenish enough tokens and are rejected. In a real system, rejected callers receive a 429 response and apply exponential backoff before retrying, allowing the bucket to refill. This prevents a single bursty client from monopolizing server capacity.

Rate limiting can operate at multiple granularities. Per-user limits prevent a single account from crowding out others. Per-API-key limits allow different products within an organization to have independent quotas. Token-based limits (counting input plus output tokens rather than requests) more accurately reflect actual resource consumption, because a 5000-token request consumes roughly 50 times as much compute as a 100-token request.

For multi-tenant deployments where different tenants have different SLAs, rate limiting alone is insufficient. You also need admission control: the ability to reject new requests when the server is at capacity, rather than accepting them into an ever-growing queue. A queue that grows without bound eventually causes every request to time out from the caller's perspective, even though the server is technically "accepting" the requests. Effective admission control keeps queue depth bounded by returning 503 (Service Unavailable) responses when the server is saturated, allowing callers to route requests to another instance or degrade gracefully.

Deployment Patterns

How you deploy a model serving system shapes its operational characteristics, including cost, reliability, and ease of management.

Single-Instance Deployment

The simplest deployment runs one model instance on one or more GPUs, serving all traffic from that single process. This is appropriate for development, low-traffic applications, and models too large to replicate economically. It eliminates load balancing complexity and makes debugging straightforward.

The obvious limitation is the lack of redundancy: if the serving process crashes or the GPU experiences a hardware fault, all service is interrupted until the instance recovers. Single-instance deployments also cannot scale horizontally to handle traffic beyond what one instance can process, though vertical scaling (adding GPUs for tensor parallelism) extends this ceiling substantially.

For development and experimentation, single-instance deployment is nearly always the right choice. The configuration simplicity means you can quickly test different models, configurations, and optimization settings without managing distributed infrastructure. Problems are easy to diagnose because there is only one log stream and one process to inspect. Move to replicated deployment when your traffic or reliability requirements demand it, not before.

Replicated Deployment

Running multiple identical instances of the model behind a load balancer provides both redundancy and horizontal scaling. Each instance handles a portion of the traffic, and a load balancer distributes requests across healthy instances. If one instance fails, its traffic routes to the remaining instances.

Stateless inference servers are well-suited to replication because there is no shared state between requests beyond the loaded model weights. Load balancers can use simple round-robin or least-connections policies because any instance can handle any request equally well. This is in contrast to stateful services where requests that belong to the same "session" must be routed to the same server.

The challenge with replicated LLM deployments is memory cost: each replica requires a full copy of the model weights in GPU memory. For large models, this makes wide replication expensive. Teams often balance the number of replicas against the tensor parallelism degree within each replica, choosing configurations that maximize served requests per dollar of GPU cost.

Consider a 70B model that requires 8 GPUs for tensor parallelism. Running 2 replicas costs 16 GPUs and provides twice the throughput with redundancy. Running a single replica with 8 GPUs costs half as much but has no redundancy. The right choice depends on your traffic volume and availability requirements. For high-availability production deployments, two replicas is a reasonable minimum: it provides both scaling headroom and tolerance for a single-instance failure.

Prefix caching across replicas adds a nuance to load balancing. If replica A has cached the KV state for a popular system prompt and replica B has not, routing a request with that system prompt to replica A avoids a prefill computation. Consistent hashing on the first N tokens of a request can direct requests with the same prefix to the same replica, preserving prefix cache hits. This is more complex than round-robin but can substantially reduce compute on workloads with common prefixes.

Blue-Green and Canary Deployments

Zero-downtime deployments use deployment patterns borrowed from general software operations.

Blue-green deployment maintains two complete environments, "blue" (current production) and "green" (new version). When deploying a model update, the green environment is spun up with the new version, validated, and then traffic is atomically switched from blue to green. If issues emerge, traffic instantly rolls back to blue. The trade-off is cost: maintaining two full environments doubles GPU expenses during the deployment window. For large models requiring many GPUs, this doubling can be substantial, making blue-green deployment economically challenging except for brief deployment windows.

Canary deployment routes a small fraction (such as 5%) of traffic to the new version while the remainder continues on the old version. Quality metrics, latency, and error rates are monitored for the canary slice. If metrics look good over a defined evaluation period, the traffic fraction increases progressively until the new version handles all traffic. This provides early failure detection with gradual exposure, balancing safety and deployment speed.

The evaluation period between traffic shifts is critical. You need enough traffic through the canary to detect regressions with statistical confidence. A 5% canary receiving 1,000 requests in its first hour gives you reasonable confidence for common failure modes, but may not expose rare edge cases. Defining clear rollout criteria before deployment (for example, "advance when p99 latency is within 10% of baseline and error rate is below 0.1%") prevents the evaluation from being subject to subjective judgment under time pressure.

Out[24]:
Visualization
Step chart showing canary traffic fraction increasing from 5% to 100% over 24 hours in a staged rollout.
Traffic allocation during a canary deployment rollout over 24 hours. Traffic to the new model version increases in stages (5%, 20%, 50%, 100%) as quality and latency metrics pass validation checks at each stage. Gray shaded bands mark evaluation windows between traffic shifts, during which metrics are assessed before advancing to the next stage. A failed evaluation at any stage would halt the rollout and maintain the current traffic split until the issue is resolved.

Autoscaling

Static deployments with a fixed number of replicas waste capacity during off-peak hours and may be insufficient during peak hours. Autoscaling dynamically adjusts the number of running instances in response to traffic.

Horizontal autoscaling (adding or removing replicas) is the standard approach for web services. For LLM serving, autoscaling has one critical challenge: the cold start problem. Adding a new replica requires loading model weights, which takes minutes for large models. By the time the new replica is ready to serve traffic, the traffic spike that triggered the scale-out may already be over. This means autoscaling for LLMs must be proactive: anticipate traffic increases before they arrive (using historical patterns or predictive signals) and scale up in advance rather than in reaction.

Minimum replica counts ensure that at least one warm instance is always ready to serve traffic. A common pattern is maintaining two replicas as a minimum even during very low traffic, so that the system can always absorb a sudden burst with one instance while the autoscaler provisions additional capacity.

Vertical scaling within a single instance (changing the tensor parallelism degree) is not typically done dynamically because it requires restarting the serving process. Horizontal scaling is therefore the primary autoscaling mechanism.

Limitations and Practical Considerations

Model serving combines software engineering, systems programming, and machine learning. Each domain contributes a distinct category of challenges.

Hardware Dependency

LLM serving is tightly coupled to GPU hardware in ways that create significant operational constraints. A model compiled with TensorRT-LLM for one GPU generation may not run on another without recompilation. Serving frameworks that use custom CUDA kernels require specific driver and toolkit versions. Upgrading hardware often means requalifying the entire serving stack. This hardware specificity contrasts with traditional web services that can run on commodity hardware with minimal configuration, and it means that organizations running LLM services must invest substantially in hardware standardization and upgrade planning.

The dependency extends to firmware and driver versions. A serving stack that works perfectly with CUDA 12.2 and driver 535 may silently produce wrong outputs with CUDA 12.4 due to kernel-level changes. Integration tests must validate model outputs against reference outputs before any infrastructure update is adopted; checking only that the server starts and responds is insufficient.

Cold Start Latency

The time required to load a large model into GPU memory, allocate KV cache structures, and complete warmup requests before the server can handle production traffic is substantial. A 70B-parameter model may require several minutes from process start to first request. This cold start latency complicates autoscaling strategies: spinning up new instances in response to a traffic spike takes too long to help with the spike itself. Serving systems must maintain warm instances proactively, accepting the cost of idle capacity to ensure rapid response to demand increases.

One mitigation is checkpoint pre-loading: caching model weights in CPU memory or on a local NVMe drive on the serving host, so that when the GPU process restarts, it can load weights from fast local storage rather than from network-attached storage. A local NVMe drive typically provides 5-7 GB/s read speed, compared to 1-2 GB/s for network storage, cutting disk read time by 3-5x for large models. When combined with GPU memory pre-allocation (reserving GPU memory before the model is fully loaded), a well-optimized cold start can bring a 70B model to ready state in under two minutes instead of five or more.

Cost Efficiency

GPU costs dominate the economics of LLM serving. A single H100 GPU costs thousands of dollars per month on cloud platforms, and serving large models requires many such GPUs. The key metric is tokens served per dollar, which depends on both raw inference speed and GPU utilization. An underutilized GPU generating few tokens is expensive; a fully utilized GPU generating many tokens is efficient.

Achieving high utilization requires workload shaping, good batching, and sometimes mixed workloads: running batch inference jobs during off-peak hours on the same infrastructure that handles interactive traffic during peak hours. Token-based pricing, where users pay per thousand tokens rather than per request, aligns incentives with efficient utilization by pricing long requests proportionally to their actual resource consumption.

The cost calculation changes significantly with model size and optimization choices. A 7B model with int8 quantization running at 80% utilization can serve far more tokens per dollar than a 70B model in float16 running at 40% utilization, even if the 70B model produces better output quality. Understanding the quality-per-dollar trade-off, not just quality-per-output, is essential for building economically sustainable LLM services.

Spot or preemptible instances from cloud providers offer discounts of 50-80% compared to on-demand pricing at the cost of potential instance interruption. For stateless inference servers, spot instances work well because any in-flight requests at interruption time can simply be retried against other instances. Systems that route around interrupted instances automatically (with proper health checking and load balancing) can run large portions of their serving infrastructure on spot, dramatically reducing costs.

Reliability Under Partial Failures

Production serving systems must handle partial failures gracefully. Individual GPUs can fail, model instances can crash, and network connectivity between components can be interrupted. A well-designed serving stack isolates these failures: a crashed instance should be removed from the load balancer pool and restarted automatically, while other instances continue serving traffic. Health check systems detect failures within seconds, and circuit breakers prevent the load balancer from sending requests to known-bad instances.

For tensor-parallel deployments where a single model spans multiple GPUs, a failure of any one GPU in the group takes down the entire instance. This is one reason teams prefer wider replication (more single-GPU or dual-GPU instances) over deeper tensor parallelism (fewer instances each using many GPUs) when reliability matters. A 4-way tensor-parallel deployment has a higher probability of GPU failure at any point in time than a 2-way deployment with twice as many replicas.

Deeper failures, such as model weight corruption or systematic errors introduced by a model update, require application-level quality monitoring rather than just infrastructure health checks. Tracking output quality metrics (answer relevance, refusal rate, format compliance) alongside infrastructure metrics allows you to detect model-level regressions that infrastructure health checks would miss. The monitoring and quality assurance capabilities covered in later chapters of this part are essential for catching these more subtle failure modes before they significantly impact users.

Dependency on Inference Optimization Techniques

Effective model serving does not stand alone: it is most powerful when combined with the inference optimization techniques covered in Part XXXI. KV caching, quantization, speculative decoding, and continuous batching all operate at the level of individual requests. Model serving adds the orchestration layer above them, making sure that requests are routed efficiently, hardware is utilized fully, and the system remains stable under load. A serving system built without these underlying optimizations will achieve a fraction of the throughput possible with them.

Speculative decoding in particular interacts with serving in non-obvious ways. The draft model used for speculation adds its own memory requirement and compute overhead. For a single-request workload, speculative decoding can reduce latency by 2-3x. For a high-throughput batched workload, the benefit shrinks because batched decode is already efficient: the KV cache reads that make decode memory-bandwidth-bound are partially offset by the larger batch. Teams deploying speculative decoding in production must validate improvements in their specific workload metrics rather than assuming the laboratory results transfer directly.

Summary

Model serving turns a trained model into a production service. The key concepts from this chapter are:

  • Model loading involves copying weights from disk to GPU memory, with large models requiring tensor parallelism across multiple GPUs. Cold start latency is high and must be managed through proactive capacity planning.

  • Memory budgeting requires accounting for model weights, activations, and KV cache. KV cache consumption grows with both sequence length and concurrent request count, making careful configuration essential. PagedAttention eliminates fragmentation and dramatically improves KV cache efficiency.

  • Request handling encompasses queuing policies (FCFS, SJF, priority), batching strategies (static, dynamic, continuous), and rate limiting to protect server capacity. Continuous batching provides the highest GPU utilization for variable-length generation workloads.

  • The prefill-decode asymmetry shapes every serving design decision. Prefill is compute-bound and spikes GPU utilization; decode is memory-bandwidth-bound and generates one token at a time. Chunked prefill and disaggregated serving architectures address this asymmetry by preventing long prefills from disrupting decode latency.

  • Serving frameworks like vLLM, TGI, and TensorRT-LLM provide production-ready implementations of these concepts with different trade-offs in performance, flexibility, and ease of use.

  • Serving configuration requires explicitly trading throughput for latency based on application requirements. Key parameters include gpu_memory_utilization, max_num_seqs, batch size limits, and chunked prefill settings. Tracking latency as a distribution (p50, p95, p99) rather than an average reveals tail behavior that average metrics hide.

  • Deployment patterns such as canary rollouts and blue-green deployments provide safe paths for model updates without service interruption. Autoscaling for LLMs requires proactive capacity management because cold start times are too long for reactive scaling.

The next chapter covers latency optimization techniques specific to serving infrastructure: profiling request latency, reducing time-to-first-token through prefill optimization, and tuning decode throughput.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about model serving.

Model Serving Quiz

Question 1 of 100 of 10 completed
A 13-billion-parameter model is stored in float16 (2 bytes per parameter). Approximately how much GPU memory do the model weights alone require?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026modelserving, author = {Michael Brenndoerfer}, title = {Model Serving: Frameworks, Loading}, year = {2026}, url = {https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Model Serving: Frameworks, Loading. Retrieved from https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration
MLAAcademic
Michael Brenndoerfer. "Model Serving: Frameworks, Loading." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration>.
CHICAGOAcademic
Michael Brenndoerfer. "Model Serving: Frameworks, Loading." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Model Serving: Frameworks, Loading'. Available at: https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Model Serving: Frameworks, Loading. https://mbrenndoerfer.com/writing/model-serving-frameworks-loading-request-handling-configuration

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.