Part of Language AI Handbook
Covers LLM serving latency from TTFT to TPOT. Topics include KV cache management, continuous batching, speculative decoding, streaming responses.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Latency Optimization
When you type a prompt and wait for a language model to respond, time passes across a surprising number of stages. There is the moment the request travels to the server, the setup the model does before it starts generating, the time spent producing each token, and finally the trip your answer takes back across the network. Each of these stages contributes to the total wall-clock time you experience, yet they have almost nothing to do with one another. Reducing one does not automatically reduce the others. That asymmetry is the central insight of latency optimization: you must understand which stage governs the user experience before you can make meaningful improvements.
Latency matters in a way that raw throughput does not. A model that produces a thousand tokens per second across a large batch might still feel sluggish if it takes five seconds before the first word appears. Conversely, a model that starts writing in under a second feels responsive even when the full completion takes longer. This chapter dissects the latency budget of a modern LLM serving system, explains the bottlenecks that arise at each stage, and covers the practical strategies, from KV-cache management to speculative decoding to streaming, that practitioners use to make large models feel fast.
What makes LLM latency optimization unusual compared to traditional software performance work is the interplay between memory, compute, and user perception. Classical web services are often network-bound or database-bound; optimize the database query and the page loads faster. LLM serving combines memory-bandwidth bottlenecks during decode, compute bottlenecks during prefill, and psychological factors around what humans perceive as "fast." A response that streams words at a natural reading pace feels interactive even when the underlying model is not unusually fast. Engineering for perceived latency is as important as engineering for measured latency.
There is also an important distinction between latency and throughput that shapes every design decision in this chapter. Throughput measures how many requests a system can serve per unit of time across a whole workload. Latency measures how long a single user waits for their answer. Many techniques that improve throughput (larger batches, longer queues) hurt individual latency. The techniques in this chapter are specifically chosen to improve latency, and we will be explicit about the throughput trade-offs they introduce.
Understanding the Latency Budget
A request to an LLM serving system travels through a pipeline of stages, and the total latency is the sum of all of them. Understanding this pipeline precisely is the first step toward optimizing it, because different optimizations target different stages.
The stages and the metrics that matter most for each one are worth enumerating clearly.
Time to First Token
Time to first token (TTFT) is the elapsed time from when a client sends a request to when the first token of the response arrives. This metric governs perceived responsiveness. If you are building a chat interface, TTFT is the most user-visible latency metric, because it determines how long the screen remains blank after the user submits a message.
TTFT is dominated by the prefill phase: the forward pass over the entire input prompt. During prefill, the model processes all input tokens simultaneously, computing keys and values for every layer and populating the KV cache. This work scales with the prompt length. A 4,000-token system prompt is dramatically more expensive to prefill than a 50-token conversational message. Improving TTFT typically means either reducing prompt length, speeding up the prefill computation (for example with tensor parallelism), or caching prefill results so they do not have to be recomputed.
One subtlety that often surprises practitioners is that TTFT also includes queue time: the interval between when a request arrives and when the model begins processing it. On a lightly loaded system, queue time is negligible. On a saturated system, a request might wait hundreds of milliseconds before prefill even begins. This means that TTFT SLOs can be violated not because prefill is slow, but because the serving system is overloaded. Optimizing prefill time on an overloaded system does nothing for user experience; adding capacity or better request scheduling is the correct lever.
Serving systems often distinguish between cold TTFT (the first request after a cold start, when nothing is cached) and warm TTFT (subsequent requests that benefit from prefix caching and warm GPU caches). In production, you care primarily about warm TTFT under steady load, but cold TTFT matters for auto-scaling scenarios where new instances are spun up to handle load spikes.
Time per Output Token
Time per output token (TPOT), also written as inter-token latency, is the time between consecutive generated tokens. It determines how fast the streamed response visually updates. If TPOT exceeds roughly 60 to 80 milliseconds per token, a human reader will notice the model "typing" one word at a time at a pace that feels slow. At 50 ms per token, the experience resembles a fast typist. At 30 ms or below, the text appears nearly instantaneous.
The decode phase that governs TPOT is fundamentally different from prefill. During decoding, the model generates tokens one at a time, each requiring a full forward pass over the model, but reusing the stored KV cache from previous steps. Each decode step processes only a single token (or a small batch), so the computation is relatively light, but it still reads the entire set of model weights from memory. On most hardware, decode is memory-bandwidth bound, not compute bound. The limiting factor is how fast the GPU can stream weights from HBM (high-bandwidth memory) into the compute cores, not how fast it can multiply matrices.
The implication is that arithmetic intensity (the ratio of floating-point operations to memory bytes transferred) during decode is very low. A transformer with 70 billion parameters at 16-bit precision requires moving roughly 140 GB of weights per decode step. If the GPU provides 2 TB/s of memory bandwidth, the raw minimum decode time is 70 ms per step per request. Batching multiple requests together increases arithmetic intensity and amortizes that memory read across many outputs simultaneously. This single observation explains most of the architecture of modern LLM serving systems: the entire engineering effort around batching, continuous scheduling, and speculative decoding is, at root, an effort to overcome the memory-bandwidth bound on decode.
TPOT is not perfectly constant across a generation. The first few decode steps after prefill may be slower because the GPU is transitioning from the compute-intensive prefill kernel to the memory-bandwidth-intensive decode kernel. Steps that trigger KV cache management operations (reordering pages, evicting less-used entries) are also slower. Production profiling typically shows a distribution of per-step TPOT values rather than a single constant, and the P99 of that distribution can be significantly higher than the median.
End-to-End Latency
End-to-end latency is the time between sending a request and receiving the complete response. It is approximately:
where is the number of output tokens generated.
When streaming is disabled (the client waits for the full completion), end-to-end latency is the only metric that matters. When streaming is enabled, TTFT and TPOT are both important, since users see each token arrive incrementally.
A useful reframing of this equation is to ask what drives end-to-end latency across different workload types. For a short conversational response of 50 tokens, TTFT may account for the majority of latency. For a long document generation of 2,000 tokens, the decode term dominates entirely. Optimizing TTFT for a long-generation workload has diminishing returns because decode time is so much larger; the reverse is true for short-answer tasks.
The Decode Bottleneck in Detail
To see why decode is memory-bandwidth bound, consider the arithmetic intensity of a single matrix-vector product, which is the dominant operation during a single-token decode step. For a weight matrix and input vector :
- FLOPs: multiply-add operations
- Bytes read: (for the weight matrix, at 2 bytes per float16 element) plus (for the input vector)
The arithmetic intensity is approximately:
Modern GPUs can perform roughly 100-300 FLOPs per byte of memory bandwidth (their compute-to-bandwidth ratio). When arithmetic intensity is only 1 FLOP/byte, the GPU's compute units sit idle while waiting for memory. Increasing batch size during decode raises arithmetic intensity because the same weight matrix serves multiple input vectors simultaneously, moving the operation closer to the hardware's compute-bound regime.
This analysis also explains why quantization has such a significant impact on decode speed. If you represent weights at int8 instead of float16, the same weight data occupies half the memory. A matrix that previously required reading 140 GB of bytes from HBM now requires only 70 GB, and the minimum decode latency halves accordingly. The arithmetic operations are still performed at higher precision after dequantization, so quality is largely preserved. This is the core reason weight-only quantization (quantizing weights but not activations) is so popular for serving: it directly attacks the memory-bandwidth bottleneck without the accuracy risks of full quantization.

The chart makes the decode bottleneck concrete. A single-request decode step achieves barely 1 FLOP per byte of memory bandwidth transferred, while the hardware can sustain 100 or more FLOPs per byte when kept busy. Prefill over a 512-token prompt sits comfortably in the compute-bound regime because all those tokens are processed in a single large matrix multiplication. The key takeaway is that even a batch of 64 concurrent decode requests barely scratches the roofline, which is why practitioners focus so heavily on maximizing batch sizes during decode.
Network and Queuing Overhead
Beyond GPU computation, requests spend time in queuing (waiting for the model to become available), serialization (converting tensors to bytes for transmission), and network transit. In collocated deployments these overheads are typically under 10 milliseconds, but in multi-region or edge deployments network round-trip time can dominate. A round trip from Europe to a US-East data center adds roughly 100 ms of irreducible latency that no amount of GPU optimization can remove. Latency profiling must instrument all stages to avoid optimizing the wrong part of the pipeline.
Queuing deserves special attention because it is often the dominant component in production systems under load. When request arrival rate approaches service rate, queuing theory predicts that average queue time grows without bound. Even at 80% utilization, a system with Poisson arrivals experiences average queue times roughly four times the average service time. This means that a system with 500 ms average service time at 80% utilization will have average queue times of 2,000 ms, making TTFT nearly 2.5 seconds regardless of how fast the GPU is. Capacity planning and admission control are therefore first-order latency concerns, not afterthoughts.
KV Cache Management
The key-value (KV) cache is the most consequential data structure in LLM serving. Every transformer layer computes key and value projections for each token during the forward pass. During autoregressive decoding, these values do not change for previously computed tokens. Storing them avoids recomputing the full attention operation from scratch at every step.
Without a KV cache, decoding token would require processing all tokens again. With the cache, decoding token processes only the new token and reads the cached keys and values for positions through . The time savings are dramatic: without caching, decode time would grow as ; with caching, it is .
To understand why the KV cache is so large and why managing it carefully matters, it helps to trace what is being stored. In each transformer layer, for each token position, the model computes a key vector and a value vector for each attention head. These vectors are the inputs to the attention operation. Once computed, they never change for that position, so storing them avoids recomputing the attention keys and values from the residual stream at every subsequent decode step. The tradeoff is memory: every token in the context requires storing vectors of dimension , and this memory grows linearly with context length.
KV Cache Size
The KV cache size for a single sequence is:
where:
- : number of transformer layers
- : number of attention heads per layer
- : dimension of each attention head
- : sequence length (prompt plus all generated tokens so far)
- The factor of 2 accounts for both keys and values
For a 70-billion parameter model with 80 layers, 64 heads, and 128 dimensions per head, storing the KV cache for a single 4,096-token context at float16 requires:
This is a substantial fraction of a GPU's HBM. Managing KV cache memory is one of the primary constraints on serving throughput and the primary reason batching cannot be increased without limit.
def kv_cache_size_gb(
num_layers, num_heads, head_dim, seq_len, bytes_per_elem=2
):
"""Compute KV cache size in GB for a single sequence."""
total_bytes = (
2 * num_layers * num_heads * head_dim * seq_len * bytes_per_elem
)
return total_bytes / (1024**3)Model 512 tokens 2K tokens 4K tokens 8K tokens --------------------------------------------------------------------------------------------- 7B (32 layers, 32 heads, 128 dim) 0.25G 1.00G 2.00G 4.00G 13B (40 layers, 40 heads, 128 dim) 0.39G 1.56G 3.12G 6.25G 70B (80 layers, 64 heads, 128 dim) 1.25G 5.00G 10.00G 20.00G
KV cache memory scales linearly with sequence length, number of layers, and number of heads. A 70B model serving a single request with an 8,192-token context uses over 20 GB of KV cache, leaving little room on an 80 GB GPU for model weights and activations, let alone other requests. This arithmetic explains why modern decoder models increasingly use grouped-query attention (GQA) or multi-query attention (MQA), both of which reduce the number of KV heads. By sharing keys and values across multiple query heads, GQA can reduce KV cache size by 4x to 8x without meaningful quality loss. We covered the architectural details of GQA in the chapter on modern decoder models; here the practical consequence is simply that models designed with serving in mind tend to have significantly smaller KV caches than their older counterparts.
Paged Attention
Traditional KV cache implementations pre-allocate a contiguous memory block for each request based on the maximum sequence length. This wastes memory because most requests do not fill their allocated space, and it makes it impossible to split a cache across non-contiguous memory regions.
PagedAttention (introduced with the vLLM system) borrows virtual memory concepts from operating systems. The KV cache is divided into fixed-size pages (typically 16 tokens per page), and a page table maps logical page indices to physical memory locations. When a request needs more cache space, it is allocated the next free physical page, which may not be contiguous with its existing pages. Attention kernels are modified to perform the scatter-gather needed to read non-contiguous pages correctly.
The gains are significant. With contiguous allocation, a request for a 1,000-token sequence that uses only 300 tokens wastes 700 tokens of memory. With paged allocation, only populated pages consume memory. System-wide memory utilization improves, allowing more requests to be batched simultaneously. Paging also enables prefix caching (discussed below) because multiple sequences can share the same physical pages for common prefixes without copying.
An additional benefit of paged allocation is that it enables copy-on-write semantics for beam search and parallel sampling. When a single prompt generates multiple completions (as in beam search), the prompt's KV pages can be shared across all beams until the beams diverge. Only the pages containing positions where the beams differ need to be copied. This reduces memory usage for multi-completion workloads substantially compared to the naive approach of allocating a full separate cache per beam.
The tradeoff of paged attention is implementation complexity. Attention kernels must support non-contiguous memory access, which requires custom CUDA kernels rather than standard matrix multiply routines. Libraries like vLLM and TensorRT-LLM have invested significant engineering effort in these kernels. For practitioners using these libraries, paged attention is transparent; for those building custom serving infrastructure, implementing it from scratch is a substantial undertaking.
Prefix Caching
Many production LLM deployments send the same system prompt or few-shot examples with every request. Without caching, each request must prefill the entire prompt, even when the first 2,000 tokens are identical to every previous request.
Prefix caching stores the KV states for common prompt prefixes in a shared cache. When a new request arrives with a prefix matching a cached entry, the serving system skips prefill for those tokens and loads their KV states directly from cache. This turns an expensive compute operation into a memory read, reducing TTFT proportionally to the fraction of the prompt that hits the cache.
The effectiveness of prefix caching depends on cache hit rate, which in turn depends on how repetitive the prompt patterns are. Retrieval-augmented generation systems, where document chunks are retrieved and inserted into prompts, can achieve high hit rates for repeatedly retrieved chunks. Multi-turn conversation systems benefit when system prompts are cached across turns.
Prefix caching works well when the shared prefix appears at the very beginning of every prompt, because token positions in a causal attention model must be consistent with the original sequence. A cached prefix for positions 1 through can only be reused if the new request's first tokens are byte-for-byte identical to those in the cache entry. Even a single token difference invalidates the match. This constraint means that serving systems typically hash the prefix tokens to identify matching entries, and any mutation of the system prompt (even adding a timestamp) defeats the cache entirely. Prompt engineering for cached deployments must be disciplined about keeping the shared prefix stable.
KV Cache Quantization
Beyond paging and prefix sharing, another approach to extending KV cache capacity is quantizing the cached tensors to lower precision. Standard KV caches store keys and values as float16 (2 bytes per element). Reducing to int8 (1 byte) halves the memory footprint; int4 reduces it to a quarter.
The accuracy impact of KV cache quantization is generally mild for int8 and acceptable for int4 in many settings, though it grows with context length. The keys and values that most affect current attention scores are those for recent tokens; distant token values can tolerate more quantization error before outputs are noticeably affected. This observation motivates mixed-precision KV caching, where recent tokens are stored at higher precision and older tokens at lower precision, though production implementations of this are still rare.
Batching Strategies
Batching is the primary technique for improving GPU utilization during decode. The core idea is simple: if the GPU must read all model weights to decode one token, it might as well decode a token for many requests simultaneously. The memory bandwidth cost is roughly the same, while the compute throughput increases linearly with batch size.
Understanding the batching strategies available, and their trade-offs, is essential because the choice of batching strategy determines both throughput and latency behavior under realistic serving conditions. No single strategy is optimal for all workloads; the best choice depends on the distribution of request lengths and the latency SLOs you need to meet.
Static Batching
The simplest approach collects a fixed set of requests, runs them all together until all sequences are complete, and then collects the next batch. The problem is that sequences in a batch complete at different times. When the first sequence finishes, its slot in the batch sits idle until all other sequences finish. This padding waste is severe when sequence lengths vary widely.
To see just how bad this can be, consider a batch of eight requests where seven generate 50 tokens and one generates 400 tokens. The batch completes only when the 400-token sequence finishes. During the final 350 decode steps, seven of the eight batch slots are doing no useful work. GPU utilization for those steps is 12.5%. This is not a pathological edge case; real-world request length distributions are highly skewed, with the majority of requests being short and a long tail of much longer ones.
Static batching does have one advantage: simplicity. The entire batch is processed with a single fixed-size matrix operation at each decode step, which maps cleanly to standard CUDA kernels without requiring custom scatter-gather logic. For workloads with very uniform sequence lengths (all requests generating exactly tokens), static batching wastes no time and achieves high throughput. In practice, however, this uniformity rarely holds outside of benchmarking scenarios.
Continuous Batching
Continuous batching (also called dynamic batching or iteration-level scheduling) eliminates padding waste by inserting new requests at the level of individual decode iterations. When one sequence in a batch completes, the scheduler immediately fills its slot with a new request. The batch composition changes from step to step, but the GPU always runs at the target batch size.
This is now the standard approach in production systems. vLLM, TGI (Hugging Face Text Generation Inference), and NVIDIA Triton Inference Server all implement variants of continuous batching. The throughput improvement over static batching on realistic workloads is often 10-50x because the GPU stays busy rather than waiting for stragglers.
The key implementation insight is that every decode step is structurally identical: the model reads its weights, attends over cached keys and values, and produces a logit distribution for the next token. The fact that different sequences in the batch are at different positions in their respective generations does not change this structure. Each sequence simply provides a different query vector to the attention operation and reads from a different slice of the KV cache. With paged attention managing those cache slices, combining sequences at arbitrary positions in their generations is straightforward.
One important consideration for continuous batching is how to handle the prefill step for new requests that join the batch mid-iteration. Prefill is a different operation from decode: it processes many tokens at once in a single large matrix multiplication. Mixing prefill work with in-flight decode work in the same GPU kernel is architecturally complex. Most serving systems handle this by interleaving: one iteration computes the prefill for a newly arrived request, and the next iteration includes that request in the decode batch. This introduces one iteration of extra latency for new requests but keeps the implementation clean.

The visualization makes the inefficiency of static batching concrete. When the longest sequence in a batch is eight times the shortest, static batching achieves only about 12% utilization during the final steps. Continuous batching achieves near-full utilization by replacing completed requests immediately.
Chunked Prefill
One tension in continuous batching is that long prefill operations can block decode steps. When a new request with a 4,000-token prompt joins the system, its prefill computation can take hundreds of milliseconds, delaying all in-flight decode requests and spiking TTFT for other users.
Chunked prefill addresses this by breaking large prompts into smaller chunks (for example, 512 tokens at a time) and interleaving them with decode steps. The prefill of a single long request is spread across many iteration steps, allowing decode requests to progress in between. This trades a modest increase in per-request TTFT (because the prefill now takes multiple iterations instead of one) for much better TTFT consistency across the system as a whole.
The chunk size is a tunable parameter. Smaller chunks reduce head-of-line blocking but increase the scheduling overhead of managing partially prefilled requests. Larger chunks reduce overhead but increase the worst-case TTFT spike for decode requests running in parallel. Production systems typically tune this parameter empirically based on observed workload characteristics.
Chunked prefill also interacts with the KV cache in an interesting way. A partially prefilled request consumes KV cache pages for the tokens processed so far, but the attention operation for those tokens is not yet useful for generation. The serving system must track the boundary between prefilled and not-yet-prefilled positions and ensure that attention is only computed over the prefilled portion until the full prompt is processed.
Speculative Decoding
The core latency bottleneck during decode is that each token requires a full forward pass through the model. Speculative decoding breaks this constraint by using a small, fast draft model to generate candidate tokens cheaply, then verifying multiple candidates with the large target model in parallel.
The intuition behind speculative decoding is straightforward. Most of the time, the next few tokens in a generation are highly predictable: function words like "the", "is", "of", high-frequency phrases, and continuations that any reasonable language model would select. A small model can predict these tokens quickly. If the large model would have chosen the same tokens, you get those tokens for free. If the small model was wrong, you fall back to the large model's choice and continue. In expectation, you generate multiple tokens per large-model forward pass rather than exactly one.
The Algorithm
Given a target model and a draft model :
-
The draft model generates candidate tokens autoregressively. This is fast because the draft model is much smaller.
-
The target model performs a single forward pass over the draft tokens simultaneously (treating them like a prefill), computing the probability distribution for each position.
-
Each draft token is accepted with probability:
where:
- : probability assigned to the draft token by the target model
- : probability assigned by the draft model
- : a rejection sampling correction that guarantees the accepted tokens follow the exact target distribution
- If a token is rejected, the target model samples a corrected token at that position, and decoding restarts from there. The key guarantee is that the final output distribution is mathematically identical to sampling directly from the target model, token by token.
The acceptance probability formula has a clean probabilistic interpretation. If the target model assigns higher probability to a draft token than the draft model does (), the token is accepted with probability 1: the small model was being conservative about a token the large model is confident in, so there is no downside to accepting. If the target model assigns lower probability (), the token is accepted only with probability : the small model was overconfident, and we accept in proportion to how plausible the target model finds the token. The rejection step, when it occurs, samples from a residual distribution that corrects for the excess probability the draft model placed on incorrect tokens.
The expected number of accepted tokens per target model call, called the acceptance length, determines the speedup. If the acceptance length is , then speculative decoding reduces the number of target model calls needed for tokens by a factor of , at the cost of running the draft model for the same tokens. The net speedup depends on the ratio of draft model cost to target model cost and the acceptance length achieved.

The heatmap reveals the regime where speculative decoding helps most: high acceptance rates (the draft model frequently matches the target) combined with a cheap draft model (low draft-to-target cost ratio). In practice, acceptance rates of 60-80% are typical when the draft model is a smaller member of the same model family, and draft models are often 10-20x cheaper than the target, placing most real systems well into the beneficial zone.
Draft Model Selection
The draft model must be fast (to minimize the cost of generating candidates) and accurate (to maximize acceptance rate). The most common approaches are:
- Smaller model from the same family: A 7B model drafting for a 70B model typically achieves higher acceptance rates than an unrelated model because they were trained on the same data and share similar inductive biases.
- Distilled draft model: A model trained specifically to mimic the target model's distribution, which can achieve high acceptance rates at very small model sizes.
- Self-speculative decoding: Drafts are generated using the first few layers of the target model and then verified by the full model. This requires no separate draft model but introduces complexity in the forward pass implementation.
- Lookahead decoding: Generates draft tokens using n-gram patterns from the current context, requiring no model at all for the draft step. Acceptance rates are lower but draft cost approaches zero.
The choice of draft model strategy depends heavily on the deployment context. If you are already running a smaller model for other purposes, using it as a draft model adds nearly zero marginal cost. If the serving system is memory-constrained (a common situation when running very large models on limited GPU counts), self-speculative decoding avoids the memory overhead of loading a separate draft model while still providing latency improvements. Lookahead decoding is most effective for tasks like code generation where n-gram repetition is high, and least effective for open-ended creative generation where context repetition is rare.
The number of draft tokens is also a tunable parameter. More draft tokens mean larger potential speedups when acceptance rates are high, but the benefit saturates because the probability of an unbroken run of accepted tokens falls as for a geometric acceptance model. A typical value of to captures most of the benefit without excessive overhead.
Correctness Guarantee
Speculative decoding is an exact algorithm: the output token distribution is identical to running the target model alone. This follows from the rejection sampling correction. When a draft token is accepted, it is a sample from (the target) because the acceptance criterion reweights the proposal to match . When a draft token is rejected, the target model samples directly from . In both cases, the accepted sequence has the same distribution as direct sampling from .
This guarantee distinguishes speculative decoding from approximations like quantization or pruning. Speculative decoding never degrades output quality, only output speed. This is a remarkable property: you get a latency reduction for free, with no quality-speed trade-off. The only cost is compute efficiency at the system level: the draft model consumes resources, and rejected draft tokens represent wasted compute. This is why speculative decoding is most attractive when a single request occupies dedicated hardware, and less attractive when the system is already at full capacity.
Speculative Decoding in Practice
Beyond the textbook algorithm, production deployments of speculative decoding require handling several practical concerns. The draft model and target model must use compatible tokenization so that draft tokens correspond to positions in the target model's vocabulary. The draft model's KV cache must be managed separately from the target model's, adding memory overhead. And the serving framework must be able to execute the draft generation and target verification steps efficiently, ideally as consecutive GPU kernel launches without returning to the CPU between steps.
Modern serving frameworks like vLLM provide first-class support for speculative decoding, handling these details transparently. For practitioners, the main decision is selecting the draft model and the speculation depth , then measuring whether the acceptance rate on your specific workload is high enough to justify the overhead. If your users are asking questions with predictable, formulaic answers, speculative decoding is likely beneficial. If they are asking for highly creative or unpredictable generations, acceptance rates may be too low for speculative decoding to help.
Streaming Responses
Streaming is the technique of sending generated tokens to the client as they are produced, rather than buffering the entire response and sending it all at once. For users, the experience difference is dramatic: instead of waiting 20 seconds for a long response and then seeing it all at once, they see words appearing progressively, making the system feel much more responsive.
The perception effect of streaming is well-documented in user research on web applications. Users consistently rate streaming interfaces as faster even when the total completion time is identical to the non-streaming version. This happens because streaming provides visual feedback that something is happening, reducing anxiety about whether the system has accepted the request, and because reading can begin before generation is complete for long responses. Streaming is therefore one of the highest-impact latency optimizations available, and it costs essentially nothing to implement in most serving frameworks.
How Streaming Works
In HTTP-based serving, streaming is typically implemented using Server-Sent Events (SSE) or chunked transfer encoding. The serving system writes each generated token (or small group of tokens) to the HTTP response body as it becomes available, and the client reads these incremental updates in a streaming loop.
Server-Sent Events use a text-based protocol where each chunk is formatted as data: <payload>\n\n. Most LLM APIs (OpenAI, Anthropic, etc.) use SSE with JSON-encoded token payloads. The client connects to the API endpoint, reads the SSE stream line by line, parses the JSON payload from each data: line, and appends each token to the display. When the model finishes generating, the server sends a special termination event (often data: [DONE]), and the client closes the connection.
import random
import time
def simulate_streaming_decode(
prompt_tokens, n_output_tokens, tpot_ms=50.0, ttft_ms=500.0
):
"""
Simulate a streaming LLM response.
Yields (token_text, elapsed_ms) tuples as each token becomes available.
"""
# Prefill phase: simulate TTFT
time.sleep(ttft_ms / 1000.0)
elapsed = ttft_ms
vocab = [
"the",
"model",
"generates",
"tokens",
"one",
"by",
"step",
"using",
"cache",
]
for i in range(n_output_tokens):
token = random.choice(vocab)
time.sleep(tpot_ms / 1000.0)
elapsed += tpot_ms
yield token, elapsedFirst token at: 350 ms (TTFT) Final token at: 1300 ms (end-to-end latency) Average TPOT: 50.0 ms Sample tokens: model the step tokens tokens generates model efficiently ...
The simulation illustrates the two-phase structure of every streaming response. The client waits for TTFT (the prefill phase), then receives tokens at the TPOT rate. End-to-end latency is the sum of both phases, but the user's perceived experience depends heavily on TTFT since that is when the response begins to appear.
Chunking and Detokenization
A subtle issue in streaming is that tokens do not always align with visible characters. A word like "unbelievable" might tokenize as "un", "believ", "able", three separate tokens. Streaming each token independently would show "un" first, then "believ", then "able", which looks unnatural in a character-based display. Many production systems buffer a small number of tokens (or characters) before flushing to the client to create a more natural visual flow.
A more significant issue is multi-byte Unicode characters. The UTF-8 encoding of many non-ASCII characters spans multiple bytes, and a byte-level tokenizer might split a character across two tokens. Sending a partial character to the client causes a decoding error. Serving systems must either use character-level streaming with proper byte buffering, or operate at the token level and handle multi-byte characters explicitly.
Special characters in Markdown and code also require careful handling in streaming contexts. If the model is generating a code block that begins with triple backticks, the client renderer may not know how to render the content until the closing backticks arrive. Some serving systems send raw tokens while others implement client-side buffering for structured content like tables, code blocks, and LaTeX equations. The trade-off is between immediate visual feedback and correct rendering of complex content.
Streaming and Latency SLOs
Streaming changes how latency service level objectives (SLOs) are defined. Without streaming, the SLO is straightforward: 95% of requests must complete in under seconds. With streaming, you typically define two SLOs:
- P95 TTFT: 95% of requests must produce their first token within ms
- P95 TPOT: 95% of inter-token intervals must be under ms
These two SLOs can be in tension. Strategies that reduce TTFT (like prioritizing prefill work) can increase TPOT because decode steps are delayed while prefill runs. The chunked prefill technique described earlier is a direct response to this tension.
Setting appropriate SLO values requires understanding your application's user experience requirements. For a chat assistant, users notice TPOT above 80 ms and TTFT above 1 second. For an API serving developers building non-interactive pipelines, end-to-end latency matters more than TTFT or TPOT individually. For a code autocomplete tool (where the model is completing partial lines as the user types), TTFT must be extremely low (under 200 ms) and TPOT is secondary because completions are usually short. The product's user experience determines the SLO structure alongside engineering constraints.
Latency Profiling
Optimizing latency without measuring it is guesswork. A profiling system that instruments every stage of the request pipeline turns optimization from art into engineering.
Profiling an LLM serving system is more challenging than profiling a typical web service because the stages of interest span multiple levels of abstraction: Python application code, CUDA kernels on the GPU, and hardware-level memory transactions. A request that appears slow might be waiting in a Python-level queue, waiting for a CUDA kernel to launch, waiting for GPU memory bandwidth, or all three simultaneously. Effective profiling instruments each level and correlates their observations into a single per-request timeline.
The most actionable profiling data combines high-level per-stage timing (easy to collect in Python with time.monotonic()) with periodic GPU utilization and memory bandwidth metrics (available through NVIDIA's DCGM telemetry and Nsight Systems). Correlating these two streams lets you distinguish between "request is slow because it is waiting to be scheduled" and "request is slow because its GPU kernels are bottlenecked on memory bandwidth."
Key Metrics to Instrument
A well-instrumented LLM serving system tracks the following for every request:
- Queue time: time spent waiting for a free GPU slot or batch slot
- Prefill time: time to compute attention and fill the KV cache for the input prompt
- Decode time: total time across all autoregressive steps
- Per-step decode time: time for each individual decode iteration (useful for detecting slowdowns from KV cache eviction or memory pressure)
- Network time: time for the response to travel from server to client
import time
from dataclasses import dataclass, field
from typing import Optional
@dataclass
class RequestTrace:
request_id: str
prompt_tokens: int
output_tokens: int = 0
queue_start: float = field(default_factory=time.monotonic)
queue_end: Optional[float] = None
prefill_start: Optional[float] = None
prefill_end: Optional[float] = None
decode_start: Optional[float] = None
decode_end: Optional[float] = None
@property
def queue_time_ms(self) -> float:
if self.queue_end is None:
return 0.0
return (self.queue_end - self.queue_start) * 1000
@property
def prefill_time_ms(self) -> float:
if self.prefill_start is None or self.prefill_end is None:
return 0.0
return (self.prefill_end - self.prefill_start) * 1000
@property
def decode_time_ms(self) -> float:
if self.decode_start is None or self.decode_end is None:
return 0.0
return (self.decode_end - self.decode_start) * 1000
@property
def ttft_ms(self) -> float:
"""Time from queue start to first token (end of prefill)."""
if self.prefill_end is None:
return 0.0
return (self.prefill_end - self.queue_start) * 1000
@property
def tpot_ms(self) -> float:
"""Average time per output token."""
if self.output_tokens <= 1 or self.decode_time_ms == 0:
return 0.0
return self.decode_time_ms / self.output_tokens=== Latency Profile Summary (N=200 simulated requests) === Time to First Token (TTFT): P50: 269.2 ms P95: 786.1 ms P99: 1012.6 ms Time per Output Token (TPOT): P50: 51.5 ms P95: 64.4 ms P99: 68.4 ms End-to-End Latency: P50: 14.70 s P95: 25.48 s P99: 29.95 s Latency Breakdown (median request): Queue: 21.8 ms (0.1%) Prefill: 235.4 ms (1.6%) Decode: 14419.8 ms (98.2%)
The profile shows how latency is distributed across stages and that different stages dominate for different workloads. Short requests with long system prompts are prefill-dominated; long generation tasks are decode-dominated; overloaded systems are queue-dominated. This three-way breakdown is the starting point for any optimization effort: there is no universal best strategy, only strategies that are best for the specific breakdown you observe.


The contrasting shapes of these two distributions capture the essential asymmetry of LLM latency. TTFT has a heavy right tail driven by long prompts, while TPOT is tightly clustered because it is hardware-bound rather than workload-bound. This asymmetry matters for SLO design: TTFT P99 can be much larger than P50, requiring the tail to be specifically optimized (for example by prioritizing short-prompt requests), while TPOT can be reliably predicted from hardware specifications.
Identifying Bottlenecks
With per-stage latency data, bottleneck identification becomes systematic. A few patterns are particularly actionable:
- High queue time relative to total latency indicates capacity shortage. The system is overloaded and either needs more capacity or better request admission control.
- Prefill time growing faster than prompt length suggests memory fragmentation or lack of prefix caching. A 2x increase in prompt tokens should produce at most a 2x increase in prefill time.
- TPOT varying with batch size confirms memory-bandwidth-bound behavior. If TPOT increases as batch size grows beyond a threshold, the GPU memory bandwidth is saturated and the batch is too large.
- Outlier TPOT values (P99 >> P50) often indicate KV cache eviction events where the system is swapping cached states to CPU memory to handle new requests.

The scatter plot confirms the linear relationship between prompt length and prefill time. The slope, roughly 0.5 ms per token, is the simulation's effective per-token prefill cost. Requests that fall well above the trend line are candidates for optimization via prefix caching or prompt compression.
Production Monitoring Infrastructure
Beyond per-request profiling, latency optimization requires continuous production monitoring. Latency regressions can be introduced silently by changes to model versions, serving configuration, hardware utilization trends, or traffic pattern shifts. A monitoring system that tracks latency percentiles over time, alerts on P95 threshold violations, and provides dashboards for capacity planning is essential infrastructure for any serious LLM deployment.
The metrics that matter most for production monitoring are organized around the percentiles that correspond to your SLOs. If your TTFT SLO is P95 under 1 second, you need a timeseries of P95 TTFT measured every minute, with alerts when it exceeds threshold. Tracking only averages is insufficient: the average TTFT can look healthy while P99 is already degrading, which is the situation that causes user complaints.
Distributed tracing (using systems like OpenTelemetry, Jaeger, or cloud-native equivalents) provides the full picture by correlating traces across all layers of the serving stack. Each request gets a unique trace identifier that propagates through all components, allowing you to reconstruct the complete timeline of a slow request including queue time, model initialization, GPU kernel launches, and network transmission. This is invaluable for diagnosing tail-latency issues that only appear on specific request types or under specific load conditions.
Putting It Together: A Latency Optimization Workflow
A principled approach to latency optimization follows a consistent sequence of steps regardless of the specific system.
The first step is measurement. Before changing anything, instrument the serving system to collect TTFT, TPOT, queue time, and per-stage breakdowns. Establish baseline percentiles (P50, P95, P99) for your target workload. This baseline is what all subsequent improvements will be compared against.
The second step is identifying the dominant bottleneck. Examine the latency breakdown. If queue time dominates, the problem is capacity, not optimization. If prefill time is high relative to prompt length, investigate whether prefix caching is effective. If TPOT is high and not improving with higher batch sizes, the model may be too large for the available hardware and quantization or model selection may be the right lever.
The third step is applying targeted optimizations. Continuous batching is almost always worth enabling if not already in place. Prefix caching is effective for workloads with repetitive system prompts. Speculative decoding provides the most benefit on workloads where the draft acceptance rate is high (conversational tasks tend to have higher acceptance rates than creative writing tasks because the vocabulary and style are more predictable). Streaming should be enabled for all user-facing applications regardless of other optimizations.
The fourth step is re-measuring and verifying that the optimization improved the right metric without degrading others. Speculative decoding, for instance, can increase per-request compute and memory usage, potentially hurting throughput. The optimization is only beneficial if the latency improvement is worth the throughput cost for your specific SLOs.
Beyond these four core steps, it is worth considering the interaction between hardware and optimization strategy. A serving stack on older hardware with lower memory bandwidth may benefit more from quantization than one on the latest A100 or H100 GPUs. A system that runs on many small GPUs connected over NVLink may see different bottlenecks from one running on a single large GPU. The optimizations described in this chapter are general principles; their relative importance varies with hardware, model size, and workload characteristics.
Limitations and Trade-offs
Latency optimization in LLM serving requires trade-offs between competing objectives, and practitioners must account for what they give up when optimizing for speed.
The most persistent tension is between latency and throughput. Batching improves throughput by amortizing the memory bandwidth cost of weight loading across many requests, but larger batches mean each individual request waits longer for its slot in the batch. Continuous batching mitigates this by filling slots immediately, but a system running at high throughput will still queue requests during load spikes. Setting the right batch size limit requires knowing your latency SLOs and your expected request arrival rate simultaneously. There is no configuration that simultaneously minimizes per-request latency and maximizes system throughput; the right balance depends on whether you are building an interactive product (where latency matters most) or a batch processing pipeline (where throughput matters most).
Speculative decoding introduces a different trade-off. It reduces latency per token but increases total compute because the draft model runs on tokens that may be rejected. Under heavy load, this extra compute can reduce throughput. Speculative decoding is most beneficial when the system is compute-light (a single user chatting with a deployed model) and least beneficial when the system is fully saturated (a batch processing job at maximum throughput). Measuring both metrics before and after enabling speculative decoding is essential because the benefit varies substantially by workload.
KV cache management presents its own constraints. Larger KV caches support longer contexts but consume memory that could otherwise support larger batch sizes. Systems with tight memory budgets must choose between supporting long contexts for fewer requests or short contexts for more requests. Paged attention and quantizing the KV cache to lower precision (int8 or even int4) extend the effective memory budget, but KV quantization introduces approximation error that can affect quality at long contexts. The error is generally larger for very long contexts where the model must aggregate information from distant positions, precisely the use case that requires the longest context windows.
Streaming, while almost universally positive for perceived responsiveness, requires the client infrastructure to support incremental delivery. HTTP/1.1 connections with slow clients can create head-of-line blocking where a slow consumer holds a server-side buffer open, consuming memory and potentially blocking other requests. Production serving systems must implement client timeouts and back-pressure mechanisms to prevent this. HTTP/2 and HTTP/3 handle streaming more gracefully through stream multiplexing, which is one reason they are preferred for high-scale LLM serving deployments.
Finally, prefix caching only helps when prompts share prefixes. A system serving highly diverse, user-specific prompts may see near-zero cache hit rates, making prefix caching a source of memory overhead with no corresponding benefit. Measuring cache hit rate is essential before committing resources to a large prefix cache. In deployments where cache hit rate is low, the memory would be better used to support larger batches during decode.
There is also an infrastructure-level trade-off between operational simplicity and latency optimality. A system with continuous batching, paged attention, prefix caching, speculative decoding, and chunked prefill is a highly capable serving stack, but it is also a complex one with many interacting parameters. Each optimization introduces potential failure modes: a prefix cache eviction bug can silently serve incorrect outputs, a speculative decoding implementation bug can subtly alter output distributions, and KV quantization errors can be context-length-dependent and hard to detect in standard benchmarks. Adopting mature, well-tested serving frameworks (vLLM, TGI, TensorRT-LLM) rather than building from scratch is usually the right call for most teams, both for reliability and because these frameworks incorporate years of hard-won optimizations.
Summary
LLM serving latency breaks down into a small number of measurable stages: queue time, prefill (which determines TTFT), and decode (which determines TPOT). Each stage has a different bottleneck and a different set of optimizations.
The key insights are:
- Prefill is compute-bound and scales with prompt length. Prefix caching eliminates redundant prefill for shared prompt prefixes. The KV cache must be carefully sized and managed, because at long contexts it dominates GPU memory usage.
- Decode is memory-bandwidth bound and improves with batching. Continuous batching keeps the GPU fully utilized by refilling completed sequence slots immediately. Weight quantization directly reduces the memory bandwidth required per decode step, often halving TPOT for int8 weights.
- Speculative decoding uses a fast draft model to generate candidate tokens, then verifies them in parallel with the target model, achieving exact sampling while reducing the number of full-model forward passes. It is most effective on predictable, conversational workloads where acceptance rates are high.
- Streaming decouples perceived responsiveness from total completion time by delivering tokens to the client as they are generated. It is the highest-ROI optimization for user-facing applications because it improves the subjective experience without changing any serving infrastructure.
- Profiling with per-stage breakdowns is the foundation of all latency optimization: different stages have different bottlenecks and require different solutions. Queue-dominated systems need more capacity; prefill-dominated systems need prefix caching; decode-dominated systems need batching and quantization.
- Trade-offs are unavoidable: latency optimizations often conflict with throughput, and every system must find the balance appropriate to its SLOs and workload.
Together, these techniques make it practical to deploy large language models with latencies that feel interactive rather than slow, transforming the user experience without sacrificing output quality or model capability. The next chapter explores the complementary challenge of throughput optimization, where the goal shifts from minimizing per-request latency to maximizing the number of requests served per second.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about LLM serving latency optimization.
Latency Optimization Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!