Caching for LLMs: Prompt, Semantic, and Invalidation

Michael BrenndoerferFebruary 11, 202652 min read

Part of Language AI Handbook

Covers LLM caching strategies including prompt caching with KV reuse, semantic caching with embedding similarity, cache invalidation policies.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Caching in LLM Production Systems

Every call to a large language model carries a cost measured in dollars, milliseconds, and compute. When users ask the same question twice, or when a system prompt appears in every request, re-running the same computation from scratch wastes all three. Caching addresses this directly: store results you have already computed, and serve them instantly on the next matching request.

Caching is not a new idea. It sits at the foundation of computer architecture, operating systems, and web infrastructure. CPU caches store recently accessed memory words because programs tend to access the same addresses repeatedly. HTTP caches store web page responses because many users request the same pages. Content delivery networks push static assets to edge nodes because geographic proximity reduces latency. In every case, the same insight applies: computation or network round-trips are expensive, and reuse is cheap.

The difficult part of LLM caching is deciding when two requests should match. A CPU cache matches exact memory addresses. A web cache matches exact URLs. An LLM cache must decide what "the same" means for natural language queries, for long shared prefixes in tokenized prompts, and for structured API calls that include model parameters, sampling configuration, and multi-turn conversation history. Each level of the system offers different granularity, different savings, and different implementation tradeoffs.

The economics also favor caching. LLM inference is expensive relative to most computational tasks. Running a large model like GPT-4 or Claude can cost hundreds of times more than a typical database query or API call. Even a modest cache hit rate of 20-30% translates directly into a proportional reduction in compute spend. At scale, the same hit rate produces a substantial reduction in inference spend.

This chapter works through the main forms of LLM caching from the inside out. We begin with prompt caching, which operates at the model level and saves computation on repeated token prefixes. We then move to semantic caching, which operates at the application level and matches queries by meaning rather than exact string equality. We cover cache invalidation, the notoriously hard problem of knowing when a cached answer is no longer valid. Finally, we examine cache hit rates, the key metric for evaluating whether a caching strategy is helping.

Prompt Caching

Prompt caching exploits the internal mechanics of transformer inference. To understand why it saves computation, you need to understand what the model does when it processes a prompt.

The KV Cache and Prefix Reuse

When a transformer processes a sequence of tokens, each attention layer computes key and value matrices for every token in the input. These matrices encode each token's representation in the context of every other token it can attend to. Computing them requires matrix multiplications that scale with both the sequence length and the model's hidden dimension, making them the dominant cost in long-context inference.

The key and value matrices for a given token depend only on that token's position and all preceding tokens. They do not depend on what comes after. This property is what makes prefix reuse possible.

The essential observation is that if two prompts share a prefix, the key and value matrices for that prefix are identical in both requests. There is no mathematical reason to recompute them. Prompt caching implements exactly this optimization: the computed KV tensors for a prefix are stored in fast memory and reused whenever a new request shares that prefix, bypassing the most expensive portion of the forward pass.

KV Cache

The KV cache stores the key and value tensors computed during attention for each token in a sequence. During autoregressive generation, it avoids recomputing attention for already-processed tokens by storing and reusing these tensors. Prompt caching extends this idea across separate requests, not just within a single generation. While the within-request KV cache is a standard optimization in all transformer inference engines, cross-request prefix caching is a newer capability offered by inference providers as an explicit feature.

The cost savings are substantial. A request whose prompt is 90% shared prefix and 10% new content requires the model to compute full attention only over the new tokens, dramatically reducing time-to-first-token (TTFT). For requests where the system prompt is 2,000 tokens and the user message is 50 tokens, prompt caching can reduce TTFT by 80-90%. The financial savings are equally large because most providers charge for input tokens processed, and cached tokens are processed at a significant discount or not charged at all.

Cache Granularity and Alignment

Not all prompts can benefit equally from prefix caching. The benefit depends on how prompts are structured relative to the caching mechanism.

Prefix caching matches on the exact sequence of token IDs from the beginning of the prompt. If a shared component appears anywhere other than the beginning, it does not help. This is why prompt engineering for production systems often includes a discipline: place stable, reusable content at the front. A typical ordering runs from most stable to least stable:

  • System prompt (instructions, persona, output format)
  • Retrieved documents or context (may vary per query but often repeated across sessions)
  • Few-shot examples
  • The user's actual question

This ordering ensures that as many leading tokens as possible are shared across requests, maximizing the prefix that can be cached. Violating this ordering, for example by placing user-specific metadata at the beginning, can eliminate all caching benefit even when 90% of the prompt content is shared.

The granularity of caching is typically at the cache block level, where each block covers a fixed number of tokens (commonly 16 or 32). A prefix must be long enough to fill at least one complete block to be eligible for caching. Short system prompts of only a few tokens may not save as much as expected because the partial final block must still be recomputed. In practice, this means that very short system prompts gain little from prefix caching even when they are consistently shared.

Minimum Prefix Length

Most providers require a minimum prompt length before activating prompt caching because the overhead of cache lookup and storage for very short prompts exceeds the savings. The cache lookup itself requires hashing the prefix, searching the cache index, and loading stored tensors from memory, all of which take time. For very short prefixes, this overhead may exceed the savings from not recomputing the prefix.

Anthropic's Claude API activates prompt caching at 1,024 tokens for Claude 3 models. OpenAI's context caching for GPT-4o activates at 1,024 tokens as well. Google's Gemini API uses a minimum of 32,768 tokens for its explicit caching feature, targeting use cases with very long shared context such as entire code repositories or long documents. These thresholds reflect each provider's judgment about where the overhead-to-savings ratio becomes favorable.

Understanding minimum lengths matters for system design. A system prompt that is 800 tokens long sits below the caching threshold and will be fully recomputed on every request regardless of how consistently it is shared. Padding to reach the threshold by adding more thorough instructions is sometimes worth doing.

Cache Lifetime and Eviction

Cached KV states are not stored indefinitely. They consume GPU memory, the scarcest resource in LLM deployments. A single cached context for a 2,000-token prompt in a large model can occupy tens of megabytes of GPU memory. Multiply this across thousands of distinct prefixes and the memory pressure becomes significant.

Providers implement time-based eviction: a cached entry is retained for a sliding window (typically 5 minutes to 1 hour) and evicted if no new request uses it within that window. The eviction policy resembles a least-recently-used (LRU) scheme at the time dimension: entries that have not been accessed recently are removed to free memory for new entries. Systems with bursty but infrequent traffic may see poor cache hit rates because the cached state is evicted between bursts.

This eviction behavior has practical implications for how you think about caching economics. A batch job that runs hourly and shares a long system prompt will likely miss the cache on every run unless you either increase the cache lifetime or keep the cache warm with periodic requests. Designing around eviction windows requires knowing your traffic pattern.

Some providers expose cache lifetime controls that allow you to trade memory usage for higher hit rates. A longer cache lifetime means higher hit rates for infrequent-but-repeated workloads at the cost of holding more GPU memory across longer periods. The right setting depends on your traffic burstiness and the ratio of cache storage cost to inference cost.

Exact vs. Fuzzy Prefix Matching

The standard prompt caching described above requires exact prefix matching at the token level. Two prompts share a cached prefix only if their token sequences are literally identical from token 0 through some boundary. A single changed character, even in whitespace, invalidates the match.

This strictness is intentional and mathematically necessary. The KV cache stores tensors that encode exact positional relationships. The key and value for a token at position ii in the attention computation depends on the token's embedding, the positional encoding for position ii, and the learned weight matrices. Reusing a cached KV state for a slightly different prefix would produce incorrect outputs because the attention patterns would be based on wrong positional information.

The practical implication is that prompt templates must be constructed with care. Dynamic content such as timestamps, user IDs, session identifiers, or request-specific metadata should go at the end of the prompt, after all shared content. Even a date injected into the system prompt ("Today's date is March 15, 2025") will prevent caching if it changes between requests, because the token "March" at position 10 will have a different KV representation than "April" at the same position.

Systems that inject per-user personalization into the system prompt must decide whether to sacrifice caching for personalization or restructure their prompts to isolate the personalizing content at the tail. Often, the best solution is to move lightweight personalization information (name, preferences) into the final human turn rather than the system prompt, preserving the shared system prompt prefix.

Measuring Prompt Cache Impact

Prompt caching benefits are visible in two metrics: TTFT and cost. TTFT is measured as the time from request submission to the first token of the model's response. Cost is measured in dollars per request.

For a prompt with PP prefix tokens and NN new tokens, where the prefix is fully cached, the effective input token count for billing is NN rather than P+NP + N. If the cached token discount is dd (e.g., 0.10 for a 90% discount on cached tokens), the cost is:

cost=N⋅cinput+P⋅cinput⋅d\text{cost} = N \cdot c_{\text{input}} + P \cdot c_{\text{input}} \cdot d

where cinputc_{\text{input}} is the per-token input price and dd is the cached token fraction (near 0 for fully cached prefixes).

For TTFT, the savings come from skipping the forward pass over the prefix. The reduction is roughly proportional to the fraction of the prompt that is cached, because the forward pass time scales linearly with the number of tokens processed.

Semantic Caching

Prompt caching operates within the model inference pipeline and is invisible to application code. Semantic caching operates at the application layer and requires explicit design decisions about what to cache, how to measure similarity, and when to return a cached answer versus generating a fresh one.

The core idea is simple: if a user asks "What is the capital of France?" and a later user asks "Which city is France's capital?", these questions have the same intent and the same answer. Why run a full LLM inference for the second request when you already have the answer? Exact-match caching would miss this case because the two strings differ. Semantic caching catches it by comparing meaning rather than text.

This capability matters because natural language is inherently paraphrastic. People express the same idea in hundreds of different ways. Customer support databases receive thousands of slight variations of the same common questions. Knowledge base assistants handle paraphrases of the same factual lookups. Without semantic caching, each variation requires a fresh model call even when the answer is already known.

Embedding-Based Similarity

The standard architecture for semantic caching uses text embeddings to measure query similarity. The intuition is that a well-trained embedding model places semantically similar texts close together in a high-dimensional vector space, regardless of surface-level phrasing differences.

When a new query arrives, you compute its embedding using an embedding model, search a vector store for the nearest cached query embedding, and return the cached response if the similarity exceeds a threshold.

The process works as follows. For each query that generates a fresh response, you store the query embedding and the response together in the cache. On subsequent requests, you compute the embedding for the new query and run a nearest-neighbor search over stored embeddings. If the closest match has a cosine similarity above your threshold, you return the cached response. Otherwise, you run inference and cache the new result.

This architecture separates the matching problem from the storage problem. The vector store handles similarity search efficiently, using approximate nearest-neighbor algorithms that can search millions of vectors in milliseconds. The cache store handles response retrieval, using a key-value store indexed by embedding identifier.

Cosine Similarity

Cosine similarity measures the angle between two vectors in high-dimensional space. For two vectors a\mathbf{a} and b\mathbf{b}, it is defined as:

sim(a,b)=a⋅b∥a∥∥b∥\text{sim}(\mathbf{a}, \mathbf{b}) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \|\mathbf{b}\|}

Two embeddings with cosine similarity near 1.0 represent texts with very similar meaning. Near 0.0 means semantically unrelated. Semantic caching typically uses a threshold between 0.85 and 0.95, depending on how much semantic variation is acceptable for a given use case.

The choice of embedding model matters significantly. You want embeddings that capture semantic meaning precisely enough to distinguish questions with different answers, while being similar enough for paraphrases of the same question. Models like text-embedding-3-small, all-MiniLM-L6-v2, and specialized task-specific embeddings are common choices. The embedding model adds latency to each cache lookup, so small, fast models are preferred for latency-sensitive applications. A model that takes 100 ms to compute an embedding makes semantic caching impractical if your target uncached latency is 150 ms.

Setting the Similarity Threshold

The similarity threshold is the most consequential hyperparameter in semantic caching. It controls the precision-recall tradeoff between cache hits and response quality.

A high threshold (say, 0.95) means only near-identical paraphrases match. You get very few false cache hits where semantically different questions receive the wrong cached answer. But the cache hit rate will be lower because slightly different phrasings do not match.

A low threshold (say, 0.70) means more queries match, increasing the hit rate. But you risk returning incorrect cached answers for queries that seemed similar but had different intent. A question about "the effects of caffeine on sleep" might incorrectly match "the effects of melatonin on sleep" at a low threshold, returning an answer about the wrong substance.

There is no universally correct threshold. The right value depends on the domain, the diversity of your query space, and the cost of a wrong answer. Factual question-answering over a narrow domain can tolerate lower thresholds because similar questions have similar answers. Creative or personalized tasks should use higher thresholds or disable semantic caching entirely.

The most rigorous way to set a threshold is empirical. Collect a labeled dataset of query pairs with human judgments of whether they have the same answer. Compute embedding similarities for all pairs. Find the similarity value that maximizes the F1 score between hit rate (recall) and precision (fraction of hits that are correct). This calibration should be redone whenever the embedding model or the knowledge domain changes significantly.

Cluster-Level vs. Query-Level Caching

Basic semantic caching stores individual queries and their responses. A more structured approach is cluster-level caching, where you identify clusters of semantically similar queries and cache a canonical response for the cluster. New queries are matched to the nearest cluster rather than to individual cached queries.

Cluster-level caching has better space efficiency: instead of storing one response per unique query, you store one response per semantic cluster. It also provides more predictable behavior because the response for each cluster is explicitly defined rather than depending on which query happened to arrive first. If the first query in a cluster was phrased unusually and received an unusually phrased response, subsequent queries in the cluster will receive that unusual response. With cluster-level caching, the canonical response is deliberately crafted or quality-reviewed.

The tradeoff is that cluster-level caching requires upfront work to define clusters and their canonical responses. It works well for structured domains like customer support, where query types are well-defined and the set of clusters is relatively stable over time. It is harder to maintain for open-domain systems where new query types emerge continuously.

A hybrid approach that works in many production systems is to run cluster-level caching for the high-frequency query types that can be identified in advance, and fall back to individual query-level caching for the long tail. The top 50 customer support question types might be cluster-cached with curated answers, while the remaining diversity of queries uses standard semantic caching.

Context-Aware Caching

Semantic caching treats queries as independent. In a conversational system, this assumption breaks down. The meaning and appropriate response to "What do you mean?" depends entirely on the preceding conversation. "Tell me more about that" is meaningless without context. "Is it expensive?" requires knowing what "it" refers to.

A semantic cache that matches these queries to cached responses from different conversations will almost certainly return something irrelevant or misleading. The embedding for "Can you clarify?" is similar in every conversation, but the correct response is completely different in each case.

Context-aware caching addresses this by including conversation history in the cache key. The simplest approach is to embed the last few turns of conversation alongside the current query and match on the concatenated representation. This prevents false matches across different conversations at the cost of a lower cache hit rate, since the combined representation is more unique.

The depth of context to include involves a tradeoff. Including only the last turn is cheap and provides some context disambiguation. Including the last five turns provides much better disambiguation but makes the embedding more expensive to compute and the cached representation less reusable (since the probability of two sessions having identical recent history is very low).

An alternative approach that avoids this tradeoff is to restrict semantic caching to the first turn of a conversation, where there is no context to contaminate the match. The first user message in a conversation is most likely to be a direct, context-free question that matches well across sessions. Subsequent turns within a conversation are handled by exact-match caching keyed on the full conversation history, or by skipping caching entirely for multi-turn context.

Vector Store Infrastructure

At small scale, you can implement semantic caching with an in-memory list of embeddings searched by brute-force cosine similarity. This works for caches with a few thousand entries and single-server deployments.

At production scale, you need a vector database that supports efficient approximate nearest-neighbor search. Options include Pinecone, Weaviate, Qdrant, and Milvus for dedicated vector databases, as well as PostgreSQL with pgvector or Redis with RedisSearch for embedding support in existing infrastructure. The right choice depends on your existing stack, your query volume, and your tolerance for managing additional infrastructure.

The latency of the vector search must fit within your overall latency budget. A vector database query typically completes in 5-20 ms for collections of millions of vectors using optimized index structures like HNSW (Hierarchical Navigable Small World graphs). Adding the embedding computation time (5-30 ms depending on model size), the total cache lookup overhead is typically 10-50 ms. This is well worth the trade for requests that would otherwise take 1-10 seconds for inference.

Caching Architecture in Practice

Understanding individual caching mechanisms is necessary but not sufficient. A production system typically operates multiple caching layers simultaneously, and the interaction between them determines overall performance.

Multi-Layer Cache Design

A well-designed LLM application cache has at least three layers, each targeting a different source of efficiency.

The first layer is exact-match caching at the request level. Before doing anything else, hash the complete request (including the full prompt text, model parameters, and any other request fields) and check a fast key-value store like Redis. If the hash matches a cached response, return it immediately. This handles repeated identical requests at near-zero latency. It is deterministic and works well when users or automated pipelines make repeated identical requests. API clients that retry failed requests, scheduled jobs that re-run with the same inputs, and automated test suites that repeat the same queries all benefit disproportionately from this layer.

The second layer is semantic caching for queries that are equivalent in meaning but not textually identical. This runs only if the exact-match cache misses. It uses the embedding and vector search approach described earlier. The latency is higher than exact matching (typically 5-20 ms for the embedding lookup plus vector search) but still much lower than a full model inference call. The semantic cache catches the paraphrase traffic that the exact-match cache misses.

The third layer is prompt caching at the model inference level. This runs automatically if the LLM provider supports it and requires no application-layer logic. It benefits all requests that share a long system prompt or common context prefix, regardless of whether they hit the application-level cache. Every request to Claude, GPT-4, or Gemini benefits from prompt caching whenever the system prompt is shared and meets the minimum length threshold.

The combination means that the most common requests (exact repeats) are handled instantly, semantically equivalent requests are handled quickly, and all requests benefit from reduced time-to-first-token through prompt caching. Each layer handles what the previous layer cannot.

Request-Level Cache Key Design

The exact-match cache key must uniquely identify a request while being stable for requests that should match. Poor key design is a common source of cache miss inflation.

A naive approach is to hash the entire HTTP request body, including all parameters. This fails if parameters are serialized in a different order on each request, or if non-semantic fields like request IDs or timestamps are included. Two requests that are semantically identical but carry different request IDs will produce different hashes and both miss the cache.

A reliable cache key includes:

  • The model identifier and version
  • The full prompt text (normalized to a canonical form)
  • Generation parameters that affect the output (temperature, max tokens, stop sequences)
  • Any sampling seed if deterministic outputs are required

Fields that should not be included: request timestamps, request IDs, tracing headers, and any session or authentication metadata that does not affect the model's output.

For responses that vary by user (personalized outputs), you must include a user identifier in the cache key. This reduces the hit rate but ensures users do not receive each other's cached responses.

Cache Warming

A cache with no entries provides no benefit. Cache warming is the practice of pre-populating a cache before traffic arrives, using knowledge of what requests are likely.

For exact-match and semantic caches, warming means making representative requests, or simulating them, during deployment so that common queries already have cached responses when real traffic arrives. For prompt caching, warming means sending requests that include the system prompt before the cache is needed, so the KV state is already stored.

Cache warming is particularly important for systems with long cache eviction windows or for deployments that restart frequently. A cache that gets cold between deploys provides inconsistent latency, which is harder to reason about than uniformly slow uncached responses. Users who happen to arrive during a cold cache period experience full inference latency; users who arrive later experience cache-accelerated latency. This inconsistency can make performance analysis difficult.

Effective warming strategies include replaying the previous day's or hour's most common requests during deployment, pre-computing responses for the top-K queries identified from query logs, and using synthetic query generation to cover the most probable queries for a given domain. The cost of warming is paid once per deployment, and the benefit accrues for the lifetime of the cache.

Monitoring and Observability

Caching without monitoring is incomplete. You need to know whether the cache is working as intended, whether quality is being maintained, and whether the cache is cost-effective to operate.

Key metrics to track include:

  • Hit rate broken down by cache layer (exact match, semantic, prompt)
  • Similarity score distribution for semantic cache hits
  • Latency percentiles (p50, p95, p99) for cached and uncached requests
  • Cache storage utilization and eviction rates
  • Cost per request with and without caching

The similarity score distribution is particularly important for semantic caches. If most hits are occurring at similarities of 0.85-0.88 and your threshold is 0.88, you are close to the quality boundary. Drifting the threshold down without reviewing hit quality can silently degrade response accuracy.

User feedback signals, such as thumbs-up/thumbs-down ratings or downstream task performance, should be correlated with cache hit/miss status. If users rate cached responses lower than fresh responses, your threshold is likely too low or your cache domain is too broad.

Cache Invalidation

Cache invalidation is famously described as one of the two hard problems in computer science, alongside naming things. In LLM systems, it is difficult because the relationship between a cached response and its continued validity is complex and often implicit.

Why Invalidation is Hard

A cached response was correct, or at least appropriate, when it was generated. It can become wrong for many reasons, and those reasons are not always visible from the cached response itself.

The underlying facts changed. If you cache the response "The current CEO of Company X is Alice Smith" and Alice Smith is replaced, the cached response is now factually wrong. External knowledge changes continuously and unpredictably. You cannot know in advance which cached responses depend on which facts.

The model changed. LLM providers update models regularly, sometimes changing behavior, safety policies, or output formatting conventions. A response cached under model version A may not be the response you would want from model version B. Model updates can change tone, verbosity, reasoning style, and safety thresholds in ways that make old cached responses look inconsistent with current behavior.

The context changed. A cached response that was appropriate for a certain user, time, or system state may not be appropriate when those conditions change. A system prompt updated to reflect a new product version should invalidate all cached responses that were generated under the old system prompt, because the instructions may now be different.

The user state changed. In personalized systems, a cached response tied to a user's preferences may become stale as the user's preferences evolve. A financial advisor bot that cached recommendations based on a user's risk tolerance should invalidate those recommendations if the user updates their risk profile.

These sources of staleness are fundamentally different from each other and require different invalidation strategies. There is no single mechanism that handles all of them correctly.

Time-Based Invalidation (TTL)

The simplest invalidation strategy is time-to-live (TTL): cached entries expire after a fixed duration. A TTL of 1 hour means no cached response is more than 1 hour old. This is easy to implement, requires no dependency tracking, and provides a clear guarantee to users about maximum response staleness.

The weakness of TTL-based invalidation is its bluntness. Stable information like mathematical definitions, historical facts, or fixed product descriptions could be cached indefinitely without staleness. Fast-changing information like stock prices, weather, or live event data should not be cached at all, or only for seconds. A fixed TTL either over-invalidates stable content (wasting cache capacity and adding unnecessary latency) or under-invalidates dynamic content (serving stale answers).

Tiered TTLs address this by assigning different expiration times to different content types. A system might use a 24-hour TTL for factual knowledge base responses, a 1-hour TTL for product information, and a 5-minute TTL for pricing or availability data. Implementing tiered TTLs requires your caching logic to know something about the content type of each response, either by tagging responses at generation time or by classifying them based on the query type.

The right TTL values are empirical. You should measure how quickly information in your domain changes, then set TTLs to be some fraction of that typical change interval. Setting a TTL based on intuition without measurement often results in either overly conservative (frequent re-computation) or overly aggressive (stale responses) caching behavior.

Event-Based Invalidation

Event-based invalidation ties cache expiration to specific events rather than elapsed time. When the system prompt changes, all cached responses generated under the old system prompt are immediately invalidated. When a knowledge base is updated, all cached responses that might have drawn on the changed content are invalidated.

Event-based invalidation is more precise than TTL and avoids serving stale responses during the window between a change event and the next TTL expiration. If you deploy an updated system prompt at noon and your TTL is 1 hour, you could serve stale responses for up to an hour. With event-based invalidation, the old responses are invalidated at noon immediately.

But it requires the caching system to track dependencies: which cached responses were generated under which system prompt version, using which knowledge base version, and so on. This dependency graph can become complex. A response might have been generated under system prompt version 3, using knowledge base documents 17, 42, and 91. Updating document 42 should invalidate that cached response, but tracking this at response granularity requires recording which documents each response depended on at generation time.

Dependency tracking adds implementation complexity that many teams find difficult to justify. One practical simplification is cache versioning: include a cache version identifier in every cache key. When the system prompt or other shared state changes, increment the version. All previous cache entries automatically become unreachable because their keys no longer match, achieving a complete cache flush for the changed component without needing to track individual dependencies.

Version-key invalidation is essentially a full flush for the affected component. Every entry generated under the old version becomes unreachable at once, and the cache must rebuild from scratch for that component. This is clean and correct, but it sacrifices the warmth of all entries that were not affected by the change. If only 5% of your responses depended on the changed document, version flushing invalidates the other 95% unnecessarily.

Partial Invalidation vs. Full Flush

When a change occurs, you have two options: invalidate only the entries likely affected by the change, or flush the entire cache and rebuild from scratch.

Full flush is simple and guaranteed correct. It is appropriate when the change is pervasive (such as a major model upgrade or a system-wide policy update) or when you cannot reason about which entries are affected. The cost is a cold cache immediately after the flush, with degraded hit rates until the cache warms up again. The warmup period can take hours or days depending on traffic volume, and during this period users experience full inference latency.

Partial invalidation is more efficient but requires careful reasoning about scope. If a specific FAQ article is updated, you might invalidate only the cached responses to questions about that article. This preserves cache warmth for unaffected queries. But if your invalidation logic is wrong and you miss some affected entries, you will serve stale responses for those queries, which is the failure mode you were trying to avoid by caching selectively.

The choice between full flush and partial invalidation should be governed by the expected scope of changes and your tolerance for serving stale responses. For infrequent, broad changes (model upgrades), full flush is safer. For frequent, narrow changes (individual document updates), partial invalidation pays off.

Invalidation in Semantic Caches

Semantic caches present an additional invalidation challenge that does not exist in exact-match caches. In an exact-match cache, you can invalidate a specific entry by its key. In a semantic cache, entries are matched by embedding similarity, not by a deterministic key. You cannot easily invalidate "all responses that might have discussed topic X" without searching the embedding space for entries near the topic's embedding.

This asymmetry creates a subtle problem. Suppose a company updates its product pricing policy. In an exact-match cache, you could search for entries whose keys contain the word "pricing" and remove them. In a semantic cache, you would need to embed a query like "product pricing policy", search for all cache entries with similarity above some threshold, and remove them. But this operation itself requires a similarity search that may be slow and may have recall errors: some relevant entries may not be found if their embeddings are below the search threshold.

One practical approach is to maintain a side index that maps semantic topics or document IDs to the cache entries that might depend on them. At cache store time, you classify the response by topic (automatically, using the LLM itself to tag the response) and add the entry ID to the topic's index. When a topic changes, you use the index to find and invalidate the relevant cache entries. Building and maintaining this index requires careful design but makes targeted invalidation tractable.

A simpler approach is to accept full flush as the invalidation strategy for semantic caches, while investing in fast cache warming to minimize the cold period. If your semantic cache can be warmed to 60% hit rate within 30 minutes of a flush (by replaying recent query logs), the cost of periodic full flushes may be acceptable.

Cache Hit Rates

A cache hit rate is the fraction of requests that were served from cache rather than by running inference. It is the primary metric for evaluating whether a caching strategy is working.

hit_rate=cache_hitscache_hits+cache_misses\text{hit\_rate} = \frac{\text{cache\_hits}}{\text{cache\_hits} + \text{cache\_misses}}

where:

  • cache_hits\text{cache\_hits}: the number of requests served from cache
  • cache_misses\text{cache\_misses}: the number of requests that required fresh inference
  • hit_rate\text{hit\_rate}: a value between 0 and 1 (or 0% to 100%)

A hit rate of 0 means the cache provided no benefit. A hit rate of 1 means every request was served from cache, and the model was never called after the initial warmup. In practice, well-designed caching systems often achieve hit rates of 20-60% for general-purpose chatbots and 70-90% for structured applications with narrow query spaces.

What Drives Hit Rates

Hit rates are driven by the repetitiveness of your workload. Some systems have inherently cacheable traffic:

  • Customer support bots where users repeatedly ask the same common questions
  • Document summarization pipelines where the same documents are processed repeatedly
  • Code generation assistants where the same boilerplate patterns appear frequently
  • Content moderation pipelines where the same, or very similar, content is checked multiple times

Other systems have traffic that is inherently diverse:

  • Open-ended creative writing assistants
  • Systems with strong personalization where every response is user-specific
  • Real-time systems that require fresh information on every request
  • Conversational systems where multi-turn context makes each request unique

Understanding which category your system falls into determines how much effort to invest in caching. Attempting to apply semantic caching to an inherently diverse workload will produce a low hit rate at the cost of cache lookup overhead on every request, a net negative.

The depth of the query space is the key structural variable. A question-answering system with 50 well-defined topics and 1,000 possible questions per topic has a query space of roughly 50,000 questions. A system handling free-form creative requests has an effectively unbounded query space. The first system can achieve high cache hit rates with a moderate-sized cache; the second system cannot, regardless of cache size.

Measuring Hit Rates Correctly

Aggregate hit rate is useful but can be misleading. A high hit rate driven by a small number of extremely common queries may coexist with very low hit rates for the long tail of queries that users care about most. Per-query-type or per-session hit rates give more actionable information.

Consider a customer support system where "reset password" accounts for 40% of all queries, and this query has a 100% cache hit rate. If the remaining 60% of queries cover hundreds of different issues and have 0% hit rate, the aggregate hit rate is 40%, which looks reasonable. But the users who came with complex issues received no caching benefit, and these users may matter disproportionately for satisfaction and cost.

For semantic caches, you should also measure the distribution of similarity scores at which cache hits occurred. A semantic cache with a high hit rate achieved mostly at similarities of 0.70-0.80 may be returning low-quality matches. A healthy distribution should show hits concentrated at high similarity (above 0.90) with the threshold positioned where the quality clearly degrades.

Another important metric is latency reduction. Even a modest hit rate can provide significant overall latency improvement if the cache hits occur on long, expensive inference calls. A 20% hit rate on a system where cached calls complete in 10 ms versus uncached calls in 3 seconds reduces average latency by nearly 70%.

To see this, let hh be the hit rate, lcl_c be the cache lookup latency, and lil_i be the inference latency. The average latency is:

lˉ=h⋅lc+(1−h)⋅li\bar{l} = h \cdot l_c + (1 - h) \cdot l_i

For h=0.20h = 0.20, lc=10 msl_c = 10\text{ ms}, li=3000 msl_i = 3000\text{ ms}:

lˉ=0.20×10+0.80×3000=2+2400=2402 ms\bar{l} = 0.20 \times 10 + 0.80 \times 3000 = 2 + 2400 = 2402\text{ ms}

Without caching, lˉ=3000\bar{l} = 3000 ms. The 20% hit rate delivers a 20% latency reduction in this scenario. Higher inference latency amplifies the benefit: a 10% hit rate on an 8-second inference call saves more absolute time than a 30% hit rate on a 500-ms inference call.

The Cost of Cache Misses with Warming

Hit rate is not the only cost metric. A cache miss followed by caching the result has a higher cost than a cache miss for a one-off query that will never be asked again. For rare queries, caching the response wastes storage and may cause evictions of more frequently used entries.

Selective caching addresses this by only caching responses that are likely to be requested again. You can estimate reuse probability from query frequency distributions, domain knowledge, or by requiring a minimum number of occurrences before caching. Frequency filtering, also called "minimum frequency caching", tracks how many times each query (or semantic cluster) has been requested and only adds it to the cache after the count exceeds a threshold. This avoids polluting the cache with one-off queries that consume capacity needed for high-frequency entries.

Building frequency estimation into your caching layer adds complexity but improves cache efficiency for large, diverse query spaces. The simplest implementation uses a count-min sketch (a probabilistic frequency estimation data structure) to track query frequencies with bounded memory. When the count for a new query crosses the threshold, the query is promoted to the cache.

The cold start problem is the flip side of selective caching. New queries need at least one uncached inference before they can be cached. During high-traffic events or product launches, many new queries arrive simultaneously, all missing the cache and all competing for inference capacity. This can create a feedback loop where high load leads to cache misses, which leads to higher inference load, which leads to even higher latency, reducing the chance that users rephrase and retry with a cacheable query. Warming before high-traffic events is the mitigation.

Zipfian Query Distributions

Natural language query distributions closely follow Zipf's law: if you rank queries by frequency, the kk-th most popular query has frequency proportional to 1/k1/k. The most popular query is asked twice as often as the second most popular, three times as often as the third, and so on.

This power-law distribution has a direct implication for caching: a small cache covers a disproportionately large fraction of traffic. If the top 1% of distinct queries accounts for 50% of all requests (a typical ratio in Zipfian distributions), then a cache large enough to hold 1% of the distinct query space will serve 50% of all traffic. This is why semantic caches deliver value even when they store only hundreds of entries against a universe of millions of possible queries.

The Zipfian property also implies diminishing returns from cache size. The gain from increasing the cache from 100 to 200 entries is much larger than the gain from increasing it from 10,000 to 10,100 entries, because the low-rank entries that the larger cache adds are individually rare. The cache size that maximizes return on investment is usually much smaller than the size that maximizes hit rate.

Code Implementation

Let's implement a semantic caching layer from scratch. We will build the core logic: embedding-based similarity search, cache storage, threshold-controlled retrieval, and hit rate tracking.

In[4]:
Code
# We will use sentence-transformers for embeddings
# uv pip install sentence-transformers
from sentence_transformers import SentenceTransformer

# Load a small, fast embedding model suitable for caching
embedding_model = SentenceTransformer("all-MiniLM-L6-v2")

The all-MiniLM-L6-v2 model produces 384-dimensional embeddings and runs in about 5 ms per query on a CPU. This is fast enough to add minimal latency to the cache lookup path.

In[5]:
Code
@dataclass
class CacheEntry:
    query: str
    response: str
    embedding: np.ndarray
    created_at: float
    hit_count: int = 0


@dataclass
class SemanticCache:
    threshold: float = 0.90
    ttl_seconds: float = 3600.0
    entries: list = field(default_factory=list)
    hits: int = 0
    misses: int = 0

    def _cosine_similarity(self, a: np.ndarray, b: np.ndarray) -> float:
        """Compute cosine similarity between two unit-normalized vectors."""
        return float(
            np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b) + 1e-10)
        )

    def _is_expired(self, entry: CacheEntry) -> bool:
        return (time.time() - entry.created_at) > self.ttl_seconds

    def lookup(self, query: str) -> Optional[str]:
        """Search cache for a semantically similar query. Returns response or None."""
        query_embedding = embedding_model.encode(
            query, normalize_embeddings=True
        )
        now = time.time()

        best_similarity = 0.0
        best_entry = None

        for entry in self.entries:
            if (now - entry.created_at) > self.ttl_seconds:
                continue
            sim = self._cosine_similarity(query_embedding, entry.embedding)
            if sim > best_similarity:
                best_similarity = sim
                best_entry = entry

        if best_entry is not None and best_similarity >= self.threshold:
            best_entry.hit_count += 1
            self.hits += 1
            return best_entry.response
        else:
            self.misses += 1
            return None

    def store(self, query: str, response: str) -> None:
        """Store a new query-response pair in the cache."""
        embedding = embedding_model.encode(query, normalize_embeddings=True)
        entry = CacheEntry(
            query=query,
            response=response,
            embedding=embedding,
            created_at=time.time(),
        )
        self.entries.append(entry)

    @property
    def hit_rate(self) -> float:
        total = self.hits + self.misses
        return self.hits / total if total > 0 else 0.0

The SemanticCache class encapsulates the full lookup-store-metric cycle. The lookup method encodes the incoming query, scans all unexpired entries for the best cosine similarity match, and returns a cached response only if the similarity exceeds the threshold. The store method encodes a query and adds it to the entry list. The hit_rate property computes the running hit fraction across all requests seen so far.

In[6]:
Code
# Simulated LLM responses for demonstration
def mock_llm(query: str) -> str:
    """Simulate an LLM call with a small delay."""
    time.sleep(0.05)  # Simulate 50ms inference time
    responses = {
        "What is the capital of France?": "The capital of France is Paris.",
        "What is the capital of Germany?": "The capital of Germany is Berlin.",
        "Explain neural networks briefly.": "Neural networks are computational models inspired by the brain, consisting of layers of interconnected nodes that learn patterns from data.",
        "What are transformers in NLP?": "Transformers are a neural network architecture that uses self-attention to process sequences in parallel, forming the basis of models like BERT and GPT.",
    }
    for key, val in responses.items():
        if any(word in query.lower() for word in key.lower().split()):
            return val
    return f"Response to: {query}"


# Populate the cache with some initial queries
cache = SemanticCache(threshold=0.88, ttl_seconds=3600)

initial_queries = [
    (
        "What is the capital of France?",
        mock_llm("What is the capital of France?"),
    ),
    (
        "Explain neural networks briefly.",
        mock_llm("Explain neural networks briefly."),
    ),
    (
        "What are transformers in NLP?",
        mock_llm("What are transformers in NLP?"),
    ),
]

for query, response in initial_queries:
    cache.store(query, response)
In[7]:
Code
# Test with queries at varying semantic distances from cached content
test_queries = [
    "What is the capital of France?",  # Exact match
    "Which city is the capital of France?",  # Paraphrase (should hit)
    "Where is France's capital located?",  # Rephrased (should hit)
    "Tell me about French neural networks.",  # Mixed topics (may hit or miss)
    "What is the capital of Japan?",  # Different country (should miss)
    "How do transformers work in deep learning?",  # Paraphrase (should hit)
]

results = []
for query in test_queries:
    start = time.time()
    cached_response = cache.lookup(query)
    latency_ms = (time.time() - start) * 1000

    if cached_response:
        source = "CACHE HIT"
        response = cached_response
    else:
        source = "CACHE MISS"
        response = mock_llm(query)
        cache.store(query, response)

    results.append(
        {
            "query": query,
            "source": source,
            "latency_ms": latency_ms,
        }
    )
Out[8]:
Console
Query                                              Result          Latency
---------------------------------------------------------------------------
What is the capital of France?                     CACHE HIT         7.9ms
Which city is the capital of France?               CACHE HIT        19.9ms
Where is France's capital located?                 CACHE HIT         7.0ms
Tell me about French neural networks.              CACHE MISS        6.9ms
What is the capital of Japan?                      CACHE MISS        7.8ms
How do transformers work in deep learning?         CACHE MISS        7.8ms

Overall hit rate: 50.0%
Total hits: 3  |  Total misses: 3
Out[9]:
Console

Cache entry usage:
Original Query                                            Hits
-----------------------------------------------------------------
What is the capital of France?                               3
Explain neural networks briefly.                             0
What are transformers in NLP?                                0
Tell me about French neural networks.                        0
What is the capital of Japan?                                0
How do transformers work in deep learning?                   0

The output shows the core behavior: exact and near-paraphrase queries return cached responses with millisecond latency, while queries about different topics miss the cache and trigger the (simulated) inference call. The hit count per entry reveals which cache entries are doing the most work, a useful metric for understanding whether your seed queries were representative of actual traffic.

Now let's examine how the hit rate changes as a function of the similarity threshold:

In[10]:
Code
# Simulate how threshold affects hit rate vs. response quality

# A set of base queries and their known paraphrases
query_groups = [
    {
        "canonical": "What is the capital of France?",
        "variants": [
            ("What is the capital of France?", 1.00),  # Exact
            ("Which city is France's capital?", 0.93),  # Near paraphrase
            ("France capital city name?", 0.87),  # Compressed form
            ("Where is France's government located?", 0.75),  # Related concept
            ("Name the French capital city.", 0.91),  # Rephrased
        ],
    },
    {
        "canonical": "Explain transformer architecture.",
        "variants": [
            ("Explain transformer architecture.", 1.00),
            ("How does the transformer model work?", 0.92),
            ("What is a transformer neural network?", 0.88),
            ("Describe the attention mechanism.", 0.72),
            ("What are key components of transformers?", 0.85),
        ],
    },
]

thresholds = np.round(np.arange(0.60, 1.01, 0.02), 2)
hit_rates = []
correct_match_rates = []

for threshold in thresholds:
    hits = 0
    correct = 0
    total = 0
    for group in query_groups:
        for variant_query, true_similarity in group["variants"]:
            total += 1
            if true_similarity >= threshold:
                hits += 1
                # A "correct" match is one where the similarity is genuinely high
                if true_similarity >= 0.85:
                    correct += 1
    hit_rates.append(hits / total)
    correct_match_rates.append(correct / total)
Out[11]:
Visualization
Line chart showing hit rate and correct-match rate versus similarity threshold, with a reference line at 0.88 and the curves first diverging at the plotted 0.74 threshold.
Cache hit rate and correct-match rate across similarity thresholds in the ten-query toy sample. Lowering the threshold from 1.0 admits more variants. At the marked 0.88 threshold, every admitted variant meets the sample's 0.85 correctness criterion. The plotted curves first diverge at 0.74, which admits a related query with 0.75 similarity as a false cache hit.

The plot shows the fundamental tradeoff. In this small, discrete sample, a threshold of 0.88 admits 60% of the variants and every admitted variant satisfies the chosen correctness criterion. Lowering the threshold to 0.84 admits all eight high-similarity variants without adding a false match. At the next relevant plotted setting, 0.74, the hit-rate curve moves above the correct-match curve because the 0.75-similarity query is now admitted. A production threshold still requires a much larger labeled validation set, but the point where these curves separate is the quantity to watch.

Now let's visualize the relationship between cache size, hit rate, and the cumulative latency savings for a realistic workload:

In[12]:
Code
# Simulate a request workload with a Zipfian query distribution
# (a small number of popular queries, a long tail of rare ones)

n_unique_queries = 500
n_total_requests = 5000
inference_latency_ms = 800
cache_lookup_ms = 15

# Zipf distribution: query i has probability proportional to 1/i
ranks = np.arange(1, n_unique_queries + 1)
zipf_probs = 1.0 / ranks
zipf_probs /= zipf_probs.sum()

# Sample request sequence
query_ids = np.random.choice(
    n_unique_queries, size=n_total_requests, p=zipf_probs
)

# Simulate caching with various cache sizes (max unique entries before eviction)
cache_sizes = [10, 25, 50, 100, 200, 500]
simulated_hit_rates = []
simulated_avg_latencies = []

for max_size in cache_sizes:
    cache_store = {}  # Maps query_id -> True (just track presence)
    cache_order = []  # LRU eviction order
    hits = 0

    for qid in query_ids:
        if qid in cache_store:
            hits += 1
            # Move to end (most recently used)
            cache_order.remove(qid)
            cache_order.append(qid)
        else:
            # Cache miss: add to cache
            if len(cache_order) >= max_size:
                # Evict least recently used
                evicted = cache_order.pop(0)
                del cache_store[evicted]
            cache_store[qid] = True
            cache_order.append(qid)

    hit_rate_sim = hits / n_total_requests
    avg_latency = hit_rate_sim * cache_lookup_ms + (1 - hit_rate_sim) * (
        inference_latency_ms + cache_lookup_ms
    )

    simulated_hit_rates.append(hit_rate_sim)
    simulated_avg_latencies.append(avg_latency)
Out[13]:
Visualization
Line chart of cache hit rate vs max cache size, showing rapid early rise then plateau
Cache hit rate as a function of maximum cache size for a simulated Zipfian query workload. Hit rate rises sharply through the first 50-100 entries, then continues with diminishing gains as the cache covers progressively rarer queries.
Line chart of average latency vs max cache size, showing rapid early fall then plateau
Average request latency as a function of cache size for the same workload. The dashed 800 ms line is the uncached inference baseline. Latency drops fastest in the early cache-size range and then approaches the 15 ms lookup cost more gradually, mirroring the diminishing hit-rate gains.
Out[14]:
Console
Cache Size | Hit Rate | Avg Latency | Latency Reduction
------------------------------------------------------------
        10 |   24.9% |        615ms |             23.1%
        25 |   40.6% |        490ms |             38.7%
        50 |   52.4% |        395ms |             50.6%
       100 |   67.1% |        278ms |             65.2%
       200 |   79.5% |        179ms |             77.6%
       500 |   90.8% |         88ms |             88.9%

The pattern in the output reflects a property of real language workloads: a Zipfian distribution means a small cache covers a disproportionately large fraction of traffic. The first 50 entries in the cache often account for 40-50% of all requests. This is why semantic caches deliver value even when they store only hundreds of entries against a universe of millions of possible queries. The diminishing returns visible in the hit rate curve confirm that there is a natural stopping point for cache investment: beyond a certain size, additional entries cover only very rare queries that contribute little to overall performance.

Key Parameters

The key parameters for the semantic caching system are:

  • threshold: The minimum cosine similarity score required for a cache hit. Higher values (0.90-0.95) give precision; lower values (0.75-0.85) give recall. Domain-specific empirical tuning is necessary.
  • ttl_seconds: How long a cached entry remains valid before expiration. Set based on how quickly the underlying information changes, not on system load.
  • max_size: The maximum number of entries the cache stores before LRU eviction begins. Tuned against the Zipfian hit-rate curve for your workload.
  • embedding_model: The model used to encode queries into vectors. Faster models reduce cache lookup latency but may produce lower-quality similarity estimates.

Limitations and Practical Implications

Caching is powerful but limited by the nature of LLM applications. The first and most persistent limitation is that not all tasks are cacheable. Creative generation, personalized content, and queries requiring fresh real-world information resist caching because no two requests have the same correct answer. Applying semantic caching to these tasks will produce responses that seem plausible but are contextually wrong, which can be more harmful than simply running inference. Before instrumenting a caching layer, you should audit your traffic to understand what fraction of requests are cacheable.

Semantic caching introduces a subtle failure mode: the embedding model and the generation model may not agree on what "similar" means. Two queries can be close in embedding space but require different responses. For example, "What is the boiling point of water at sea level?" and "What is the boiling point of water at high altitude?" are semantically similar (both about water's boiling point) but have different answers (100 degrees Celsius vs. roughly 90 degrees Celsius). An embedding model trained for general semantic similarity may assign these queries a high cosine similarity, causing the cache to return the wrong answer. Domain-specific embedding models and higher thresholds reduce but do not eliminate this risk. For safety-critical or high-stakes applications, you should validate the cache against a labeled dataset of known-similar and known-distinct query pairs before deploying it.

Prompt caching has a different set of limitations. It requires that the shared prefix be truly identical at the token level. Systems that inject dynamic content into system prompts, such as timestamps, request IDs, or per-user personalization strings, lose prompt caching benefits unless the dynamic content is carefully isolated at the end of the prompt. The discipline of ordering prompt components from most stable to least stable is not always compatible with other design goals, particularly in highly personalized systems where the personalization is conceptually part of the "setup" that comes before the user's question.

Cache invalidation remains an open engineering problem. The strategies described in this chapter, TTL, event-based invalidation, and version keys, all involve tradeoffs between freshness, complexity, and performance. Systems that serve frequently updated information must choose between serving stale cached responses and paying the full cost of inference on every request. There is no free lunch: the only way to guarantee freshness is to not cache, which means giving up the performance and cost benefits entirely. Good production systems typically accept a defined maximum staleness, communicate it to users where relevant, and design their cache TTLs to keep staleness within that bound.

A practical concern that often goes unaddressed in initial caching implementations is the effect of caching on model output diversity. Many LLM applications benefit from some variability in responses: slightly different wordings feel more natural to users and reduce the sensation of talking to a machine. Aggressive caching reduces this variability. If the same response is returned for every paraphrase of "what is the weather like?", conversations start to feel repetitive. For applications where conversational naturalness matters, you can introduce a randomization layer that occasionally bypasses the cache even for high-similarity matches, preserving some response diversity at the cost of a slightly lower effective hit rate.

Scalability is a final concern. An in-process Python cache of a few thousand entries (as implemented above) works for low-traffic prototypes. A production system with millions of cached entries distributed across many replicas requires a different architecture: a centralized vector database, a distributed key-value store for responses, and careful cache key partitioning to avoid hot spots. The embedding computation itself scales to dedicated inference endpoints if the volume requires it. As caching infrastructure becomes more complex, the operational overhead of maintaining it must be weighed against the cost savings it delivers.

Despite these limitations, caching consistently delivers some of the best return on investment of any latency optimization in LLM systems. In systems where even 10-20% of traffic can be served from cache, the cost and latency reductions are significant enough to change the economics of deployment. The engineering effort required to implement a basic exact-match or semantic cache is modest compared to the gains, making caching one of the first optimizations to add when moving from prototype to production.

Summary

Caching in LLM production systems operates at multiple levels, each targeting a different source of repeated computation.

Prompt caching works at the model level by reusing pre-computed KV tensors for shared prompt prefixes. It is transparent to application code and can reduce time-to-first-token by 80-90% for requests with long shared system prompts. The key requirement is placing stable content at the beginning of the prompt and keeping it exactly consistent across requests. Providers enforce minimum prefix lengths (typically 1,024 tokens) and use time-based eviction with windows of 5 minutes to 1 hour.

Semantic caching works at the application level by matching queries on semantic similarity rather than exact text equality. It uses text embeddings and vector similarity search to find cached responses for queries that are paraphrases of previously answered questions. The similarity threshold is the critical hyperparameter, balancing hit rate against the risk of returning incorrect cached answers. Context-aware caching handles the multi-turn case by including conversation history in the matching key.

Cache invalidation is the hardest part of caching. Time-based TTL is simple but blunt. Event-based invalidation is precise but requires dependency tracking. Version keys provide a clean mechanism for flushing stale entries when shared state changes. Semantic caches require special handling for invalidation because entries cannot be identified by deterministic keys, often requiring side indices or full-flush strategies.

Cache hit rates are the primary metric for evaluating caching effectiveness. Workloads following a Zipfian distribution, as most natural language workloads do, deliver disproportionate hit rates from small caches. Hit rate, combined with latency reduction and cost savings, provides the full picture of caching value. Aggregate hit rate alone can be misleading; per-query-type and similarity-score distributions give a sharper view of cache health.

The practical approach for production systems is to layer these mechanisms: exact-match caching for identical requests, semantic caching for paraphrases, and prompt caching at the model level for all requests. Together, these layers cover the most common sources of redundant computation and deliver meaningful improvements to cost, latency, and throughput. Monitoring, threshold calibration, and cache warming are the ongoing operational work that keeps each layer performing well as traffic patterns and content evolve.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about caching in LLM production systems.

LLM Caching Quiz

Question 1 of 100 of 10 completed
What does prompt caching reuse across separate API requests?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026cachingllms, author = {Michael Brenndoerfer}, title = {Caching for LLMs: Prompt, Semantic, and Invalidation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Caching for LLMs: Prompt, Semantic, and Invalidation. Retrieved from https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm
MLAAcademic
Michael Brenndoerfer. "Caching for LLMs: Prompt, Semantic, and Invalidation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm>.
CHICAGOAcademic
Michael Brenndoerfer. "Caching for LLMs: Prompt, Semantic, and Invalidation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Caching for LLMs: Prompt, Semantic, and Invalidation'. Available at: https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Caching for LLMs: Prompt, Semantic, and Invalidation. https://mbrenndoerfer.com/writing/caching-prompt-semantic-invalidation-hit-rates-llm

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.