Part of Language AI Handbook
Explains why the first tokens in transformer sequences absorb excess attention weight, how this causes streaming inference failures.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Attention Sinks
When you deploy a language model for streaming generation, something strange happens. After processing a few thousand tokens, the model's output quality degrades. It starts repeating itself, loses coherence, and eventually produces gibberish. This failure mode isn't about running out of memory or hitting a hard context limit. It's about a subtle architectural quirk: the first few tokens in a sequence absorb a disproportionate amount of attention weight, regardless of their semantic relevance.
These overloaded positions are called attention sinks. Understanding why they exist and how to work around them makes streaming inference practical: models can generate coherent text indefinitely without quality degradation.
To appreciate why this quirk matters, consider how self-attention works. Every token in a transformer attends to every previous token, and the resulting attention weights must sum to exactly 1 due to the softmax normalization. When a model processes a 4,000-token document and generates the 4,001st token, it computes a probability distribution over all 4,000 prior positions. In theory, most of that weight should flow to semantically relevant positions. In practice, a substantial fraction always ends up on the very first token, even when that token is something as inert as a formatting character or a system prompt header. The model has no choice: softmax cannot assign zero weight to any position, so it routes "spare" probability mass to positions it has learned are safe to absorb it.
Think of attention sinks as pressure-release valves in the attention mechanism. When a hydraulic system has excess pressure, it routes the surplus through a designated low-resistance path rather than letting pressure build unpredictably. Attention sinks play the same role. They give the model a predictable place to deposit attention weight that has nowhere useful to go, keeping the remaining distribution clean and semantically meaningful. Remove the valve and the pressure distributes chaotically through paths that weren't designed to handle it.
This phenomenon was first systematically documented by Xiao et al. in their 2023 paper introducing StreamingLLM. Their key observation was deceptively simple: preserving just four initial tokens alongside a sliding window of recent tokens allowed language models to generate coherent text across sequences of arbitrary length. The memory footprint stays constant. The generation quality stays stable. The trick works because it preserves exactly the tokens the model's attention mechanism needs, even when those tokens carry no semantic content that would survive a reading comprehension test.
Understanding attention sinks requires working through three distinct questions. First, why does the attention mechanism develop these sink positions in the first place? Second, what goes wrong when you naively remove them during streaming? Third, how does StreamingLLM's cache design provide a principled solution that costs essentially nothing in memory? This chapter answers all three, building from the mathematics of softmax through a complete implementation of a streaming inference cache.
The Discovery: Why First Tokens Hoard Attention
The attention sink phenomenon was first systematically analyzed in the StreamingLLM paper by Xiao et al. (2023). Researchers observed that in autoregressive language models, the very first tokens in a sequence consistently receive high attention weights from all later positions, even when those initial tokens carry no special semantic meaning.
An attention sink is a token position that accumulates disproportionately high attention weights across many query positions, typically regardless of the token's actual content or relevance. In autoregressive transformers, the first few tokens often serve as attention sinks due to softmax normalization requirements.
Consider what happens when a model processes a long sequence. For each new token, the self-attention mechanism computes attention weights over all previous positions. These weights must sum to 1 due to the softmax normalization. When the model encounters positions that aren't particularly relevant to the current prediction, it still needs to distribute some probability mass. The first tokens become a convenient "dump" for this excess attention.
This behavior emerges from training dynamics, not explicit design. During pretraining, models learn that attending to initial tokens is "safe" because they're always present and their representations stabilize early. The first token position essentially becomes a learned bias term that absorbs attention weight that would otherwise need to be spread across irrelevant positions.
Notice that the content of those first tokens is largely irrelevant to the sink function. The Xiao et al. study found that adding a dedicated newline character, a period, or even a blank token at position 0 produced an effective sink, because the model's key vectors at that position adapt during training to maximize their "catchall" property rather than to represent any particular linguistic meaning. The sink is a role, not a token type. The first position gets pressed into service because it is the position with the longest and most consistent training signal: every sequence in the pretraining corpus starts at position 0, so the key vector at position 0 accumulates gradient updates from more contexts than any other position.
The phenomenon extends across the model. Attention sinks are not confined to a single transformer layer or a single attention head. Visualizing attention maps across all layers shows the same pattern repeating: the first few positions receive high attention weights at virtually every layer and in the majority of heads. This cross-layer consistency indicates that attention sinks are woven into the model's overall information processing strategy, not a surface-level artifact of the final prediction head. The sink pattern stabilizes the hidden representations that flow between layers, because downstream layers receive slightly diluted but predictable signals from a stable upstream attractor.
The observation that transformers concentrate attention on early positions predates StreamingLLM. Researchers studying model interpretability noted "attention head specialization" as far back as 2019, with some heads appearing to copy information and others appearing to aggregate over long spans. What was missing was the practical insight: instead of viewing sinks as an interpretability curiosity, Xiao et al. asked whether they could be exploited for memory-efficient streaming. That reframing, from quirk to feature, produced one of the simplest and most practically impactful LLM inference optimizations published in 2023.

The visualization shows the characteristic shape: a sharp spike at the beginning of the sequence (the sink tokens), low uniform attention across the middle positions, and slightly higher attention weights on recent tokens that provide immediate context.
To understand just how much attention concentrates in the first few positions, let's look at the cumulative attention distribution. By summing the attention weights from left to right, we can see what fraction of total attention is captured by the first tokens.

The cumulative view highlights the sink phenomenon. The first 4 tokens alone capture about 29% of all attention weight, despite representing only 4% of the sequence positions. This concentration explains why removing these tokens causes such dramatic failures: you're removing the positions that the model relies on most heavily.
Why Removing Initial Tokens Breaks Generation
This puzzle shows the importance of attention sinks. Suppose you want to do streaming inference: process a very long document one chunk at a time, discarding old tokens to save memory. A natural approach is to keep a sliding window of the most recent tokens.
But this fails catastrophically. When you remove the first tokens, the model loses its attention sinks. The excess attention weight that previously flowed to position 0 must now go somewhere else. The model wasn't trained with this attention distribution, so it produces outputs that don't match any pattern it learned during pretraining.
The key insight is that the model's behavior at inference time is always conditioned on the distribution of attention it saw during training. Every attention head learned its weights, every layer learned its transformations, under the assumption that a certain fraction of attention would flow to the sink positions. When you remove those positions, you are not simply asking the model to "forget" those tokens. You are scrambling the normalization constant inside every softmax operation. What looked like 30% of attention going to semantically useful tokens now becomes the model's entire distribution, normalized over a strictly smaller set, producing inflated weights on positions that were never meant to receive them.
In practice, the failure mode follows a characteristic arc. For the first few tokens after the sinks are removed, the model may still produce plausible output, because the recent context window provides enough signal to overcome the distribution mismatch. But as more tokens are generated without sinks, the accumulated error in attention distribution corrupts the hidden states layer by layer. The model begins generating repetitive phrases, then grammatical non-sentences, then complete gibberish. This progression is not random. It reflects the model searching for new "safe" positions to serve as sinks, failing to find candidates with the right key-vector geometry, and progressively destabilizing its internal representations.
The failure is also not recoverable by longer windows. Even if you give the model a window of 10,000 recent tokens but remove the first few, quality degrades. The window size addresses a different problem (how much local context the model can use) than the sink problem (where the model routes excess attention weight). You need to solve both simultaneously to achieve stable streaming generation.


The contrast is clear. With sink tokens present (left), attention follows the pattern the model learned during training: high weight on the initial positions, with the rest distributed sensibly across recent context. Without sink tokens (right), the attention distribution becomes erratic. The model tries to find new sinks among the remaining tokens, but these positions weren't trained to serve that role.
In practice, this manifests as:
- Increased perplexity on downstream tokens
- Repetitive or looping generation
- Loss of coherence over long generations
- Complete degeneration into nonsense after enough tokens
The StreamingLLM Solution
StreamingLLM proposes a simple fix: instead of a pure sliding window, always keep the first few tokens. The elegance of this solution is that it requires no retraining, no changes to model weights, and no modifications to the attention computation itself. It is purely a cache management strategy. You change which tokens you store and present to the attention mechanism, and the model's existing behavior takes care of the rest.
The observation that leads to the fix is worth stating explicitly. The problem with naive sliding windows is not that old tokens are discarded. It is that the specific tokens at positions 0 through 3 are discarded. Tokens at positions 50, 100, or 500 can be discarded without meaningful harm, because the model's attention mechanism was not specifically trained to rely on their presence at a particular position. The sink tokens at the very beginning, however, have developed key-vector representations that act as attractors for excess attention weight across the entire model. They play a structural role in keeping attention distributions normalized and predictable.
The attention mechanism context becomes:
By preserving just 1 to 4 initial tokens (the attention sinks), plus a sliding window of recent tokens, you maintain the attention distribution the model expects while bounding memory usage. The generation can continue indefinitely without quality degradation.
In practice, this means the key-value cache that a transformer maintains during generation has two regions instead of one. The first region is a small fixed buffer holding the key and value vectors of the sink tokens. This buffer never changes once the initial tokens are processed. The second region is a circular buffer holding the most recent key-value pairs, which shifts by one position each time a new token is generated. The total cache size is bounded by regardless of how many tokens have been generated in total, where is the number of sink tokens.
Think of it like a news anchor keeping permanent notes about the show's topic and guest background (the sink tokens) alongside a rolling transcript of the last few minutes of conversation (the sliding window). The notes are never thrown away because they anchor the context. The transcript slides forward as new dialogue arrives. Together they give the anchor enough to respond coherently to the current moment without needing to store every word ever spoken on the show.

The key insight is that you don't need the entire history. You need:
- The attention sinks to absorb excess attention weight
- Recent context for actual language modeling
Everything in between can be safely discarded without affecting generation quality.
Mathematical Analysis
To understand why attention sinks emerge and why they're essential for streaming inference, we need to examine the mathematics of attention itself. This walkthrough from the basic attention formula to the StreamingLLM solution will show a core tension in transformer design: the softmax function that makes attention work also creates the problem that sinks must solve.
The analysis proceeds in three stages. First, we examine the standard attention formula to identify exactly where the softmax constraint forces excess attention to accumulate somewhere. Second, we trace the training dynamics that direct that excess toward the first positions. Third, we follow the mathematics through a sink removal event to see precisely why and how the model breaks when sinks are absent. Each stage builds directly on the previous one, so by the end you will have a complete mechanistic account of the phenomenon from first principles.
One thing to keep in mind as you read through the formulas: the key tension is not unique to transformers. Any mechanism that must produce a normalized probability distribution over a set of candidates faces this pressure. The attention mechanism simply makes the problem particularly visible because the number of candidates (sequence positions) can be thousands deep, making the "spare" probability mass substantial even when the model is making perfectly good predictions about the relevant positions.
The Attention Mechanism: A Probability Distribution Problem
Self-attention is a weighted averaging mechanism. For each position in a sequence, the model asks: "Which previous positions should I draw information from, and how much from each?" The answer comes as a set of weights that sum to 1, forming a probability distribution over the context.
In a standard autoregressive transformer, the attention computation at position produces:
where:
- : the query vector at position , representing what information this position is looking for
- : the key matrix containing key vectors from positions 1 through , where each key represents what information that position offers
- : the value matrix containing value vectors from positions 1 through , holding the actual content to be aggregated
- : the dimension of the key vectors, used to scale the dot products and prevent them from growing too large
The softmax function is the critical piece. It takes the raw attention scores (dot products between queries and keys) and converts them into a probability distribution:
where each individual weight is:
where:
- : the attention weight from query position to key position , indicating how much position attends to position
- : the dot product between query and key , measuring their compatibility
- : the exponential function, which ensures all values are positive
- The denominator sums over all positions, normalizing the weights to form a valid probability distribution
This is where the tension arises. The exponential function has a mathematical property that creates the sink phenomenon: for all finite . No matter how irrelevant a position is, the model cannot assign it exactly zero attention. The softmax constraint forces the model to distribute some probability mass everywhere.
When the current token needs information from only a handful of positions, what happens to the attention weight that must go to the other 995 tokens in a 1000-token sequence? The model needs somewhere to put this "excess" attention, somewhere that won't disrupt the computation. During training, the first tokens become that destination.
How Sink Tokens Emerge During Training
The emergence of sink tokens is not designed but learned. During pretraining on billions of tokens, the model discovers that attending to early positions is "safe" for several reasons:
- They're always present: Position 0 exists in every training example, so the model can reliably use it
- Their representations stabilize early: By the time later positions attend to them, their hidden states have passed through many layers
- Attending to them causes minimal harm: Mixing in a small amount of irrelevant but stable information is better than spreading that attention across many varying positions
Let denote the key vector at position 0 after training. This vector develops a distinctive property: for most query vectors , the dot product is relatively high compared to random positions. The model has learned to make a universal "attractor" in key space.
To see why this specialization emerges, consider the gradient dynamics during training. Whenever the model generates a token that benefits from "dumping" excess attention, the gradient updates reinforce key-query alignment at position 0. Because this pattern occurs consistently across diverse training examples (position 0 is always present; the content at other positions varies wildly), the key vector experiences a stronger and more consistent gradient pressure than any other position. Over millions of training steps, this consistently pushes toward a direction in key space that maximizes dot products with a wide variety of queries.
Notice that this is a self-reinforcing dynamic. Once starts absorbing excess attention during training, the gradient updates make it even better at absorbing excess attention, which causes it to absorb even more, and so on. The model converges to a configuration where position 0 is a reliable sink precisely because the training process rewards that configuration. It is an emergent coordination mechanism: there is no explicit objective that says "make position 0 a sink," but the model discovers that arrangement independently as the globally stable solution to the softmax pressure problem.
The same mechanism applies, to a lesser extent, to positions 1 and 2. Each subsequent early position sees a consistent gradient pressure from all later positions, diminishing as we move further from the start. This explains the characteristic shape of the attention spike: highest at position 0, decaying quickly across the next few positions, then dropping to near-uniform levels for middle positions.

The heatmap shows that sink behavior is not an artifact of any single layer. Across all transformer layers, the first few positions consistently attract high attention weights. This cross-layer consistency suggests that sink tokens are part of how the model processes information, rather than a quirk of one particular layer's learned weights.
We can write the attention weight on position 0 using simplified notation:
where:
- : the attention weight from query position to position 0 (the first token)
- : the scaled attention score between query and key
- : the numerator, which grows exponentially with how well query matches key 0
- : the normalization constant summing over all previous positions
For to be consistently high across different queries and contexts, must be consistently large. The key vector evolves during training to project strongly onto the directions that queries typically occupy, regardless of the actual content at position 0.
The Mathematics of Failure: What Happens When Sinks Disappear
Understanding why removal of sink tokens breaks the model requires following the math carefully. When you remove position 0 from the key-value cache, the attention computation changes in a specific and damaging way.
The new attention weights become:
where:
- : the modified attention weight after removing position 0
- The denominator now sums only over positions , excluding the removed sink
The critical change is in the denominator. By removing from the sum, we've made the denominator smaller. Let's trace through what this means:
- Numerators stay the same: For any remaining position , its numerator hasn't changed
- Denominator shrinks: The sum that divides the numerator is now missing the term
- All weights increase: Since numerator/smaller-denominator > numerator/larger-denominator, every remaining position receives more attention
The probability mass that previously went to position 0 must redistribute across the remaining positions. If the sink absorbed 30% of attention (), that 30% now spreads across positions that weren't trained to receive it.
This redistribution causes two compounding problems:
-
Distribution mismatch: The new attention distribution doesn't match what the model saw during training. Each attention head has learned to produce useful representations given a specific expected distribution of weights. When that distribution changes, the representations become unreliable, leading to out-of-distribution hidden states that propagate through subsequent layers.
-
Cascade of new sinks: The model may try to use other early positions as sinks, but they lack the specialized key representation of position 0. As generations continue, the model keeps shifting which positions absorb excess attention, never stabilizing into a pattern it can use effectively.
The StreamingLLM Solution: Preserving the Distribution
StreamingLLM's insight is that you don't need the full history to maintain the attention distribution the model expects. You need exactly two things: the sink tokens to absorb excess attention, and recent context for actual language modeling. Everything in between can be discarded.
Formally, we maintain positions as sinks plus a sliding window for recent context. The attention weight for any position in this combined set becomes:
where:
- : the attention weight from query position to position
- : the attention score between query and key
- : the set of preserved sink token positions
- : the sliding window of the most recent positions
- The denominator sums over both sets. This makes the attention weights form a valid probability distribution
The key insight is that the denominator structure matches what the model saw during training. The sink tokens contribute their terms to the normalization, absorbing their usual share of attention weight. The window tokens receive the context-relevant attention for actual language modeling. Because the sink tokens are present, the model operates in its training distribution rather than an out-of-distribution state.
This is why StreamingLLM works: it doesn't try to eliminate the sink phenomenon or work around it. Instead, it embraces the sink as a necessary component of how the model has learned to distribute attention, and simply ensures that component is always present.
Worked Example: Tracing One Eviction Step
A concrete walkthrough of a single cache eviction step makes the mechanics tangible. Suppose you are running a streaming inference session with 4 sink tokens and a window of size 8, giving a maximum cache size of 12. You have already processed 12 tokens, filling the cache. The cache currently holds positions 0, 1, 2, 3 (sinks) and positions 4 through 11 (the window). Now a 13th token arrives.
The cache is full, so something must be evicted. The eviction rule is: preserve all sink tokens, and drop the oldest window token. Position 4 is the oldest window token, so it gets evicted. After eviction, the cache holds positions 0, 1, 2, 3 (sinks) and positions 5 through 12 (the updated window). The cache is still 12 entries, and position 4 is gone permanently.
Here is what happens to the attention distribution during this step. Before eviction, the query for token 13 would compute attention scores against cache positions {0, 1, 2, 3, 4, 5, ..., 11}. The denominator in the softmax sums the exponentials of all 12 scores. After eviction, the query computes attention against {0, 1, 2, 3, 5, 6, ..., 12}: the sink tokens are present, and the new token (12) has been appended, but position 4 is absent. The denominator changes, but the sink tokens' contribution to it remains. The model behaves just as it was trained to: the sink positions absorb excess attention, the recent window provides local context, and the new token's generation is stable.
Now consider what would have happened with a pure sliding window (no sinks). The cache before eviction holds positions 2 through 13 (12 tokens). After eviction, it holds positions 3 through 13, and the positions 0 and 1 that once served as sinks are long gone. The attention denominator no longer includes any of the specialized sink positions. The probability mass that those positions used to absorb now pushes up attention weights on positions 3, 4, and 5, which were never trained to handle it. Hidden states computed from this distorted distribution propagate forward and corrupt the next computation. By token 100 of streaming, the error has compounded across many generations and the output is incoherent.
The worked example makes one more thing clear: the information content of the sink tokens is not what matters. Position 4, which we evicted, might have held a semantically rich word like "transformer" or "architecture." But we can discard it because the model can still make good predictions from positions 5 through 12 combined with the sink anchors at 0 through 3. By contrast, position 0 might hold something entirely unremarkable, like a BOS token or a formatting character, but we must keep it because its key vector has been shaped by training to stabilize the entire attention distribution.
Implementation
With the mathematical foundation in place, let's translate these concepts into working code. The implementation centers on a key insight from our analysis: the cache must always contain the sink tokens, so we need a data structure that preserves specific positions while sliding others.
Our StreamingLLMCache class manages this by treating the cache as two distinct regions: a fixed sink region at the beginning that never changes, and a sliding window region that shifts as new tokens arrive. When the cache fills up, we evict the oldest window tokens while the sink tokens remain untouched.
import torch
class StreamingLLMCache:
"""
KV cache for streaming inference that preserves attention sinks.
The cache maintains:
- First `num_sink_tokens` positions (attention sinks)
- Most recent `window_size` positions (sliding window)
"""
def __init__(
self,
num_sink_tokens=4,
window_size=1020,
num_layers=12,
num_heads=12,
head_dim=64,
):
self.num_sink_tokens = num_sink_tokens
self.window_size = window_size
self.max_cache_size = num_sink_tokens + window_size
self.num_layers = num_layers
self.num_heads = num_heads
self.head_dim = head_dim
# Initialize empty caches for keys and values
# Shape: (num_layers, batch_size, num_heads, cache_len, head_dim)
self.k_cache = None
self.v_cache = None
self.cache_len = 0
def update(self, layer_idx, new_k, new_v):
"""
Add new key-value pairs to the cache, evicting old non-sink tokens if needed.
Args:
layer_idx: Which transformer layer this update is for
new_k: New key tensor of shape (batch, heads, new_len, head_dim)
new_v: New value tensor of shape (batch, heads, new_len, head_dim)
Returns:
Combined key and value tensors for attention computation
"""
batch_size = new_k.shape[0]
new_len = new_k.shape[2]
# Initialize cache on first update
if self.k_cache is None:
self.k_cache = torch.zeros(
self.num_layers,
batch_size,
self.num_heads,
self.max_cache_size,
self.head_dim,
device=new_k.device,
dtype=new_k.dtype,
)
self.v_cache = torch.zeros_like(self.k_cache)
# If cache has room, simply append
if self.cache_len + new_len <= self.max_cache_size:
self.k_cache[
layer_idx, :, :, self.cache_len : self.cache_len + new_len
] = new_k
self.v_cache[
layer_idx, :, :, self.cache_len : self.cache_len + new_len
] = new_v
if layer_idx == self.num_layers - 1:
self.cache_len += new_len
else:
# Cache is full: preserve sinks, slide window
# Keep first num_sink_tokens, evict oldest non-sink tokens
sink_k = self.k_cache[layer_idx, :, :, : self.num_sink_tokens]
sink_v = self.v_cache[layer_idx, :, :, : self.num_sink_tokens]
# Calculate how many window tokens we can keep
tokens_to_keep = self.window_size - new_len
if tokens_to_keep > 0:
# Keep recent window tokens
window_k = self.k_cache[layer_idx, :, :, -tokens_to_keep:]
window_v = self.v_cache[layer_idx, :, :, -tokens_to_keep:]
# Reconstruct cache: sinks + kept window + new tokens
self.k_cache[layer_idx, :, :, : self.num_sink_tokens] = sink_k
self.k_cache[
layer_idx,
:,
:,
self.num_sink_tokens : self.num_sink_tokens
+ tokens_to_keep,
] = window_k
self.k_cache[
layer_idx,
:,
:,
self.num_sink_tokens + tokens_to_keep : self.num_sink_tokens
+ tokens_to_keep
+ new_len,
] = new_k
self.v_cache[layer_idx, :, :, : self.num_sink_tokens] = sink_v
self.v_cache[
layer_idx,
:,
:,
self.num_sink_tokens : self.num_sink_tokens
+ tokens_to_keep,
] = window_v
self.v_cache[
layer_idx,
:,
:,
self.num_sink_tokens + tokens_to_keep : self.num_sink_tokens
+ tokens_to_keep
+ new_len,
] = new_v
else:
# Window is smaller than new tokens, just keep sinks and new
self.k_cache[layer_idx, :, :, : self.num_sink_tokens] = sink_k
self.k_cache[
layer_idx,
:,
:,
self.num_sink_tokens : self.num_sink_tokens + new_len,
] = new_k
self.v_cache[layer_idx, :, :, : self.num_sink_tokens] = sink_v
self.v_cache[
layer_idx,
:,
:,
self.num_sink_tokens : self.num_sink_tokens + new_len,
] = new_v
if layer_idx == self.num_layers - 1:
self.cache_len = (
self.num_sink_tokens + max(0, tokens_to_keep) + new_len
)
# Return the active portion of the cache for this layer
return (
self.k_cache[layer_idx, :, :, : self.cache_len],
self.v_cache[layer_idx, :, :, : self.cache_len],
)
def get_cache_length(self):
return self.cache_lenThe cache management is straightforward: when adding new tokens would exceed the maximum cache size, we preserve the sink tokens, keep as many recent tokens as possible, and append the new tokens. Let's test the cache behavior by adding tokens in batches:
# Test the cache behavior
cache = StreamingLLMCache(
num_sink_tokens=4, window_size=12, num_layers=2, num_heads=4, head_dim=32
)
# Simulate adding tokens in batches
batch_size = 1
cache_lengths = []
for step in range(5):
new_k = torch.randn(batch_size, 4, 4, 32) # 4 new tokens per step
new_v = torch.randn(batch_size, 4, 4, 32)
for layer in range(2):
k, v = cache.update(layer, new_k, new_v)
cache_lengths.append(cache.get_cache_length())StreamingLLM Cache Demonstration ================================================== Sink tokens: 4 Window size: 12 Max cache size: 16 Step 1: Added 4 tokens, cache length = 4 Step 2: Added 4 tokens, cache length = 8 Step 3: Added 4 tokens, cache length = 12 Step 4: Added 4 tokens, cache length = 16 Step 5: Added 4 tokens, cache length = 16
The cache grows from 4 to 16 tokens over the first four steps, then stabilizes at 16 (the maximum cache size of 4 sink tokens + 12 window tokens). Once the cache is full, the window slides forward while always preserving the sink tokens.
Now let's implement the attention mechanism that uses this cache:
import torch.nn.functional as F
def streaming_attention(query, cache, layer_idx, new_k, new_v, scale=None):
"""
Compute attention with StreamingLLM cache.
Args:
query: Query tensor of shape (batch, heads, 1, head_dim) for single token
cache: StreamingLLMCache instance
layer_idx: Current layer index
new_k: Key for new token(s)
new_v: Value for new token(s)
scale: Attention scale factor (default: 1/sqrt(head_dim))
Returns:
Attention output and updated cache
"""
head_dim = query.shape[-1]
if scale is None:
scale = 1.0 / (head_dim**0.5)
# Update cache and get all keys/values
keys, values = cache.update(layer_idx, new_k, new_v)
# Compute attention scores: (batch, heads, query_len, cache_len)
scores = torch.matmul(query, keys.transpose(-2, -1)) * scale
# Apply softmax to get attention weights
attn_weights = F.softmax(scores, dim=-1)
# Compute output
output = torch.matmul(attn_weights, values)
return output, attn_weights# Demonstrate attention weight distribution
torch.manual_seed(42)
cache = StreamingLLMCache(
num_sink_tokens=4, window_size=20, num_layers=1, num_heads=1, head_dim=64
)
# Fill cache with some initial tokens
initial_k = torch.randn(1, 1, 20, 64)
initial_v = torch.randn(1, 1, 20, 64)
# Make all 4 sink tokens have distinctive key patterns (simulating trained sinks)
initial_k[:, :, 0, :] = initial_k[:, :, 0, :] * 3.0 + 1.0 # Sink token 0
initial_k[:, :, 1, :] = initial_k[:, :, 1, :] * 2.5 + 0.8 # Sink token 1
initial_k[:, :, 2, :] = initial_k[:, :, 2, :] * 2.0 + 0.6 # Sink token 2
initial_k[:, :, 3, :] = initial_k[:, :, 3, :] * 1.5 + 0.4 # Sink token 3
cache.update(0, initial_k, initial_v)
# Now add a new query token
query = torch.randn(1, 1, 1, 64)
new_k = torch.randn(1, 1, 1, 64)
new_v = torch.randn(1, 1, 1, 64)
output, attn_weights = streaming_attention(query, cache, 0, new_k, new_v)
The attention weights show the expected pattern: higher attention on the first few positions (the sink tokens) with the remaining attention distributed across the window tokens.
Designing Effective Sink Tokens
Not all initial tokens make equally good sinks. The effectiveness of an attention sink depends on what token occupies that position during training. A token becomes a good sink when its key vector, after training, projects strongly onto a wide variety of query vectors. That property is acquired through repeated exposure: the more consistently a token appears at position 0 across diverse training examples, the more thoroughly the training process shapes its key representation into a universal attractor.
This has practical consequences for how you choose which tokens to preserve during streaming inference. The StreamingLLM paper tested several choices and found that the model's own BOS token is almost always the most effective single sink. Supplementary sinks can come from system prompt tokens if the model was instruction-tuned on conversations, or from learned sink tokens if the model was trained with explicit sink support. In all cases, the quality criterion is the same: how well does this token's key vector absorb excess attention across diverse query contexts?
Several strategies produce effective sinks, each with different implementation requirements:
Beginning-of-Sequence Token
Most tokenizers include a dedicated <BOS> (beginning-of-sequence) or <s> token that always appears at position 0. Because this token appears at the start of every training sequence, it develops a strong sink representation. The model learns that this token is always present and always available to absorb excess attention.
System Prompt Tokens
For instruction-tuned models, the system prompt often provides natural sink tokens. Phrases like "You are a helpful assistant" appear at the start of every conversation, making them reliable candidates for sink positions. The model has seen these tokens in position 0-20 millions of times, so their key representations have thoroughly adapted to the sink role.
Dedicated Sink Tokens
Some models are trained with explicit sink tokens. These are special tokens inserted at the beginning of sequences specifically to serve as attention sinks. During training, the model learns to route excess attention to these tokens rather than developing the behavior organically.
def count_sink_tokens_needed(model_perplexity_data, threshold=1.05):
"""
Determine the optimal number of sink tokens based on perplexity analysis.
When removing initial tokens, perplexity increases. We find the minimum
number of tokens to preserve such that perplexity stays within threshold.
Args:
model_perplexity_data: Dict mapping num_preserved_tokens -> perplexity
threshold: Maximum acceptable perplexity ratio vs. baseline
Returns:
Recommended number of sink tokens to preserve
"""
if not model_perplexity_data:
return 4 # Default recommendation from StreamingLLM paper
baseline_perplexity = max(model_perplexity_data.values())
for num_tokens in sorted(model_perplexity_data.keys()):
ppl = model_perplexity_data[num_tokens]
if ppl <= baseline_perplexity * threshold:
return num_tokens
return max(model_perplexity_data.keys())# Simulated perplexity data showing effect of sink token count
simulated_data = {
0: 45.2, # No sinks: very high perplexity
1: 18.7, # 1 sink: still elevated
2: 12.3, # 2 sinks: getting better
3: 10.1, # 3 sinks: close to baseline
4: 9.8, # 4 sinks: matches baseline
8: 9.7, # 8 sinks: marginal improvement
}
recommended = count_sink_tokens_needed(simulated_data)Perplexity vs Number of Preserved Sink Tokens ============================================= 0 sink tokens: perplexity = 45.2 <-- recommended 1 sink tokens: perplexity = 18.7 2 sink tokens: perplexity = 12.3 3 sink tokens: perplexity = 10.1 4 sink tokens: perplexity = 9.8 8 sink tokens: perplexity = 9.7
The data shows that perplexity drops sharply when preserving the first few tokens, then plateaus. With zero sink tokens, perplexity explodes to 45.2, indicating severe model degradation. Adding just one sink token cuts perplexity by more than half. By four sink tokens, perplexity reaches 9.8, nearly matching baseline performance. Additional sinks beyond four provide diminishing returns, matching the StreamingLLM paper's recommendation.
Streaming Inference Pipeline
Let's put everything together into a complete streaming inference pipeline:
class StreamingInferenceEngine:
"""
Complete streaming inference engine with attention sink preservation.
"""
def __init__(self, model, tokenizer, num_sink_tokens=4, window_size=2044):
self.model = model
self.tokenizer = tokenizer
self.num_sink_tokens = num_sink_tokens
self.window_size = window_size
# Extract model configuration
config = model.config
self.num_layers = config.num_hidden_layers
self.num_heads = config.num_attention_heads
self.head_dim = config.hidden_size // config.num_attention_heads
# Initialize cache
self.cache = None
self.total_tokens_generated = 0
def reset(self):
"""Reset the cache for a new generation session."""
self.cache = None
self.total_tokens_generated = 0
def _init_cache(self, batch_size, device, dtype):
"""Initialize the streaming cache."""
self.cache = StreamingLLMCache(
num_sink_tokens=self.num_sink_tokens,
window_size=self.window_size,
num_layers=self.num_layers,
num_heads=self.num_heads,
head_dim=self.head_dim,
)
def generate_token(self, input_ids, past_key_values=None):
"""
Generate a single token using the streaming cache.
In a real implementation, this would integrate with the model's
forward pass to use the streaming cache instead of standard KV cache.
"""
# This is a simplified demonstration
# Real implementation would modify the model's attention layers
passLet's demonstrate the memory savings achieved by StreamingLLM compared to a full KV cache:
# Simulate a long generation scenario
total_tokens = 10000
sink_tokens = 4
window_size = 2044
max_cache_size = sink_tokens + window_size
# Calculate memory usage (assuming 768-dim embeddings, 12 layers, float32)
# KV cache stores keys and values (2x) for each layer
full_cache_memory = total_tokens * 768 * 12 * 2 * 4 / (1024**2) # MB
streaming_cache_memory = max_cache_size * 768 * 12 * 2 * 4 / (1024**2) # MB
memory_saved = full_cache_memory - streaming_cache_memory
savings_percent = (1 - streaming_cache_memory / full_cache_memory) * 100Streaming Inference Demonstration ================================================== Generating 10,000 tokens with: - 4 sink tokens (always preserved) - 2,044 window tokens (sliding) - Maximum cache size: 2,048 tokens Memory usage comparison: Full KV cache: 703.1 MB StreamingLLM cache: 144.0 MB Memory saved: 559.1 MB (79.5%)
The streaming approach enables generation of arbitrarily long sequences while using fixed memory. In this example, generating 10,000 tokens with a full cache would require over 700 MB, but StreamingLLM uses only about 145 MB, an 80% reduction.

The visualization makes the scaling difference clear. Full KV caching follows a diagonal line that quickly exceeds typical GPU memory limits. At 100,000 tokens, a full cache would require nearly 7 GB. StreamingLLM's flat line at the bottom shows constant memory usage, enabling generation of arbitrarily long sequences without hitting memory limits.
To put the numbers in context: for a GPT-2-medium sized model with 768-dimensional hidden states and 12 layers, each cached token consumes 768 dimensions 12 layers 2 (keys and values) 4 bytes = roughly 72 KB per token. A StreamingLLM cache of 4 sinks + 2044 window tokens uses about 147 MB total, regardless of how many tokens have been generated. A full cache for a 100,000-token sequence would require over 7 GB. For larger models like the 7B-parameter class with 4096 dimensions and 32 layers, each cached token uses about 1 MB, so StreamingLLM's 2048-token cache costs about 2 GB while a 100,000-token full cache would cost roughly 100 GB. At that scale, full caching is simply not feasible on any single GPU, so streaming approaches become necessary rather than merely convenient.
Empirical Validation
The mathematical argument for StreamingLLM explains why the method works, but empirical confirmation is equally important. Two questions matter most in practice: how many sink tokens do you need, and how large does the window need to be? The answers determine the memory footprint of your deployment. The original paper's findings on these questions have been replicated across many model families.
The key metric for evaluating streaming generation quality is perplexity: the model's average uncertainty per token, measured on a held-out evaluation set. Lower perplexity means the model assigns higher probability to the correct next token, indicating better prediction quality. When we vary the number of sink tokens or the window size and measure how perplexity changes, we get a direct empirical measurement of how much each parameter affects generation quality.
Let's examine how perplexity changes as we vary the number of sink tokens and window size:


The left plot confirms that preserving 4 sink tokens is sufficient for most models. Additional sinks provide diminishing returns. The right plot shows that window size matters too: larger windows give the model more recent context to work with, improving perplexity until it approaches the full-context baseline.
Beyond StreamingLLM: Learnable Sink Tokens
The original StreamingLLM approach preserves the first tokens from the input sequence. A more principled approach is to train models with explicit, learnable sink tokens. These tokens are:
- Added to the vocabulary as special tokens
- Always prepended to the input sequence during training
- Optimized end-to-end to serve as effective attention sinks
Models trained this way develop more efficient sink representations. The sink tokens learn key vectors that are optimally positioned to capture excess attention, rather than relying on whatever tokens happen to appear at the start of the sequence.
Learnable sink tokens are special tokens added to a model's vocabulary and trained end-to-end to serve as attention sinks. Unlike implicit sinks that emerge from training on natural text, learnable sinks are explicitly optimized to absorb excess attention weight.
The training procedure is straightforward:
- Add sink tokens to the tokenizer vocabulary
- Prepend these tokens to every training sequence
- Train the model normally, allowing the sink token embeddings to learn
During inference, you prepend the same sink tokens to the input, and they immediately serve their purpose without needing to preserve specific input tokens.
The advantage of learnable sink tokens over naturally occurring ones is consistency and efficiency. When a model relies on its BOS token as a sink, it is repurposing a token that was also trained to carry other information, such as signaling the start of a sequence to the model's positional encoding. The BOS token's key vector has to serve two masters: representing "start of sequence" for the generation process and serving as an attention attractor for all subsequent positions. These goals are not always aligned. A dedicated learnable sink token has a single purpose, so its key vector can be optimized entirely for the attractor role. The result is a more compact, more reliable sink that often requires fewer sink tokens to achieve the same generation quality.
This design choice is increasingly common in recently developed models. If you are building a model from scratch, or fine-tuning a base model for production streaming, adding one or two learnable sink tokens to the vocabulary is a low-cost, high-value investment. The extra tokens add negligible computation per step while providing a principled foundation for unlimited streaming generation.
def prepare_input_with_sink_tokens(tokenizer, text, num_sink_tokens=4):
"""
Prepare input with explicit sink tokens prepended.
Args:
tokenizer: Tokenizer with sink tokens in vocabulary
text: Input text to tokenize
num_sink_tokens: Number of sink tokens to prepend
Returns:
Token IDs with sink tokens at the beginning
"""
# Assume sink tokens are named [SINK_0], [SINK_1], etc.
sink_token_ids = [
tokenizer.convert_tokens_to_ids(f"[SINK_{i}]")
for i in range(num_sink_tokens)
]
# Tokenize the actual text
text_tokens = tokenizer.encode(text, add_special_tokens=False)
# Combine: sinks first, then text
return sink_token_ids + text_tokensIn Practice: Deploying Streaming Inference
The gap between the algorithm as described in the paper and a working production deployment is narrower than it might seem, but a few practical considerations determine whether streaming inference works reliably. Understanding them saves you from debugging sessions that feel like the model is randomly failing when there is in fact a systematic cause.
The first practical consideration is integration with the model's forward pass. The StreamingLLM cache is a drop-in replacement for the standard key-value cache that modern transformer implementations already maintain for fast autoregressive generation. Most inference frameworks (HuggingFace Transformers, vLLM, llama.cpp) expose a past_key_values argument that carries the key-value cache across generation steps. Replacing the standard cache management logic with the two-region (sinks plus sliding window) management is sufficient. You do not need to modify attention weights, layer normalization, or any other part of the forward pass.
The second consideration is batching. In single-sequence generation, the cache management is straightforward because there is one sequence with one set of sink tokens. In batched generation, you need independent cache management per sequence, because different sequences in a batch may be at different positions and have different eviction histories. Sharing a cache across sequences would corrupt the sink-token invariant for most of them.
The third consideration is integration with generation strategies like beam search. Standard beam search maintains multiple candidate sequences in parallel, expanding and pruning beams at each step. Streaming inference is most natural for greedy decoding or top-k/top-p sampling, where each step produces exactly one token per sequence. Beam search with StreamingLLM requires one StreamingLLM cache per beam, which multiplies memory usage by the beam width. For most conversational and streaming applications, greedy or sampling-based generation is appropriate, so this is rarely a limitation in practice.
Finally, a word on when to use StreamingLLM versus alternative long-context approaches. StreamingLLM is the right choice when you need unbounded generation length with fixed memory and are willing to accept that the model cannot recall content from beyond the window boundary. It is not the right choice when precise recall of specific facts from throughout a long document is required. For that use case, retrieval-augmented generation (which we cover in a separate chapter) provides a better tradeoff: the model can access any part of the document on demand, at the cost of a retrieval step. StreamingLLM and RAG are complementary technologies: RAG for deep long-range recall, streaming for fluent continuous generation.
Key Parameters
When implementing StreamingLLM, four parameters govern the tradeoff between generation quality, memory usage, and task suitability. Getting these parameters right means understanding what each one controls and why it affects streaming behavior in the way it does. Fortunately, the empirical guidance from the original paper is clear enough that you can start with sensible defaults and tune conservatively from there.
The fundamental tradeoff is that larger parameter values (more sinks, larger window) produce better quality at the cost of more memory, while smaller values reduce memory at the cost of some quality degradation. For most practical deployments, the right balance lands in a relatively narrow range that the original paper characterized well.
The key parameters are:
-
num_sink_tokens: Number of initial tokens to preserve as attention sinks. The default of 4 works well for most models, but you may need fewer (1-2) for smaller models or more (8-16) for models with unusual attention patterns. Start with 4 and adjust based on perplexity measurements.
-
window_size: Number of recent tokens to keep in the sliding window. Larger windows capture more context but require more memory. Common values range from 512 to 2048. The right choice depends on the typical dependency length in your generation task: conversational AI may need 512-1024, while document summarization benefits from 2048+.
-
max_cache_size: Total cache size, computed as
num_sink_tokens + window_size. This determines the fixed memory footprint. For a 7B parameter model with 32 layers and 4096 head dimension, each token in the cache uses approximately 1 MB, so a cache of 2048 tokens requires about 2 GB. -
position_encoding_strategy: How to handle position encodings for window tokens. Options include keeping original positions (works with RoPE), resetting positions when sliding (may cause distribution mismatch), or using relative encodings (which are less sensitive to this mismatch). The choice depends on the base model's positional encoding scheme.
The position encoding strategy deserves special attention because it is the parameter most likely to cause subtle failures that are hard to diagnose. When a token at original position 5000 slides into the "first window slot" after many evictions, it still carries its position-5000 encoding if you are using absolute position embeddings. The model was trained to interpret position 5000 as a middle-of-document location, not as a near-past location. For models using RoPE (Rotary Position Embedding), keeping original positions works reasonably well because RoPE encodes relative distance rather than absolute position in a way that degrades gracefully. For models using absolute sinusoidal embeddings, you may need to recompute positions after each eviction to avoid systematic distribution mismatch.
Limitations and Considerations
While attention sinks and StreamingLLM enable streaming inference, they come with limitations you should understand before choosing this approach for a production system. These limitations are not reasons to avoid StreamingLLM, but they are constraints that define where the technique excels and where it falls short.
The fundamental limitation is that StreamingLLM trades context for memory. When you slide the window forward, you permanently lose access to information in the discarded tokens. If the model needs to recall a specific detail from 5,000 tokens ago but that token is no longer in the cache, the information is simply gone. For tasks requiring precise long-range recall (like answering questions about specific facts from early in a document), StreamingLLM will struggle compared to full-context approaches. This is not a subtle degradation. It is a hard information-theoretic boundary: if the relevant tokens are not in the cache, the model cannot retrieve them regardless of how much it "tries." The model may confabulate plausible-sounding but incorrect details to fill the gap, which is arguably worse than simply saying it does not know.
The attention sink phenomenon is also somewhat architecture-dependent. While it appears consistently across decoder-only transformer models trained autoregressively, the exact number of sink tokens needed and their effectiveness varies. Models with different training procedures, position encodings, or attention patterns may exhibit different sink behaviors. The 4-sink recommendation from the original paper is a good starting point, but you may need to tune this for your specific model. In particular, models that use group-query attention (GQA) or multi-query attention (MQA) have fewer key-value heads than query heads, which can alter the sink dynamics. Models fine-tuned heavily on instruction-following data may have shifted their sink behavior relative to the base model.
There is also a subtle issue with positional encoding. When tokens slide out of the window, the remaining tokens do not automatically adjust their position encodings. A token at original position 5000 keeps its position-5000 encoding even when it becomes the "first" window token. This works reasonably well with relative position encodings like RoPE, but can cause issues with absolute position encodings. Some implementations recompute position encodings for the active cache, but this adds complexity and may not match the model's training distribution. The severity of this issue depends on whether the model was trained on sequences longer than the expected streaming context, and on how the position encoding interacts with the attention computation.
StreamingLLM also introduces a quality gap compared to full-context inference, even when the relevant tokens are within the window. The model knows, in some sense, that information is missing: the "middle" region between the sink tokens and the current window was once populated but is now absent. The attention denominator in the softmax still sums over only the sinks and the window, so the normalization is not identical to what the model saw during pretraining on full sequences. This introduces a mild distribution mismatch that is smaller than the mismatch from removing sinks entirely, but not zero. For most generation tasks, this mismatch is imperceptible. For tasks that require very precise probability calibration (such as scoring or ranking candidate continuations), it may matter.
Finally, StreamingLLM does not help with the initial context window. You still cannot attend to more than your window allows at any single step. If understanding a passage requires simultaneously seeing tokens 1-1000 and tokens 5000-6000, you are limited by what fits in the window plus sinks. StreamingLLM enables infinite generation, not infinite context. These are different problems with different solutions. Infinite generation is what StreamingLLM solves: maintaining quality across arbitrarily long outputs. Infinite context, meaning the ability to reason over arbitrarily long inputs, requires different architectural innovations such as hierarchical attention, memory-augmented transformers, or retrieval-augmented generation.
Summary
Attention sinks are an emergent property of autoregressive transformers: the first few tokens in a sequence learn to absorb excess attention weight that would otherwise need to be distributed across irrelevant positions. This behavior emerges from the softmax normalization requirement that attention weights sum to 1, combined with training dynamics that consistently reward the use of early positions as "safe" targets for excess attention.
StreamingLLM exploits this phenomenon for practical benefit. By preserving a small number of sink tokens alongside a sliding window of recent context, you can generate text of arbitrary length with fixed memory usage. The key insights are:
- Attention sinks are essential: Removing the first tokens breaks the attention distribution the model learned during training, causing quality degradation
- A few sinks suffice: Four sink tokens typically achieve baseline-equivalent performance, with diminishing returns from additional sinks
- Window size matters: Larger windows capture more local context, improving generation quality until the window is large enough to cover typical dependency ranges
- Memory is constant: Unlike full KV caching that grows linearly with sequence length, StreamingLLM uses fixed memory regardless of generation length
In practice, you can deploy language models for continuous generation tasks (chatbots, streaming summaries, real-time translation) without worrying about memory growth. The model maintains coherent output quality for thousands or millions of tokens, constrained only by compute time rather than memory.
The implementation is also remarkably lightweight. You do not need to retrain a model, modify its architecture, or change its attention computation. You need only a custom key-value cache with two regions: a fixed sink buffer and a sliding window buffer. The total memory footprint is bounded by the sum of these two regions regardless of generation length, and the implementation integrates cleanly with existing transformer inference frameworks through the standard past_key_values interface.
Understanding attention sinks also provides insight into transformer behavior more broadly. The fact that models spontaneously learn to use certain positions as "attention dumps" shows how they manage the constraint that attention weights must form a probability distribution. This has implications for model interpretability, efficient architecture design, and the development of even longer-context language models. The sink phenomenon is one of several ways that transformers develop internal coordination strategies not specified in their training objective, strategies that only become visible when you probe the model's behavior at scale. It is a reminder that large models are not simply scaling up small models: they develop qualitatively different internal organization that both creates new capabilities and imposes new constraints on inference-time deployment.
The next chapter builds on this foundation by examining sparse attention patterns and sliding window attention architectures, which address the computational cost of full self-attention over long sequences in a complementary way. Where StreamingLLM solves the memory problem of streaming generation, sparse attention solves the quadratic cost of attention over long inputs, and the two techniques can be combined in models designed for extremely long-context applications.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about attention sinks and StreamingLLM.
Attention Sinks Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!