Part of Language AI Handbook
Global tokens create communication hubs in sparse attention, reducing path length and helping efficient transformers exchange information across long documents.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Global Tokens
Sparse attention patterns like sliding windows solve the quadratic complexity problem but introduce a new challenge: how does information travel across the full sequence? If each token only attends to nearby neighbors, a token at position 0 cannot directly communicate with a token at position 1000. Information must hop through many intermediate windows, risking degradation along the way. Global tokens solve this by designating certain positions as communication hubs that can attend to and be attended by every position in the sequence.
To understand why this matters, think about what attention is really doing. In standard self-attention, every token can directly ask "what information anywhere in the sequence is relevant to me?" That direct query-answer relationship is precisely what gives transformers their power over recurrent models, which had to pass information through a long chain of hidden states. Sliding window attention sacrifices that directness in exchange for efficiency, and global tokens are the mechanism that wins it back, at least partially.
The key insight is that you do not need every pair of tokens to communicate directly. You need every token to be reachable from every other token within a small, fixed number of steps. If you can guarantee that, the model can still reason about the full sequence across multiple layers, even if no single layer has global receptive fields for all tokens. Global tokens provide exactly those shortcuts. Think of them as express stops on a subway network: you might have to take a local train one stop to reach the nearest express station, but from there you can travel anywhere in the system in one more hop.
The mathematical elegance of this approach is that the cost grows linearly rather than quadratically. Adding global tokens to a sequence of length introduces additional attention connections on top of the sliding window connections, where is the window size. As long as stays small relative to , the overall complexity remains linear. In practice, even a single global token is often sufficient for many tasks, and even using 64 global tokens for a sequence of 4096 tokens adds negligible overhead.
This chapter builds directly on the sliding window and sparse attention mechanisms covered in earlier chapters. If you have been following the progression of efficient attention techniques, you will recognize global tokens as a targeted surgical fix rather than a wholesale redesign. The underlying sliding window machinery is left intact. Global tokens are added on top, requiring only a modest change to the attention mask and the computation that handles global positions separately from local ones.
We will explore how global tokens evolved from BERT's classification token, how different tasks motivate different placement strategies, how learned global tokens discovered by training differ from task-designated ones, and how all of this interacts with computational complexity. By the end, you will understand the mechanics of global tokens and why they represent one of the cleanest solutions to the long-range problem in efficient transformers.
The Information Bottleneck in Local Attention
Sliding window attention limits each token to a local neighborhood. While efficient, this creates a path length problem reminiscent of RNNs. To transfer information from the start of a document to the end, data must flow through overlapping windows one hop at a time.
The problem appears in practice. A document understanding model reading a legal contract needs tokens near the end of the document, where the signature block appears, to be influenced by the definitions established at the beginning. A summarization model processing a news article needs the concluding paragraph to be shaped by the topic sentence in the lede. A question answering model reading a scientific paper needs the answer token, which may appear in the methods section, to be directly aware of the question about experimental design. In all these cases, the long-range dependency is not optional: it is the core of the task.
Consider a sequence of 4096 tokens with a window size of 512. A token at position 0 can directly see tokens up to position 256 (half the window). To reach position 4000, information must traverse roughly 16 hops through successive windows. Each hop introduces potential for information loss or distortion.
This 16-hop path can cause problems in neural networks. Each attention layer transforms representations, and there is no guarantee that a specific piece of information introduced at position 0 will survive intact through 16 layers of transformation as it propagates toward position 4000. In an RNN, this was called the vanishing gradient problem in a different guise. In sliding window attention, it manifests as effective forgetting: the further the semantic distance, the more attenuated the information signal becomes across successive layers.
Path length measures the number of attention operations needed for one position to influence another. Standard self-attention has path length 1 (direct connection). Sliding window attention has path length proportional to sequence length divided by window size. The formal bound is:
where is the distance between positions and is the window size. For a 4096-token sequence with window size 512, two positions at opposite ends of the sequence have a path length of 16.
Global tokens restore direct connections by serving as intermediaries. Any token can communicate with any other token by passing through a global token, reducing the maximum path length to 2.
The path length of 2 deserves careful attention, because it provides a stronger guarantee than it first appears. With full self-attention, every pair of tokens communicates in exactly 1 hop: direct attention. With sliding window attention alone, the maximum path is , which grows with sequence length. With global tokens added, the bound becomes 2 regardless of sequence length. This means that even for a million-token sequence, any two tokens are always just one global token hop away from each other. The global token architecture achieves a property that is qualitatively closer to full attention than to pure sliding window attention, even though its complexity profile resembles the latter.
import numpy as np
def compute_path_lengths(seq_len, window_size, num_global=0):
"""
Compute maximum path lengths for different attention patterns.
Path length = minimum number of attention operations for info to flow
between any two positions.
"""
# Sliding window: path length is ceil(distance / (window_size / 2))
max_distance = seq_len - 1
window_path = int(np.ceil(max_distance / (window_size // 2)))
# With global tokens: max path is 2 (token -> global -> token)
global_path = 2 if num_global > 0 else window_path
return {
"full_attention": 1,
"sliding_window": window_path,
"with_global_tokens": global_path,
}
# Example: 4096 token sequence
seq_len = 4096
window_size = 512
paths = compute_path_lengths(seq_len, window_size, num_global=1)
The difference is stark. Sliding window attention requires 16 hops to connect distant positions, while adding just one global token reduces this to 2. This improvement in connectivity comes with minimal computational overhead since only a few tokens become global.
What the chart also shows is that the three approaches divide naturally into two qualitative tiers. Full attention and global-token attention both achieve constant-bounded path lengths (1 and 2 respectively), while pure sliding window attention has a path length that grows with the sequence. That qualitative difference has real consequences during training: models with bounded path lengths can learn long-range dependencies through gradient backpropagation much more reliably than models where the gradient signal must travel through 16 or more attention layers to propagate between distant positions.
CLS Token as Global Attention
The idea of global tokens predates efficient transformers. BERT's [CLS] token, placed at the start of every sequence, was designed to aggregate sentence-level information. During fine-tuning, the [CLS] representation feeds into classification heads because it has learned to summarize the entire input.
The [CLS] token was introduced in the original BERT paper (Devlin et al., 2018) primarily as a bookkeeping convenience. Because BERT was trained on pairs of sentences, the model needed a position whose representation could be used to classify whether the two sentences were consecutive in the original text (the next sentence prediction task). The [CLS] token, attending to both sentences through standard full self-attention, naturally aggregated information from both halves of the input. Researchers later discovered that this aggregation property made it ideal as a sentence embedding: the final-layer [CLS] representation often captured the overall semantic content of the input well enough to drive downstream classification tasks with a simple linear layer on top. Longformer's designers, faced with the problem of extending BERT to long documents, recognized that the [CLS] token's existing role as a sequence-level aggregator mapped perfectly onto the global attention hub concept. Rather than inventing a new mechanism, they preserved the familiar token but gave it the architectural privilege it had always semantically deserved.
Longformer and BigBird formalized this intuition by giving the [CLS] token bidirectional global attention. In standard BERT, [CLS] only receives information through the normal attention mechanism. In Longformer, [CLS] can attend to every token and every token can attend to [CLS].
The distinction between unidirectional and bidirectional global attention is worth dwelling on. In standard BERT's self-attention, [CLS] does attend to all tokens in the sequence, because BERT uses full quadratic attention. What changes in Longformer is the context: when using sliding window attention for all other tokens, [CLS] would normally only see a small local window if treated like any other position. Making it globally attending means [CLS] retains its BERT-like property of seeing the full sequence even in a sparse attention framework. The bidirectional aspect, that every token also attends back to [CLS], is what maps it from a mere aggregator into a true communication hub. Without bidirectionality, local tokens could push information to [CLS] but could not pull information back from the global context it has accumulated.
Think of it this way: a non-bidirectional [CLS] is like a town bulletin board where everyone can post notices, but nobody reads them. A bidirectional global [CLS] is like a town hall meeting where everyone speaks and everyone listens, concentrated into a single representative who then broadcasts the collective understanding back to the whole community.
def create_longformer_attention_mask(seq_len, window_size, global_positions):
"""
Create an attention mask for Longformer-style attention.
Returns:
mask: Boolean array where True means position i can attend to position j
"""
mask = np.zeros((seq_len, seq_len), dtype=bool)
# Local window attention for all positions
for i in range(seq_len):
start = max(0, i - window_size // 2)
end = min(seq_len, i + window_size // 2 + 1)
mask[i, start:end] = True
# Global positions can attend to everything and be attended by everything
for g in global_positions:
mask[g, :] = True # Global attends to all
mask[:, g] = True # All attend to global
return mask
# Example: 16 tokens, window of 4, CLS token at position 0
seq_len = 16
window_size = 4
global_positions = [0] # CLS token
mask = create_longformer_attention_mask(seq_len, window_size, global_positions)
The attention mask reveals the structure clearly. The cross pattern emanating from position 0 shows the CLS token's global connectivity. The diagonal band represents local window attention for all other positions. This combination ensures that every token is at most 2 hops from any other token: local token to CLS, then CLS to distant local token.
In practice, this mask is sparse enough to provide major computational savings while being connected enough to preserve the essential long-range reasoning capabilities. The fraction of enabled attention connections in this 16-token example is small, and for realistic sequences of thousands of tokens, the fraction enabled by the global token grows even smaller relative to the baseline of full attention.
Task-Specific Global Tokens
While CLS provides classification-focused global attention, different tasks benefit from different global token configurations. The key insight here is that the optimal placement of global tokens is a function of the task's information requirements, and those requirements are often apparent from the task structure itself.
Question answering, for instance, naturally has two distinct segments: the question and the context passage. The question defines what information is relevant; the context contains candidate answer spans. Every word in the context must, in some sense, evaluate itself against the question: does this token or phrase answer what was asked? That evaluation is precisely what attention enables. Making all question tokens global ensures that every context token can directly attend to the question in a single hop, without having to route through local windows that may not include any question tokens.
This is a fundamentally different design philosophy from the single-CLS approach. Rather than designating one universal aggregator, we designate an entire semantically meaningful segment as globally visible. The architectural choice reflects the task structure: in QA, the question is the global context, and the context passage is the local content being searched. Making question tokens global encodes that semantic relationship directly into the attention mask.
Named entity recognition and coreference resolution benefit similarly. In coreference resolution, a pronoun at position 3000 in a document might refer to a noun phrase at position 100. If the noun phrase tokens are made globally visible, the pronoun token can directly attend to them regardless of the distance. The alternative, waiting for the sliding window to propagate the noun phrase's representation to position 3000, requires many hops and risks losing the specific lexical information needed to resolve the reference.
def create_qa_attention_mask(question_len, context_len, window_size):
"""
Create attention mask for question answering.
Question tokens are global, context tokens use sliding window.
"""
total_len = question_len + context_len
mask = np.zeros((total_len, total_len), dtype=bool)
# Question tokens are global (positions 0 to question_len-1)
for i in range(question_len):
mask[i, :] = True # Question attends to all
mask[:, i] = True # All attend to question
# Context tokens use sliding window among themselves
for i in range(question_len, total_len):
start = max(question_len, i - window_size // 2)
end = min(total_len, i + window_size // 2 + 1)
mask[i, start:end] = True
return mask
# Example: 4 question tokens, 12 context tokens
question_len = 4
context_len = 12
window_size = 4
qa_mask = create_qa_attention_mask(question_len, context_len, window_size)
The mask shows a block pattern. The upper-left quadrant is fully connected (question tokens attending to each other). The left columns and top rows extend this connectivity to context tokens. The lower-right shows the sliding window pattern among context tokens.
This task-specific design makes semantic sense. When answering a question, every context word should be able to consider the question directly. Without global question tokens, a context word might need to hop through multiple windows before accessing question information.
Notice also what happens to the question tokens themselves. They attend to each other fully (the upper-left block), and they attend to all context tokens (the top rows). This means question tokens can aggregate information from the entire context passage, not just from nearby positions. A question token representing "where" can attend directly to the answer span wherever it appears in the context, pulling its representation into the question token's key-value space. The result is that question tokens and context tokens both have mutual global visibility, which is exactly the right structure for extractive QA.
Learned Global Tokens
CLS and task-specific tokens are fixed by design, meaning a human engineer must decide which positions serve as global hubs before training begins. An alternative approach introduces learnable global tokens that the model discovers during training. These tokens do not correspond to any input words but serve purely as memory and communication buffers.
This shift from designated to learned global tokens represents a deeper change in perspective. With designated tokens, we are imposing prior knowledge about the task structure onto the architecture: we believe the question tokens are globally important, so we encode that belief in the mask. With learned global tokens, we step back and say: we do not know exactly what global information the task needs, so let the model figure it out. The learned tokens become a trainable bottleneck through which global information flows, and the model optimizes how to use them to minimize task loss.
Perceiver and Perceiver IO use this idea extensively. They introduce a small set of latent tokens, often 256 to 512, that cross-attend to a much longer input sequence. The latent tokens then self-attend among themselves efficiently, and finally cross-attend back to produce outputs. This three-stage process, cross-attend in, self-attend internally, cross-attend out, allows the model to process inputs of arbitrary length while keeping the expensive self-attention computation confined to the small latent space.
The elegance of learned global tokens is their generality. The same mechanism works whether the input is text, images, audio, or structured data. The latent tokens do not know or care about the modality of the input. They simply learn to extract whatever information is relevant for the task from whatever input is presented. This is why Perceiver IO succeeded at processing multimodal inputs: the latent tokens served as a modality-agnostic interface between the raw input stream and the task-specific output.
Think of learned global tokens as empty vessels at initialization. At the start of training, their embeddings are random. This provides no useful information. As training proceeds, the gradient signal shapes them to carry exactly the information that other tokens need to perform the task. By the end of training, each latent token has a specialized role in the model's information routing scheme, even though that role was never explicitly programmed.
def create_latent_global_mask(input_len, num_latents, window_size):
"""
Create attention mask with learned latent global tokens.
Latent tokens attend globally; input tokens attend locally + to latents.
"""
total_len = num_latents + input_len
mask = np.zeros((total_len, total_len), dtype=bool)
# Latent tokens (positions 0 to num_latents-1) have global attention
for i in range(num_latents):
mask[i, :] = True
mask[:, i] = True
# Input tokens use sliding window among themselves
for i in range(num_latents, total_len):
input_pos = i - num_latents
# Local window in input space
start = num_latents + max(0, input_pos - window_size // 2)
end = num_latents + min(input_len, input_pos + window_size // 2 + 1)
mask[i, start:end] = True
return mask
# Example: 4 latent tokens, 16 input tokens
num_latents = 4
input_len = 16
window_size = 4
latent_mask = create_latent_global_mask(input_len, num_latents, window_size)
The advantage of learned latents is flexibility. The model can discover what information to store in global memory rather than being constrained to predefined positions. The latent tokens act as a compressed representation of the full sequence, enabling efficient long-range communication.
In practice, learned global tokens tend to specialize organically. Probing studies on Perceiver models have found that different latent tokens attend preferentially to different parts of the input or different semantic categories of information, even though no such specialization was explicitly programmed. One latent might concentrate its attention on entities, another on relational phrases, another on discourse markers. This emergent specialization is a sign that the model is efficiently using its global token budget, allocating different aspects of sequence-level understanding to different latent positions.
Global Token Count
How many global tokens do you need? This depends on the task and sequence length. Too few global tokens create an information bottleneck. Too many defeat the purpose of sparse attention by reintroducing quadratic interactions.
The right way to think about this trade-off is in terms of information-theoretic capacity. A single global token with embedding dimension has at most real-valued numbers to represent the global context of a sequence that may contain thousands of tokens. For a simple classification task where the label depends on one or two facts from the document, a single global token may be sufficient: it just needs to encode those facts and suppress everything else. For a complex task like multi-hop question answering, where the answer requires combining three or four facts distributed across the document, a single global token becomes a bottleneck. Each fact competes for the limited representational capacity, and some may be lost.
The operational rule of thumb is to start with one global token and add more only when task-specific evaluation metrics show that performance has plateaued. A useful diagnostic is to inspect the attention patterns of the global token after training: if it shows very diffuse, nearly uniform attention over the entire sequence, it is probably overloaded and struggling to concentrate on the most important information. If it shows sharply peaked attention at a few positions, it may have sufficient capacity and adding more tokens may not help much.
For classification tasks with a single CLS token, one global token often suffices. The CLS token aggregates a summary representation, and the classification head only needs this single vector.
For tasks requiring fine-grained outputs, such as question answering, named entity recognition, or span extraction, more global tokens help. Making all question tokens global in QA provides richer context for every answer span candidate. Each question token carries part of the question's semantics, so different context tokens can attend to whichever question token is most relevant to their local content.
def compute_attention_complexity(seq_len, window_size, num_global):
"""
Compute number of attention connections for different configurations.
"""
# Full attention: n^2 connections
full = seq_len * seq_len
# Sliding window: approximately n * window_size
local = seq_len * window_size
# Global token connections: each global attends to all (n)
# and is attended by all (n), minus overlap
# Approximately: 2 * num_global * seq_len
global_connections = 2 * num_global * seq_len - num_global * num_global
# Total sparse: local + global (with some overlap correction)
sparse = local + global_connections
return {
"full_attention": full,
"local_only": local,
"local_plus_global": sparse,
"savings_percent": 100 * (1 - sparse / full),
}
# Test with different global token counts
seq_len = 4096
window_size = 512
results = []
for num_global in [1, 4, 16, 64]:
complexity = compute_attention_complexity(seq_len, window_size, num_global)
results.append(
{
"num_global": num_global,
"connections": complexity["local_plus_global"],
"savings": complexity["savings_percent"],
}
)

Even with 64 global tokens, we maintain over 84% savings compared to full attention. The marginal cost of additional global tokens is linear in sequence length, not quadratic. This means you can afford to be generous with global tokens for complex tasks without sacrificing efficiency.
To understand why this holds, recall that the total number of connections in global-local attention is approximately:
where:
- : sequence length
- : window size
- : number of global tokens
The first term, , comes from the sliding window: each of the positions attends to neighbors. The second term, , comes from global tokens: each global token attends to all positions (adding connections) and each of the positions attends to all global tokens (adding another connections). Both terms are linear in , while full attention costs . So the savings relative to full attention are:
For , , and : savings . As grows, the savings percentage approaches 100% for any fixed and .
Worked Example: Tracing Information Through a Document
To make the mechanics concrete, let us trace how a specific piece of information travels through a Longformer-style model processing a 20-token document. We will use a single global CLS token and a window size of 4.
Consider this document (indexed 0-19): "The cat [CLS] sat on the mat near the old red barn where the white horse lived yesterday."
Wait, that framing is slightly off: the CLS token is always at position 0 in Longformer, and the actual text tokens follow. Let us say the document is: [CLS] The cat sat on the mat near the old red barn where the white horse lived yesterday . (positions 0 through 17).
Suppose we want to understand how the final token yesterday at position 17 can be influenced by the word cat at position 2. With a window size of 4, yesterday (position 17) directly sees positions 15 through 17 and the CLS token at position 0. It cannot directly see position 2.
In layer 1, cat at position 2 attends to positions 0 through 4 (its local window) plus the CLS token. It contributes to the CLS token's representation by being in the CLS token's attention scope (since CLS attends to all positions). So after layer 1, the CLS token's representation has absorbed information from cat.
In layer 2, yesterday at position 17 attends to its local window (positions 15-17) and the CLS token. The CLS token now carries information derived from cat. So yesterday's representation after layer 2 is shaped by cat's information, mediated through the CLS token. The path is: cat (layer 1) to CLS, then CLS (layer 2) to yesterday. Two hops, two layers.
Let us verify this computationally:
# Worked example: trace information flow in a 18-token document
doc_len = 18
doc_window = 4
doc_global = [0] # CLS at position 0
cat_pos = 2
yesterday_pos = 17
# Layer 1 mask
layer1_mask = create_longformer_attention_mask(doc_len, doc_window, doc_global)
# Can 'cat' reach 'CLS' in layer 1?
cat_can_see_cls_l1 = layer1_mask[cat_pos, 0]
cls_can_see_cat_l1 = layer1_mask[0, cat_pos] # CLS has global attention
# Can 'yesterday' reach 'CLS' in layer 1?
yesterday_can_see_cls_l1 = layer1_mask[yesterday_pos, 0]
# After layer 2: yesterday has seen CLS, which has seen cat
# So information path: cat -> CLS (layer 1) -> yesterday (layer 2)
path_exists_in_2_layers = cls_can_see_cat_l1 and yesterday_can_see_cls_l1=== Worked Example: 18-token document, window=4, CLS at position 0 === Target path: 'cat' (pos 2) -> 'yesterday' (pos 17) Direct distance: 15 positions Layer 1 direct connections: cat (pos 2) can see CLS: True CLS can see cat (pos 2): True yesterday (pos 17) can see CLS: True After 2 layers, 'cat' can reach 'yesterday': True Positions visible to each key token in layer 1: 'cat' (pos 2) attends to: [0, 1, 2, 3, 4] 'yesterday' (pos 17) attends to: [0, 15, 16, 17] 'CLS' (pos 0) attends to: all 18 positions
The output confirms the 2-hop path. In layer 1, cat is in the CLS token's global attention scope. This contributes to CLS's updated representation. In layer 2, yesterday attends to CLS, which now encodes information from cat. The 15-position gap between cat and yesterday is bridged in exactly 2 layers.
The guarantee has a practical consequence: a Longformer model with only 2 transformer layers could, in principle, model dependencies between any two tokens regardless of distance, provided those dependencies route through the global token. In practice, models use many more layers to build rich, multi-hop reasoning chains through the global tokens.

Global-Local Attention Mixing
The interplay between global and local attention creates interesting information flow patterns. In a single attention layer, information moves in two streams. The local stream propagates through sliding windows, carrying fine-grained positional information about neighboring context. The global stream broadcasts through designated tokens. This provides sequence-wide context without the positional specificity of local attention.
These two streams are not independent. Local tokens contribute to the global token's representation at each layer, and the global token's representation then flows back to local tokens. Over multiple layers, this creates a feedback loop: local context informs the global summary, which then enriches local representations, which then further refine the global summary. The process converges to a state where both local and global representations are mutually informed by each other, a property that pure sliding window models cannot achieve.
The information flow through layers is cumulative in a precise sense. A token's effective receptive field, meaning the set of all positions that can influence its representation, grows with each layer. After one layer with a window of size and a single global token, a token's receptive field includes its -neighborhood plus the global token's information (which encompasses the full sequence). After two layers, the token can also see through the global token's two-layer receptive field, and through the local neighbors' one-layer receptive fields. The receptive field does not grow linearly: it jumps to near-full coverage after the first layer due to the global token's presence, then deepens in quality over subsequent layers.
def simulate_information_flow(
seq_len, window_size, global_positions, num_layers
):
"""
Simulate how information spreads through global-local attention.
Returns reachability matrix: which positions can influence which after L layers.
"""
# Start with direct connections from attention mask
mask = create_longformer_attention_mask(
seq_len, window_size, global_positions
)
reachable = mask.copy()
layer_snapshots = [reachable.copy()]
# Propagate through layers
for _ in range(num_layers - 1):
# Position j can reach position i if there's any intermediate k
# where j reaches k and k reaches i
new_reachable = reachable @ reachable > 0
reachable = reachable | new_reachable
layer_snapshots.append(reachable.copy())
return layer_snapshots
# Simulate: 32 tokens, window 8, CLS at 0
seq_len = 32
window_size = 8
global_positions = [0]
snapshots = simulate_information_flow(
seq_len, window_size, global_positions, num_layers=3
)


The visualization reveals how quickly information propagates. After one layer, connectivity is limited to local neighborhoods plus the global token. By layer 2, the bidirectional global token has served as a relay and every position can reach every other position. Layer 3 remains fully connected; its value is representational refinement, not additional graph reachability.
This suggests that even with aggressive sparsity, a few layers of global-local attention can match the connectivity of full attention. The key insight is that global tokens provide "shortcuts" that reduce the effective diameter of the attention graph.
The jump between layer 1 and layer 2 connectivity is particularly striking. After one layer, only about 34% of position pairs can reach each other. After two layers, that number reaches 100%. This discontinuous jump is a direct consequence of the bidirectional global token. Without it, the window would expand only incrementally and still miss distant pairs. The global token converts slow growth in connectivity into full two-hop coverage.
Implementation Strategies
Implementing global tokens efficiently requires careful attention to the attention computation itself. The naive approach computes global and local attention separately, then merges results. A more efficient approach uses a single sparse attention kernel that handles both patterns in one pass.
The reason efficiency matters here is not just academic. The whole point of global tokens is to add minimal overhead to an already-efficient sliding window attention. If the implementation of global attention adds significant constant-factor overhead, it can undermine the efficiency gains of the sparse pattern. Production implementations in frameworks like HuggingFace Transformers handle this by computing global attention using dense matrix operations for the small set of global positions, while using specialized CUDA kernels for the local window attention for the remaining positions. The two results are then combined.
For a pure-Python reference implementation, the dense approach is simpler and still instructive. The key step is applying the attention mask before softmax: setting non-attended positions to negative infinity ensures that those positions contribute zero weight in the softmax output, effectively excluding them from the weighted average.
def global_local_attention(
query, key, value, window_size, global_positions, d_k
):
"""
Compute global-local attention efficiently.
Args:
query, key, value: Arrays of shape (seq_len, d_model)
window_size: Size of local attention window
global_positions: List of positions with global attention
d_k: Dimension for scaling
Returns:
output: Attended values of shape (seq_len, d_model)
attention_weights: Sparse attention matrix
"""
seq_len = query.shape[0]
# Create attention mask
mask = create_longformer_attention_mask(
seq_len, window_size, global_positions
)
# Compute raw attention scores
scores = query @ key.T / np.sqrt(d_k)
# Apply mask: set non-attended positions to -inf
masked_scores = np.where(mask, scores, -np.inf)
# Softmax (row-wise)
exp_scores = np.exp(
masked_scores - masked_scores.max(axis=1, keepdims=True)
)
# Handle -inf: exp(-inf) = 0
exp_scores = np.nan_to_num(exp_scores, nan=0.0, posinf=0.0, neginf=0.0)
attention_weights = exp_scores / (
exp_scores.sum(axis=1, keepdims=True) + 1e-9
)
# Compute output
output = attention_weights @ value
return output, attention_weightsThe scoring step computes the full matrix of dot products, then zeroes out the positions that do not appear in the mask. The formula is:
where:
- : the attention score from query position to key position
- : the query vector at position
- : the key vector at position
- : the key dimensionality (used for scaling to prevent vanishing gradients)
- : 1 if position is allowed to attend to position , 0 otherwise
Setting masked positions to before the row-wise softmax is the standard trick: since , those positions contribute zero to the softmax denominator and produce zero attention weight. The numerical stability step of subtracting the row maximum before exponentiation (the log-sum-exp trick) prevents overflow for the non-masked positions.
The mask determines which attention scores to compute. Setting masked positions to negative infinity before softmax effectively zeros their contribution to the weighted average. The global tokens see the full sequence in their rows, while local tokens see only their windows plus global positions.
Let us verify the implementation with a small example:
# Test the implementation
seq_len = 16
d_model = 64
d_k = 64
window_size = 4
global_positions = [0, 8] # CLS and mid-sequence global token
# Random queries, keys, values
Q = np.random.randn(seq_len, d_model)
K = np.random.randn(seq_len, d_model)
V = np.random.randn(seq_len, d_model)
output, weights = global_local_attention(
Q, K, V, window_size, global_positions, d_k
)Input shape: (16, 64) Output shape: (16, 64) Attention weights shape: (16, 16) Active connections per position: Position 0 (global): 16 connections Position 1 (local ): 5 connections Position 2 (local ): 6 connections Position 3 (local ): 7 connections Position 4 (local ): 7 connections Position 5 (local ): 7 connections Position 6 (local ): 6 connections Position 7 (local ): 6 connections Position 8 (global): 16 connections Position 9 (local ): 6 connections Position 10 (local ): 6 connections Position 11 (local ): 7 connections Position 12 (local ): 7 connections Position 13 (local ): 7 connections Position 14 (local ): 6 connections Position 15 (local ): 5 connections
Global positions connect to all 16 tokens. Local positions connect to their window (4-5 tokens) plus the 2 global tokens, totaling around 6-7 connections each.
Notice the asymmetry in connection counts. Global tokens at positions 0 and 8 have 16 connections each, one per token. Local tokens have 6-7 connections. This asymmetry is intentional and algorithmically motivated: the global tokens need to see everything to serve as effective relays, while local tokens only need to see their neighborhood and the global hubs. The total number of connections is much smaller than the required by full attention.

The attention weight heatmap shows the characteristic pattern. The global token rows (0 and 8) have distributed attention across all positions. The local token rows show concentrated attention within their windows, with visible bumps at columns 0 and 8 where they attend to global tokens.
In practice, the attention weights in the global rows are not uniform. The global token uses the standard scaled dot-product attention mechanism to decide how much to attend to each position, so some positions receive more attention than others based on their semantic relevance. A global CLS token does not treat every word equally: it attends more to content words and salient phrases than to function words and punctuation. This learned non-uniformity is what makes global tokens more powerful than a simple average pooling operation would be.
Comparison: Different Global Token Strategies
Different architectures use global tokens in distinct ways. Understanding these variations helps when choosing or designing efficient attention mechanisms for a specific task. The variety of approaches reflects the fact that "which tokens should be globally visible" is fundamentally a domain knowledge question, and different tasks come with different structural priors about which positions carry globally relevant information.
The design space can be organized along two axes: whether the global token set is determined by position (positional strategies) or by semantics (semantic strategies), and whether it is fixed before training (static strategies) or discovered during training (learned strategies). CLS-only is positional and static. Task-specific tokens like question tokens are semantic and static. Periodic tokens are positional and static. Learned latents are semantic and dynamic, because their effective content, though not their positions, is determined by training.
| Strategy | Global Tokens | Use Case | Advantages |
|---|---|---|---|
| CLS-only (Longformer) | First token | Classification | Simple, single aggregation point |
| Task-specific (QA) | Question tokens | Extractive tasks | Semantic alignment with task structure |
| Periodic (BigBird) | Every k-th token | General | Uniform coverage, no special tokens needed |
| Learned latents (Perceiver) | Separate buffer | Multimodal, long sequences | Flexible, task-agnostic |
Each strategy makes trade-offs. CLS-only is simplest but creates a single bottleneck. Task-specific requires knowing the task structure. Periodic global tokens offer even coverage but may waste capacity on uninformative positions. Learned latents are most flexible but add trainable parameters and require the model to discover the global structure from scratch, which can slow convergence on small datasets.
The periodic strategy, used in BigBird, deserves particular attention because it offers a theoretical guarantee absent from the others. BigBird proved that its combination of random and local attention, along with global attention, is a universal approximator of sequence functions in the sense that it can approximate any function on sequences that a full-attention transformer can compute. The periodic global tokens in BigBird act as the "global" component in this theoretical framework. This ensures that there are always positions that can see the full sequence regardless of where the random attention heads happen to land.
def compare_global_strategies(seq_len, window_size):
"""Compare connectivity of different global token strategies."""
strategies = {
"CLS only": [0],
"CLS + SEP": [0, seq_len - 1],
"Periodic (every 8)": list(range(0, seq_len, 8)),
"First 4 tokens": list(range(4)),
}
results = []
for name, global_pos in strategies.items():
mask = create_longformer_attention_mask(
seq_len, window_size, global_pos
)
density = mask.sum() / (seq_len * seq_len)
complexity = compute_attention_complexity(
seq_len, window_size, len(global_pos)
)
results.append(
{
"strategy": name,
"num_global": len(global_pos),
"density": density * 100,
"savings": complexity["savings_percent"],
}
)
return results
# Compare strategies
seq_len = 64
window_size = 8
comparison = compare_global_strategies(seq_len, window_size)Comparison: seq_len=64, window_size=8 Strategy # Global Density Savings -------------------------------------------------------- CLS only 1 16.5% 84.4% CLS + SEP 2 19.3% 81.3% Periodic (every 8) 8 33.9% 64.1% First 4 tokens 4 24.8% 75.4%
All strategies maintain significant savings compared to full attention. The periodic strategy with 8 global tokens has higher density (more connections) but still achieves over 80% savings. The choice depends on whether you need uniform global coverage (periodic) or task-aligned global positions (CLS, question tokens).
The density column is the most informative for understanding the attention pattern's sparsity. A density of 15% means that only 15% of all possible token pairs attend to each other. This sparsity is what makes the attention computable in linear time: the 85% of pairs that do not attend need not be computed at all. Notice that even the densest strategy in the comparison (Periodic with every-8th token as global) achieves over 80% savings, confirming that the linear cost of global tokens is small relative to the quadratic cost of full attention.
In Practice: Deployment Considerations
Understanding global tokens conceptually is one thing; deploying them in production systems raises additional concerns that are worth addressing directly.
The most common practical question is whether to use global tokens at all, or to simply use a larger sliding window. The answer depends on the nature of the long-range dependencies in your data. If dependencies are predominantly local with occasional long-range exceptions, a larger window may be simpler and sufficient. If the task requires global context, such as document classification, where the final label depends on the document's overall topic rather than any local phrase, global tokens are the right choice. The failure mode of a too-small window is subtle and hard to diagnose: the model will appear to work but will systematically fail on documents where the key information happens to fall outside the window of the tokens being queried.
A second practical consideration is where to place global tokens in the sequence. Most implementations follow BERT's convention of placing the CLS token at position 0. This is fine for classification, but for generation tasks it may be better to place the global token at the end of the input sequence (position ), since decoder-style models typically generate left-to-right and the final position has access to all preceding context through causal attention anyway.
For fine-tuning pre-trained models on long-document tasks, the standard approach is to extend a BERT-like model by replacing its full self-attention with Longformer-style global-local attention. The sliding window attention for local positions can be initialized from the original attention weights, since attending to a nearby window is a special case of attending to the full sequence. The global attention for the CLS token is initialized similarly: at the start of fine-tuning, the global attention weights are identical to the original full-attention weights, and fine-tuning gradually specializes them for the long-document task.
Memory is another practical concern. Even though global-local attention has linear computational complexity, naive dense matrix implementations still allocate memory. Production implementations use specialized sparse attention kernels, such as the sliding window attention kernel in the HuggingFace Longformer implementation, to achieve the memory footprint that the theoretical complexity implies. Without these kernels, you may find that your GPU runs out of memory for long sequences even though the computation itself would be fast enough.
Finally, gradient flow deserves attention during training. Because the global token positions receive gradients from all other positions in the sequence, their parameter updates can be much larger in magnitude than updates to local token parameters. This imbalance can destabilize training if left unaddressed. The standard mitigation is to use a lower learning rate for global attention parameters, or to clip gradients separately for global and local attention components. Alternatively, using a warmup schedule where the model starts with a larger effective window and gradually reduces it to the target window size can help the global attention mechanism learn its role before being exposed to extremely long-range dependencies.
Limitations and Impact
Global tokens solve the long-range dependency problem in sparse attention, but they come with limitations that become more pronounced as sequence lengths and task complexities grow.
The primary limitation is the potential information bottleneck. A single CLS token must compress an entire document's worth of information into one vector of dimension , typically 768 or 1024 in standard BERT-scale models. For a short paragraph, this is entirely feasible: the CLS token can capture the main topic and sentiment, plus key entities, without strain. For a 10,000-word document covering multiple topics, the same fixed-size vector must somehow encode everything that might be globally relevant, and something will inevitably be lost.
This bottleneck manifests in a specific failure mode: global tokens tend to saturate on frequent, high-salience features and lose less prominent information. In a legal document, the CLS token might learn to encode the parties, the date, and the type of agreement, but lose track of specific clause exceptions buried in the body text. Those exceptions may be precisely what a downstream model needs to answer a specific legal question. The solution, adding more global tokens, works but increases the number of parameters receiving full-sequence gradients, which in turn increases the effective computation.
Global tokens also introduce asymmetry in the attention pattern. The CLS token sees everything, while other tokens see only windows plus global positions. This asymmetry can lead to uneven gradient flow during training, with global token parameters receiving disproportionately large updates. In a transformer with many layers, this can cause the global token's representation to drift toward patterns that are easy to learn but not maximally informative for the task. Careful initialization and learning rate scheduling help mitigate this issue, but they add complexity to the training pipeline.
A third limitation is that global tokens require the practitioner to make a design decision before training: how many global tokens, and where? For well-studied tasks like classification and extractive QA, the answer is relatively clear. For novel tasks or tasks where the relevant global context is not known in advance, choosing the wrong global token configuration can hurt performance. Learned global tokens side-step this problem but require more data and longer training to discover good configurations.
Despite these limitations, global tokens work remarkably well in practice. Longformer achieved state-of-the-art results on long-document tasks like WikiHop and TriviaQA when it was introduced in 2020. The ability to process 4096 tokens efficiently, compared to BERT's maximum of 512, opened new applications in document understanding, long-form question answering, and summarization. BigBird extended this further to 4096 and beyond, with theoretical guarantees about universal approximation that provided a principled foundation for the empirical successes of sparse global-local attention.
The conceptual impact may be even more significant than the benchmark numbers. Global tokens demonstrate that carefully designed sparse patterns can match or exceed the performance of full attention while being much more efficient. This principle has influenced subsequent architectures, from Perceiver's latent arrays to modern retrieval-augmented models that retrieve relevant context rather than attending to everything. The insight that "not everything needs to attend to everything, but some things need to attend to everything" is simple to state but took the field several years and multiple architectures to fully internalize and exploit. It continues to shape how researchers design attention mechanisms for long-context language models today.
Key Parameters
When implementing global-local attention, several parameters control the trade-off between efficiency and expressiveness:
-
window_size: The number of tokens each position can attend to locally. Larger windows capture more local context but increase computation. Typical values range from 256 to 512 for long-document models. The window should be large enough to capture phrase-level dependencies. Setting the window too small forces the model to rely heavily on global tokens for even moderately long-range dependencies, which can overwhelm the global token capacity. -
global_positions: A list of token indices designated as global. Common strategies include:- First position only (CLS token) for classification
- First and last positions (CLS + SEP) for sequence-pair tasks
- All question tokens for extractive QA
- Every k-th position for uniform coverage
-
num_global: The count of global tokens. More global tokens reduce the information bottleneck but add connections where is the global count and is sequence length. Start with 1-4 global tokens and increase if task performance plateaus. -
d_k: The dimension used for scaling attention scores. Standard practice sets this equal to the model's head dimension (typically 64). Scaling by keeps attention logits in a stable range regardless of dimensionality.
The relationship between window size and number of global tokens defines the fundamental trade-off in global-local attention design. A large window with few global tokens provides rich local context and minimal global routing capacity. A small window with many global tokens provides tight local focus and rich global communication bandwidth. Most practical architectures land somewhere in the middle: a moderately large window (256-512) with a small number of global tokens (1-8), sized to cover the task's specific global information requirements.
Summary
Global tokens bridge efficient local attention and effective long-range modeling. By designating certain positions as communication hubs with full attention span, they reduce the maximum path length between any two positions to just 2 hops, achieving near-full-attention connectivity at linear computational cost.
Key takeaways:
- The bottleneck problem: Sliding window attention limits direct communication to local neighborhoods, requiring many hops for long-range information flow.
- Global tokens as hubs: Designated positions that can attend to and be attended by all positions, serving as relay points for sequence-wide communication.
- CLS token attention: BERT's classification token naturally fits the global attention role, aggregating sequence information for downstream tasks.
- Task-specific globals: Question tokens in QA, separator tokens in sentence pairs, or any semantically meaningful positions can serve as global tokens.
- Learned latents: Trainable global tokens that discover what information to aggregate. This gives flexibility at the cost of additional parameters.
- Efficient scaling: Global token count scales linearly, so even many global tokens preserve the efficiency gains of sparse attention.
- Reduced path length: With global tokens, maximum path length drops from to 2, where is the sequence length and is the window size. This matches full attention's connectivity properties.
- Deployment considerations: Production use requires sparse attention kernels for true linear memory, careful gradient management for global token parameters, and thoughtful decisions about global token placement and count for the specific task.
The next chapter examines Longformer, which combines sliding window attention with global tokens into a complete architecture. We will see how these components work together to achieve strong performance on document-level NLP tasks while maintaining linear complexity in sequence length.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about global tokens in efficient transformers.
Global Tokens Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!