Part of Language AI Handbook
Explains how Longformer combines sliding window and global attention to process documents of 4,096+ tokens with O(n) complexity instead of O(n²).
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Longformer
Standard transformer attention scales quadratically with sequence length, making it prohibitively expensive for documents that span thousands of tokens. If you want to process an entire research paper, a legal contract, or a book chapter, vanilla attention becomes a memory and compute bottleneck long before you reach the interesting parts of the text.
To see why this is a problem in practice, consider what happens when you increase the sequence length from 512 tokens to 4,096 tokens. With full attention, each of the 4,096 tokens must compute an attention score against every other token, producing over 16 million attention weights per layer. On a GPU with 16 GB of memory, this easily saturates available resources before you even add the residual connections, feedforward layers, and batch dimension. Standard BERT was trained with 512-token sequences for exactly this reason: it was the longest sequence that remained feasible given typical hardware constraints in 2018.
This limitation matters enormously for real-world text. A legal contract averages 10,000 words. A scientific paper runs 8,000 to 12,000 words. A court opinion might span 20,000 words. When you truncate these documents to 512 tokens, you throw away roughly 90% of the content. A model analyzing whether a contract clause conflicts with another clause mentioned 8,000 tokens earlier simply cannot make that connection if both sections are chopped off.
Longformer addresses this by combining two complementary attention patterns: local sliding window attention for capturing nearby context, and global attention for aggregating information across the entire sequence. This hybrid approach reduces complexity from to , where is the sequence length. In practical terms, doubling the sequence length with standard attention quadruples the cost, but with Longformer it only doubles the cost. The result is a transformer that can handle sequences of 4,096 tokens or more without the memory explosion that would occur with full attention.
The key insight is that the quadratic cost of full attention is largely wasted. Not every pair of tokens needs a direct attention connection. Language has structure: meaning flows locally through sentences and paragraphs, and only a small number of special tokens (like document-level question markers or classification anchors) need to see the entire sequence. By designing an attention pattern that reflects this structure, Longformer recovers efficiency without sacrificing the representational power needed for long-document understanding.
BERT's 512-token limit was not a fundamental architectural choice but a practical constraint imposed by the quadratic cost of self-attention. Researchers working on scientific literature analysis, legal document review, and clinical note processing routinely encountered documents that exceeded this limit and resorted to heuristics such as chunking documents into overlapping segments and ensembling predictions across chunks. This produced brittle systems that could not reason across chunk boundaries. Longformer, introduced by Beltagy and Peters along with Cohan at the Allen Institute for AI in 2020, was motivated directly by the needs of scientific literature analysis and addressed the long-document limitation in a principled way that required no chunking heuristics.
The Core Insight: Local + Global Attention
Longformer's key innovation is recognizing that not every token needs to attend to every other token. Most language understanding happens locally. When reading a sentence, the meaning of a word primarily depends on its immediate neighbors. The subject of a sentence matters for its verb, but a word in paragraph 15 rarely affects the interpretation of a word in paragraph 3.
Think of it this way: when you read the word "therefore" in a paragraph, you look left to understand the cause and right to understand the conclusion. You do not simultaneously scan back to the title page or forward to the appendix. Human reading is overwhelmingly local. Longformer formalizes this observation into an attention pattern where each token has a fixed-width spotlight that illuminates only its immediate neighborhood, rather than broadcasting attention to every other token in the document.
However, some tokens need global reach. A question token in question answering must see the entire context to find the answer. A classification token needs to aggregate information from the whole document. Longformer solves this by treating these global tokens differently from the rest of the sequence. Rather than giving all tokens equal capabilities, it partitions them into two classes with different attention behaviors tailored to their roles.
Notice that this design closely mirrors how humans process documents. When reading a research paper, you read each sentence in context with its neighbors. But certain anchoring concepts, like the research question stated in the abstract, remain globally relevant throughout. Longformer's global tokens play the same role as these anchoring concepts: they provide document-level context that every local token can access.
Longformer is a transformer model designed for long documents that combines sliding window attention (local context) with global attention (full sequence access) to achieve linear complexity in sequence length while maintaining the ability to model long-range dependencies.
The architecture defines two types of attention:
- Sliding window attention: Each token attends only to tokens within a fixed window around it. With window size , token attends to positions
- Global attention: Selected tokens attend to all tokens in the sequence, and all tokens attend back to them
This combination allows Longformer to handle documents that would overwhelm standard transformers, while still capturing the relationships that matter for understanding. The elegance of the design is that neither component requires the other to function: you could use sliding window attention alone for tasks with only local dependencies, or use global attention alone for tasks requiring full-sequence reasoning. The hybrid approach earns the best of both worlds at a cost close to the cheaper option.
Mathematical Formulation
To understand why Longformer works, we need to think carefully about what attention computes and why the standard approach becomes expensive. This journey from intuition to formal mathematics will reveal Longformer's mechanics and the deeper insight that makes efficient attention possible.
Understanding the math is not just academic. The formulas encode the design decisions that determine what relationships the model can and cannot learn. Once you see the equations, the tradeoffs in window size selection, global token placement, and the need for separate projection matrices all become transparent.
The Problem: Why Standard Attention Breaks Down
Recall that standard transformer attention computes, for each token, a weighted average of all other tokens in the sequence. For a token at position , we ask: "How relevant is every other position to understanding this token?" The answer comes as a set of attention weights that sum to 1.
The issue is that this question scales poorly. If you have tokens, each token must compute attention scores, giving us total computations. For a 512-token sequence, that's about 260,000 attention weights per layer. For a 4,096-token document, it explodes to over 16 million. This quadratic scaling is what makes long documents prohibitively expensive.
In practice, memory is often the binding constraint before compute. Attention weights must be stored during the forward pass to allow gradient computation during backpropagation. A single attention head over a 4,096-token sequence stores a matrix of float32 values, requiring 64 MB of GPU memory. A 12-layer model with 12 heads per layer requires around 9 GB for attention weights alone, leaving almost no headroom for activations, gradients, or optimizer states. Reducing the number of stored attention weights is therefore the most direct path to longer sequences.

Most of these computations are wasted. When you read the word "cat" in a sentence, you don't need to simultaneously consider a word from three paragraphs away to understand its meaning. Most language understanding happens locally. Longformer exploits this observation by restricting which tokens each position can attend to.
Dividing the Sequence: Local and Global Tokens
Longformer partitions the sequence into two distinct sets:
- Local tokens : The vast majority of tokens, which only need to see their immediate neighborhood
- Global tokens : A small number of special tokens that need to aggregate information from the entire sequence
For a sequence of length , we typically have local tokens and global tokens. This asymmetry is the source of Longformer's efficiency: we do expensive full-sequence attention for only a handful of tokens, while the rest use a much cheaper local pattern.
Sliding Window Attention: The Local Pattern
For local tokens, we define a window of size centered on each position. Instead of attending to all tokens, position only attends to positions within the range . Think of it as a spotlight that follows each token, illuminating only its immediate neighborhood.
This restriction reflects a property of language: most syntactic and semantic relationships are local. Subjects tend to be near their verbs. Adjectives modify nearby nouns. Pronouns typically refer to recently mentioned entities. By focusing attention on a fixed-size window, we capture the dependencies that matter most while ignoring distant positions that rarely contribute meaningful information.
In practice, the window size is a critical hyperparameter. The original Longformer paper uses for most experiments, matching the maximum sequence length of standard BERT. This means each token can see up to 256 positions to its left and 256 to its right, which is sufficient to capture most sentence-level and paragraph-level syntactic dependencies. Smaller windows (like 128 or 256) save memory and compute but may miss longer-range dependencies within a paragraph. Larger windows (like 1024) capture more context but reduce the sparsity gains that make Longformer efficient. The right window size depends on your task and the typical length of relevant local contexts within your documents.
The formal computation for a local token at position is:
Let's unpack each component:
-
: The query vector for position . This is what the token is "asking for" from other positions. We compute it by projecting the input embedding through a learned weight matrix .
-
: The key matrix containing only keys for tokens in the local window. This is a matrix of shape where each row is the key vector for one position in the window. Keys represent what each position "offers" to other tokens.
-
: The value matrix for the same window, also shape . Values contain the actual information that gets aggregated. The attention weights determine how much of each value to include in the output.
-
: The dimension of key and query vectors, typically 64 in base transformer models.
-
: A scaling factor that prevents attention scores from growing too large. Without this, dot products between high-dimensional vectors can become very large, causing softmax to produce nearly one-hot distributions that are hard to learn from.
The computation proceeds through three carefully designed steps:
-
Score computation: We compute , which produces a vector of scores. Each score measures how well the query at position matches each key in the window. Geometrically, this is a dot product: vectors pointing in similar directions produce high scores.
-
Normalization: We divide by to stabilize the magnitude, then apply softmax to convert raw scores into a probability distribution. The weights now sum to 1, telling us what fraction of attention to allocate to each position.
-
Aggregation: We compute a weighted average of the value vectors using these probabilities. Positions with high attention weights contribute more to the output; positions with low weights contribute less.
The critical complexity reduction comes from step 1: instead of computing scores, we compute only scores. Since is a constant (typically 256 or 512), the cost per token is instead of . Across all tokens, the total complexity is , which is linear in sequence length.
Global Attention: Bridging Distant Positions
Sliding window attention solves the efficiency problem, but it creates a new one: how does information travel between distant parts of the sequence? If token 1 can only see tokens 1-256, and token 4000 can only see tokens 3744-4000, how do they ever communicate?
In a standard transformer, information propagates through repeated local interactions across layers. After layers with window size , a token's effective receptive field covers positions. With 8 layers and window size 512, a token can be influenced by content up to 4,096 positions away, but this requires routing information through many intermediate tokens before it arrives. The signal may degrade or get mixed with irrelevant context along the way.
The answer in Longformer is global tokens. These are special positions that break the local-only rule. A global token attends to every position in the sequence, and every position attends back to it. Think of global tokens as relay stations: information from any part of the sequence can flow through them to reach any other part in a single attention step rather than through many layers of local propagation.
Global tokens are task-specific designations, not architectural constants. For a classification task, you designate the [CLS] token as global because it needs to aggregate the entire document before predicting a label. For question answering, you designate all question tokens as global because the model must scan the entire passage to locate the answer. For coreference resolution, you might designate entity mentions as global to allow the model to directly compare them regardless of how far apart they appear. This flexibility is one of Longformer's most practical advantages over methods that fix the sparse attention pattern at architecture design time.
For a global token at position , the attention computation spans the entire sequence:
where:
- : The query vector for the global token, computed using a separate projection matrix
- : The key matrix for all tokens, shape
- : The value matrix for all tokens, shape
Global tokens use different projection matrices (, , ) than local tokens (, , ) because local and global attention serve fundamentally different purposes.
When a local token computes attention, it is asking: "Which of my neighbors help me understand my immediate context?" When a global token computes attention, it is asking: "What are the most important pieces of information across the entire document?" These are different questions, and the model benefits from learning different representations to answer them.
The separation also matters for the tokens attending to global positions. When a local token attends to a global token, it should receive a representation tailored for aggregation, not one optimized for local context understanding. The separate projection matrices allow this flexibility. Without them, the model would need a single set of projections that performs well at both local syntactic reasoning and global semantic aggregation simultaneously. In practice, sharing projections forces a compromise that reduces performance on both tasks. The separate projections roughly double the parameter count in attention layers, but this cost is worthwhile for the quality improvement they provide.
The separate global projections also help with training stability. Because global tokens see vastly more positions than local tokens, the gradients flowing back through them have a different character: they average information from across the entire sequence rather than from a small local neighborhood. Using separate parameters allows the optimizer to calibrate learning rates and weight updates appropriately for each regime without the two objectives interfering with each other.
The Symmetry Property
Global attention is bidirectional: global tokens attend to all positions, and all positions attend back to global tokens. This symmetry is essential for information flow. It is worth pausing to understand exactly what "bidirectional" means here and why it matters.
If global attention were one-directional, say global tokens attend to all positions but local tokens cannot attend back to global tokens, then local tokens would be unaware that a global token exists. A paragraph in the middle of a document would compute its local representations without incorporating any document-level signal. The classification token at position 0 would aggregate information from the full document, but the words in paragraph 7 would not have access to what the classification token "knows" about the document structure.
With bidirectional global attention, every local token's output representation includes a weighted combination of the global token representations as part of its attended context. This is the mechanism by which document-level information percolates down to even the most local tokens. A sentence in paragraph 7 can effectively "know" what the document's main question is because it attends to the [CLS] or question tokens that encode that information.

Consider document classification with a [CLS] token at position 0. The [CLS] token needs to see the entire document to make a classification decision, so it receives global attention. But equally important, every token in the document can "see" the [CLS] token and incorporate that global context into its own representation. This bidirectional flow ensures that even local tokens have access to document-level information, just mediated through the global tokens rather than computed directly.
The symmetry also provides a useful mental model for thinking about information paths in the network. Without global tokens, information takes a long path: it travels layer by layer through local windows, like a game of telephone where each intermediary can only talk to its neighbors. With global tokens, any token can communicate with any other token in two steps: first to a global token, then from that global token to the destination. The maximum information distance collapses from hops (proportional to sequence length divided by window size) to just 2 hops when a global token mediates the connection.
Combining the Patterns: The Complete Picture
Now we can see how Longformer attention works as a unified system:
-
Most tokens use sliding window attention, computing relevance only within their local neighborhood. This keeps the bulk of computation cheap.
-
A few tokens (typically [CLS], question tokens, or task-specific markers) use global attention, seeing the entire sequence. This maintains the ability to aggregate long-range information.
-
All tokens can attend to global tokens, even if they're outside the local window. This ensures information can flow from any position to any other position through the global intermediaries.
The result is a sparse attention pattern that covers the essential dependencies while skipping the redundant ones. Local relationships are captured directly. Global relationships are captured through designated relay tokens.



Complexity Analysis: Why It Works
Let's verify that this design achieves linear complexity. The total attention computation involves three components:
-
Local attention: Each of the tokens computes attention over a window of size . This contributes to the total cost.
-
Global tokens attending to all: Each of the global tokens attends to all positions. This contributes .
-
All tokens attending to global tokens: Each token includes global positions in its attention set. This cost is already absorbed into the local window computation (global tokens simply become part of the attended set).
The total complexity is:
where:
- : the sequence length, the variable that grows with document size
- : the sliding window size, a constant chosen at model design time (typically 512)
- : the number of global tokens, a constant chosen per task (typically 1 to 10)
The key observation is that is a constant. It doesn't grow with sequence length. This means the overall complexity is , linear in sequence length.
Compare this to standard attention's complexity. When you double the sequence length:
- Standard attention: cost quadruples ()
- Longformer attention: cost only doubles ()
This difference is what enables Longformer to process documents of 4,096 tokens or more on hardware that would run out of memory with standard attention at just 512 tokens.
Visualizing the Attention Pattern
The Longformer attention pattern creates a distinctive sparse structure when visualized as a matrix. The sparsity pattern is not random. It has a specific geometry: a diagonal band of width (the local window) with full rows and columns at the global token positions. When you see this pattern, you can immediately read off the design intent of the model. The density of the diagonal band tells you how much local context each token accesses. The number of full rows tells you how many global aggregation points the task requires. Building intuition for this visual representation will help you debug attention configurations and understand why certain configurations work better than others.
Let's build a visualization to understand this pattern in detail.
import numpy as np
def create_longformer_attention_mask(seq_len, window_size, global_positions):
"""
Create a Longformer-style attention mask.
Args:
seq_len: Total sequence length
window_size: Size of sliding window (must be even)
global_positions: List of positions that have global attention
Returns:
Attention mask of shape (seq_len, seq_len)
"""
mask = np.zeros((seq_len, seq_len))
half_window = window_size // 2
# Add sliding window attention for all positions
for i in range(seq_len):
start = max(0, i - half_window)
end = min(seq_len, i + half_window + 1)
mask[i, start:end] = 1
# Add global attention: global tokens attend to all, all attend to global
for g in global_positions:
mask[g, :] = 1 # Global token attends to all
mask[:, g] = 1 # All tokens attend to global token
return mask
The visualization reveals the structure of Longformer attention. The diagonal band represents sliding window attention: each position attends to its local neighborhood. The vertical and horizontal stripes at positions 0 and 32 show global attention: these tokens have bidirectional access to the entire sequence.
Notice how sparse this pattern is compared to full attention. Full attention would fill the entire square, requiring storage for attention weights (where is our sequence length in this example). The Longformer pattern only needs storage for the non-zero entries, which scales linearly with sequence length.
Comparing Attention Sparsity
Let's quantify the sparsity achieved by Longformer compared to full attention.
def compute_attention_density(seq_len, window_size, num_global):
"""Compute the fraction of non-zero attention weights."""
# Sliding window: each of n tokens attends to w tokens
sliding_window_entries = seq_len * window_size
# Global attention: each global token adds (n-1) new connections
# (excluding the window entries we already counted)
global_entries = num_global * (seq_len - window_size) * 2
# Total possible entries
total_possible = seq_len**2
# Approximate density (some overlap, but gives the right idea)
density = min(
1.0, (sliding_window_entries + global_entries) / total_possible
)
return density
# Compare densities at different sequence lengths
seq_lengths = [512, 1024, 2048, 4096, 8192, 16384]
window_size = 512
num_global = 2
densities = [
compute_attention_density(n, window_size, num_global) for n in seq_lengths
]
full_attention_memory = [n**2 for n in seq_lengths]
longformer_memory = [n * (window_size + 2 * num_global) for n in seq_lengths]Sequence Length | Attention Density | Memory Ratio (Longformer/Full)
-----------------------------------------------------------------
512 | 100.0% | 100.78%
1024 | 50.2% | 50.39%
2048 | 25.1% | 25.20%
4096 | 12.6% | 12.60%
8192 | 6.3% | 6.30%
16384 | 3.1% | 3.15%The numbers show the difference. At 512 tokens, Longformer uses about the same memory as full attention. But as sequence length grows, the savings become dramatic. At 4,096 tokens, Longformer uses only about 12% of the memory. At 16,384 tokens, it drops to around 3%. This is the power of linear versus quadratic scaling.
Window Size and Its Effect on Coverage
The window size is the single most consequential hyperparameter in Longformer. Choosing it well requires understanding the trade-off between local coverage, memory usage, and the number of layers needed for full-sequence reasoning. The following visualization illustrates how different window sizes produce different attention density and coverage profiles.

Notice that as window size increases, attention density rises and the number of layers needed for full-sequence coverage falls. A window of 64 tokens covers only 1.6% of a 4,096-token sequence per layer and would require 64 layers for a token to be influenced by the full document. A window of 512 tokens covers 12.5% per layer and requires only 8 layers for full propagation. A window of 2,048 tokens covers 50% but uses half as much memory as full attention, making it worth considering only for very dense tasks where far more than just local context matters.
Worked Example: Tracing Attention Through a Document
To make the abstract mechanics concrete, let's trace what happens when Longformer processes a short legal document. Suppose the document has 20 tokens representing the sentence: "[CLS] The buyer shall pay the seller within 30 days of delivery as specified in clause 5 of this agreement [SEP]". We use a window size of 4 and designate position 0 ([CLS]) as the only global token.
The first thing to establish is which positions each token can attend to. Token 0 ([CLS]) is global, so it attends to all 20 positions. Every other token also attends back to position 0, even if it is outside their local window. Token 1 ("The") can attend to positions 0 through 3 (the local window plus the out-of-window global token at position 0 is already covered since the window extends to position 0 here). Token 5 ("pay") can attend to positions 3 through 7 under its local window, plus position 0 because it is global. Token 14 ("clause") can attend to positions 12 through 16 locally, plus position 0 globally.
Now consider a question: does the word "days" (token 10) have any connection to the phrase "clause 5" (tokens 14-15)? Under pure sliding window attention with window size 4, token 10 can only see positions 8 through 12, so it has no direct view of tokens 14-15. However, both token 10 and tokens 14-15 attend to the [CLS] token at position 0. In the first layer, the [CLS] token aggregates a weighted combination of all 20 tokens. In the second layer, when token 10 attends to the [CLS] token, it receives a representation that has already absorbed information from "clause 5". The long-range connection is established in two attention hops rather than requiring the sequence to be short enough for direct attention.
This two-hop communication is the mechanism behind Longformer's ability to model long-range dependencies despite using a local attention window for most tokens. Notice that the quality of this connection depends heavily on the [CLS] token learning to preserve the relevant information about "clause 5" in its aggregated representation. This is why training with global tokens on appropriate tasks matters: the model must learn to use global tokens as effective information relay stations, which happens through exposure to tasks that require cross-document reasoning.
In practice, real documents are far longer than 20 tokens, and the two-hop path becomes even more valuable. A 4,096-token legal contract might require reasoning between a definition in section 1 and its application in section 8. Under pure local attention with window size 512, reaching from position 100 to position 3,800 would require at least 7 layers of propagation. With global attention on section headers or the [CLS] token, that connection becomes a two-hop path available from the very first layer.
Implementing Longformer Attention
With the mathematical foundation in place, let's translate these ideas into working code. Building the implementation from scratch solidifies understanding and reveals the practical choices that make Longformer work. We'll start with the simpler sliding window mechanism, then layer on global attention support.
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
class LongformerSlidingWindowAttention(nn.Module):
"""
Sliding window attention for Longformer.
Each position attends only to tokens within a fixed window.
"""
def __init__(self, embed_dim, num_heads, window_size):
super().__init__()
self.embed_dim = embed_dim
self.num_heads = num_heads
self.head_dim = embed_dim // num_heads
self.window_size = window_size
assert embed_dim % num_heads == 0, (
"embed_dim must be divisible by num_heads"
)
# Projections for local attention
self.q_proj = nn.Linear(embed_dim, embed_dim)
self.k_proj = nn.Linear(embed_dim, embed_dim)
self.v_proj = nn.Linear(embed_dim, embed_dim)
self.out_proj = nn.Linear(embed_dim, embed_dim)
def forward(self, x, attention_mask=None):
"""
Args:
x: Input tensor of shape (batch, seq_len, embed_dim)
attention_mask: Optional mask for padding tokens
Returns:
Output tensor of shape (batch, seq_len, embed_dim)
"""
batch_size, seq_len, _ = x.shape
# Project queries, keys, and values
q = self.q_proj(x)
k = self.k_proj(x)
v = self.v_proj(x)
# Reshape for multi-head attention
q = q.view(
batch_size, seq_len, self.num_heads, self.head_dim
).transpose(1, 2)
k = k.view(
batch_size, seq_len, self.num_heads, self.head_dim
).transpose(1, 2)
v = v.view(
batch_size, seq_len, self.num_heads, self.head_dim
).transpose(1, 2)
# For simplicity, we'll use a loop-based implementation
# Production code would use more efficient sparse operations
half_window = self.window_size // 2
scale = 1.0 / math.sqrt(self.head_dim)
outputs = []
for i in range(seq_len):
start = max(0, i - half_window)
end = min(seq_len, i + half_window + 1)
# Get local keys and values
local_k = k[:, :, start:end, :] # (batch, heads, window, head_dim)
local_v = v[:, :, start:end, :]
# Compute attention scores
query = q[:, :, i : i + 1, :] # (batch, heads, 1, head_dim)
scores = torch.matmul(query, local_k.transpose(-2, -1)) * scale
# Apply softmax and compute weighted sum
attn_weights = F.softmax(scores, dim=-1)
output = torch.matmul(attn_weights, local_v)
outputs.append(output)
# Concatenate outputs
output = torch.cat(outputs, dim=2) # (batch, heads, seq_len, head_dim)
output = (
output.transpose(1, 2)
.contiguous()
.view(batch_size, seq_len, self.embed_dim)
)
return self.out_proj(output)This implementation demonstrates the core sliding window mechanism. Notice how the loop over positions extracts only the local window of keys and values for each query. In production code, this loop would be replaced with efficient sparse matrix operations, but the conceptual structure remains the same: each position computes attention only over its local window, keeping memory usage bounded regardless of sequence length.
The key insight from the code is that the window size half_window determines how far each token can "see." Positions near the edges of the sequence have smaller effective windows (they can't attend to positions that don't exist), but this is handled gracefully by the max and min bounds.
Now let's extend this to the full Longformer attention with global token support:
class LongformerAttention(nn.Module):
"""
Full Longformer attention with both sliding window and global attention.
"""
def __init__(self, embed_dim, num_heads, window_size):
super().__init__()
self.embed_dim = embed_dim
self.num_heads = num_heads
self.head_dim = embed_dim // num_heads
self.window_size = window_size
# Local attention projections
self.q_proj = nn.Linear(embed_dim, embed_dim)
self.k_proj = nn.Linear(embed_dim, embed_dim)
self.v_proj = nn.Linear(embed_dim, embed_dim)
# Separate projections for global attention
self.q_proj_global = nn.Linear(embed_dim, embed_dim)
self.k_proj_global = nn.Linear(embed_dim, embed_dim)
self.v_proj_global = nn.Linear(embed_dim, embed_dim)
self.out_proj = nn.Linear(embed_dim, embed_dim)
def forward(self, x, global_attention_mask=None):
"""
Args:
x: Input tensor of shape (batch, seq_len, embed_dim)
global_attention_mask: Boolean mask where True indicates global attention
Shape: (batch, seq_len)
Returns:
Output tensor of shape (batch, seq_len, embed_dim)
"""
batch_size, seq_len, _ = x.shape
# Compute local projections
q_local = self.q_proj(x)
k_local = self.k_proj(x)
v_local = self.v_proj(x)
# Compute global projections
q_global = self.q_proj_global(x)
k_global = self.k_proj_global(x)
v_global = self.v_proj_global(x)
# Reshape for multi-head attention
def reshape(t):
return t.view(
batch_size, seq_len, self.num_heads, self.head_dim
).transpose(1, 2)
q_local = reshape(q_local)
k_local = reshape(k_local)
v_local = reshape(v_local)
q_global = reshape(q_global)
k_global = reshape(k_global)
v_global = reshape(v_global)
scale = 1.0 / math.sqrt(self.head_dim)
half_window = self.window_size // 2
# Determine which positions have global attention
if global_attention_mask is None:
global_positions = []
else:
global_positions = (
global_attention_mask[0].nonzero(as_tuple=True)[0].tolist()
)
outputs = torch.zeros_like(q_local)
for i in range(seq_len):
is_global = i in global_positions
if is_global:
# Global token: attend to entire sequence using global projections
query = q_global[:, :, i : i + 1, :]
scores = torch.matmul(query, k_global.transpose(-2, -1)) * scale
attn_weights = F.softmax(scores, dim=-1)
output = torch.matmul(attn_weights, v_global)
else:
# Local token: sliding window + attend to global tokens
start = max(0, i - half_window)
end = min(seq_len, i + half_window + 1)
# Gather local and global keys/values
local_indices = list(range(start, end))
attend_indices = list(set(local_indices + global_positions))
attend_indices.sort()
local_k = k_local[:, :, attend_indices, :]
local_v = v_local[:, :, attend_indices, :]
# Use global k,v for global positions within the attended set
for g in global_positions:
if g in attend_indices:
idx = attend_indices.index(g)
local_k[:, :, idx, :] = k_global[:, :, g, :]
local_v[:, :, idx, :] = v_global[:, :, g, :]
query = q_local[:, :, i : i + 1, :]
scores = torch.matmul(query, local_k.transpose(-2, -1)) * scale
attn_weights = F.softmax(scores, dim=-1)
output = torch.matmul(attn_weights, local_v)
outputs[:, :, i : i + 1, :] = output
# Reshape and project output
outputs = (
outputs.transpose(1, 2)
.contiguous()
.view(batch_size, seq_len, self.embed_dim)
)
return self.out_proj(outputs)The implementation reveals several important design decisions:
-
Separate projection matrices: We maintain two complete sets of Q, K, V projections. The local projections (
q_proj,k_proj,v_proj) are used for sliding window attention, while the global projections (q_proj_global,k_proj_global,v_proj_global) are used when global tokens are involved. -
Asymmetric handling: Global tokens use global projections for everything. Local tokens use local projections for local attention, but when attending to global tokens, they use the global key and value vectors. This asymmetry ensures consistent representations regardless of which token is doing the attending.
-
Dynamic attention sets: For local tokens, we compute the union of the local window and global positions. This ensures that even if a global token is outside the local window, every token can still attend to it.
Let's verify our implementation works correctly:
# Create a small deterministic test using the chapter-level plot seed
batch_size = 2
seq_len = 32
embed_dim = 64
num_heads = 4
window_size = 8
model = LongformerAttention(embed_dim, num_heads, window_size)
x = torch.randn(batch_size, seq_len, embed_dim)
# Mark positions 0 and 16 as global
global_mask = torch.zeros(batch_size, seq_len, dtype=torch.bool)
global_mask[:, 0] = True
global_mask[:, 16] = True
output = model(x, global_attention_mask=global_mask)Input shape: (2, 32, 64) Output shape: (2, 32, 64) Global positions: [0, 16] Window size: 8 Embedding dimension: 64 Number of attention heads: 4
The implementation correctly handles both local sliding window attention and global attention for designated tokens. The output maintains the same shape as the input (batch size 2, sequence length 32, embedding dimension 64), confirming that our attention mechanism preserves dimensionality while routing attention based on the global attention mask. The two global positions (0 and 16) receive full sequence attention, while all other positions use the sliding window of size 8.
Using Longformer from Hugging Face
For practical applications, you'll want to use the optimized Longformer implementation from Hugging Face. The production implementation differs substantially from our educational implementation in one critical way: it uses a dedicated CUDA kernel for sliding window attention that operates on chunked, diagonal-band sparse matrices rather than a Python loop over positions. This kernel achieves attention in time proportional to the number of non-zero entries rather than the full grid, which delivers the speed and memory benefits in practice. Our loop-based implementation demonstrates the correct logic but would not be fast enough for real workloads.
Let's see how to apply the Hugging Face Longformer to a document processing task. The interface closely mirrors standard transformer models, with one additional parameter: the global_attention_mask that specifies which tokens receive global attention.
from transformers import LongformerModel, LongformerTokenizer
# Load pre-trained Longformer
tokenizer = LongformerTokenizer.from_pretrained("allenai/longformer-base-4096")
model = LongformerModel.from_pretrained("allenai/longformer-base-4096")
# Sample long document (concatenated for demonstration)
document = (
"""
Machine learning has transformed how we process and understand text.
Traditional approaches relied on hand-crafted features and statistical methods.
Modern neural networks learn representations directly from data, enabling
unprecedented performance on tasks from translation to summarization.
The transformer architecture, introduced in 2017, revolutionized the field.
Self-attention allows models to capture relationships between any positions
in a sequence, overcoming the sequential limitations of recurrent networks.
However, the quadratic complexity of attention limited sequence lengths.
Longformer addresses this limitation through a hybrid attention pattern.
By combining local sliding window attention with strategic global attention,
it achieves linear complexity while maintaining the ability to model
long-range dependencies essential for document understanding.
"""
* 10
) # Repeat to create a longer document# Tokenize the document
inputs = tokenizer(
document,
return_tensors="pt",
max_length=4096,
truncation=True,
padding="max_length",
)
# Create global attention mask - set first token ([CLS]) to have global attention
global_attention_mask = torch.zeros_like(inputs["input_ids"])
global_attention_mask[:, 0] = 1 # [CLS] token has global attention
# Forward pass
with torch.no_grad():
outputs = model(
input_ids=inputs["input_ids"],
attention_mask=inputs["attention_mask"],
global_attention_mask=global_attention_mask,
)
# Get the [CLS] token representation for classification
cls_representation = outputs.last_hidden_state[:, 0, :]Document tokens: 1,522 Max sequence length: 4,096 [CLS] representation dimension: 768 Global attention positions: 1 token(s)
The document contains over 1,000 tokens, padded to Longformer's maximum sequence length of 4,096. Despite processing this long sequence, only a single token (the [CLS] token at position 0) has global attention. This token's 768-dimensional representation now contains information aggregated from the entire document, enabling downstream tasks like classification or similarity computation without the quadratic memory cost of full attention.
Configuring Global Attention for Different Tasks
The power of Longformer comes from its flexibility in configuring global attention. Different tasks benefit from different global attention patterns. This configurability is not a minor implementation detail; it is a fundamental design philosophy. Rather than hardcoding a fixed sparse attention structure, Longformer treats the global attention pattern as a task-specific hyperparameter that the practitioner controls. Getting this right has a significant impact on downstream performance.
The guiding principle for choosing global tokens is to ask: which tokens in my input absolutely must have access to the entire document to perform their role in the task? These are the tokens that should receive global attention. For tasks where the answer is "no specific tokens need full-sequence access," you can rely on sliding window attention alone and use Longformer as a more efficient BERT replacement. For tasks where the answer is "many tokens need full-sequence access," the computational advantages of Longformer diminish as global attention becomes more prevalent. The sweet spot is tasks where a small number of designated tokens (typically 1 to 20) need global reach while the majority of the sequence processes locally.
Document Classification
For classification, you typically only need global attention on the [CLS] token:
def create_classification_global_mask(input_ids, cls_token_id):
"""Global attention only on [CLS] token."""
global_mask = torch.zeros_like(input_ids)
global_mask[:, 0] = 1 # [CLS] is always at position 0
return global_mask
# Example usage
cls_mask = create_classification_global_mask(
inputs["input_ids"], tokenizer.cls_token_id
)Question Answering
For question answering, the question tokens need to see the entire context to find the answer:
def create_qa_global_mask(input_ids, question_end_position):
"""Global attention on question tokens (positions 0 to question_end)."""
global_mask = torch.zeros_like(input_ids)
global_mask[:, :question_end_position] = 1
return global_mask
# Example: question ends at position 20
qa_mask = create_qa_global_mask(inputs["input_ids"], question_end_position=20)Global attention configurations: Classification: 1 global token ([CLS]) Question Answering: 20 global tokens (question) Memory comparison at 4,096 tokens: Full attention: 16,777,216 attention weights Classification Longformer: ~2,105,344 weights (12.5%) QA Longformer: ~2,260,992 weights (13.5%)
The difference in global token count between classification and QA tasks has minimal impact on memory efficiency. Classification uses just 12.5% of full attention memory, while QA with 20 global question tokens uses only 13.5%. Both represent massive savings compared to the 16.7 million attention weights required by full attention at 4,096 tokens.
Named Entity Recognition
For token-level tasks like NER, you might want global attention on sentence boundaries or special delimiter tokens:
def create_ner_global_mask(input_ids, sep_token_id):
"""Global attention on [SEP] tokens (sentence boundaries)."""
global_mask = (input_ids == sep_token_id).long()
# Also include [CLS]
global_mask[:, 0] = 1
return global_maskThe flexibility to configure global attention per task is one of Longformer's key advantages. You control exactly which tokens have global reach, optimizing the trade-off between computational cost and model capability.
Beyond the three task patterns shown here, there are more advanced configurations worth knowing. For multi-label document classification where you predict multiple categories simultaneously, you might use one global token per category label, with each global token learning to aggregate evidence for its associated category from the full document. For structured prediction tasks like relation extraction, you might assign global attention to entity spans so that each entity can directly compare itself with every other entity in the document, regardless of distance. For dialogue systems processing multi-turn conversation histories, you might use global attention on turn separators to allow the model to track conversational state transitions across arbitrarily long histories.
One practical caution: if you add global attention to tokens that do not semantically need it, you typically see minor performance improvements while increasing memory usage substantially. The model can always choose to assign near-zero attention weights to irrelevant global tokens, so the expressiveness never hurts performance, but the memory cost of storing and computing the additional full-sequence attention rows is real. Always benchmark the minimal global attention configuration before adding more global tokens.
Memory and Speed Analysis
Understanding memory usage in precise terms helps you make informed decisions about batch size, sequence length, and hardware requirements when deploying Longformer. The theoretical complexity analysis tells us attention memory scales linearly, but the actual megabytes on a GPU depend on the specific configuration. Let's measure the savings achieved by Longformer compared to full attention across a range of sequence lengths.
def measure_attention_memory(
seq_lengths, embed_dim=768, num_heads=12, window_size=512
):
"""Estimate memory usage for attention matrices."""
results = []
for seq_len in seq_lengths:
# Full attention: n x n attention matrix per head per layer
full_attention_size = seq_len * seq_len * num_heads
# Longformer: n x w for local + n x g for global (approximate)
# Assuming 1 global token
num_global = 1
longformer_size = seq_len * (window_size + 2 * num_global) * num_heads
# Convert to MB (assuming float32)
bytes_per_float = 4
full_mb = full_attention_size * bytes_per_float / (1024**2)
longformer_mb = longformer_size * bytes_per_float / (1024**2)
results.append(
{
"seq_len": seq_len,
"full_attention_mb": full_mb,
"longformer_mb": longformer_mb,
"savings_ratio": longformer_mb / full_mb,
}
)
return results
seq_lengths = [512, 1024, 2048, 4096, 8192]
memory_results = measure_attention_memory(seq_lengths)

The memory comparison reveals the dramatic difference between full attention and Longformer. At 512 tokens, Longformer uses nearly as much memory as full attention because the window size is comparable to the sequence length. But as sequences grow, the gap widens exponentially. At 8,192 tokens, Longformer uses only about 6% of the memory that full attention would require.
This difference is why Longformer can process 4,096-token documents on hardware that would run out of memory with standard BERT-style attention at just 512 tokens.
Practical Applications
Longformer shines on tasks that require understanding long documents where context from distant parts matters. The common thread across these applications is that standard 512-token models require either truncation (losing information) or chunking (breaking the ability to reason across chunk boundaries). Longformer processes the full document in one pass, enabling reasoning that previously required multi-step pipelines or task-specific architectures.
Scientific Paper Analysis
Research papers often span 3,000 to 5,000 tokens. A model analyzing citations needs to connect references in the text to the bibliography at the end. The introduction summarizes findings that appear in detail in the results section. Longformer's global attention on section headers and key sentences enables these long-range connections.
The Allen Institute for AI, which developed Longformer, applied it extensively to scientific literature tasks. On SciRepEval, a benchmark for scientific paper representations, Longformer-based models substantially outperform BERT-based models that truncate papers at 512 tokens. The performance gap is not surprising: a paper's abstract, which typically falls within the first 512 tokens, does not contain all the information needed to classify the paper's methodology, evaluate its contribution to a specific subfield, or identify which prior results it contradicts.
Legal Document Review
Contracts frequently cross-reference clauses defined elsewhere in the document. A clause on page 15 might modify conditions stated on page 3. With Longformer, you can designate clause markers as global tokens, allowing the model to track these dependencies across the entire document.
Legal applications also benefit from Longformer's ability to process entire contracts as unified inputs for information extraction. Extracting all obligations from a service agreement, for instance, requires seeing both the general terms in section 2 and the specific conditions in appendix A that modify them. A chunking approach might process these sections independently and miss that the appendix overrides the general terms. Longformer processes the full document, allowing the model to learn the precedence relationships directly from the training data.
Book Summarization
Summarizing a book chapter requires understanding themes that develop over thousands of words. Characters introduced early affect events that happen much later. Longformer's ability to process long sequences while maintaining global tokens for key narrative elements enables coherent summarization.
For summarization tasks specifically, the choice of global tokens requires more thought than for classification. There is no single obvious "summary token" equivalent to the [CLS] token. One effective approach is to prepend the question "What is this document about?" and give global attention to all question tokens. The model then learns to interpret the full document through the lens of this question, and the question token representations aggregate document-level meaning. Another approach uses the beginning-of-output position as a global token, signaling to the model that this position should accumulate the information needed to generate the summary.
Multi-Document Question Answering
When answering questions that require reasoning across multiple documents, you can concatenate them and use global attention on the question tokens. The model can then search across all provided evidence to find the answer.
This concatenation approach has been applied successfully to tasks like WikiHop, where answering a question requires following a chain of reasoning across multiple Wikipedia paragraphs. Rather than retrieving and re-ranking candidate passages separately, Longformer processes the entire set of retrieved passages as one long sequence, allowing the model to directly compare supporting evidence from different paragraphs in a single forward pass. This eliminates the error-prone re-ranking stage and lets the model reason across documents directly.
Limitations and Trade-offs
Longformer represents a significant advance for long document processing, but it comes with trade-offs you should understand before adopting it.
The sliding window creates an information bottleneck for tokens far from any global token. If important information lies 1,000 positions away from the nearest global token and outside any token's window, the model must rely on the stacking of attention layers to propagate that information. Deep transformer stacks help mitigate this through repeated local interactions that gradually spread information, but the effective receptive field is still limited compared to full attention. For tasks where any token might need to attend to any other token equally, Longformer's sparse pattern might lose important connections.
This bottleneck becomes apparent in tasks involving fine-grained comparisons between distant tokens. Consider a coreference resolution task where "the defendant" in paragraph 1 must be resolved to "John Smith, age 34" in paragraph 6, which is 800 tokens away. Without a global token near both mentions, the model must propagate this connection through many layers of local attention. If the transformer has only 6 layers and the window size is 256, the effective receptive field at layer 6 is only 1,536 tokens, which might be sufficient. But for a 12-layer model processing a 16,000-token document with no global tokens near the relevant mentions, the connection may simply not form reliably.
The separate projection matrices for global attention increase parameter count by approximately 50% for the attention layers. For models with billions of parameters, this overhead is significant. You are trading memory during inference for additional model weights that must be stored and loaded. The global projections also add complexity to fine-tuning: you need to carefully initialize them, and learning dynamics can differ between local and global components.
In practice, the parameter overhead is manageable for most applications. A Longformer-base model has 149 million parameters compared to BERT-base's 110 million, an increase of about 35%. This is small compared to the functional gains for long-document tasks. However, if you are deploying on resource-constrained hardware or trying to maximize performance per parameter, the additional weights are a real cost. One mitigation is to initialize the global projection matrices as copies of the local projection matrices at the start of fine-tuning, so the model begins from a reasonable starting point before the two sets of projections diverge to serve their distinct roles.
Choosing which tokens should have global attention requires task-specific knowledge. For classification, the [CLS] token is an obvious choice. For question answering, the question tokens work well. But for open-ended tasks like summarization or dialogue, the optimal global attention configuration is less clear. Getting this wrong can significantly hurt performance, and finding the right configuration often requires experimentation.
The ambiguity around global token selection is especially challenging when you are applying Longformer to a new domain without existing research to guide you. You may find yourself running ablations to compare single-CLS-global against multi-token-global configurations, or testing whether global attention on domain-specific markers (like section headings in legal documents) improves over CLS-only. This experimentation cost is real and should be factored into project timelines when adopting Longformer for novel tasks.
Longformer's linear complexity assumes the window size and number of global tokens remain constant as sequence length grows. If your task requires global attention on a growing fraction of tokens (say, one global token per paragraph), complexity can approach again. The linear scaling benefit only materializes when global attention is sparse.
There is also a subtle inference-time consideration: while Longformer's memory footprint scales linearly with sequence length, it does not match the raw speed of full attention on shorter sequences. The sparse attention implementation has higher constant-factor overhead compared to the dense matrix multiplication used by full attention, which is heavily optimized in standard linear algebra libraries. For sequences under 1,024 tokens, standard BERT with full attention is often faster in wall-clock time than Longformer with sparse attention. Longformer's efficiency advantages become meaningful starting around 2,048 tokens, and grow substantially beyond that. If your workload is dominated by short documents, Longformer may not be the right tool.
Finally, Longformer was pretrained on relatively long documents from scientific literature, web text, and books. If your target domain involves a very different kind of long-form text, such as code, tabular data embedded in prose, or structured reports with consistent formatting, the pretrained representations may not transfer as effectively as they do for narrative prose. In such cases, domain-adaptive pretraining on your target corpus before fine-tuning can improve results.
Key Parameters
When working with Longformer, the following parameters have the greatest impact on model behavior and performance:
-
window_size: Controls how many neighboring tokens each position can attend to. Larger windows capture more context but increase memory usage linearly. The default of 512 works well for most document tasks, but you might reduce it to 256 for memory-constrained environments or increase it for tasks requiring broader local context.
-
global_attention_mask: A binary tensor indicating which tokens have global attention. Set to 1 for tokens that need to see the entire sequence (e.g., [CLS] for classification, question tokens for QA). Keep the number of global tokens small to maintain linear complexity.
-
max_length: Maximum sequence length the model can process. Longformer-base supports 4,096 tokens by default. Longer sequences require more memory and compute, but the linear scaling makes lengths up to 16,384 feasible.
-
attention_mode (in custom implementations): Determines whether to use sliding window only, global only, or the hybrid pattern. The hybrid pattern is standard for most tasks.
When fine-tuning Longformer, pay special attention to the global attention configuration. The pretrained model learns both local and global projection matrices, so changing which tokens receive global attention during fine-tuning is straightforward. However, adding too many global tokens can negate the efficiency benefits that make Longformer attractive for long documents.
One aspect of the window size parameter that is easy to overlook is its relationship to the number of transformer layers. A model with 12 layers and window size 512 has an effective receptive field of positions, comfortably covering a 4,096-token sequence even without any global tokens. However, a model with only 6 layers and window size 256 has an effective receptive field of just 1,536 positions, meaning that tokens near the end of a 4,096-token document would have a strongly impoverished view of the document's beginning even after all layers. When choosing window size, factor in the number of layers in your model and the maximum sequence length you intend to process. The rule of thumb is that the product of layers times window size should comfortably exceed the maximum sequence length for purely local tasks. If it does not, you need more global tokens to bridge the remaining gap.
The attention_window parameter in Hugging Face's Longformer implementation allows you to specify a different window size per layer. This is an advanced configuration used in some research work where lower layers benefit from smaller windows (capturing syntactic structure) while higher layers use larger windows (capturing broader semantic context). In most practical applications, a uniform window size across all layers performs well and is much simpler to reason about.
Summary
Longformer tackles the quadratic attention bottleneck through an elegant combination of local and global attention patterns. The key ideas are:
-
Sliding window attention reduces per-token complexity from to , where is the window size. Most tokens only need local context to understand their meaning.
-
Global attention preserves the ability to aggregate information across the entire sequence. By designating specific tokens as global, you maintain long-range dependencies where they matter most.
-
Separate projections for local and global attention allow the model to learn distinct representations for different types of context aggregation.
-
Linear complexity enables processing sequences of 4,096 tokens or more on hardware that would fail with standard attention at 512 tokens.
The practical impact is substantial. Research papers, legal documents, and book chapters that were previously too long for transformer processing are now accessible. Tasks requiring reasoning across thousands of tokens become feasible without the memory explosion of full attention.
Longformer represents a broader trend in efficient attention: recognizing that the dense attention pattern of standard transformers is often more than necessary. By designing sparse patterns that match the actual information flow needed for a task, we can dramatically reduce computational costs while maintaining model quality.
When you use Longformer, the most important decisions you make are not about the architecture itself but about how you configure the global attention mask. The model architecture is fixed after pretraining, but the global attention pattern is a task-specific parameter you control at inference time. Investing time in understanding which tokens need full-sequence access for your task will pay dividends in both performance and efficiency. Use the [CLS] token for classification, the question tokens for question answering, and entity or structural markers for tasks requiring cross-document reasoning.
Looking ahead, the ideas in Longformer influenced a wave of subsequent work on efficient attention for long documents. BigBird extended the sparse attention idea by adding random attention alongside the local and global patterns. This provides theoretical guarantees that the resulting attention graph is a connected expander with favorable information-mixing properties. Hierarchical approaches like Longformer Encoder-Decoder (LED) adapted the Longformer attention mechanism for generation tasks like summarization, where the decoder must attend to a long encoded document. The design choices made in Longformer, particularly the task-specific global attention configuration and the separate projection matrices, became templates that later efficient attention models either adopted or explicitly chose to deviate from. Understanding Longformer deeply gives you the vocabulary to evaluate these successor models and understand why they make the choices they do.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Longformer's efficient attention mechanism.
Longformer Attention Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!