Part of Language AI Handbook
Explains how ALiBi encodes position through linear attention biases instead of embeddings. Topics include head-specific slopes, extrapolation properties.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
ALiBi: Attention with Linear Biases
In the previous chapters, we explored various approaches to encoding position information: sinusoidal encodings that add fixed patterns to embeddings, learned position embeddings that train position vectors from scratch, relative position encodings that capture pairwise distances, and RoPE that rotates embeddings based on position. Each method has trade-offs involving complexity, performance, and the ability to handle sequences longer than those seen during training. We have seen that some methods are elegant but fragile at long range, and others are flexible but costly to implement or interpret.
ALiBi (Attention with Linear Biases) takes a radically different approach. Instead of modifying embeddings or inventing clever rotation schemes, ALiBi simply subtracts a value from attention scores based on the distance between tokens. The farther apart two tokens are, the larger the penalty. This remarkably simple idea, introduced by Press and Smith along with Lewis in 2022, achieves strong performance while enabling something the other methods struggle with: extrapolation to sequences far longer than anything seen during training.
Think of ALiBi as adding a proximity premium to attention. Every pair of tokens starts with a content-based score. This reflects how relevant their representations are to each other. ALiBi then applies a distance tax: nearby tokens pay very little, while distant tokens pay heavily. The model never needs to learn what position 512 means, because position is never represented as a vector. Instead, the model learns to work with attention patterns where nearby context is systematically favored. That behavioral pattern stays consistent no matter how long the sequence grows.
The key insight is that position information can enter the model through the mechanism of attention itself, rather than through the token representations. Every prior method we examined grafted position onto the input before it reached the attention layers. ALiBi asks a different question: why not inject position directly where positional relationships matter? Attention is where the model decides which tokens to look at, so shaping attention by distance is a direct and transparent way to express locality.
This design choice has several consequences. First, the model's token embeddings remain pure semantic representations, uncontaminated by position. Second, the position bias is fully interpretable: you can look at any attention head and immediately understand its locality preference by inspecting its slope. Third, and most importantly, the bias depends only on relative distance, which means the model has never seen a "position 1024" encoding and so does not need one during inference at length 2048.
In this chapter, we will work through ALiBi from the ground up. We will develop the mathematical formulation step by step, implement it in code, visualize how different heads behave, and understand exactly why extrapolation works. We will also compare ALiBi to RoPE, discuss practical trade-offs, and examine which real-world models have adopted it.
The Extrapolation Problem
Before understanding ALiBi's solution, we need to appreciate the problem it solves. Transformers trained on sequences of length often fail dramatically when given sequences of length or beyond. This is not a subtle degradation: performance can collapse entirely, with the model producing incoherent outputs for inputs only slightly longer than the training window.
Length extrapolation refers to a model's ability to process sequences longer than those encountered during training while maintaining reasonable performance. Many position encoding schemes fail this test because they produce position representations the model has never learned to interpret.
With sinusoidal encodings, the encoding at position 1025 is a specific vector of cosines and sines that the model was never trained to use. The model's linear and attention layers have developed weights that respond appropriately to positions 0 through 1024, but position 1025 looks foreign, and the model has no basis for interpreting it correctly. With learned position embeddings, the situation is even worse: positions beyond the training length simply do not exist in the embedding table. The model cannot process them at all without modification.
RoPE handles this more gracefully because its rotations are defined by a mathematical formula rather than learned weights. In theory, you can compute a rotation for any position. In practice, however, the model's layers have been tuned on rotation patterns corresponding to positions 0 through , and rotation angles at positions far beyond can fall into regions of the complex plane that the model has never been trained to interpret. Research teams working with RoPE models often need additional techniques like position interpolation or NTK-aware scaling to extend the training context length.
Why does extrapolation matter so much in practice? First, training on very long sequences is extraordinarily expensive. Attention's quadratic complexity means that doubling the sequence length quadruples the attention computation. If you could train on sequences of 1,024 tokens and reliably deploy on sequences of 4,096 or 8,192 tokens, you would save enormous training costs. Second, real-world inputs are unpredictable. A conversational AI might encounter a long document, a code model might see a large file, and a summarization system might receive a lengthy article. A model that degrades gracefully on longer inputs is far more practically valuable than one that requires the input to stay within a hard length ceiling.
The extrapolation problem is essentially a generalization problem across a dimension that most machine learning practitioners don't spend much time thinking about: the sequence length dimension. We routinely generalize across input content, but generalizing across input structure, specifically the sequential length structure, requires position encodings that work by relative relationship rather than absolute coordinate.
The Core Idea: Penalizing Distance
ALiBi's insight is elegant in its simplicity: do not encode position in the embeddings at all. Instead, modify the attention mechanism itself to prefer nearby tokens over distant ones. Embeddings carry semantic content; let attention carry positional structure.
Think of it as a gravitational field. Every token exerts an attentional "pull" on every other token based on content compatibility. ALiBi adds a distance decay to this pull: the further two tokens are from each other, the weaker the pull, regardless of content. Nearby tokens always have a structural advantage. This is not unlike how we read: we naturally weight recent context more heavily than distant context when interpreting the meaning of a word.
Consider what attention scores represent. They measure compatibility between a query and a key. After softmax, these scores determine how much each position contributes to the output. ALiBi introduces a simple bias that subtracts from the score based on distance. To compute the modified attention score between position (as the query) and position (as the key), we start with the standard dot product and then impose the distance penalty:
where:
- : the modified attention score between positions and
- : the original dot product between query and key , capturing content-based compatibility
- : a slope parameter that controls how aggressively to penalize distance (different for each attention head)
- : the absolute distance between the two positions in the sequence
The subtraction means distant tokens receive lower scores. A token 10 positions away gets penalized by , while an adjacent token gets penalized by just . After softmax normalization, this translates to nearby tokens receiving higher attention weights. The key insight is that the penalty is linear in distance: every additional step away costs exactly more in score, regardless of whether you are stepping from position 1 to position 2 or from position 1000 to position 1001.
ALiBi adds a linear bias to attention scores based on the distance between query and key positions. The bias is always non-positive, penalizing distant positions. The penalty grows linearly with distance, controlled by a slope parameter that is fixed (not learned) and specific to each attention head.
This is position encoding without position embeddings. The model learns nothing about absolute position during pretraining of its embeddings. Position information enters only through the attention bias, and only at the moment of computing attention scores. The embeddings remain clean semantic representations of token content, and the relative positional structure is layered on top at attention time.
Why does the penalty have to be linear? The authors experimented with other functional forms, including concave penalties that taper off quickly and convex penalties that grow steeply. Linear penalties produced the best combination of in-distribution performance and out-of-distribution (long-sequence) extrapolation. The linear form also has a pleasing mathematical property: it represents a constant cost per unit of distance, which creates a uniform "discount rate" on attention over distance. This discount rate is consistent regardless of absolute position, which is the property that makes extrapolation work.
The Mathematical Formulation
Now that we understand ALiBi's core insight, let's develop the complete mathematical framework. We will build from standard attention to ALiBi-augmented attention, showing exactly where and how the position bias enters the computation.
Standard attention and ALiBi attention are nearly identical. The only difference is one additional term added to the score matrix before the softmax. This surgical simplicity is part of what makes ALiBi attractive: it requires minimal changes to an existing attention implementation and adds almost no computational overhead.
Starting Point: Standard Scaled Dot-Product Attention
Recall that standard self-attention operates on three matrices derived from the input sequence. For a sequence of tokens, the attention mechanism computes:
where:
- : the query matrix containing query vectors for all positions
- : the key matrix containing key vectors for all positions
- : the value matrix containing value vectors for all positions
- : the dimension of queries and keys
- : the raw attention score matrix where entry is the dot product between query and key
- : scaling factor that prevents dot products from growing too large in high dimensions, which would push the softmax into regions of very small gradients
The matrix captures content-based similarity: how much each query "wants" to attend to each key based purely on their learned representations. This matrix is completely blind to position. Token 1 attending to token 2 produces the same content-based score whether they appear next to each other or 500 positions apart. This is by design for the core attention mechanism, which wants to capture global relationships. But it means we need another mechanism to introduce positional awareness. That mechanism is the ALiBi bias.
Why does ALiBi add bias at this stage rather than modifying the queries and keys? Adding to the pre-softmax scores is the most direct possible intervention. The scores are the raw votes about which tokens matter. By adjusting the scores before normalization, we influence the attention distribution in a clear and simple way. Modifying queries and keys (as RoPE does) achieves a similar effect but through a more complex indirect path involving embedding geometry.
Injecting Position: The Bias Matrix
ALiBi's modification is surgical. Rather than changing how queries, keys, or values are computed, it adds a single term to the attention scores. The complete ALiBi-augmented attention reads:
The bias matrix is added directly to the scaled scores before softmax. This is the only change to standard attention. The matrix encodes position information through a remarkably simple rule: each entry is determined entirely by the distance between the corresponding positions. Specifically,
where:
- : the bias added to the attention score between query position and key position
- : the slope parameter for this attention head, controlling the strength of the distance penalty
- : the absolute distance between positions and in the token sequence
Let's unpack why this formula works. The absolute distance measures how far apart two positions are in the sequence. Adjacent tokens have distance 1, tokens separated by 10 positions have distance 10. Multiplying by a positive slope and negating creates a penalty that grows linearly with distance. A token that is close gets a small penalty; a token that is far gets a large penalty.
Consider what happens for a query at position 5 attending to various keys:
- Attending to position 5 (itself): (no penalty, full score preserved)
- Attending to position 4 (adjacent): (small penalty)
- Attending to position 3 (two steps away): (moderate penalty)
- Attending to position 0 (distant): (large penalty)
The negative bias reduces the attention score, and larger distances produce more negative biases. After softmax normalization, this translates to lower attention weights for distant tokens. The content-based similarity in still matters completely: a very semantically relevant distant token can overcome its distance penalty if its dot product score is large enough. But now content has to compete against a structural penalty that favors proximity.
Why does this formula make sense? Notice that the formula simply implements a universal assumption that language models hold implicitly: nearby context tends to be more relevant than distant context. Rather than forcing the model to learn this bias from data (which is expensive and may lead to overfitting on long-range spurious correlations), ALiBi builds it in as an inductive bias. The model is free to override it with strong content signals, but the default behavior is local attention. This is similar to how convolutional neural networks build in a translation-invariance inductive bias, rather than forcing the network to learn that from scratch.
The Causal Case: Masking the Future
For decoder-style models that generate text left-to-right, we need causal masking: position can only attend to positions . The combined bias matrix integrates both the ALiBi distance penalties and the causal mask:
The structure reveals two distinct components working together. The lower triangle, including the diagonal, contains the ALiBi distance penalties. Entry with holds the value , which is zero on the diagonal and grows more negative as you move left along any row. The upper triangle holds values from the causal mask.
The entries ensure that after applying inside the softmax, future positions contribute nothing to the weighted sum. The linear penalties in the lower triangle then shape how attention flows among the allowed past positions. These two components operate in entirely different regimes: the causal mask is a hard prohibition (zero weight, always), while the ALiBi penalty is a soft preference (lower weight, but nonzero).
Notice that the causal mask interacts naturally with ALiBi. For a model processing position , the only tokens it can attend to are positions 0 through . These are exactly the tokens for which , so the distance is always , always positive, and always measured in the backward direction. The ALiBi penalty is therefore always a backward-looking decay, consistently favoring the most recent tokens in the context.
Building the Bias Matrix: Step-by-Step Implementation
Let's translate this mathematics into code, building up the bias matrix piece by piece. The core insight behind the implementation is that we can compute all pairwise distances simultaneously using broadcasting, rather than iterating over all pairs with nested loops.
import numpy as np
def create_alibi_bias(seq_len, slope):
"""
Create the ALiBi bias matrix for a single attention head.
Args:
seq_len: Length of the sequence
slope: The slope parameter m for this head
Returns:
Bias matrix of shape (seq_len, seq_len)
"""
# Create position indices
positions = np.arange(seq_len)
# Compute pairwise distances: |i - j|
# positions[:, None] is (seq_len, 1), positions[None, :] is (1, seq_len)
# Broadcasting gives (seq_len, seq_len)
distances = np.abs(positions[:, None] - positions[None, :])
# Apply linear penalty
bias = -slope * distances
return biasThe key implementation insight is using NumPy broadcasting to compute all pairwise distances at once. By reshaping the position array into a column vector positions[:, None] and a row vector positions[None, :], subtraction produces an matrix where entry is . Taking the absolute value gives us the distance matrix. Multiplying by the slope and negating gives the penalty matrix. This vectorized approach is both faster and more readable than explicit loops.
Let's see what this produces for a small sequence:
ALiBi bias matrix for sequence length 6, slope 0.5: [[-0. -0.5 -1. -1.5 -2. -2.5] [-0.5 -0. -0.5 -1. -1.5 -2. ] [-1. -0.5 -0. -0.5 -1. -1.5] [-1.5 -1. -0.5 -0. -0.5 -1. ] [-2. -1.5 -1. -0.5 -0. -0.5] [-2.5 -2. -1.5 -1. -0.5 -0. ]]
Reading this matrix: the diagonal is zero because a token attending to itself has distance zero and receives no penalty. Moving away from the diagonal in either direction, penalties grow linearly. Position 5 (row 5) attending to position 0 (column 0) shows , which is exactly as expected. The matrix is symmetric because distance is symmetric: . In a causal model we will only use the lower triangle, but the full symmetric matrix is the natural output of the distance computation.
For decoder models, we overlay the causal mask on top of the distance penalties:
def create_causal_alibi_bias(seq_len, slope):
"""
Create ALiBi bias with causal masking.
Args:
seq_len: Length of the sequence
slope: The slope parameter m
Returns:
Bias matrix with -inf for future positions
"""
# Start with the distance-based bias
bias = create_alibi_bias(seq_len, slope)
# Create causal mask (upper triangle should be -inf)
causal_mask = np.triu(np.ones((seq_len, seq_len)), k=1)
# Apply causal mask: future positions get -inf
bias = np.where(causal_mask == 1, -np.inf, bias)
return biasThe np.triu function creates an upper triangular matrix of ones, which we use to identify positions that should be masked. The k=1 argument excludes the diagonal, since a token should be able to attend to itself. The np.where call then replaces those positions with negative infinity, which becomes effectively zero after the exponential in softmax.
Causal ALiBi bias matrix: [[-0. -inf -inf -inf -inf -inf] [-0.5 -0. -inf -inf -inf -inf] [-1. -0.5 -0. -inf -inf -inf] [-1.5 -1. -0.5 -0. -inf -inf] [-2. -1.5 -1. -0.5 -0. -inf] [-2.5 -2. -1.5 -1. -0.5 -0. ]]
Now the upper triangle shows inf (NumPy's display for ), while the lower triangle retains the distance penalties. Row 5 can attend to all previous positions with penalties for positions 0 through 5 respectively. Row 0 can only attend to itself with penalty 0, since there are no previous positions available.
Let's visualize both the distance matrix and the resulting causal ALiBi bias matrix side by side to make the structure visually clear:


The left matrix shows raw distances: the diagonal is 0 (same position), adjacent cells are 1, and values grow as positions diverge. The right matrix shows what happens after applying ALiBi: distances become negative penalties (scaled by the slope), and the upper triangle is masked to enforce causality. This visual makes clear that the bias matrix is nothing more than a scaled, negated distance matrix with causal masking applied on top.
Head-Specific Slopes
A single slope would force all attention heads to have the same locality preference. That would be limiting: different aspects of language operate at different scales, and a model benefits from heads that specialize in different scopes. Some heads might focus on very local context (steep slope, harsh penalties for distance), while others maintain broader receptive fields (gentle slope, mild penalties).
Think of it like a camera system with multiple lenses. A telephoto lens captures fine detail up close but misses the broader scene. A wide-angle lens captures the full picture but loses local detail. You want both working together. ALiBi gives each attention head a different "focal length" by assigning it a different slope. The ensemble of heads covers a wide spectrum of temporal scales, from immediate local syntax to long-range topic coherence.
The slopes are not learned but fixed according to a geometric sequence. The decision to fix them (rather than learn them) is deliberate: it eliminates hyperparameters from the position encoding, makes the encoding fully deterministic and reproducible, and avoids the risk of the model learning pathological slope values during training. For a model with attention heads, the slope for head is:
where:
- : the slope parameter for attention head
- : the total number of attention heads in the model
- : the head index, ranging from 1 to
- : a scaling factor that ensures slopes span a consistent range regardless of how many heads the model has
- : the denominator that grows exponentially with head index, making later heads have gentler slopes
Why does this formula make sense? Notice that the exponent ranges from (head 1) to (head ). The denominator is therefore always through , independent of . This ensures that the range of slopes remains consistent across model architectures with different head counts. Whether you have 4 heads or 32 heads, the steepest head always has a slope near and the gentlest head always has a slope near .
The base of 8 in the exponent was chosen empirically by the ALiBi authors. With 8 heads, the exponent simplifies to just , giving slopes of , which equals . With other head counts, the formula interpolates appropriately.
def get_alibi_slopes(num_heads):
"""
Compute ALiBi slopes for each attention head.
The slopes follow a geometric sequence, with the first head
having the steepest slope (most local attention) and the last
head having the gentlest slope (broadest attention).
Args:
num_heads: Number of attention heads
Returns:
Array of slopes, one per head
"""
# Compute the ratio for the geometric sequence
ratio = 2 ** (8 / num_heads)
# Generate slopes: 1/ratio, 1/ratio^2, ..., 1/ratio^num_heads
slopes = 1.0 / (ratio ** np.arange(1, num_heads + 1))
return slopes4 heads: slopes = [0.25 0.0625 0.015625 0.003906] 8 heads: slopes = [0.5 0.25 0.125 0.0625 0.03125 0.015625 0.007812 0.003906] 16 heads: slopes = [0.707107 0.5 0.353553 0.25 0.176777 0.125 0.088388 0.0625 0.044194 0.03125 0.022097 0.015625 0.011049 0.007812 0.005524 0.003906]
The geometric progression ensures that slopes span several orders of magnitude. With 8 heads, the steepest slope (0.5) penalizes a distance of 10 by 5 logits, effectively eliminating distant tokens from consideration after softmax. The gentlest slope (roughly 0.004) penalizes the same distance by only 0.04 logits, allowing the head to attend broadly across hundreds of positions. This spread is what gives ALiBi its multi-scale character.
Notice that the formula produces a different set of slopes for 4, 8, and 16 heads, but in each case, the slopes span roughly the same overall range from steep to gentle. With 4 heads, you get fewer granular distinctions but the extremes remain similar. With 16 heads, you get finer granularity, more heads at intermediate slopes, but still roughly the same endpoints. This consistency is the purpose of the scaling factor.

Visualizing the Attention Bias
The slope values are informative in the abstract, but seeing what they produce as attention bias matrices makes the effect concrete. Let's visualize how ALiBi biases shape attention patterns for the steepest and gentlest heads:


The contrast is stark and instructive. Head 1's bias matrix shows deep red (strongly negative) values just a few positions from the diagonal. By position 10, the penalty exceeds -5 logits, making those tokens nearly invisible after softmax unless their content score is extraordinarily strong. The effective attention window for this head is perhaps 3 to 5 tokens wide. Head 8, in contrast, shows nearly uniform mild penalties throughout. Even at distance 19, the penalty is less than 0.1 logits, allowing meaningful attention to distant tokens across essentially the entire sequence.
This division of labor is intentional and reflects something real about natural language. Linguistic phenomena operate at different scales. Adjacent tokens matter for subject-verb agreement, determiner-noun matching, and immediate syntactic structure. Tokens a few steps away matter for multi-word expressions, phrasal dependencies, and clause-level coherence. Tokens many steps away matter for coreference resolution (a pronoun referring back to a noun introduced many sentences earlier), topic consistency, and document-level discourse structure. By giving different heads different locality preferences, ALiBi enables the model to capture phenomena at multiple scales simultaneously without requiring the model to learn these scales from scratch.
ALiBi in Attention: Complete Implementation
Now we have all the pieces to implement ALiBi-augmented attention from scratch. The implementation closely mirrors standard attention. We compute and scale the score matrix, add the bias, apply the optional causal mask, take the softmax, and use the result to weight the values.
The only new element is the construction and addition of the ALiBi bias tensor. For multi-head attention, we need one bias matrix per head, which we stack into a three-dimensional tensor of shape (num_heads, seq_len, seq_len). Each slice along the first dimension is the bias matrix for one head, with that head's specific slope.
def softmax(x, axis=-1):
"""Numerically stable softmax."""
x_max = np.max(x, axis=axis, keepdims=True)
exp_x = np.exp(x - x_max)
return exp_x / np.sum(exp_x, axis=axis, keepdims=True)
def alibi_attention(Q, K, V, slopes, causal=True):
"""
Compute attention with ALiBi position encoding.
Args:
Q: Query matrix of shape (num_heads, seq_len, d_k)
K: Key matrix of shape (num_heads, seq_len, d_k)
V: Value matrix of shape (num_heads, seq_len, d_v)
slopes: ALiBi slopes, one per head, shape (num_heads,)
causal: Whether to apply causal masking
Returns:
Output of shape (num_heads, seq_len, d_v)
Attention weights of shape (num_heads, seq_len, seq_len)
"""
num_heads, seq_len, d_k = Q.shape
# Compute raw attention scores: Q @ K^T
# Shape: (num_heads, seq_len, seq_len)
scores = np.matmul(Q, K.transpose(0, 2, 1))
# Scale by sqrt(d_k)
scores = scores / np.sqrt(d_k)
# Create ALiBi biases for each head
# Shape: (num_heads, seq_len, seq_len)
positions = np.arange(seq_len)
distances = np.abs(positions[:, None] - positions[None, :])
# Broadcast slopes: (num_heads, 1, 1) * (seq_len, seq_len)
alibi_bias = -slopes[:, None, None] * distances[None, :, :]
# Apply causal mask if needed
if causal:
causal_mask = np.triu(np.ones((seq_len, seq_len)), k=1)
alibi_bias = np.where(causal_mask == 1, -np.inf, alibi_bias)
# Add ALiBi bias to scores
scores = scores + alibi_bias
# Apply softmax to get attention weights
attention_weights = softmax(scores, axis=-1)
# Handle NaN from -inf (all masked positions)
attention_weights = np.nan_to_num(attention_weights, nan=0.0)
# Compute output: weighted sum of values
output = np.matmul(attention_weights, V)
return output, attention_weightsThe broadcasting step deserves attention. We have slopes with shape (num_heads,). Reshaping to (num_heads, 1, 1) and multiplying by distances[None, :, :] with shape (1, seq_len, seq_len) produces a tensor of shape (num_heads, seq_len, seq_len). Each head's slice is the distance matrix multiplied by that head's slope. This is a single vectorized operation that replaces what would otherwise be a loop over heads.
Let's test this implementation with a concrete example to verify the shapes and inspect the attention patterns:
# Create a simple example
np.random.seed(42)
num_heads = 4
seq_len = 8
d_k = 16
d_v = 16
# Random Q, K, V matrices
Q = np.random.randn(num_heads, seq_len, d_k) * 0.5
K = np.random.randn(num_heads, seq_len, d_k) * 0.5
V = np.random.randn(num_heads, seq_len, d_v) * 0.5
# Get ALiBi slopes
slopes = get_alibi_slopes(num_heads)
# Compute ALiBi attention
output, attention_weights = alibi_attention(Q, K, V, slopes, causal=True)Output shape: (4, 8, 16) Attention weights shape: (4, 8, 8) Attention weights for head 1 (steep slope = 0.250): [[1. 0. 0. 0. 0. 0. 0. 0. ] [0.45 0.55 0. 0. 0. 0. 0. 0. ] [0.233 0.37 0.397 0. 0. 0. 0. 0. ] [0.227 0.205 0.262 0.306 0. 0. 0. 0. ] [0.118 0.068 0.121 0.279 0.414 0. 0. 0. ] [0.083 0.086 0.13 0.201 0.176 0.324 0. 0. ] [0.065 0.089 0.092 0.127 0.137 0.272 0.218 0. ] [0.025 0.038 0.057 0.073 0.136 0.214 0.233 0.224]] Attention weights for head 4 (gentle slope = 0.003906): [[1. 0. 0. 0. 0. 0. 0. 0. ] [0.562 0.438 0. 0. 0. 0. 0. 0. ] [0.36 0.453 0.187 0. 0. 0. 0. 0. ] [0.344 0.23 0.245 0.181 0. 0. 0. 0. ] [0.184 0.232 0.181 0.169 0.233 0. 0. 0. ] [0.121 0.125 0.286 0.214 0.096 0.158 0. 0. ] [0.104 0.124 0.171 0.176 0.08 0.175 0.169 0. ] [0.109 0.137 0.063 0.124 0.158 0.147 0.163 0.099]]
Compare the attention patterns between heads. Head 1 with its steep slope concentrates attention heavily on recent positions, with weights dropping rapidly as distance increases. The rows show a clear diagonal structure: most of the weight is on the most recent 1 to 3 tokens, with essentially nothing beyond 5 positions. Head 4 with its gentle slope distributes attention more evenly across the available context. The rows are more uniform, with distant positions receiving non-negligible weight.
To understand the impact of ALiBi more directly, let's compare attention patterns with and without the position bias. We'll compute attention for the same queries and keys, once with standard attention (no position encoding) and once with ALiBi:


The difference is striking. Standard attention distributes weight based purely on content similarity, sometimes attending strongly to distant positions because the content happens to match well. ALiBi reshapes this pattern, pulling attention toward recent tokens while still allowing content to influence the final distribution. This locality bias emerges from a single matrix addition, requiring no learned position embeddings and no modifications to the token representations themselves.


Worked Example: Tracing Through ALiBi Step by Step
Let's work through a concrete numerical example to solidify the mechanics. We'll use a very small sequence so every number is legible.
Consider a sequence of 4 tokens with 2 attention heads and . We'll use simple numbers so you can verify each step by hand.
# Small worked example with traceable numbers
np.random.seed(0)
n = 4 # sequence length
h = 2 # number of heads
dk = 4 # key/query dimension
# Simple Q and K with small values
Q_ex = np.array(
[
[
[1.0, 0.0, 0.5, 0.0],
[0.0, 1.0, 0.0, 0.5],
[0.5, 0.0, 1.0, 0.0],
[0.0, 0.5, 0.0, 1.0],
],
[
[1.0, 0.0, 0.5, 0.0],
[0.0, 1.0, 0.0, 0.5],
[0.5, 0.0, 1.0, 0.0],
[0.0, 0.5, 0.0, 1.0],
],
]
) # shape (2, 4, 4)
K_ex = Q_ex.copy() # Use K = Q for simplicity
V_ex = np.random.randn(h, n, dk) * 0.1
slopes_ex = get_alibi_slopes(h) # slopes for 2 headsStep 1: Raw dot-product scores QK^T (head 1): [[1.25 0. 1. 0. ] [0. 1.25 0. 1. ] [1. 0. 1.25 0. ] [0. 1. 0. 1.25]] Step 2: Scaled scores (divided by sqrt(4) = 2.000) (head 1): [[0.625 0. 0.5 0. ] [0. 0.625 0. 0.5 ] [0.5 0. 0.625 0. ] [0. 0.5 0. 0.625]] Step 3: ALiBi bias matrix (head 1, slope = 0.062): [[-0. -0.062 -0.125 -0.188] [-0.062 -0. -0.062 -0.125] [-0.125 -0.062 -0. -0.062] [-0.188 -0.125 -0.062 -0. ]] Step 4: ALiBi bias with causal mask applied (head 1): [[-0. -inf -inf -inf] [-0.062 -0. -inf -inf] [-0.125 -0.062 -0. -inf] [-0.188 -0.125 -0.062 -0. ]] Step 5: Final scores (scaled + bias) (head 1): [[ 0.625 -inf -inf -inf] [-0.062 0.625 -inf -inf] [ 0.375 -0.062 0.625 -inf] [-0.188 0.375 -0.062 0.625]] Step 6: Attention weights after softmax (head 1): [[1. 0. 0. 0. ] [0.335 0.665 0. 0. ] [0.341 0.22 0.438 0. ] [0.163 0.286 0.184 0.367]]
Let's read through this trace carefully. In step 1, the raw dot-product scores capture content similarity. Since we used with identity-like structure, each token scores highest against itself. In step 2, we scale by , which shrinks all values without changing their relative ordering. In step 3, the ALiBi bias is purely a function of distance: entry is where is the head's slope. For head 1 with the steeper slope, the penalties at distance 1, 2, and 3 are meaningful numbers. In step 4, the upper triangle is set to to enforce causality. In step 5, we add the bias to the scaled scores. Notice how the off-diagonal entries drop: the content similarity gets penalized by distance, pulling the distribution toward the diagonal. In step 6, softmax turns these into probabilities. The diagonal entries (attending to self) dominate strongly for head 1 because the ALiBi penalty has suppressed the off-diagonal scores.
This step-by-step trace illustrates the key property of ALiBi: it does not change what the model "wants" to attend to (the content signal), but it imposes a systematic cost for looking farther away. When content signals are strong, they can dominate. When content signals are weak or comparable, the distance penalty tips the balance toward nearby tokens.
Why ALiBi Extrapolates
The key to ALiBi's extrapolation ability lies in what the model learns during training. With other position encoding schemes, the model learns to interpret specific position representations. A sinusoidal encoding for position 512 produces a particular vector that the model has seen and learned to use. A learned position embedding for position 512 is a specific trained vector. When you encounter position 1025, this is a novel representation the model has not seen, and its layers have no well-defined response to it.
ALiBi sidesteps this problem entirely. The model never learns position representations because there are none to learn. Instead, it learns to work with relative attention patterns shaped by the linear bias. During training on sequences of length 1024, the model sees attention patterns where nearby tokens are favored and distant tokens are penalized. This is true at position 10, position 500, and position 1000. At every position in the training range, the model observes the same structural pattern: recent tokens have higher weight, distant tokens have lower weight.
When inference extends to position 2048, the same principle applies. The local neighborhood still receives favorable bias. Tokens 10 positions away still get penalized by the same amount. The absolute positions are larger, but the relative structure is completely unchanged. The model has learned to extract information from attention patterns that favor locality, and those patterns remain consistent regardless of sequence length.
ALiBi extrapolates because it encodes relative distance, not absolute position. The linear penalty for distance 10 is whether you are at position 50 or position 5000. The model learns to work with distance-biased attention patterns, which remain structurally consistent across all sequence lengths. This is the fundamental distinction from methods that encode absolute position as a vector.
There is a subtle but important point here. When a model processes a long sequence at inference time, tokens near the beginning of the sequence will be far from the current query position. ALiBi will penalize them heavily, potentially so heavily that they receive near-zero attention weight. This is not a failure: it is the expected behavior. The model is effectively operating with a soft window: it pays close attention to recent context and discounts the distant past. For sequences significantly longer than the training length, the very first tokens may be practically invisible to later queries. Whether this hurts performance depends on the task: for many language modeling and generation tasks, the most relevant context is recent, so this graceful distance falloff can improve performance.
Let's visualize the consistency property that makes extrapolation work:

The curves overlap perfectly because ALiBi's bias depends only on distance, never on absolute position. Whether the query is at token 20 or token 5000, the bias for a key 30 positions back is the same: . This is the mathematical basis for length extrapolation, and it is also the mathematical reason why it could not work for sinusoidal or learned position embeddings: those methods embed absolute position, so the representation at position 5000 is fundamentally different from (and trained less on than) the representation at position 20.
ALiBi vs. RoPE: A Comparison
Both ALiBi and RoPE are widely used in modern language models, and both address relative position. But they take fundamentally different approaches that lead to different trade-offs in practice.
RoPE encodes position by rotating query and key vectors. When computing dot products between rotated vectors, the rotation angles combine such that the result depends on relative position rather than absolute position. This is mathematically elegant: the dot product after rotation naturally depends on rather than on and separately. RoPE modifies the embedding space itself: the vectors that participate in dot products carry rotational information baked into their components.
ALiBi's approach is different in character: it does not modify the vectors at all. The dot products remain the raw content-based scores. ALiBi then overlays a separate, explicit position signal on top of those scores. The two components are additive and fully transparent. You can look at a score and immediately decompose it into a content part (the dot product) and a position part (the bias).
| Aspect | ALiBi | RoPE |
|---|---|---|
| Position encoding location | Attention score bias | Query and key embeddings |
| Mechanism | Subtracts linear penalty from attention scores | Rotates Q and K vectors by position-dependent angles |
| Parameters | Fixed slopes, no learned position parameters | No additional parameters |
| Computation | Simple matrix addition to score matrix | Complex number arithmetic or 2D rotation matrices |
| Extrapolation | Strong out-of-the-box at long sequences | Requires additional techniques (NTK scaling, interpolation) |
| Interpretability | Fully transparent: slope directly controls window width | Less direct: rotation angle interplay is harder to interpret |
| Impact on embeddings | None: token embeddings are pure semantic vectors | Embeds absolute rotation into token representations |
The choice between ALiBi and RoPE often comes down to empirical performance on the target task. RoPE has shown excellent results on many benchmarks and is used by prominent models including LLaMA and Mistral, as well as Falcon. ALiBi has shown strong extrapolation and is used by BLOOM and MPT. Newer research (including models like LLaMA 3 and Mixtral) tends to favor RoPE with context extension techniques, suggesting that RoPE's expressiveness may ultimately win out as practitioners develop better tools for extending its context window. However, ALiBi remains a practical choice when implementation simplicity and extrapolation matter, especially when minimizing implementation risk.
Let's compare the implementation complexity concretely, which is where ALiBi's advantage is clearest:
def rope_attention(Q, K, V, seq_len, d_k, causal=True):
"""
Simplified RoPE attention for comparison.
This is a sketch showing the additional complexity.
"""
# RoPE requires computing rotation matrices or using complex numbers
# For each position, apply rotation to Q and K before computing attention
# Step 1: Compute rotation angles for each position and dimension
positions = np.arange(seq_len)
dim_indices = np.arange(d_k // 2)
# Frequency for each dimension pair (simplified)
freqs = 1.0 / (10000 ** (2 * dim_indices / d_k))
# Angle matrix: (seq_len, d_k/2)
angles = positions[:, None] * freqs[None, :]
# Step 2: Apply rotation to Q and K (complex number approach)
# This involves reshaping Q and K, applying cos/sin transformations...
# (full implementation omitted for brevity)
# The key point: RoPE modifies the embeddings themselves
# before attention computation
pass
def alibi_attention_simple(Q, K, V, slopes, causal=True):
"""
ALiBi attention for comparison.
"""
# Step 1: Standard attention scores
scores = np.matmul(Q, K.transpose(0, 2, 1)) / np.sqrt(Q.shape[-1])
# Step 2: Add distance-based bias (one line!)
seq_len = Q.shape[1]
distances = np.abs(
np.arange(seq_len)[:, None] - np.arange(seq_len)[None, :]
)
scores = scores + (-slopes[:, None, None] * distances)
# That's it. Apply softmax and compute output.
return scoresThe ALiBi bias is a single line of code added to standard attention. The entire position encoding logic lives in one array operation. RoPE requires restructuring how queries and keys are computed, introducing trigonometric functions and careful handling of dimension pairs. Both work, but ALiBi's simplicity is a practical advantage: fewer lines of code means fewer places for bugs to hide, and the position bias is always directly visible and inspectable.
Historical Context
ALiBi was introduced in the paper "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation" by Ofir Press, Noah A. Smith, and Mike Lewis, published in 2022 at ICLR. The title captures the core contribution succinctly: train on short sequences, test (inference) on long sequences, enabled by the linear bias design.
The paper was motivated by the observation that all existing position encoding methods (sinusoidal, learned, T5 relative biases, and RoPE) failed to generalize beyond training sequence length. The authors systematically evaluated length extrapolation and found that ALiBi dramatically outperformed all competitors on this metric while remaining competitive on standard perplexity benchmarks.
ALiBi was adopted by the BLOOM project (the largest open multilingual language model at the time of its release in 2022 with 176 billion parameters) and by MosaicML's MPT family of models. This adoption showed that ALiBi could support production-scale language models. The method influenced subsequent thinking about position encoding design and contributed to the recognition that inductive biases favoring locality are valuable, not limiting.
Practical Considerations
When deploying ALiBi in your own models, several practical considerations affect performance and behavior. Understanding these factors helps you make informed decisions about whether ALiBi is the right choice and how to configure it.
ALiBi's computational overhead relative to standard attention is minimal. The bias matrix computation is , matching the overall cost of attention, and the actual operation (a matrix add) is extremely fast on modern hardware. There are no additional parameters to initialize, no embedding lookups, and no trigonometric operations. In a typical transformer implementation, adding ALiBi increases the attention computation time by perhaps 1 to 2 percent, well within noise.
The fixed slopes, while convenient, raise a natural question: should they be learned instead? Several research papers have explored learned ALiBi slopes. The findings are mixed: on standard benchmarks within the training length, learned slopes sometimes improve performance slightly. But on out-of-distribution sequence lengths (the key use case for ALiBi), learned slopes may not extrapolate as reliably as the fixed formula. The fixed slopes guarantee consistent behavior across all sequence lengths by design. Learned slopes might overfit to the training length range. If you decide to experiment with learned slopes, treat them as a carefully monitored hyperparameter rather than a drop-in replacement.
The causal versus bidirectional distinction matters for deployment. Decoder-only language models (for text generation) use causal ALiBi, where the upper triangle of the bias matrix is . Encoder models (for classification, understanding) use bidirectional ALiBi, where the full symmetric bias matrix applies and both left and right context receive distance penalties. Encoder-decoder models (for translation, summarization) typically use bidirectional ALiBi in the encoder and causal ALiBi in the decoder. The mathematics remains the same in all cases; only which entries of the bias matrix are live changes.

The effective context window visualization reveals a striking difference across heads. Head 1 drops below the 10 percent relative weight threshold very quickly, around distance 5, because its steep slope () imposes an exponentially decaying effective weight. Head 4 with slope 0.0625 reaches the 10 percent threshold around distance 37. Head 8 with slope 0.0039 barely crosses it even at distance 100. When all heads work together, the model simultaneously has access to very local structure (from head 1), medium-range context (from heads 2 through 7), and long-range dependencies (from head 8). This is arguably more structured than what emerges organically in models without explicit position encoding.
Limitations
ALiBi's simplicity is both its strength and its limitation. Understanding where the method falls short is essential for making informed architectural decisions.
The linear penalty assumes that relevance decreases monotonically with distance, which is a reasonable prior for many language tasks but not universally true. Consider structured data like code: a function definition might appear 500 tokens before a call site, but understanding the call requires attending back to the definition with full weight. The ALiBi penalty at distance 500 is , which with a steep slope effectively zeroes out that attention. The gentle-slope heads may still capture it, but there is no mechanism to override the penalty when the task demands strong long-range attention. Models using ALiBi compensate by learning strong content-based signals that overcome the bias, but this requires more capacity dedicated to fighting the prior.
Similarly, some programming languages and formal documents use nested structure (parentheses, XML tags, logical blocks) where the relevant matching token could be arbitrarily far away. The distance penalty is agnostic to this structure: it cannot know that an opening brace needs to attend to its closing brace regardless of the distance. Tasks that require this kind of distance-invariant matching are challenging for ALiBi. RoPE, being content-neutral in its position encoding (the rotation is applied to all dimensions uniformly), has less of this structural disadvantage.
The fixed slopes also create an implicit ceiling on what the gentlest head can do. Even with slope , a distance of 1000 tokens incurs a penalty of logits. After softmax, this is still a significant suppression. For tasks that require integration across truly long ranges (thousands of tokens of prior context), ALiBi's gentlest head may still fall short. This was one motivation for context window extension techniques like sliding window attention and retrieval augmentation in systems that use ALiBi.
Another limitation is the lack of asymmetry. The penalty is symmetric: attending backward 10 steps is penalized the same as attending forward 10 steps (in bidirectional models). Natural language is often asymmetric: we typically interpret a word in light of what came before more than what comes after. While the causal mask handles this in decoder models, bidirectional ALiBi (for encoders) applies equal penalties in both directions. A refinement would be to use different slopes for forward and backward attention, but this is not part of the original ALiBi design.
Despite these limitations, ALiBi has proven remarkably effective. The BLOOM family of models, including the 176-billion parameter BLOOM-176B, uses ALiBi. So does the MPT family from MosaicML. These models demonstrate that ALiBi scales to the largest parameter counts and handles remarkably diverse tasks. The limitations manifest primarily in edge cases and specialized tasks, not in the general language modeling and generation use cases that dominate production deployment.
The impact of ALiBi extends beyond its direct use. It demonstrated that position encoding can be far simpler than previously thought. The original Transformer's sinusoidal encodings were ingenious but perhaps overengineered for the task. ALiBi showed that a linear penalty on distance, applied at attention time, is sufficient for strong performance and far superior for extrapolation. This insight has influenced subsequent work on efficient transformers, long-context models, and the general question of how much structure should be built into the architecture versus learned from data.
Key Parameters
When implementing ALiBi in your own models, the following parameters control its behavior:
-
num_heads: The number of attention heads in your model. ALiBi automatically computes slopes for each head using the geometric sequence formula. More heads create finer granularity in locality preferences: with 16 heads you get smoother coverage of the slope range than with 4 heads. -
slope(m): The penalty strength for each head. Steeper slopes (larger values like 0.5) create strong locality bias where attention concentrates on nearby tokens. Gentler slopes (smaller values like 0.004) allow broader attention across the sequence. These are fixed by the formula rather than tuned, which eliminates a hyperparameter while potentially leaving some performance on the table for specialized tasks. -
causal: Whether to apply causal masking. Set toTruefor decoder-style autoregressive models where tokens can only attend to previous positions. Set toFalsefor encoder-style bidirectional attention, where both preceding and following tokens are available. -
Base value (8): The constant in the slope formula that controls the range of slopes. The original ALiBi paper uses 8, which ensures slopes span several orders of magnitude regardless of head count. Increasing this base would compress all slopes toward smaller values (gentler biases), while decreasing it would shift toward steeper slopes. The value 8 was selected empirically and is almost never modified in practice.
Summary
ALiBi offers a refreshingly simple approach to position encoding that achieves strong performance and excellent length extrapolation through a single design insight: replace position embeddings with a direct penalty on attention score based on distance.
The key ideas in this chapter are:
-
No position embeddings. Position enters only through attention biases, not through modifications to token representations. This keeps embeddings clean and makes position encoding completely separable from content encoding.
-
Linear distance penalty. Attention scores are reduced by , where is a head-specific slope and is the distance between query position and key position . Nearby tokens are favored, distant tokens are penalized, and the penalty is always proportional to distance.
-
Geometric slopes. Different attention heads use different slopes, creating multi-scale attention. Some heads focus locally on syntax and immediate context, others attend broadly to long-range dependencies. The slopes are fixed by formula, not learned.
-
Strong extrapolation. Because only relative distance matters (not absolute position), models trained on short sequences can process longer sequences at inference time. The bias structure is exactly the same at position 100 as at position 5000.
-
Minimal overhead. ALiBi adds one matrix addition to attention computation. There are no additional parameters, no complex rotation arithmetic, and no changes to the embedding pipeline.
The next chapter will compare the position encoding methods we have covered, from sinusoidal and learned encodings to relative approaches such as RoPE and ALiBi. You will see how each handles key challenges like extrapolation, computational cost, and representational power, and when to choose each method for real-world applications.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about ALiBi (Attention with Linear Biases).
ALiBi Position Encoding Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!