ALiBi: Attention with Linear Biases for Position Encoding

Michael BrenndoerferUpdated May 30, 202552 min read

Part of Language AI Handbook

Explains how ALiBi encodes position through linear attention biases instead of embeddings. Topics include head-specific slopes, extrapolation properties.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

ALiBi: Attention with Linear Biases

In the previous chapters, we explored various approaches to encoding position information: sinusoidal encodings that add fixed patterns to embeddings, learned position embeddings that train position vectors from scratch, relative position encodings that capture pairwise distances, and RoPE that rotates embeddings based on position. Each method has trade-offs involving complexity, performance, and the ability to handle sequences longer than those seen during training. We have seen that some methods are elegant but fragile at long range, and others are flexible but costly to implement or interpret.

ALiBi (Attention with Linear Biases) takes a radically different approach. Instead of modifying embeddings or inventing clever rotation schemes, ALiBi simply subtracts a value from attention scores based on the distance between tokens. The farther apart two tokens are, the larger the penalty. This remarkably simple idea, introduced by Press and Smith along with Lewis in 2022, achieves strong performance while enabling something the other methods struggle with: extrapolation to sequences far longer than anything seen during training.

Think of ALiBi as adding a proximity premium to attention. Every pair of tokens starts with a content-based score. This reflects how relevant their representations are to each other. ALiBi then applies a distance tax: nearby tokens pay very little, while distant tokens pay heavily. The model never needs to learn what position 512 means, because position is never represented as a vector. Instead, the model learns to work with attention patterns where nearby context is systematically favored. That behavioral pattern stays consistent no matter how long the sequence grows.

The key insight is that position information can enter the model through the mechanism of attention itself, rather than through the token representations. Every prior method we examined grafted position onto the input before it reached the attention layers. ALiBi asks a different question: why not inject position directly where positional relationships matter? Attention is where the model decides which tokens to look at, so shaping attention by distance is a direct and transparent way to express locality.

This design choice has several consequences. First, the model's token embeddings remain pure semantic representations, uncontaminated by position. Second, the position bias is fully interpretable: you can look at any attention head and immediately understand its locality preference by inspecting its slope. Third, and most importantly, the bias depends only on relative distance, which means the model has never seen a "position 1024" encoding and so does not need one during inference at length 2048.

In this chapter, we will work through ALiBi from the ground up. We will develop the mathematical formulation step by step, implement it in code, visualize how different heads behave, and understand exactly why extrapolation works. We will also compare ALiBi to RoPE, discuss practical trade-offs, and examine which real-world models have adopted it.

The Extrapolation Problem

Before understanding ALiBi's solution, we need to appreciate the problem it solves. Transformers trained on sequences of length LL often fail dramatically when given sequences of length 2L2L or beyond. This is not a subtle degradation: performance can collapse entirely, with the model producing incoherent outputs for inputs only slightly longer than the training window.

Length Extrapolation

Length extrapolation refers to a model's ability to process sequences longer than those encountered during training while maintaining reasonable performance. Many position encoding schemes fail this test because they produce position representations the model has never learned to interpret.

With sinusoidal encodings, the encoding at position 1025 is a specific vector of cosines and sines that the model was never trained to use. The model's linear and attention layers have developed weights that respond appropriately to positions 0 through 1024, but position 1025 looks foreign, and the model has no basis for interpreting it correctly. With learned position embeddings, the situation is even worse: positions beyond the training length simply do not exist in the embedding table. The model cannot process them at all without modification.

RoPE handles this more gracefully because its rotations are defined by a mathematical formula rather than learned weights. In theory, you can compute a rotation for any position. In practice, however, the model's layers have been tuned on rotation patterns corresponding to positions 0 through LL, and rotation angles at positions far beyond LL can fall into regions of the complex plane that the model has never been trained to interpret. Research teams working with RoPE models often need additional techniques like position interpolation or NTK-aware scaling to extend the training context length.

Why does extrapolation matter so much in practice? First, training on very long sequences is extraordinarily expensive. Attention's quadratic complexity means that doubling the sequence length quadruples the attention computation. If you could train on sequences of 1,024 tokens and reliably deploy on sequences of 4,096 or 8,192 tokens, you would save enormous training costs. Second, real-world inputs are unpredictable. A conversational AI might encounter a long document, a code model might see a large file, and a summarization system might receive a lengthy article. A model that degrades gracefully on longer inputs is far more practically valuable than one that requires the input to stay within a hard length ceiling.

The extrapolation problem is essentially a generalization problem across a dimension that most machine learning practitioners don't spend much time thinking about: the sequence length dimension. We routinely generalize across input content, but generalizing across input structure, specifically the sequential length structure, requires position encodings that work by relative relationship rather than absolute coordinate.

The Core Idea: Penalizing Distance

ALiBi's insight is elegant in its simplicity: do not encode position in the embeddings at all. Instead, modify the attention mechanism itself to prefer nearby tokens over distant ones. Embeddings carry semantic content; let attention carry positional structure.

Think of it as a gravitational field. Every token exerts an attentional "pull" on every other token based on content compatibility. ALiBi adds a distance decay to this pull: the further two tokens are from each other, the weaker the pull, regardless of content. Nearby tokens always have a structural advantage. This is not unlike how we read: we naturally weight recent context more heavily than distant context when interpreting the meaning of a word.

Consider what attention scores represent. They measure compatibility between a query and a key. After softmax, these scores determine how much each position contributes to the output. ALiBi introduces a simple bias that subtracts from the score based on distance. To compute the modified attention score between position ii (as the query) and position jj (as the key), we start with the standard dot product and then impose the distance penalty:

scoreij=qi⋅kj−m⋅∣i−j∣\text{score}_{ij} = \mathbf{q}_i \cdot \mathbf{k}_j - m \cdot |i - j|

where:

  • scoreij\text{score}_{ij}: the modified attention score between positions ii and jj
  • qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j: the original dot product between query ii and key jj, capturing content-based compatibility
  • mm: a slope parameter that controls how aggressively to penalize distance (different for each attention head)
  • ∣i−j∣|i - j|: the absolute distance between the two positions in the sequence

The subtraction means distant tokens receive lower scores. A token 10 positions away gets penalized by 10m10m, while an adjacent token gets penalized by just mm. After softmax normalization, this translates to nearby tokens receiving higher attention weights. The key insight is that the penalty is linear in distance: every additional step away costs exactly mm more in score, regardless of whether you are stepping from position 1 to position 2 or from position 1000 to position 1001.

ALiBi Bias

ALiBi adds a linear bias to attention scores based on the distance between query and key positions. The bias is always non-positive, penalizing distant positions. The penalty grows linearly with distance, controlled by a slope parameter mm that is fixed (not learned) and specific to each attention head.

This is position encoding without position embeddings. The model learns nothing about absolute position during pretraining of its embeddings. Position information enters only through the attention bias, and only at the moment of computing attention scores. The embeddings remain clean semantic representations of token content, and the relative positional structure is layered on top at attention time.

Why does the penalty have to be linear? The authors experimented with other functional forms, including concave penalties that taper off quickly and convex penalties that grow steeply. Linear penalties produced the best combination of in-distribution performance and out-of-distribution (long-sequence) extrapolation. The linear form also has a pleasing mathematical property: it represents a constant cost per unit of distance, which creates a uniform "discount rate" on attention over distance. This discount rate is consistent regardless of absolute position, which is the property that makes extrapolation work.

The Mathematical Formulation

Now that we understand ALiBi's core insight, let's develop the complete mathematical framework. We will build from standard attention to ALiBi-augmented attention, showing exactly where and how the position bias enters the computation.

Standard attention and ALiBi attention are nearly identical. The only difference is one additional term added to the score matrix before the softmax. This surgical simplicity is part of what makes ALiBi attractive: it requires minimal changes to an existing attention implementation and adds almost no computational overhead.

Starting Point: Standard Scaled Dot-Product Attention

Recall that standard self-attention operates on three matrices derived from the input sequence. For a sequence of nn tokens, the attention mechanism computes:

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

where:

  • Q∈Rn×dkQ \in \mathbb{R}^{n \times d_k}: the query matrix containing query vectors for all nn positions
  • K∈Rn×dkK \in \mathbb{R}^{n \times d_k}: the key matrix containing key vectors for all positions
  • V∈Rn×dvV \in \mathbb{R}^{n \times d_v}: the value matrix containing value vectors for all positions
  • dkd_k: the dimension of queries and keys
  • QKT∈Rn×nQK^T \in \mathbb{R}^{n \times n}: the raw attention score matrix where entry (i,j)(i, j) is the dot product between query ii and key jj
  • dk\sqrt{d_k}: scaling factor that prevents dot products from growing too large in high dimensions, which would push the softmax into regions of very small gradients

The matrix QKTQK^T captures content-based similarity: how much each query "wants" to attend to each key based purely on their learned representations. This matrix is completely blind to position. Token 1 attending to token 2 produces the same content-based score whether they appear next to each other or 500 positions apart. This is by design for the core attention mechanism, which wants to capture global relationships. But it means we need another mechanism to introduce positional awareness. That mechanism is the ALiBi bias.

Why does ALiBi add bias at this stage rather than modifying the queries and keys? Adding to the pre-softmax scores is the most direct possible intervention. The scores are the raw votes about which tokens matter. By adjusting the scores before normalization, we influence the attention distribution in a clear and simple way. Modifying queries and keys (as RoPE does) achieves a similar effect but through a more complex indirect path involving embedding geometry.

Injecting Position: The Bias Matrix

ALiBi's modification is surgical. Rather than changing how queries, keys, or values are computed, it adds a single term to the attention scores. The complete ALiBi-augmented attention reads:

Attention(Q,K,V)=softmax(QKTdk+B)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}} + B\right)V

The bias matrix B∈Rn×nB \in \mathbb{R}^{n \times n} is added directly to the scaled scores before softmax. This is the only change to standard attention. The matrix BB encodes position information through a remarkably simple rule: each entry is determined entirely by the distance between the corresponding positions. Specifically,

Bij=−m⋅∣i−j∣B_{ij} = -m \cdot |i - j|

where:

  • BijB_{ij}: the bias added to the attention score between query position ii and key position jj
  • mm: the slope parameter for this attention head, controlling the strength of the distance penalty
  • ∣i−j∣|i - j|: the absolute distance between positions ii and jj in the token sequence

Let's unpack why this formula works. The absolute distance ∣i−j∣|i - j| measures how far apart two positions are in the sequence. Adjacent tokens have distance 1, tokens separated by 10 positions have distance 10. Multiplying by a positive slope mm and negating creates a penalty that grows linearly with distance. A token that is close gets a small penalty; a token that is far gets a large penalty.

Consider what happens for a query at position 5 attending to various keys:

  • Attending to position 5 (itself): B55=−m⋅0=0B_{55} = -m \cdot 0 = 0 (no penalty, full score preserved)
  • Attending to position 4 (adjacent): B54=−m⋅1=−mB_{54} = -m \cdot 1 = -m (small penalty)
  • Attending to position 3 (two steps away): B53=−m⋅2=−2mB_{53} = -m \cdot 2 = -2m (moderate penalty)
  • Attending to position 0 (distant): B50=−m⋅5=−5mB_{50} = -m \cdot 5 = -5m (large penalty)

The negative bias reduces the attention score, and larger distances produce more negative biases. After softmax normalization, this translates to lower attention weights for distant tokens. The content-based similarity in QKTQK^T still matters completely: a very semantically relevant distant token can overcome its distance penalty if its dot product score is large enough. But now content has to compete against a structural penalty that favors proximity.

Why does this formula make sense? Notice that the formula simply implements a universal assumption that language models hold implicitly: nearby context tends to be more relevant than distant context. Rather than forcing the model to learn this bias from data (which is expensive and may lead to overfitting on long-range spurious correlations), ALiBi builds it in as an inductive bias. The model is free to override it with strong content signals, but the default behavior is local attention. This is similar to how convolutional neural networks build in a translation-invariance inductive bias, rather than forcing the network to learn that from scratch.

The Causal Case: Masking the Future

For decoder-style models that generate text left-to-right, we need causal masking: position ii can only attend to positions j≤ij \le i. The combined bias matrix integrates both the ALiBi distance penalties and the causal mask:

B=[0−∞−∞−∞⋯−m0−∞−∞⋯−2m−m0−∞⋯−3m−2m−m0⋯⋮⋮⋮⋮⋱]B = \begin{bmatrix} 0 & -\infty & -\infty & -\infty & \cdots \\ -m & 0 & -\infty & -\infty & \cdots \\ -2m & -m & 0 & -\infty & \cdots \\ -3m & -2m & -m & 0 & \cdots \\ \vdots & \vdots & \vdots & \vdots & \ddots \end{bmatrix}

The structure reveals two distinct components working together. The lower triangle, including the diagonal, contains the ALiBi distance penalties. Entry (i,j)(i, j) with j≤ij \le i holds the value −m⋅(i−j)-m \cdot (i - j), which is zero on the diagonal and grows more negative as you move left along any row. The upper triangle holds −∞-\infty values from the causal mask.

The −∞-\infty entries ensure that after applying exp⁡(−∞)=0\exp(-\infty) = 0 inside the softmax, future positions contribute nothing to the weighted sum. The linear penalties in the lower triangle then shape how attention flows among the allowed past positions. These two components operate in entirely different regimes: the causal mask is a hard prohibition (zero weight, always), while the ALiBi penalty is a soft preference (lower weight, but nonzero).

Notice that the causal mask interacts naturally with ALiBi. For a model processing position ii, the only tokens it can attend to are positions 0 through ii. These are exactly the tokens for which j≤ij \le i, so the distance is always i−ji - j, always positive, and always measured in the backward direction. The ALiBi penalty is therefore always a backward-looking decay, consistently favoring the most recent tokens in the context.

Building the Bias Matrix: Step-by-Step Implementation

Let's translate this mathematics into code, building up the bias matrix piece by piece. The core insight behind the implementation is that we can compute all pairwise distances simultaneously using broadcasting, rather than iterating over all pairs with nested loops.

In[3]:
Code
import numpy as np


def create_alibi_bias(seq_len, slope):
    """
    Create the ALiBi bias matrix for a single attention head.

    Args:
        seq_len: Length of the sequence
        slope: The slope parameter m for this head

    Returns:
        Bias matrix of shape (seq_len, seq_len)
    """
    # Create position indices
    positions = np.arange(seq_len)

    # Compute pairwise distances: |i - j|
    # positions[:, None] is (seq_len, 1), positions[None, :] is (1, seq_len)
    # Broadcasting gives (seq_len, seq_len)
    distances = np.abs(positions[:, None] - positions[None, :])

    # Apply linear penalty
    bias = -slope * distances

    return bias

The key implementation insight is using NumPy broadcasting to compute all pairwise distances at once. By reshaping the position array into a column vector positions[:, None] and a row vector positions[None, :], subtraction produces an n×nn \times n matrix where entry (i,j)(i, j) is i−ji - j. Taking the absolute value gives us the distance matrix. Multiplying by the slope and negating gives the penalty matrix. This vectorized approach is both faster and more readable than explicit loops.

Let's see what this produces for a small sequence:

Out[4]:
Console
ALiBi bias matrix for sequence length 6, slope 0.5:
[[-0.  -0.5 -1.  -1.5 -2.  -2.5]
 [-0.5 -0.  -0.5 -1.  -1.5 -2. ]
 [-1.  -0.5 -0.  -0.5 -1.  -1.5]
 [-1.5 -1.  -0.5 -0.  -0.5 -1. ]
 [-2.  -1.5 -1.  -0.5 -0.  -0.5]
 [-2.5 -2.  -1.5 -1.  -0.5 -0. ]]

Reading this matrix: the diagonal is zero because a token attending to itself has distance zero and receives no penalty. Moving away from the diagonal in either direction, penalties grow linearly. Position 5 (row 5) attending to position 0 (column 0) shows −2.5-2.5, which is exactly −0.5×5-0.5 \times 5 as expected. The matrix is symmetric because distance is symmetric: ∣i−j∣=∣j−i∣|i - j| = |j - i|. In a causal model we will only use the lower triangle, but the full symmetric matrix is the natural output of the distance computation.

For decoder models, we overlay the causal mask on top of the distance penalties:

In[5]:
Code
def create_causal_alibi_bias(seq_len, slope):
    """
    Create ALiBi bias with causal masking.

    Args:
        seq_len: Length of the sequence
        slope: The slope parameter m

    Returns:
        Bias matrix with -inf for future positions
    """
    # Start with the distance-based bias
    bias = create_alibi_bias(seq_len, slope)

    # Create causal mask (upper triangle should be -inf)
    causal_mask = np.triu(np.ones((seq_len, seq_len)), k=1)

    # Apply causal mask: future positions get -inf
    bias = np.where(causal_mask == 1, -np.inf, bias)

    return bias

The np.triu function creates an upper triangular matrix of ones, which we use to identify positions that should be masked. The k=1 argument excludes the diagonal, since a token should be able to attend to itself. The np.where call then replaces those positions with negative infinity, which becomes effectively zero after the exponential in softmax.

Out[6]:
Console
Causal ALiBi bias matrix:
[[-0.  -inf -inf -inf -inf -inf]
 [-0.5 -0.  -inf -inf -inf -inf]
 [-1.  -0.5 -0.  -inf -inf -inf]
 [-1.5 -1.  -0.5 -0.  -inf -inf]
 [-2.  -1.5 -1.  -0.5 -0.  -inf]
 [-2.5 -2.  -1.5 -1.  -0.5 -0. ]]

Now the upper triangle shows inf (NumPy's display for −∞-\infty), while the lower triangle retains the distance penalties. Row 5 can attend to all previous positions with penalties −2.5,−2.0,−1.5,−1.0,−0.5,0.0-2.5, -2.0, -1.5, -1.0, -0.5, 0.0 for positions 0 through 5 respectively. Row 0 can only attend to itself with penalty 0, since there are no previous positions available.

Let's visualize both the distance matrix and the resulting causal ALiBi bias matrix side by side to make the structure visually clear:

Out[7]:
Visualization
Heatmap of distance matrix with values increasing symmetrically away from the diagonal.
Distance matrix showing absolute distances between all position pairs. Entry (i, j) contains |i - j|. The matrix is symmetric with zeros on the diagonal and values increasing as positions diverge.
Heatmap of causal ALiBi bias matrix showing negative values in the lower triangle and masked upper triangle.
Causal ALiBi bias matrix with slope 0.5. The distance matrix is negated and scaled to produce penalties, then the upper triangle is masked with negative infinity to prevent attending to future positions.

The left matrix shows raw distances: the diagonal is 0 (same position), adjacent cells are 1, and values grow as positions diverge. The right matrix shows what happens after applying ALiBi: distances become negative penalties (scaled by the slope), and the upper triangle is masked to enforce causality. This visual makes clear that the bias matrix is nothing more than a scaled, negated distance matrix with causal masking applied on top.

Head-Specific Slopes

A single slope would force all attention heads to have the same locality preference. That would be limiting: different aspects of language operate at different scales, and a model benefits from heads that specialize in different scopes. Some heads might focus on very local context (steep slope, harsh penalties for distance), while others maintain broader receptive fields (gentle slope, mild penalties).

Think of it like a camera system with multiple lenses. A telephoto lens captures fine detail up close but misses the broader scene. A wide-angle lens captures the full picture but loses local detail. You want both working together. ALiBi gives each attention head a different "focal length" by assigning it a different slope. The ensemble of heads covers a wide spectrum of temporal scales, from immediate local syntax to long-range topic coherence.

The slopes are not learned but fixed according to a geometric sequence. The decision to fix them (rather than learn them) is deliberate: it eliminates hyperparameters from the position encoding, makes the encoding fully deterministic and reproducible, and avoids the risk of the model learning pathological slope values during training. For a model with hh attention heads, the slope for head ii is:

mi=128h⋅ifor i=1,2,…,hm_i = \frac{1}{2^{\frac{8}{h} \cdot i}} \quad \text{for } i = 1, 2, \ldots, h

where:

  • mim_i: the slope parameter for attention head ii
  • hh: the total number of attention heads in the model
  • ii: the head index, ranging from 1 to hh
  • 8h\frac{8}{h}: a scaling factor that ensures slopes span a consistent range regardless of how many heads the model has
  • 28h⋅i2^{\frac{8}{h} \cdot i}: the denominator that grows exponentially with head index, making later heads have gentler slopes

Why does this formula make sense? Notice that the exponent 8h⋅i\frac{8}{h} \cdot i ranges from 8h\frac{8}{h} (head 1) to 88 (head hh). The denominator is therefore always 28h2^{\frac{8}{h}} through 28=2562^8 = 256, independent of hh. This ensures that the range of slopes remains consistent across model architectures with different head counts. Whether you have 4 heads or 32 heads, the steepest head always has a slope near 1/28/h1/2^{8/h} and the gentlest head always has a slope near 1/256≈0.0041/256 \approx 0.004.

The base of 8 in the exponent was chosen empirically by the ALiBi authors. With 8 heads, the exponent simplifies to just ii, giving slopes of 121,122,…,128\frac{1}{2^1}, \frac{1}{2^2}, \ldots, \frac{1}{2^8}, which equals 0.5,0.25,0.125,…,0.003906250.5, 0.25, 0.125, \ldots, 0.00390625. With other head counts, the formula interpolates appropriately.

In[8]:
Code
def get_alibi_slopes(num_heads):
    """
    Compute ALiBi slopes for each attention head.

    The slopes follow a geometric sequence, with the first head
    having the steepest slope (most local attention) and the last
    head having the gentlest slope (broadest attention).

    Args:
        num_heads: Number of attention heads

    Returns:
        Array of slopes, one per head
    """
    # Compute the ratio for the geometric sequence
    ratio = 2 ** (8 / num_heads)

    # Generate slopes: 1/ratio, 1/ratio^2, ..., 1/ratio^num_heads
    slopes = 1.0 / (ratio ** np.arange(1, num_heads + 1))

    return slopes
Out[9]:
Console
4 heads: slopes = [0.25     0.0625   0.015625 0.003906]
8 heads: slopes = [0.5      0.25     0.125    0.0625   0.03125  0.015625 0.007812 0.003906]
16 heads: slopes = [0.707107 0.5      0.353553 0.25     0.176777 0.125    0.088388 0.0625
 0.044194 0.03125  0.022097 0.015625 0.011049 0.007812 0.005524 0.003906]

The geometric progression ensures that slopes span several orders of magnitude. With 8 heads, the steepest slope (0.5) penalizes a distance of 10 by 5 logits, effectively eliminating distant tokens from consideration after softmax. The gentlest slope (roughly 0.004) penalizes the same distance by only 0.04 logits, allowing the head to attend broadly across hundreds of positions. This spread is what gives ALiBi its multi-scale character.

Notice that the formula produces a different set of slopes for 4, 8, and 16 heads, but in each case, the slopes span roughly the same overall range from steep to gentle. With 4 heads, you get fewer granular distinctions but the extremes remain similar. With 16 heads, you get finer granularity, more heads at intermediate slopes, but still roughly the same endpoints. This consistency is the purpose of the 8/h8/h scaling factor.

Out[10]:
Visualization
Bar chart showing ALiBi slopes for 8 attention heads, with values decreasing geometrically from 0.5 for head 1 to near zero for head 8.
ALiBi slopes across 8 attention heads, shown on a linear scale. The geometric progression creates a rapid decay from head 1 (slope 0.5, strong locality bias) to head 8 (slope near 0.004, broad attention). Each head effectively operates on a different spatial scale, allowing the model to capture both local and distant dependencies simultaneously.

Visualizing the Attention Bias

The slope values are informative in the abstract, but seeing what they produce as attention bias matrices makes the effect concrete. Let's visualize how ALiBi biases shape attention patterns for the steepest and gentlest heads:

Out[11]:
Visualization
Heatmap showing ALiBi bias for head 1 with steep slope, displaying strong negative values far from the diagonal.
ALiBi bias matrix for head 1 (slope 0.5, steepest). Strong negative values appear just a few positions from the diagonal, creating a narrow attention window. A query at position 18 attending to position 8 (ten steps away) faces a bias of -5.0, effectively eliminating that token from consideration after softmax.
Out[12]:
Visualization
Heatmap showing ALiBi bias for head 8 with gentle slope, displaying mild penalties across all distances.
ALiBi bias matrix for head 8 (slope 0.0039, gentlest). The penalties are so mild that even tokens 19 positions away face a bias of only -0.074. This head can attend broadly across the entire available context with minimal positional discounting.

The contrast is stark and instructive. Head 1's bias matrix shows deep red (strongly negative) values just a few positions from the diagonal. By position 10, the penalty exceeds -5 logits, making those tokens nearly invisible after softmax unless their content score is extraordinarily strong. The effective attention window for this head is perhaps 3 to 5 tokens wide. Head 8, in contrast, shows nearly uniform mild penalties throughout. Even at distance 19, the penalty is less than 0.1 logits, allowing meaningful attention to distant tokens across essentially the entire sequence.

This division of labor is intentional and reflects something real about natural language. Linguistic phenomena operate at different scales. Adjacent tokens matter for subject-verb agreement, determiner-noun matching, and immediate syntactic structure. Tokens a few steps away matter for multi-word expressions, phrasal dependencies, and clause-level coherence. Tokens many steps away matter for coreference resolution (a pronoun referring back to a noun introduced many sentences earlier), topic consistency, and document-level discourse structure. By giving different heads different locality preferences, ALiBi enables the model to capture phenomena at multiple scales simultaneously without requiring the model to learn these scales from scratch.

ALiBi in Attention: Complete Implementation

Now we have all the pieces to implement ALiBi-augmented attention from scratch. The implementation closely mirrors standard attention. We compute and scale the score matrix, add the bias, apply the optional causal mask, take the softmax, and use the result to weight the values.

The only new element is the construction and addition of the ALiBi bias tensor. For multi-head attention, we need one bias matrix per head, which we stack into a three-dimensional tensor of shape (num_heads, seq_len, seq_len). Each slice along the first dimension is the bias matrix for one head, with that head's specific slope.

In[13]:
Code
def softmax(x, axis=-1):
    """Numerically stable softmax."""
    x_max = np.max(x, axis=axis, keepdims=True)
    exp_x = np.exp(x - x_max)
    return exp_x / np.sum(exp_x, axis=axis, keepdims=True)


def alibi_attention(Q, K, V, slopes, causal=True):
    """
    Compute attention with ALiBi position encoding.

    Args:
        Q: Query matrix of shape (num_heads, seq_len, d_k)
        K: Key matrix of shape (num_heads, seq_len, d_k)
        V: Value matrix of shape (num_heads, seq_len, d_v)
        slopes: ALiBi slopes, one per head, shape (num_heads,)
        causal: Whether to apply causal masking

    Returns:
        Output of shape (num_heads, seq_len, d_v)
        Attention weights of shape (num_heads, seq_len, seq_len)
    """
    num_heads, seq_len, d_k = Q.shape

    # Compute raw attention scores: Q @ K^T
    # Shape: (num_heads, seq_len, seq_len)
    scores = np.matmul(Q, K.transpose(0, 2, 1))

    # Scale by sqrt(d_k)
    scores = scores / np.sqrt(d_k)

    # Create ALiBi biases for each head
    # Shape: (num_heads, seq_len, seq_len)
    positions = np.arange(seq_len)
    distances = np.abs(positions[:, None] - positions[None, :])

    # Broadcast slopes: (num_heads, 1, 1) * (seq_len, seq_len)
    alibi_bias = -slopes[:, None, None] * distances[None, :, :]

    # Apply causal mask if needed
    if causal:
        causal_mask = np.triu(np.ones((seq_len, seq_len)), k=1)
        alibi_bias = np.where(causal_mask == 1, -np.inf, alibi_bias)

    # Add ALiBi bias to scores
    scores = scores + alibi_bias

    # Apply softmax to get attention weights
    attention_weights = softmax(scores, axis=-1)

    # Handle NaN from -inf (all masked positions)
    attention_weights = np.nan_to_num(attention_weights, nan=0.0)

    # Compute output: weighted sum of values
    output = np.matmul(attention_weights, V)

    return output, attention_weights

The broadcasting step deserves attention. We have slopes with shape (num_heads,). Reshaping to (num_heads, 1, 1) and multiplying by distances[None, :, :] with shape (1, seq_len, seq_len) produces a tensor of shape (num_heads, seq_len, seq_len). Each head's slice is the distance matrix multiplied by that head's slope. This is a single vectorized operation that replaces what would otherwise be a loop over heads.

Let's test this implementation with a concrete example to verify the shapes and inspect the attention patterns:

In[14]:
Code
# Create a simple example
np.random.seed(42)

num_heads = 4
seq_len = 8
d_k = 16
d_v = 16

# Random Q, K, V matrices
Q = np.random.randn(num_heads, seq_len, d_k) * 0.5
K = np.random.randn(num_heads, seq_len, d_k) * 0.5
V = np.random.randn(num_heads, seq_len, d_v) * 0.5

# Get ALiBi slopes
slopes = get_alibi_slopes(num_heads)

# Compute ALiBi attention
output, attention_weights = alibi_attention(Q, K, V, slopes, causal=True)
Out[15]:
Console
Output shape: (4, 8, 16)
Attention weights shape: (4, 8, 8)

Attention weights for head 1 (steep slope = 0.250):
[[1.    0.    0.    0.    0.    0.    0.    0.   ]
 [0.45  0.55  0.    0.    0.    0.    0.    0.   ]
 [0.233 0.37  0.397 0.    0.    0.    0.    0.   ]
 [0.227 0.205 0.262 0.306 0.    0.    0.    0.   ]
 [0.118 0.068 0.121 0.279 0.414 0.    0.    0.   ]
 [0.083 0.086 0.13  0.201 0.176 0.324 0.    0.   ]
 [0.065 0.089 0.092 0.127 0.137 0.272 0.218 0.   ]
 [0.025 0.038 0.057 0.073 0.136 0.214 0.233 0.224]]

Attention weights for head 4 (gentle slope = 0.003906):
[[1.    0.    0.    0.    0.    0.    0.    0.   ]
 [0.562 0.438 0.    0.    0.    0.    0.    0.   ]
 [0.36  0.453 0.187 0.    0.    0.    0.    0.   ]
 [0.344 0.23  0.245 0.181 0.    0.    0.    0.   ]
 [0.184 0.232 0.181 0.169 0.233 0.    0.    0.   ]
 [0.121 0.125 0.286 0.214 0.096 0.158 0.    0.   ]
 [0.104 0.124 0.171 0.176 0.08  0.175 0.169 0.   ]
 [0.109 0.137 0.063 0.124 0.158 0.147 0.163 0.099]]

Compare the attention patterns between heads. Head 1 with its steep slope concentrates attention heavily on recent positions, with weights dropping rapidly as distance increases. The rows show a clear diagonal structure: most of the weight is on the most recent 1 to 3 tokens, with essentially nothing beyond 5 positions. Head 4 with its gentle slope distributes attention more evenly across the available context. The rows are more uniform, with distant positions receiving non-negligible weight.

To understand the impact of ALiBi more directly, let's compare attention patterns with and without the position bias. We'll compute attention for the same queries and keys, once with standard attention (no position encoding) and once with ALiBi:

Out[16]:
Visualization
Heatmap showing attention weights without position encoding, with relatively uniform distribution.
Standard attention without position encoding. Attention weights depend only on content similarity between queries and keys. The pattern is irregular and non-local, with no systematic preference for nearby tokens.
Heatmap showing attention weights with ALiBi, showing stronger diagonal pattern due to locality bias.
ALiBi attention with the same queries and keys. The position bias systematically pulls weight toward recent tokens, creating a diagonal structure. Content signals can still override the bias for strongly relevant distant tokens.

The difference is striking. Standard attention distributes weight based purely on content similarity, sometimes attending strongly to distant positions because the content happens to match well. ALiBi reshapes this pattern, pulling attention toward recent tokens while still allowing content to influence the final distribution. This locality bias emerges from a single matrix addition, requiring no learned position embeddings and no modifications to the token representations themselves.

Out[17]:
Visualization
Heatmap of attention weights for head 1 showing strong diagonal pattern with rapid falloff.
Attention weights for head 1 (steep slope 0.5). Each row shows how one query position distributes attention across key positions. The strong diagonal pattern confirms that attention concentrates sharply on nearby tokens, with distant positions receiving negligible weight despite their content.
Out[18]:
Visualization
Heatmap of attention weights for head 4 showing broader attention distribution.
Attention weights for head 4 (slope 0.0625). Attention spreads more broadly across all available positions, with no strong locality bias. This head can pick up on long-range semantic relationships that the steep-slope heads miss.

Worked Example: Tracing Through ALiBi Step by Step

Let's work through a concrete numerical example to solidify the mechanics. We'll use a very small sequence so every number is legible.

Consider a sequence of 4 tokens with 2 attention heads and dk=4d_k = 4. We'll use simple numbers so you can verify each step by hand.

In[19]:
Code
# Small worked example with traceable numbers
np.random.seed(0)
n = 4  # sequence length
h = 2  # number of heads
dk = 4  # key/query dimension

# Simple Q and K with small values
Q_ex = np.array(
    [
        [
            [1.0, 0.0, 0.5, 0.0],
            [0.0, 1.0, 0.0, 0.5],
            [0.5, 0.0, 1.0, 0.0],
            [0.0, 0.5, 0.0, 1.0],
        ],
        [
            [1.0, 0.0, 0.5, 0.0],
            [0.0, 1.0, 0.0, 0.5],
            [0.5, 0.0, 1.0, 0.0],
            [0.0, 0.5, 0.0, 1.0],
        ],
    ]
)  # shape (2, 4, 4)

K_ex = Q_ex.copy()  # Use K = Q for simplicity
V_ex = np.random.randn(h, n, dk) * 0.1

slopes_ex = get_alibi_slopes(h)  # slopes for 2 heads
Out[20]:
Console
Step 1: Raw dot-product scores QK^T (head 1):
[[1.25 0.   1.   0.  ]
 [0.   1.25 0.   1.  ]
 [1.   0.   1.25 0.  ]
 [0.   1.   0.   1.25]]

Step 2: Scaled scores (divided by sqrt(4) = 2.000) (head 1):
[[0.625 0.    0.5   0.   ]
 [0.    0.625 0.    0.5  ]
 [0.5   0.    0.625 0.   ]
 [0.    0.5   0.    0.625]]

Step 3: ALiBi bias matrix (head 1, slope = 0.062):
[[-0.    -0.062 -0.125 -0.188]
 [-0.062 -0.    -0.062 -0.125]
 [-0.125 -0.062 -0.    -0.062]
 [-0.188 -0.125 -0.062 -0.   ]]

Step 4: ALiBi bias with causal mask applied (head 1):
[[-0.      -inf   -inf   -inf]
 [-0.062 -0.      -inf   -inf]
 [-0.125 -0.062 -0.      -inf]
 [-0.188 -0.125 -0.062 -0.   ]]

Step 5: Final scores (scaled + bias) (head 1):
[[ 0.625   -inf   -inf   -inf]
 [-0.062  0.625   -inf   -inf]
 [ 0.375 -0.062  0.625   -inf]
 [-0.188  0.375 -0.062  0.625]]

Step 6: Attention weights after softmax (head 1):
[[1.    0.    0.    0.   ]
 [0.335 0.665 0.    0.   ]
 [0.341 0.22  0.438 0.   ]
 [0.163 0.286 0.184 0.367]]

Let's read through this trace carefully. In step 1, the raw dot-product scores capture content similarity. Since we used Q=KQ = K with identity-like structure, each token scores highest against itself. In step 2, we scale by 4=2\sqrt{4} = 2, which shrinks all values without changing their relative ordering. In step 3, the ALiBi bias is purely a function of distance: entry (i,j)(i, j) is −m⋅∣i−j∣-m \cdot |i-j| where mm is the head's slope. For head 1 with the steeper slope, the penalties at distance 1, 2, and 3 are meaningful numbers. In step 4, the upper triangle is set to −∞-\infty to enforce causality. In step 5, we add the bias to the scaled scores. Notice how the off-diagonal entries drop: the content similarity gets penalized by distance, pulling the distribution toward the diagonal. In step 6, softmax turns these into probabilities. The diagonal entries (attending to self) dominate strongly for head 1 because the ALiBi penalty has suppressed the off-diagonal scores.

This step-by-step trace illustrates the key property of ALiBi: it does not change what the model "wants" to attend to (the content signal), but it imposes a systematic cost for looking farther away. When content signals are strong, they can dominate. When content signals are weak or comparable, the distance penalty tips the balance toward nearby tokens.

Why ALiBi Extrapolates

The key to ALiBi's extrapolation ability lies in what the model learns during training. With other position encoding schemes, the model learns to interpret specific position representations. A sinusoidal encoding for position 512 produces a particular vector that the model has seen and learned to use. A learned position embedding for position 512 is a specific trained vector. When you encounter position 1025, this is a novel representation the model has not seen, and its layers have no well-defined response to it.

ALiBi sidesteps this problem entirely. The model never learns position representations because there are none to learn. Instead, it learns to work with relative attention patterns shaped by the linear bias. During training on sequences of length 1024, the model sees attention patterns where nearby tokens are favored and distant tokens are penalized. This is true at position 10, position 500, and position 1000. At every position in the training range, the model observes the same structural pattern: recent tokens have higher weight, distant tokens have lower weight.

When inference extends to position 2048, the same principle applies. The local neighborhood still receives favorable bias. Tokens 10 positions away still get penalized by the same amount. The absolute positions are larger, but the relative structure is completely unchanged. The model has learned to extract information from attention patterns that favor locality, and those patterns remain consistent regardless of sequence length.

Extrapolation Mechanism

ALiBi extrapolates because it encodes relative distance, not absolute position. The linear penalty for distance 10 is −10m-10m whether you are at position 50 or position 5000. The model learns to work with distance-biased attention patterns, which remain structurally consistent across all sequence lengths. This is the fundamental distinction from methods that encode absolute position as a vector.

There is a subtle but important point here. When a model processes a long sequence at inference time, tokens near the beginning of the sequence will be far from the current query position. ALiBi will penalize them heavily, potentially so heavily that they receive near-zero attention weight. This is not a failure: it is the expected behavior. The model is effectively operating with a soft window: it pays close attention to recent context and discounts the distant past. For sequences significantly longer than the training length, the very first tokens may be practically invisible to later queries. Whether this hurts performance depends on the task: for many language modeling and generation tasks, the most relevant context is recent, so this graceful distance falloff can improve performance.

Let's visualize the consistency property that makes extrapolation work:

Out[21]:
Visualization
Line plot showing ALiBi bias as a function of distance for three different query positions, all following the same linear penalty curve.
ALiBi bias as a function of distance for three query positions (20, 100, and 500). All three curves are identical because the bias depends only on distance, not on absolute position. This structural consistency is what allows models trained on short sequences to process longer ones without degradation.

The curves overlap perfectly because ALiBi's bias depends only on distance, never on absolute position. Whether the query is at token 20 or token 5000, the bias for a key 30 positions back is the same: −30m-30m. This is the mathematical basis for length extrapolation, and it is also the mathematical reason why it could not work for sinusoidal or learned position embeddings: those methods embed absolute position, so the representation at position 5000 is fundamentally different from (and trained less on than) the representation at position 20.

ALiBi vs. RoPE: A Comparison

Both ALiBi and RoPE are widely used in modern language models, and both address relative position. But they take fundamentally different approaches that lead to different trade-offs in practice.

RoPE encodes position by rotating query and key vectors. When computing dot products between rotated vectors, the rotation angles combine such that the result depends on relative position rather than absolute position. This is mathematically elegant: the dot product qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j after rotation naturally depends on i−ji - j rather than on ii and jj separately. RoPE modifies the embedding space itself: the vectors that participate in dot products carry rotational information baked into their components.

ALiBi's approach is different in character: it does not modify the vectors at all. The dot products qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j remain the raw content-based scores. ALiBi then overlays a separate, explicit position signal on top of those scores. The two components are additive and fully transparent. You can look at a score and immediately decompose it into a content part (the dot product) and a position part (the bias).

ALiBi vs. RoPE comparison across key dimensions.
AspectALiBiRoPE
Position encoding locationAttention score biasQuery and key embeddings
MechanismSubtracts linear penalty from attention scoresRotates Q and K vectors by position-dependent angles
ParametersFixed slopes, no learned position parametersNo additional parameters
ComputationSimple matrix addition to score matrixComplex number arithmetic or 2D rotation matrices
ExtrapolationStrong out-of-the-box at long sequencesRequires additional techniques (NTK scaling, interpolation)
InterpretabilityFully transparent: slope directly controls window widthLess direct: rotation angle interplay is harder to interpret
Impact on embeddingsNone: token embeddings are pure semantic vectorsEmbeds absolute rotation into token representations

The choice between ALiBi and RoPE often comes down to empirical performance on the target task. RoPE has shown excellent results on many benchmarks and is used by prominent models including LLaMA and Mistral, as well as Falcon. ALiBi has shown strong extrapolation and is used by BLOOM and MPT. Newer research (including models like LLaMA 3 and Mixtral) tends to favor RoPE with context extension techniques, suggesting that RoPE's expressiveness may ultimately win out as practitioners develop better tools for extending its context window. However, ALiBi remains a practical choice when implementation simplicity and extrapolation matter, especially when minimizing implementation risk.

Let's compare the implementation complexity concretely, which is where ALiBi's advantage is clearest:

In[22]:
Code
def rope_attention(Q, K, V, seq_len, d_k, causal=True):
    """
    Simplified RoPE attention for comparison.
    This is a sketch showing the additional complexity.
    """
    # RoPE requires computing rotation matrices or using complex numbers
    # For each position, apply rotation to Q and K before computing attention

    # Step 1: Compute rotation angles for each position and dimension
    positions = np.arange(seq_len)
    dim_indices = np.arange(d_k // 2)

    # Frequency for each dimension pair (simplified)
    freqs = 1.0 / (10000 ** (2 * dim_indices / d_k))

    # Angle matrix: (seq_len, d_k/2)
    angles = positions[:, None] * freqs[None, :]

    # Step 2: Apply rotation to Q and K (complex number approach)
    # This involves reshaping Q and K, applying cos/sin transformations...
    # (full implementation omitted for brevity)

    # The key point: RoPE modifies the embeddings themselves
    # before attention computation
    pass


def alibi_attention_simple(Q, K, V, slopes, causal=True):
    """
    ALiBi attention for comparison.
    """
    # Step 1: Standard attention scores
    scores = np.matmul(Q, K.transpose(0, 2, 1)) / np.sqrt(Q.shape[-1])

    # Step 2: Add distance-based bias (one line!)
    seq_len = Q.shape[1]
    distances = np.abs(
        np.arange(seq_len)[:, None] - np.arange(seq_len)[None, :]
    )
    scores = scores + (-slopes[:, None, None] * distances)

    # That's it. Apply softmax and compute output.
    return scores

The ALiBi bias is a single line of code added to standard attention. The entire position encoding logic lives in one array operation. RoPE requires restructuring how queries and keys are computed, introducing trigonometric functions and careful handling of dimension pairs. Both work, but ALiBi's simplicity is a practical advantage: fewer lines of code means fewer places for bugs to hide, and the position bias is always directly visible and inspectable.

Historical Context

Historical Context

ALiBi was introduced in the paper "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation" by Ofir Press, Noah A. Smith, and Mike Lewis, published in 2022 at ICLR. The title captures the core contribution succinctly: train on short sequences, test (inference) on long sequences, enabled by the linear bias design.

The paper was motivated by the observation that all existing position encoding methods (sinusoidal, learned, T5 relative biases, and RoPE) failed to generalize beyond training sequence length. The authors systematically evaluated length extrapolation and found that ALiBi dramatically outperformed all competitors on this metric while remaining competitive on standard perplexity benchmarks.

ALiBi was adopted by the BLOOM project (the largest open multilingual language model at the time of its release in 2022 with 176 billion parameters) and by MosaicML's MPT family of models. This adoption showed that ALiBi could support production-scale language models. The method influenced subsequent thinking about position encoding design and contributed to the recognition that inductive biases favoring locality are valuable, not limiting.

Practical Considerations

When deploying ALiBi in your own models, several practical considerations affect performance and behavior. Understanding these factors helps you make informed decisions about whether ALiBi is the right choice and how to configure it.

ALiBi's computational overhead relative to standard attention is minimal. The bias matrix computation is O(n2)O(n^2), matching the overall cost of attention, and the actual operation (a matrix add) is extremely fast on modern hardware. There are no additional parameters to initialize, no embedding lookups, and no trigonometric operations. In a typical transformer implementation, adding ALiBi increases the attention computation time by perhaps 1 to 2 percent, well within noise.

The fixed slopes, while convenient, raise a natural question: should they be learned instead? Several research papers have explored learned ALiBi slopes. The findings are mixed: on standard benchmarks within the training length, learned slopes sometimes improve performance slightly. But on out-of-distribution sequence lengths (the key use case for ALiBi), learned slopes may not extrapolate as reliably as the fixed formula. The fixed slopes guarantee consistent behavior across all sequence lengths by design. Learned slopes might overfit to the training length range. If you decide to experiment with learned slopes, treat them as a carefully monitored hyperparameter rather than a drop-in replacement.

The causal versus bidirectional distinction matters for deployment. Decoder-only language models (for text generation) use causal ALiBi, where the upper triangle of the bias matrix is −∞-\infty. Encoder models (for classification, understanding) use bidirectional ALiBi, where the full symmetric bias matrix applies and both left and right context receive distance penalties. Encoder-decoder models (for translation, summarization) typically use bidirectional ALiBi in the encoder and causal ALiBi in the decoder. The mathematics remains the same in all cases; only which entries of the bias matrix are live changes.

Out[23]:
Visualization
Line plot showing attention weight decay with distance for different ALiBi slopes. This shows varying effective context windows.
Relative attention weight as a function of distance for three ALiBi heads (steep, medium, and gentle). Head 1 drops below 10 percent at around distance 5, effectively capping its context window. Head 8 maintains above 60 percent weight at distance 100, allowing it to integrate very distant context. This multi-scale coverage is why the ensemble of heads outperforms any single slope.

The effective context window visualization reveals a striking difference across heads. Head 1 drops below the 10 percent relative weight threshold very quickly, around distance 5, because its steep slope (m=0.5m = 0.5) imposes an exponentially decaying effective weight. Head 4 with slope 0.0625 reaches the 10 percent threshold around distance 37. Head 8 with slope 0.0039 barely crosses it even at distance 100. When all heads work together, the model simultaneously has access to very local structure (from head 1), medium-range context (from heads 2 through 7), and long-range dependencies (from head 8). This is arguably more structured than what emerges organically in models without explicit position encoding.

Limitations

ALiBi's simplicity is both its strength and its limitation. Understanding where the method falls short is essential for making informed architectural decisions.

The linear penalty assumes that relevance decreases monotonically with distance, which is a reasonable prior for many language tasks but not universally true. Consider structured data like code: a function definition might appear 500 tokens before a call site, but understanding the call requires attending back to the definition with full weight. The ALiBi penalty at distance 500 is −500m-500m, which with a steep slope effectively zeroes out that attention. The gentle-slope heads may still capture it, but there is no mechanism to override the penalty when the task demands strong long-range attention. Models using ALiBi compensate by learning strong content-based signals that overcome the bias, but this requires more capacity dedicated to fighting the prior.

Similarly, some programming languages and formal documents use nested structure (parentheses, XML tags, logical blocks) where the relevant matching token could be arbitrarily far away. The distance penalty is agnostic to this structure: it cannot know that an opening brace needs to attend to its closing brace regardless of the distance. Tasks that require this kind of distance-invariant matching are challenging for ALiBi. RoPE, being content-neutral in its position encoding (the rotation is applied to all dimensions uniformly), has less of this structural disadvantage.

The fixed slopes also create an implicit ceiling on what the gentlest head can do. Even with slope m=0.004m = 0.004, a distance of 1000 tokens incurs a penalty of −4-4 logits. After softmax, this is still a significant suppression. For tasks that require integration across truly long ranges (thousands of tokens of prior context), ALiBi's gentlest head may still fall short. This was one motivation for context window extension techniques like sliding window attention and retrieval augmentation in systems that use ALiBi.

Another limitation is the lack of asymmetry. The penalty −m⋅∣i−j∣-m \cdot |i - j| is symmetric: attending backward 10 steps is penalized the same as attending forward 10 steps (in bidirectional models). Natural language is often asymmetric: we typically interpret a word in light of what came before more than what comes after. While the causal mask handles this in decoder models, bidirectional ALiBi (for encoders) applies equal penalties in both directions. A refinement would be to use different slopes for forward and backward attention, but this is not part of the original ALiBi design.

Despite these limitations, ALiBi has proven remarkably effective. The BLOOM family of models, including the 176-billion parameter BLOOM-176B, uses ALiBi. So does the MPT family from MosaicML. These models demonstrate that ALiBi scales to the largest parameter counts and handles remarkably diverse tasks. The limitations manifest primarily in edge cases and specialized tasks, not in the general language modeling and generation use cases that dominate production deployment.

The impact of ALiBi extends beyond its direct use. It demonstrated that position encoding can be far simpler than previously thought. The original Transformer's sinusoidal encodings were ingenious but perhaps overengineered for the task. ALiBi showed that a linear penalty on distance, applied at attention time, is sufficient for strong performance and far superior for extrapolation. This insight has influenced subsequent work on efficient transformers, long-context models, and the general question of how much structure should be built into the architecture versus learned from data.

Key Parameters

When implementing ALiBi in your own models, the following parameters control its behavior:

  • num_heads: The number of attention heads in your model. ALiBi automatically computes slopes for each head using the geometric sequence formula. More heads create finer granularity in locality preferences: with 16 heads you get smoother coverage of the slope range than with 4 heads.

  • slope (m): The penalty strength for each head. Steeper slopes (larger values like 0.5) create strong locality bias where attention concentrates on nearby tokens. Gentler slopes (smaller values like 0.004) allow broader attention across the sequence. These are fixed by the formula rather than tuned, which eliminates a hyperparameter while potentially leaving some performance on the table for specialized tasks.

  • causal: Whether to apply causal masking. Set to True for decoder-style autoregressive models where tokens can only attend to previous positions. Set to False for encoder-style bidirectional attention, where both preceding and following tokens are available.

  • Base value (8): The constant in the slope formula mi=1/2(8/h)⋅im_i = 1/2^{(8/h) \cdot i} that controls the range of slopes. The original ALiBi paper uses 8, which ensures slopes span several orders of magnitude regardless of head count. Increasing this base would compress all slopes toward smaller values (gentler biases), while decreasing it would shift toward steeper slopes. The value 8 was selected empirically and is almost never modified in practice.

Summary

ALiBi offers a refreshingly simple approach to position encoding that achieves strong performance and excellent length extrapolation through a single design insight: replace position embeddings with a direct penalty on attention score based on distance.

The key ideas in this chapter are:

  • No position embeddings. Position enters only through attention biases, not through modifications to token representations. This keeps embeddings clean and makes position encoding completely separable from content encoding.

  • Linear distance penalty. Attention scores are reduced by m⋅∣i−j∣m \cdot |i - j|, where mm is a head-specific slope and ∣i−j∣|i - j| is the distance between query position ii and key position jj. Nearby tokens are favored, distant tokens are penalized, and the penalty is always proportional to distance.

  • Geometric slopes. Different attention heads use different slopes, creating multi-scale attention. Some heads focus locally on syntax and immediate context, others attend broadly to long-range dependencies. The slopes are fixed by formula, not learned.

  • Strong extrapolation. Because only relative distance matters (not absolute position), models trained on short sequences can process longer sequences at inference time. The bias structure is exactly the same at position 100 as at position 5000.

  • Minimal overhead. ALiBi adds one matrix addition to attention computation. There are no additional parameters, no complex rotation arithmetic, and no changes to the embedding pipeline.

The next chapter will compare the position encoding methods we have covered, from sinusoidal and learned encodings to relative approaches such as RoPE and ALiBi. You will see how each handles key challenges like extrapolation, computational cost, and representational power, and when to choose each method for real-world applications.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about ALiBi (Attention with Linear Biases).

ALiBi Position Encoding Quiz

Question 1 of 100 of 10 completed
How does ALiBi encode position information in transformers?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025alibiattention, author = {Michael Brenndoerfer}, title = {ALiBi: Attention with Linear Biases for Position Encoding}, year = {2025}, url = {https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). ALiBi: Attention with Linear Biases for Position Encoding. Retrieved from https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding
MLAAcademic
Michael Brenndoerfer. "ALiBi: Attention with Linear Biases for Position Encoding." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding>.
CHICAGOAcademic
Michael Brenndoerfer. "ALiBi: Attention with Linear Biases for Position Encoding." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding.
HARVARDAcademic
Michael Brenndoerfer (2025) 'ALiBi: Attention with Linear Biases for Position Encoding'. Available at: https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). ALiBi: Attention with Linear Biases for Position Encoding. https://mbrenndoerfer.com/writing/alibi-attention-linear-biases-position-encoding

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.