Relative Position Encoding

Michael BrenndoerferUpdated June 2, 202549 min read

Part of Language AI Handbook

Explains how relative position encoding improves transformer generalization by encoding token distances rather than absolute positions.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Relative Position Encoding

Previous chapters explored absolute position encoding: each position receives a fixed representation based on its index in the sequence. Whether through sinusoidal functions or learned embeddings, the model learns that "position 5 means something specific." But language doesn't work this way. When you read "the cat sat on the mat," the relationship between "cat" and "sat" depends on their relative distance (one word apart), not on whether they appear at positions 2 and 3 versus positions 102 and 103. The same grammatical relationship holds whether the sentence opens a document or appears buried on page fifty.

Relative position encoding shifts the focus from "where am I?" to "how far apart are we?" This distinction sounds subtle, but it changes how models generalize to new sequences and new lengths. A model trained on short sequences may encounter position pairs during inference that it has never seen during training. With absolute encodings, this creates a generalization gap. With relative encodings, the model only needs to have seen a given distance during training, and that knowledge applies everywhere that same distance appears.

This chapter explores why that distinction matters, how to formulate relative attention mathematically, and how Shaw et al.'s (2018) influential approach brought relative positions into transformer self-attention in a practical, learnable way. The Shaw et al. paper is one of the most widely cited works in positional encoding because it introduced a clean, trainable formulation that could be dropped into existing transformer architectures with minimal modification.

Understanding relative position encoding also sets up important intuitions for the chapters ahead. Rotary Position Embeddings (RoPE) and ALiBi, two methods that power many of today's largest language models, both take the "encode distance, not index" philosophy and push it further with different mathematical formulations. By working through Shaw et al.'s approach first, you'll see both the core insight and the limitations that motivated those later innovations.

The material builds naturally on what you've already learned about self-attention and absolute position encodings. If you're comfortable with the query-key-value attention mechanism and understand why position information matters at all, you have everything you need to follow the derivation and implementation in this chapter.

Why Relative Position Matters

Consider the sentence "The cat that I saw yesterday sat on the mat." The verb "sat" needs to find its subject "cat." With absolute position encoding, the model must learn that "position 5 attending to position 2" captures a subject-verb relationship. But if we move the same phrase later in the document, it becomes "position 105 attending to position 102." The model must learn these patterns for every possible absolute position pair, treating what is essentially the same linguistic relationship as hundreds of distinct cases.

The Generalization Problem

Absolute position encodings create position-specific patterns. A model trained on short sequences may never see "position 500 attending to position 497," even though the relationship is identical to "position 5 attending to position 2": a distance of 3 positions.

Relative position encoding solves this by focusing on the offset between positions. The distance from position 5 to position 2 is −3-3 (looking backward three positions). The distance from position 105 to position 102 is also −3-3. By encoding this offset rather than absolute positions, the model learns a single pattern for "three positions back" that generalizes across the entire sequence.

This shift has practical consequences that go beyond test-set metrics. Language has many distance-dependent patterns: adjectives typically appear one position before their nouns, determiners precede noun phrases by a few positions, and verbs often follow their subjects within a local window. These patterns are not anchored to absolute positions; they describe structural relationships that can appear anywhere in a document. Relative encoding captures these patterns directly, letting a single learned representation cover all occurrences of a given structural relationship regardless of where that relationship appears in the input.

The generalization benefit becomes especially important for long-context applications. A model trained on sequences of 512 tokens and then applied to sequences of 2048 tokens will encounter absolute positions it has never seen. Absolute encodings have no principled way to handle this: the embeddings for positions 513 through 2048 were never trained. Relative encodings, by contrast, are anchored to distances rather than indices. As long as the model has seen sequences where tokens are kk positions apart, it can apply that knowledge at any absolute location where that same distance appears.

In[3]:
Code
# Demonstrate the generalization problem with absolute positions
def count_position_pairs(max_len_train, max_len_test):
    """Count how many position pairs in test were never seen during training."""
    seen_pairs = set()
    for i in range(max_len_train):
        for j in range(max_len_train):
            seen_pairs.add((i, j))

    unseen = 0
    total = 0
    for i in range(max_len_test):
        for j in range(max_len_test):
            total += 1
            if (i, j) not in seen_pairs:
                unseen += 1

    return unseen, total


# Training on sequences up to length 128, testing on 256
unseen_abs, total_abs = count_position_pairs(128, 256)
Out[4]:
Console
Absolute Position Generalization Problem
=============================================
Training sequence length: 128
Test sequence length:     256
Total position pairs at test time: 65,536
Unseen position pairs:            49,152
Percentage unseen:                 75.0%

With relative encoding, a distance of -3 seen during training
applies equally to positions (5, 2) and (105, 102).
Out[5]:
Visualization
Heatmap showing seen vs unseen position pairs, with a dark square in the bottom-left corner and light L-shaped region indicating unseen pairs.
Visualization of the absolute position generalization problem. The dark region (bottom-left) represents position pairs seen during training on 128-token sequences. The light region shows unseen pairs when testing on 256 tokens. Nearly 75% of test pairs were never seen during training, meaning a model relying purely on absolute positions encounters mostly unfamiliar combinations at inference time.

Nearly 75% of position pairs at test time were never seen during training. A model relying on absolute positions must somehow generalize to these unseen combinations, either by interpolating from nearby positions or by pattern-matching to semantically similar content. Neither approach is principled. With relative encoding, every distance seen during training transfers to all positions at that distance, turning a difficult extrapolation problem into a simple reuse of already-learned representations.

From Absolute to Relative: The Key Insight

Recall the standard self-attention formula. Given queries Q\mathbf{Q}, keys K\mathbf{K}, and values V\mathbf{V}, we compute:

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^T}{\sqrt{d_k}}\right)\mathbf{V}

where:

  • Q∈Rn×dk\mathbf{Q} \in \mathbb{R}^{n \times d_k}: query matrix containing all query vectors
  • K∈Rn×dk\mathbf{K} \in \mathbb{R}^{n \times d_k}: key matrix containing all key vectors
  • V∈Rn×dv\mathbf{V} \in \mathbb{R}^{n \times d_v}: value matrix containing all value vectors
  • dkd_k: dimension of queries and keys
  • dk\sqrt{d_k}: scaling factor to prevent dot products from growing too large
  • nn: sequence length

With absolute position encoding, each token's embedding includes position information before the QKV projections. Position is baked into the input: xi=ei+pi\mathbf{x}_i = \mathbf{e}_i + \mathbf{p}_i, where ei\mathbf{e}_i is the token embedding and pi\mathbf{p}_i is the position encoding. The attention scores then implicitly depend on absolute positions through the projected queries and keys. The model sees position information, but only indirectly: it must disentangle "what is this token?" from "where is this token?" inside the projection matrices themselves.

Relative position encoding takes a different approach: inject position information directly into the attention computation. Instead of adding position to the input, we modify how attention scores are calculated to explicitly account for the offset j−ij - i between positions ii and jj. This is a cleaner separation of concerns. Content similarity handles semantic matching; a separate position-dependent term handles distance-based preferences. The model does not need to learn to disentangle them because they are never mixed in the first place.

The core idea is to add a position-dependent term to the attention score. To give the model a sense of why tokens at certain distances should receive more or less attention, we write:

scoreij=qi⋅kj+relative_bias(i,j)\text{score}_{ij} = \mathbf{q}_i \cdot \mathbf{k}_j + \text{relative\_bias}(i, j)

where:

  • qi\mathbf{q}_i: the query vector at position ii
  • kj\mathbf{k}_j: the key vector at position jj
  • relative_bias(i,j)\text{relative\_bias}(i, j): a term that depends only on the offset j−ij - i

This formulation separates content matching (qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j) from position matching (the relative bias). The model can learn that certain offsets are generally more or less important, regardless of absolute position. Notice that the relative bias only depends on the difference j−ij - i, not on ii and jj individually. This is the key mathematical property that enables generalization: two position pairs with the same offset get the same positional treatment, no matter where they appear in the sequence.

Historical Context: Toward Relative Attention

The idea of encoding relative rather than absolute positions has roots in cognitive science research from the 1990s, which showed that humans process distance relationships more reliably than absolute locations. In NLP, early sequence models like recurrent networks encoded positions implicitly through the sequential processing order, which is inherently relative. Transformer models broke that implicit ordering with parallel attention, creating the need to explicitly reinject position information. Shaw et al. (2018) were among the first to show that this could be done at the attention-score level rather than the input level, a distinction that turned out to be both theoretically cleaner and empirically better. Their work influenced a wave of relative position schemes including Transformer-XL's segment-level recurrence, T5's relative bias buckets, and ultimately RoPE and ALiBi. The common thread across all of these is the insight from this section: position information belongs in the attention function, not just in the embeddings.

Shaw et al. Relative Position Representations

The most influential formulation of relative positions in self-attention comes from Shaw et al. (2018). Their approach asks a fundamental question: if we want attention to be distance-aware, where exactly should we inject position information? Rather than modifying the input embeddings, Shaw et al. discovered that injecting relative positions directly into the attention computation gives the model finer control over how distance affects both which tokens get attention and what information flows.

To understand their approach, we need to build up three connected ideas: what we want the model to learn about distance, how to represent those distances as learnable parameters, and how to integrate those representations into the attention mechanism itself.

The Design Question: What Should Distance Encode?

Consider what happens when position ii attends to position jj in standard self-attention. The model computes a score based on content similarity, then uses that score to weight how much of jj's information flows to ii. But we want distance to influence both of these operations independently.

Think of it this way: a verb attending to possible subjects has two separate questions to answer. First, which position is the most likely subject based on both content and distance? Second, once the subject is found, how much of its information should flow into the verb's representation? Both questions depend on distance, but in different ways. The model might want to assign high attention to a subject that is two positions back, and to weight that subject's contribution with a "subject-attended-from-two-back" color that differs from "subject-attended-from-ten-back."

Shaw et al. address both needs by introducing two sets of learnable embeddings: one that modifies attention scores (controlling who gets attention), and one that modifies value aggregation (controlling what information flows). This separation gives the model independent control over "who to attend to" versus "what to gather."

  1. Which positions get attention: A verb might prefer to attend to positions 1-3 positions back (where subjects typically appear), regardless of content similarity.

  2. What information flows: When aggregating information, we might want "nearby context" to carry different weight than "distant context," even when attention weights are equal.

Relative Position Embeddings

For each possible offset r=j−ir = j - i between positions, Shaw et al. define learnable vectors that the model adjusts during training just like any other parameter:

  • arK∈Rdk\mathbf{a}^K_r \in \mathbb{R}^{d_k}: a relative position embedding that influences attention scores (key-side)
  • arV∈Rdv\mathbf{a}^V_r \in \mathbb{R}^{d_v}: a relative position embedding that modifies aggregated values (value-side)

where:

  • r=j−ir = j - i: the relative offset (positive means looking forward, negative means looking backward)
  • dkd_k: dimension of query and key vectors
  • dvd_v: dimension of value vectors

These embeddings are learned during training, just like the QKV projection matrices. The critical insight is that the number of distinct offsets is bounded by the sequence length (from −(n−1)-(n-1) to +(n−1)+(n-1)), not by the number of position pairs (which is n2n^2). This means the model learns a single pattern for "three positions back" that applies universally, rather than learning separate patterns for every pair of absolute positions. The parameter count scales linearly with the clipping distance kk, not quadratically with sequence length.

The key-side embeddings (aK\mathbf{a}^K) and value-side embeddings (aV\mathbf{a}^V) are separate parameter matrices, not tied to each other. This gives the model full flexibility: the same distance can produce different effects on the routing decision (score computation) versus the aggregation decision (value weighting).

Modifying Attention Scores

Let's trace through how relative positions enter the attention computation. In standard self-attention, the score measuring how much position ii should attend to position jj captures only content similarity:

eij=qi⋅kjdke_{ij} = \frac{\mathbf{q}_i \cdot \mathbf{k}_j}{\sqrt{d_k}}

where:

  • eije_{ij}: the raw attention score from query position ii to key position jj
  • qi∈Rdk\mathbf{q}_i \in \mathbb{R}^{d_k}: the query vector at position ii
  • kj∈Rdk\mathbf{k}_j \in \mathbb{R}^{d_k}: the key vector at position jj
  • dkd_k: the dimension of query and key vectors

This score depends only on the content at each position. Two identical words at different positions produce identical scores, regardless of their distance. The model cannot distinguish "I saw a cat (right here, adjacent)" from "I saw a cat (twenty words ago)."

Shaw et al. modify this by adding the relative position embedding to the key before computing the dot product:

eij=qi⋅(kj+aj−iK)dke_{ij} = \frac{\mathbf{q}_i \cdot (\mathbf{k}_j + \mathbf{a}^K_{j-i})}{\sqrt{d_k}}

Why add to the key rather than the query? The key represents "what this position offers," so adding position information to it means the model can learn that positions at certain distances offer something different, even if their content is identical. The query remains focused on "what I'm looking for" without distance context; the key communicates "here is what I have, adjusted for how far away I am from the query."

Expanding the dot product reveals the two-component structure:

eij=qi⋅kj+qi⋅aj−iKdke_{ij} = \frac{\mathbf{q}_i \cdot \mathbf{k}_j + \mathbf{q}_i \cdot \mathbf{a}^K_{j-i}}{\sqrt{d_k}}

where:

  • qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j: the content term, measuring semantic compatibility between positions
  • qi⋅aj−iK\mathbf{q}_i \cdot \mathbf{a}^K_{j-i}: the position term, encoding how relative distance affects the score
  • aj−iK∈Rdk\mathbf{a}^K_{j-i} \in \mathbb{R}^{d_k}: the learned relative position embedding for offset (j−i)(j-i)

The position term is a dot product between the query and the relative position embedding. This means different queries can respond differently to the same distance. A query learned to find subjects might have a high dot product with a−2K\mathbf{a}^K_{-2} (two positions back), while a query learned to find objects might prefer a+1K\mathbf{a}^K_{+1} (one position forward). The model learns these preferences end-to-end from training data, so the resulting position patterns reflect actual linguistic structure rather than any hand-designed prior.

Modifying Value Aggregation

Shaw et al. apply the same principle to value aggregation. After computing attention weights αij=softmax(eij)\alpha_{ij} = \text{softmax}(e_{ij}), standard attention aggregates values as a weighted sum. We want to add relative position information here too, so that the output at each position carries the weighted content of attended tokens and a record of where that content came from.

The modified value aggregation is:

oi=∑j=1nαij(vj+aj−iV)\mathbf{o}_i = \sum_{j=1}^{n} \alpha_{ij} (\mathbf{v}_j + \mathbf{a}^V_{j-i})

where:

  • oi∈Rdv\mathbf{o}_i \in \mathbb{R}^{d_v}: output vector at position ii, now enriched with position-aware context
  • αij\alpha_{ij}: attention weight from position ii to position jj (sums to 1 across jj)
  • vj∈Rdv\mathbf{v}_j \in \mathbb{R}^{d_v}: value vector at position jj, carrying the content to aggregate
  • aj−iV∈Rdv\mathbf{a}^V_{j-i} \in \mathbb{R}^{d_v}: relative position embedding for values at offset (j−i)(j-i)
  • nn: sequence length

This modification means the output at each position incorporates the weighted content from other positions and position-dependent information about where that content came from. Even if two positions have identical values and identical attention weights, their contributions differ based on their relative distances. Think of it as: "I got this information from a token that was three positions back, and that proximity matters for how I should interpret it."

The two sets of embeddings (aK\mathbf{a}^K and aV\mathbf{a}^V) work together but serve distinct roles. The key embeddings affect which positions receive attention (the "routing" decision). The value embeddings affect what information flows once routing is decided (the "content" decision). This separation allows the model to learn, for example, that adjacent positions should receive high attention (key embedding effect) while also learning that adjacent context should contribute a specific "locality" signal to the output (value embedding effect).

In practice, many implementations omit the value-side embeddings (aV\mathbf{a}^V) for efficiency reasons. The key-side modification tends to deliver most of the benefit, and removing the value modification cuts the number of relative position parameters in half. The T5 model, for instance, uses only key-side biases. Shaw et al.'s original paper includes both, but their ablation experiments showed that omitting the value-side embeddings costs only a small amount of task performance.

Clipping Relative Positions

The formulation above assumes we have a separate embedding for every possible relative position. But this creates a practical problem: for a sequence of length nn, relative positions range from −(n−1)-(n-1) to (n−1)(n-1), requiring 2n−12n-1 embeddings. For a sequence of length 512, that's 1023 different embeddings per dimension, for both keys and values.

This approach has two flaws. First, memory grows linearly with sequence length. Second, distant positions appear infrequently during training. In a 128-token sequence, only one position pair has offset +127+127 (position 0 to position 127), while thousands of pairs have offset +1+1. The embedding for offset +127+127 would be severely undertrained: the gradient updates for that embedding come from only one pair per sequence, while the embedding for offset +1+1 receives gradient contributions from 127 pairs. The result is that distant-position embeddings never become reliable.

Shaw et al. address both problems with a simple insight: fine-grained distance matters more locally than globally. The difference between "3 positions back" and "4 positions back" may be linguistically significant, as it could distinguish an adjective-noun relationship from a determiner-noun relationship. But the difference between "100 positions back" and "101 positions back" rarely carries meaning. Both are simply "far away." The model can usefully represent "very distant" as a single category without losing important information.

This motivates clipping: relative positions beyond a maximum distance kk are clamped to ±k\pm k:

clip(j−i,−k,k)\text{clip}(j - i, -k, k)

where:

  • j−ij - i: the raw relative position (offset from position ii to position jj)
  • kk: the maximum relative position to distinguish (a hyperparameter)
  • clip(x,a,b)\text{clip}(x, a, b): clamps xx to the range [a,b][a, b], returning aa if x<ax < a, bb if x>bx > b, and xx otherwise

With clipping, we need only 2k+12k + 1 relative position embeddings regardless of sequence length. Position pairs further than kk apart share the same "far away" embedding, either a−kK\mathbf{a}^K_{-k} for "far backward" or a+kK\mathbf{a}^K_{+k} for "far forward."

The choice of kk reflects a trade-off between expressiveness and learnability. Typical values range from 16 to 128, depending on the expected locality of relevant patterns in the data. For sentence-level tasks where the longest dependencies span 10-20 tokens, k=16k = 16 often suffices. For document-level tasks with paragraph-scale dependencies, k=64k = 64 or k=128k = 128 gives the model more distance resolution. In practice, the performance difference between k=32k = 32 and k=64k = 64 is often small, but the difference between k=4k = 4 and k=16k = 16 can be significant for tasks involving long-range grammatical dependencies.

In[6]:
Code
def clip_relative_position(j, i, max_dist):
    """
    Compute clipped relative position.

    Args:
        j: Key/value position
        i: Query position
        max_dist: Maximum relative distance to distinguish

    Returns:
        Clipped relative position in range [-max_dist, max_dist]
    """
    rel_pos = j - i
    return max(-max_dist, min(max_dist, rel_pos))


# Demonstrate clipping
max_k = 4
positions = list(range(10))
query_pos = 5
Out[7]:
Console
Query position: 5
Maximum relative distance (k): 4

Position | Raw Offset | Clipped Offset
------------------------------------------
    0    |     -5     |      -4
    1    |     -4     |      -4
    2    |     -3     |      -3
    3    |     -2     |      -2
    4    |     -1     |      -1
    5    |     +0     |      +0
    6    |     +1     |      +1
    7    |     +2     |      +2
    8    |     +3     |      +3
    9    |     +4     |      +4

The table reveals how clipping works in practice. Positions 0 and 1 have raw offsets of -5 and -4, but both clip to -4 because they exceed our maximum distance of 4. Position 9 has raw offset +4, which equals the maximum so no clipping occurs. All positions beyond the clipping threshold share the same "far away" embedding, which captures the intuition that distinguishing between "very far" and "extremely far" rarely matters for language understanding.

Out[8]:
Visualization
Heatmap showing clipped relative positions for a 10x10 matrix with k=4, displaying values from -4 to +4.
Effect of clipping on the relative position matrix for a 10-token sequence with k=4. Each cell shows the clipped relative offset used to index the embedding table. The diagonal band of distinct values (from -4 to +4) widens as k increases, while the saturated corners (darkest blue and darkest red) show positions that share the same embedding because they are further than k steps apart.

Notice the diagonal band structure: each anti-diagonal contains the same relative position. The corners, where positions are far apart, are clipped to the boundary values ±k\pm k. This Toeplitz-like structure is key to efficient implementation because it means position biases repeat in a predictable pattern that can be computed once and reused.

Worked Example: Step-by-Step Score Computation

Let's walk through a concrete, minimal example to make the math tangible. Consider a four-token sequence with one-dimensional embeddings, maximum clipping distance k=2k = 2, and very simple relative position embeddings. The goal is to show exactly where the relative bias comes from and how it modifies the attention scores.

Suppose we have queries q1=[1.0]\mathbf{q}_1 = [1.0], q2=[0.5]\mathbf{q}_2 = [0.5], q3=[0.8]\mathbf{q}_3 = [0.8], q4=[0.3]\mathbf{q}_4 = [0.3] and keys k1=[0.9]\mathbf{k}_1 = [0.9], k2=[0.7]\mathbf{k}_2 = [0.7], k3=[0.4]\mathbf{k}_3 = [0.4], k4=[1.2]\mathbf{k}_4 = [1.2]. The key-side relative position embeddings (after clipping to k=2k = 2) are:

  • a−2K=[−0.3]\mathbf{a}^K_{-2} = [-0.3] (two positions back)
  • a−1K=[+0.4]\mathbf{a}^K_{-1} = [+0.4] (one position back)
  • a0K=[+0.1]\mathbf{a}^K_{0} = [+0.1] (same position)
  • a+1K=[−0.2]\mathbf{a}^K_{+1} = [-0.2] (one position forward)
  • a+2K=[−0.5]\mathbf{a}^K_{+2} = [-0.5] (two positions forward)

Now let's compute the score from query position 1 to all key positions. With dk=1d_k = 1 so dk=1\sqrt{d_k} = 1:

For position pair (i=1,j=1)(i=1, j=1): offset is 00, content term is 1.0×0.9=0.91.0 \times 0.9 = 0.9, position term is 1.0×0.1=0.11.0 \times 0.1 = 0.1. Combined score: 1.01.0.

For position pair (i=1,j=2)(i=1, j=2): offset is +1+1, content term is 1.0×0.7=0.71.0 \times 0.7 = 0.7, position term is 1.0×(−0.2)=−0.21.0 \times (-0.2) = -0.2. Combined score: 0.50.5.

For position pair (i=1,j=3)(i=1, j=3): offset is +2+2, content term is 1.0×0.4=0.41.0 \times 0.4 = 0.4, position term is 1.0×(−0.5)=−0.51.0 \times (-0.5) = -0.5. Combined score: −0.1-0.1.

For position pair (i=1,j=4)(i=1, j=4): offset is +3+3, which clips to +2+2, content term is 1.0×1.2=1.21.0 \times 1.2 = 1.2, position term is 1.0×(−0.5)=−0.51.0 \times (-0.5) = -0.5. Combined score: 0.70.7.

The key insight is visible immediately: position 4, despite having the highest content similarity (q1⋅k4=1.2\mathbf{q}_1 \cdot \mathbf{k}_4 = 1.2), is penalized by the relative position embedding because it is far away. Position 1 (same position) receives a slight boost. In a trained model, these biases would reflect linguistic preferences learned from data rather than arbitrary initialization values, but the mechanics are identical.

After applying softmax to the four scores [1.0,0.5,−0.1,0.7][1.0, 0.5, -0.1, 0.7], we get attention weights that favor the nearby positions more than pure content similarity would suggest. This is exactly the behavior we want: locality is naturally preferred, but strong content signals can still override the distance penalty when a truly relevant token appears far away.

Implementation: Relative Position in Self-Attention

Now that we understand the mathematical formulation and the worked example, let's translate it into code. We'll build up the implementation in stages: first the relative position embeddings, then the modified attention score computation, then the modified value aggregation, and finally the complete layer.

Building the Relative Position Embedding Table

The foundation of Shaw-style attention is a table of learned embeddings indexed by relative position. Since relative positions range from −k-k to +k+k after clipping, we need 2k+12k + 1 embeddings. The key implementation detail is converting from relative positions (which can be negative) to array indices (which must be non-negative).

In[9]:
Code
import numpy as np


class RelativePositionEmbedding:
    """
    Learnable embeddings for relative positions.
    """

    def __init__(self, max_distance, d_k, d_v, seed=None):
        """
        Initialize relative position embeddings.

        Args:
            max_distance: Maximum relative position to distinguish (k)
            d_k: Dimension for key-side embeddings
            d_v: Dimension for value-side embeddings
            seed: Random seed for reproducibility
        """
        if seed is not None:
            np.random.seed(seed)

        # Number of unique relative positions: from -k to +k
        n_positions = 2 * max_distance + 1

        # Initialize embeddings with small random values
        scale_k = np.sqrt(2.0 / (n_positions + d_k))
        scale_v = np.sqrt(2.0 / (n_positions + d_v))

        self.a_K = np.random.randn(n_positions, d_k) * scale_k
        self.a_V = np.random.randn(n_positions, d_v) * scale_v
        self.max_distance = max_distance

    def get_key_embedding(self, rel_pos):
        """Get the key-side relative position embedding for an offset."""
        clipped = max(-self.max_distance, min(self.max_distance, rel_pos))
        # Convert from [-k, k] to [0, 2k] for indexing
        idx = clipped + self.max_distance
        return self.a_K[idx]

    def get_value_embedding(self, rel_pos):
        """Get the value-side relative position embedding for an offset."""
        clipped = max(-self.max_distance, min(self.max_distance, rel_pos))
        idx = clipped + self.max_distance
        return self.a_V[idx]

The index conversion is the subtle but essential detail. Relative positions range from −k-k to +k+k, but array indices must be non-negative. Adding kk shifts the range to [0,2k][0, 2k]:

  • Offset −k-k (maximum backward distance) maps to index 0
  • Offset 0 (same position) maps to index kk
  • Offset +k+k (maximum forward distance) maps to index 2k2k

This simple arithmetic lets us use relative positions as array indices without any conditional logic. The same conversion applies uniformly across all offsets, making the lookup fast and predictable.

Computing Attention Scores with Relative Position

With the embedding table in place, we can implement the modified attention score computation. Recall the formula: eij=(qi⋅kj+qi⋅aj−iK)/dke_{ij} = (\mathbf{q}_i \cdot \mathbf{k}_j + \mathbf{q}_i \cdot \mathbf{a}^K_{j-i}) / \sqrt{d_k}. For each query-key pair, we compute both the content term and the position term, then sum them before scaling.

In[10]:
Code
def relative_attention_scores(Q, K, rel_embed):
    """
    Compute attention scores with relative position information.

    Args:
        Q: Query matrix of shape (seq_len, d_k)
        K: Key matrix of shape (seq_len, d_k)
        rel_embed: RelativePositionEmbedding instance

    Returns:
        Attention scores of shape (seq_len, seq_len)
    """
    seq_len, d_k = Q.shape
    scores = np.zeros((seq_len, seq_len))

    for i in range(seq_len):  # Query position
        for j in range(seq_len):  # Key position
            # Content term: q_i · k_j
            content_score = np.dot(Q[i], K[j])

            # Position term: q_i · a^K_{j-i}
            rel_pos = j - i
            a_K = rel_embed.get_key_embedding(rel_pos)
            position_score = np.dot(Q[i], a_K)

            # Combined score
            scores[i, j] = content_score + position_score

    # Scale by sqrt(d_k)
    scores = scores / np.sqrt(d_k)

    return scores

This implementation iterates over all position pairs, making the logic transparent. For each pair (i,j)(i, j), we compute the content score (qi⋅kj\mathbf{q}_i \cdot \mathbf{k}_j), look up the relative position embedding for offset j−ij - i, compute the position score (qi⋅aj−iK\mathbf{q}_i \cdot \mathbf{a}^K_{j-i}), and sum them. The final division by dk\sqrt{d_k} applies the same scaling as standard attention to prevent score magnitudes from growing with dimension.

Computing Output with Relative Position in Values

The value aggregation follows the same pattern. For each output position ii, we aggregate contributions from all positions jj, but we add the relative position embedding to each value before weighting:

In[11]:
Code
def relative_attention_output(attention_weights, V, rel_embed):
    """
    Compute attention output with relative position in values.

    Args:
        attention_weights: Softmax attention weights (seq_len, seq_len)
        V: Value matrix of shape (seq_len, d_v)
        rel_embed: RelativePositionEmbedding instance

    Returns:
        Output of shape (seq_len, d_v)
    """
    seq_len, d_v = V.shape
    output = np.zeros((seq_len, d_v))

    for i in range(seq_len):  # Output position
        for j in range(seq_len):  # Value position
            # Get value with relative position embedding
            rel_pos = j - i
            a_V = rel_embed.get_value_embedding(rel_pos)
            value_with_pos = V[j] + a_V

            # Weighted contribution
            output[i] += attention_weights[i, j] * value_with_pos

    return output

The value modification works similarly to the key modification. For each contributing position jj, we add its relative position embedding aj−iV\mathbf{a}^V_{j-i} to its value vector before applying the attention weight. This means the output at position ii contains weighted content and weighted position information about where that content originated. The model can learn to tag information with its source distance, which helps downstream layers reason about the structure of the input.

The Complete Relative Self-Attention Layer

With both components in place, we can assemble the complete layer. This class wraps the QKV projections, relative position embeddings, and the modified attention computation into a single forward pass:

In[12]:
Code
class RelativeSelfAttention:
    """
    Self-attention with relative position representations (Shaw et al.).
    """

    def __init__(self, embed_dim, d_k, d_v, max_distance=16, seed=None):
        """
        Initialize relative self-attention layer.

        Args:
            embed_dim: Dimension of input embeddings
            d_k: Query/key dimension
            d_v: Value dimension
            max_distance: Maximum relative position to distinguish
            seed: Random seed
        """
        if seed is not None:
            np.random.seed(seed)

        # QKV projection matrices
        scale_qk = np.sqrt(2.0 / (embed_dim + d_k))
        scale_v = np.sqrt(2.0 / (embed_dim + d_v))

        self.W_q = np.random.randn(embed_dim, d_k) * scale_qk
        self.W_k = np.random.randn(embed_dim, d_k) * scale_qk
        self.W_v = np.random.randn(embed_dim, d_v) * scale_v

        # Relative position embeddings
        self.rel_embed = RelativePositionEmbedding(max_distance, d_k, d_v, seed)

        self.d_k = d_k

    def forward(self, X):
        """
        Compute relative self-attention.

        Args:
            X: Input embeddings of shape (seq_len, embed_dim)

        Returns:
            output: Attention output of shape (seq_len, d_v)
            attention_weights: Weights of shape (seq_len, seq_len)
        """
        # Project to Q, K, V
        Q = X @ self.W_q
        K = X @ self.W_k
        V = X @ self.W_v

        # Compute attention scores with relative positions
        scores = relative_attention_scores(Q, K, self.rel_embed)

        # Softmax
        scores_stable = scores - scores.max(axis=1, keepdims=True)
        exp_scores = np.exp(scores_stable)
        attention_weights = exp_scores / exp_scores.sum(axis=1, keepdims=True)

        # Compute output with relative positions in values
        output = relative_attention_output(attention_weights, V, self.rel_embed)

        return output, attention_weights

The forward pass follows the standard self-attention pattern with two modifications. After projecting input embeddings to QKV, we use our modified relative_attention_scores function instead of the standard dot product, and we use relative_attention_output instead of the standard weighted sum. The softmax step remains unchanged, as the relative position terms are already incorporated into the scores before normalization.

Testing the Implementation

Let's verify our implementation produces sensible results:

In[13]:
Code
# Create a test sequence
np.random.seed(42)
seq_len = 6
embed_dim = 8
d_k = d_v = 4
max_dist = 3

# Random input embeddings
X = np.random.randn(seq_len, embed_dim)

# Create relative attention layer
rel_attn = RelativeSelfAttention(
    embed_dim, d_k, d_v, max_distance=max_dist, seed=123
)
output, weights = rel_attn.forward(X)
Out[14]:
Console
Relative Self-Attention Test
=============================================
Sequence length: 6
Embedding dim:   8
Q/K dimension:   4
V dimension:     4
Max distance:    3

Input shape:     (6, 8)
Output shape:    (6, 4)
Weights shape:   (6, 6)

Attention weights (rows sum to 1):
[[0.008 0.028 0.001 0.12  0.62  0.223]
 [0.26  0.098 0.35  0.157 0.052 0.083]
 [0.794 0.002 0.077 0.122 0.002 0.002]
 [0.016 0.394 0.025 0.108 0.356 0.101]
 [0.475 0.023 0.002 0.13  0.069 0.301]
 [0.002 0.227 0.001 0.014 0.66  0.097]]

The attention weights look similar to standard self-attention, but they now incorporate relative position information. Each query-key compatibility is biased by the relative distance between positions, even at random initialization. In a trained model, these biases become meaningful: heads specialize in different distance profiles, with some heads prioritizing local context and others focusing on longer-range dependencies.

Out[15]:
Visualization
Heatmap showing relative position embedding values with offsets on y-axis and dimensions on x-axis.
Relative position embedding matrix (key-side) after random initialization with max distance k=3. Each row corresponds to a relative offset from -3 to +3, and each column is an embedding dimension. The distinct patterns across rows show that each offset is represented differently from the start, and training will push these representations toward linguistically meaningful distance patterns.

The embedding matrix shows how each relative offset is encoded as a different vector. Offset 0 (same position) has a distinct pattern from offset -3 (three positions back) or offset +3 (three positions forward). These patterns are learned during training to capture which distances are relevant for the task. After training on a language modeling task, you would typically see the nearby offsets (±1\pm 1, ±2\pm 2) develop sharper, more structured patterns while the boundary offsets (±k\pm k) develop softer patterns reflecting their shared "far away" role.

Visualizing Relative Position Effects

Let's visualize how relative position embeddings influence attention. We'll examine the position term qi⋅aj−iK\mathbf{q}_i \cdot \mathbf{a}^K_{j-i} separately from the content term. Isolating these two components lets us see how much of the final attention pattern comes from what tokens mean versus where they are.

In[16]:
Code
# Extract position-only attention biases
def compute_position_biases(Q, rel_embed):
    """
    Compute attention score contributions from relative positions only.
    """
    seq_len, d_k = Q.shape
    biases = np.zeros((seq_len, seq_len))

    for i in range(seq_len):
        for j in range(seq_len):
            rel_pos = j - i
            a_K = rel_embed.get_key_embedding(rel_pos)
            biases[i, j] = np.dot(Q[i], a_K) / np.sqrt(d_k)

    return biases


# Compute position biases for our test case
Q_test = X @ rel_attn.W_q
position_biases = compute_position_biases(Q_test, rel_attn.rel_embed)

# Also compute content-only scores for comparison
K_test = X @ rel_attn.W_k
content_scores = (Q_test @ K_test.T) / np.sqrt(d_k)
Out[17]:
Visualization
Heatmap of content-based attention scores between six positions showing irregular patterns based on embedding similarity.
Content scores computed from query-key dot products alone. These irregular patterns reflect purely semantic compatibility between tokens and carry no distance information.
Heatmap of position-based attention biases showing diagonal band structure reflecting relative distances.
Position biases from the relative position embeddings. Each diagonal stripe corresponds to a fixed offset, revealing that tokens at the same relative distance always receive the same positional treatment regardless of their absolute positions.

The content scores show irregular patterns depending on the input embeddings. The position biases show a distinctive diagonal structure: each anti-diagonal corresponds to positions at a fixed relative distance. For example, the main diagonal (offset 0) shows self-attention bias, while diagonals above and below show biases for looking forward and backward. Notice that in the position biases panel, all cells on the same anti-diagonal have identical values. This is because they share the same relative offset and therefore the same position embedding. This structural constraint is what makes relative encoding powerful: every occurrence of the same distance gets the same treatment.

Let's examine how the combined scores compare to content-only attention:

In[18]:
Code
# Compute attention weights with and without position
def softmax_rows(x):
    exp_x = np.exp(x - x.max(axis=1, keepdims=True))
    return exp_x / exp_x.sum(axis=1, keepdims=True)


content_only_weights = softmax_rows(content_scores)
combined_scores = content_scores + position_biases
combined_weights = softmax_rows(combined_scores)
Out[19]:
Visualization
Heatmap of attention weights using only content similarity between positions.
Content-only attention weights based purely on query-key similarity. The spread of weights across positions reflects only semantic compatibility between tokens.
Heatmap of attention weights combining content and relative position, showing modified attention patterns.
Attention weights combining content and relative position biases. The relative position term shifts weights toward or away from certain distances, introducing a structured distance preference on top of the semantic signal.

The position biases subtly shift attention patterns. In this random example, the differences are modest because the embeddings are random. In trained models, relative position embeddings learn linguistically meaningful patterns: verbs attending strongly to positions 1-2 back (typical subject distance), determiners attending 1-2 forward (to their nouns), and so on. The combined plot differs from a sum of the two individual plots because the softmax nonlinearity mixes them: a moderate positive position bias for a nearby token can dramatically amplify its weight if the content score is also high, while a large positive bias for a distant token with low content score may still produce a negligible final weight.

Efficient Matrix Implementation

The loop-based implementation above is clear but slow. In practice, we compute relative attention efficiently using matrix operations. The key insight is that position biases form a Toeplitz-like structure: all entries on each diagonal share the same relative distance.

The naive loop runs in O(n2dk)O(n^2 d_k) time, with a Python-level loop over all n2n^2 position pairs. The efficient version replaces the inner position term computation with a single matrix multiplication, reducing the Python overhead from O(n2)O(n^2) loop iterations to O(1)O(1). The total asymptotic complexity remains O(n2dk)O(n^2 d_k), but the constant factor drops dramatically because NumPy's matrix operations run in compiled code.

In[20]:
Code
def efficient_relative_scores(Q, K, rel_embed):
    """
    Compute relative attention scores using efficient matrix operations.

    Args:
        Q: Query matrix of shape (seq_len, d_k)
        K: Key matrix of shape (seq_len, d_k)
        rel_embed: RelativePositionEmbedding instance

    Returns:
        Attention scores of shape (seq_len, seq_len)
    """
    seq_len, d_k = Q.shape

    # Content scores: standard Q @ K.T
    content_scores = Q @ K.T

    # Build relative position bias matrix
    # Each entry [i, j] needs Q[i] · a^K_{j-i}

    # Step 1: Collect all relative position embeddings into a matrix
    # Shape: (2*max_dist + 1, d_k)
    all_a_K = rel_embed.a_K
    max_dist = rel_embed.max_distance

    # Step 2: Compute Q @ all_a_K.T to get all possible position scores
    # Shape: (seq_len, 2*max_dist + 1)
    all_position_scores = Q @ all_a_K.T

    # Step 3: Index to build the (seq_len, seq_len) position bias matrix
    position_biases = np.zeros((seq_len, seq_len))
    for i in range(seq_len):
        for j in range(seq_len):
            rel_pos = j - i
            clipped = max(-max_dist, min(max_dist, rel_pos))
            idx = clipped + max_dist
            position_biases[i, j] = all_position_scores[i, idx]

    # Combined scores
    scores = (content_scores + position_biases) / np.sqrt(d_k)

    return scores


# Verify the efficient implementation matches the naive one
scores_naive = relative_attention_scores(Q_test, K_test, rel_attn.rel_embed)
scores_efficient = efficient_relative_scores(Q_test, K_test, rel_attn.rel_embed)
Out[21]:
Console
Maximum difference between naive and efficient: 8.88e-16
Implementations are equivalent!

The efficient version first computes all possible query-position dot products in a single (n,2k+1)(n, 2k+1) matrix multiply, then indexes into this precomputed matrix. This reduces the inner loop computation from O(dk)O(d_k) per entry to O(1)O(1), with the upfront cost of a single (n,2k+1)(n, 2k+1) matrix multiplication. The Toeplitz structure of the position bias matrix is what makes this possible: once you know how every query interacts with every possible offset embedding, you only need to map indices to offsets, which is a pure indexing operation.

Real-world implementations in PyTorch or JAX go further. The indexing step can be replaced with a gather operation or einsum that exploits the diagonal structure, bringing the full bias computation down to a few vectorized calls with no Python loops at all. This is important for GPU utilization: Python loops are sequential, but GPU kernels can parallelize across the entire (n,n)(n, n) matrix simultaneously.

Comparison with Absolute Position Encoding

Let's directly compare relative and absolute position encodings on a simple task: detecting whether a token is attending to its immediate neighbors versus distant positions.

In[22]:
Code
def analyze_position_bias(attention_weights, max_offset=5):
    """
    Analyze average attention by relative offset.
    """
    seq_len = attention_weights.shape[0]
    offset_weights = {}

    for offset in range(-max_offset, max_offset + 1):
        weights_at_offset = []
        for i in range(seq_len):
            j = i + offset
            if 0 <= j < seq_len:
                weights_at_offset.append(attention_weights[i, j])
        if weights_at_offset:
            offset_weights[offset] = np.mean(weights_at_offset)

    return offset_weights


# Analyze position patterns in our relative attention
offset_analysis = analyze_position_bias(combined_weights)
Out[23]:
Visualization
Bar chart showing average attention weight for relative offsets from -5 to +5, with the self-attention bar at offset 0 highlighted in red below the uniform baseline.
Average attention weight by relative offset for the test sequence. The red bar at offset 0 marks self-attention. The dashed line marks the uniform baseline (equal weight to all positions). Bars above the baseline indicate offsets that receive more attention than chance, while bars below indicate less. In trained models, this profile takes on a characteristic shape that reflects the model's learned distance preferences for a given task.

The profile shows how attention is distributed by relative distance. In this random initialization, patterns are noisy, but trained models develop clear preferences: immediate neighbors typically receive more attention, with the pattern decaying for distant positions. The key observation is that this profile is the same at every absolute position in the sequence. A token at position 50 and a token at position 500 both use the same distance-indexed biases. Absolute position encoding cannot achieve this: the bias for "position 50 attending to position 52" has no relationship to the bias for "position 500 attending to position 502," because the absolute embedding at position 52 is entirely independent from the absolute embedding at position 502.

In Practice: Adopting Relative Position Encoding

Shaw et al.'s formulation has been adopted in several widely-used models, and understanding how those adoptions differ from the original paper helps calibrate when and how to use relative position encoding in your own work.

BERT and GPT variants originally used absolute position embeddings (either sinusoidal or learned), but researchers have since explored retrofitting relative position attention to these architectures. The main finding is that relative encoding consistently outperforms absolute encoding on tasks requiring cross-sentence reasoning or document-level understanding, where relevant information can appear at variable distances.

T5 uses a simplified form of relative position encoding that omits the value-side embeddings entirely and uses a smaller number of learned biases shared across all heads. This shared-bias design reduces the parameter count and works well for text-to-text generation tasks. T5's position biases are learned per-layer but shared across heads, a further simplification that incurs little performance cost.

Transformer-XL introduced segment-level recurrence to handle very long sequences, and paired it with a relative position encoding scheme based on sinusoidal functions rather than learned embeddings. Their formulation is more complex than Shaw et al.'s but avoids the need to re-initialize position embeddings when extending to new sequence lengths.

When implementing relative position encoding in practice, several configuration decisions matter:

  • Clipping distance kk: For sentence-level tasks, k=16k = 16 or k=32k = 32 is usually sufficient. For document-level tasks or tasks with long-range dependencies (coreference resolution, discourse structure), larger values like k=64k = 64 or k=128k = 128 are more appropriate.

  • Value-side embeddings: Omitting aV\mathbf{a}^V reduces parameters by half with modest performance loss. For memory-constrained settings or when training data is limited, this is a reasonable trade-off.

  • Initialization scale: Relative position embeddings should be initialized with smaller values than QKV projections. A scale of 1/dk1/\sqrt{d_k} or the Xavier formula used in the code above prevents position biases from dominating attention scores at the start of training.

  • Multi-head sharing: Some implementations share position embeddings across all heads (T5's approach), while others use per-head embeddings (Shaw et al.'s approach). Shared embeddings reduce parameters but limit each head's ability to specialize in different distance profiles.

The most important practical consideration is batching. In standard attention, you batch multiple sequences together and process them in parallel. With relative position encoding, all sequences in a batch share the same position embedding table, but the offset matrix for each sequence needs to be computed based on that sequence's length. This is straightforward for fixed-length sequences (pad to the same length) but requires careful handling for variable-length sequences with attention masking.

Learned Position Patterns

To illustrate what relative position embeddings might learn, let's simulate embeddings that encode common linguistic patterns:

In[24]:
Code
# Simulate learned relative position embeddings
# Pattern: slight preference for immediately adjacent positions

d_k_sim = 4
max_dist_sim = 5

# Create embeddings that encode locality preference
np.random.seed(42)
a_K_sim = np.random.randn(2 * max_dist_sim + 1, d_k_sim) * 0.1

# Add locality bias: positions closer to 0 have higher dot product potential
for idx in range(2 * max_dist_sim + 1):
    offset = idx - max_dist_sim  # Convert back to relative position
    locality_factor = np.exp(-0.5 * (offset**2))  # Gaussian decay
    a_K_sim[idx, 0] += locality_factor  # Boost first dimension

# Create a uniform query that will reveal the position pattern
uniform_query = np.ones(d_k_sim) / np.sqrt(d_k_sim)
position_profile = uniform_query @ a_K_sim.T
Out[25]:
Visualization
Line plot showing position score peaking at offset 0 and decaying symmetrically for more distant positions.
Simulated learned relative position profile showing a locality preference. The Gaussian-shaped curve peaks at offset 0 (self-attention) and decays for more distant positions. Real models learn more asymmetric patterns reflecting directional linguistic dependencies, such as verbs preferring to look back for subjects and forward for objects.

This simulated pattern shows a Gaussian-like preference for nearby positions. Real models learn more complex patterns that capture linguistic structure: determiners attending to following nouns, verbs attending to preceding subjects, and so on. A trained model's position profiles are also typically asymmetric. This reflects the fact that English (and many other languages) has predominantly left-to-right dependency structure. Subjects appear before verbs, modifiers often precede their heads, and complementizers introduce subordinate clauses that follow the main verb. These directional biases show up clearly when you visualize position profiles from models trained on real text.

Key Parameters

When implementing relative position encoding, several parameters control the behavior and efficiency of the mechanism. Getting these right matters more for relative encoding than for absolute encoding, because the position embeddings are learned and their training dynamics depend on initialization and capacity choices.

The key parameters are:

  • max_distance (k): Maximum relative position to distinguish. Positions beyond this distance share the same "far away" embedding. Typical values range from 16 to 128. Smaller values reduce memory but lose fine-grained distance information for distant positions. Larger values preserve more detail but require more parameters and training data. The right value depends on the task: sentence-level tasks rarely need k>32k > 32, while document-level tasks may benefit from k=128k = 128.

  • d_k: Dimension of query and key vectors, and key-side relative position embeddings. Must match between queries and keys for dot product compatibility. Common values are 64 for single-head attention or embed_dim / num_heads for multi-head attention. Larger dkd_k gives the position embeddings more capacity to encode complex distance patterns.

  • d_v: Dimension of value vectors and value-side relative position embeddings. Determines the output dimension of each attention head. Often set equal to d_k for simplicity, though they can differ.

  • Embedding initialization scale: Relative position embeddings are typically initialized with small random values. Xavier/Glorot initialization scales by 2/(n_positions+d)\sqrt{2/(n\_\text{positions} + d)} to maintain stable gradient magnitudes during training. If initialization is too large, position biases dominate the attention scores from the start and the model struggles to learn content-based patterns. If too small, position information takes many epochs to develop.

The total number of relative position parameters is (2k+1)×(dk+dv)(2k + 1) \times (d_k + d_v). With k=64k = 64 and dk=dv=64d_k = d_v = 64, this adds 16,512 parameters per attention head, which is modest compared to the QKV projection matrices (which have 3×dmodel×dk3 \times d_{\text{model}} \times d_k parameters). In a 12-head transformer with dmodel=768d_{\text{model}} = 768, the QKV projections total about 1.77M parameters per layer, while relative position embeddings add only about 200K. The overhead is real but small.

Limitations and Practical Considerations

Relative position encoding solves the generalization problem but introduces its own challenges. Understanding these limitations is important both for choosing between absolute and relative encoding in a new project and for anticipating where Shaw et al.'s approach will fall short.

The most significant limitation is computational complexity. Standard absolute position encoding adds to embeddings once at the input, with no additional cost during attention. Relative position encoding modifies every attention score computation. For a sequence of length nn, we compute n2n^2 relative position biases per attention layer. With multiple layers and attention heads, this overhead accumulates. Shaw et al. note that their approach adds roughly 5-10% to training time, which is acceptable for most applications. However, for very long sequences (4096+ tokens), even a 10% overhead becomes significant because attention is already the bottleneck, and any additional computation in the O(n2)O(n^2) inner loop compounds quickly.

Memory usage also increases. We store (2k+1)(2k+1) relative position embeddings per dimension, for both keys and values. With dk=dv=64d_k = d_v = 64 and k=64k = 64, this adds about 16K parameters per attention head. This is modest compared to the projection matrices, but the position bias matrix computed during the forward pass has shape (n,n)(n, n), which for n=4096n = 4096 requires 64 MB per layer just for the bias storage (in float32). For very long sequences, this memory footprint rivals the attention weight matrix itself.

The clipping distance kk introduces a hyperparameter choice with no universally correct answer. Too small, and the model loses fine-grained distance information for medium-range dependencies. Too large, and distant embeddings have insufficient training signal, particularly early in training when most gradients flow through frequent short-distance pairs. Shaw et al. used k=16k = 16 in their original experiments on the WMT translation benchmark, but this value is too small for tasks involving long-range dependencies in English prose. In practice, kk is often set to match the typical context length during training, such as 64 or 128 for sentence-level tasks.

A subtler limitation is that Shaw et al.'s formulation encodes position at a single point in the transformer architecture: the attention score computation. This means position information can only influence the model through the routing decision (which tokens get attention) and the value aggregation (what information flows). It cannot influence the feedforward sublayers, layer norm operations, or any other component. More recent approaches like RoPE embed position information in a way that is more deeply integrated into the computation, allowing it to influence behavior throughout the network.

Finally, Shaw et al.'s relative position embeddings are fully learned, which means they depend on having sufficient training data to cover the full range of distance patterns the model needs to handle. For tasks with very limited data, the distant-position embeddings may remain undertrained throughout fine-tuning, creating a different kind of generalization problem than the one we set out to solve. In such cases, a fixed relative position scheme (like the sinusoidal approach used in Transformer-XL) may generalize better.

Despite these limitations, relative position encoding's impact on transformer research has been substantial. It demonstrated that position information can be injected at the attention level rather than the embedding level, opening the door to more sophisticated schemes like Rotary Position Embeddings and ALiBi that we'll explore in subsequent chapters. Every major language model architecture developed since 2019 has incorporated some form of relative position encoding, and the foundational ideas in this chapter appear in all of them.

Summary

Relative position encoding shifts the focus from "where am I?" to "how far apart are we?" By encoding the offset between positions rather than absolute indices, models learn patterns that generalize across all positions at the same distance.

Key takeaways from this chapter:

  • Absolute positions limit generalization: A model trained on short sequences may never see position pairs that appear in longer sequences. Relative encoding ensures that distance patterns transfer regardless of absolute position. Nearly 75% of position pairs in a 2x-length test sequence are unseen during training on the shorter length.

  • Shaw et al. formulation: Modify attention scores by adding qi⋅aj−iK\mathbf{q}_i \cdot \mathbf{a}^K_{j-i} to the content term. This separates semantic compatibility from distance-based compatibility. The position term is a dot product between the query and the relative position embedding, allowing different query types to respond differently to the same distance.

  • Clipping bounds memory: Relative positions beyond distance kk share the same embedding. This limits the number of parameters to 2k+12k + 1 regardless of sequence length, and also prevents distant embeddings from being undertrained due to infrequent occurrence.

  • Two components to modify: Shaw et al. add relative embeddings to both keys (affecting which positions get attention) and values (affecting what content flows). Both contribute to the learned distance patterns, though many implementations omit the value-side for efficiency.

  • Diagonal structure in attention: Position biases form a Toeplitz-like pattern where each anti-diagonal corresponds to a fixed relative distance. This structure enables efficient computation by precomputing all query-offset dot products in a single matrix multiply.

  • Trade-off between complexity and generalization: Relative position encoding adds computational overhead (roughly 5-10% at training time) but provides better generalization to unseen position pairs and longer sequences. The overhead grows with sequence length, which has motivated subsequent methods designed for long-context settings.

In the next chapter, we'll explore Rotary Position Embeddings (RoPE), a method that encodes relative positions through rotation rather than additive biases. RoPE achieves similar generalization benefits with a more elegant mathematical formulation that integrates naturally with the attention computation, and it has become the dominant approach in large language models such as LLaMA and Mistral, along with GPT-NeoX.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about relative position encoding.

Relative Position Encoding

Question 1 of 80 of 8 completed
What is the main advantage of relative position encoding over absolute position encoding?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025relativeposition, author = {Michael Brenndoerfer}, title = {Relative Position Encoding}, year = {2025}, url = {https://mbrenndoerfer.com/writing/relative-position-encoding-transformers}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-27} }
APAAcademic
Michael Brenndoerfer (2025). Relative Position Encoding. Retrieved from https://mbrenndoerfer.com/writing/relative-position-encoding-transformers
MLAAcademic
Michael Brenndoerfer. "Relative Position Encoding." 2026. Web. September 27, 2026. <https://mbrenndoerfer.com/writing/relative-position-encoding-transformers>.
CHICAGOAcademic
Michael Brenndoerfer. "Relative Position Encoding." Accessed September 27, 2026. https://mbrenndoerfer.com/writing/relative-position-encoding-transformers.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Relative Position Encoding'. Available at: https://mbrenndoerfer.com/writing/relative-position-encoding-transformers (Accessed: September 27, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Relative Position Encoding. https://mbrenndoerfer.com/writing/relative-position-encoding-transformers

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.