The Position Problem

Michael BrenndoerferUpdated June 5, 202547 min read

Part of Language AI Handbook

Explains why self-attention is blind to word order and what properties positional encodings need.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

The Position Problem: Why Self-Attention Is Blind to Word Order

In the previous chapters, we explored self-attention and the query-key-value mechanism in detail. We saw how each token can attend to every other token in a sequence, computing rich contextual representations that capture relationships regardless of distance. A token at position 1 can attend to a token at position 100 just as easily as to its immediate neighbor, which is one of self-attention's most powerful properties. But there is a fundamental limitation hiding in plain sight: self-attention has no notion of order. It treats the input sequence as an unordered bag of tokens, completely ignoring the positions where those tokens appear.

Think of it this way. Imagine you receive a bag of Scrabble tiles spelling out a message. You can see every letter and every pair of letters, but the bag has no concept of left-to-right sequence. You know which letters co-occur with which other letters, but you cannot reconstruct the original order from that information alone. Self-attention faces the same problem. It sees every pair of tokens and computes their compatibility, but the positions those tokens occupy leave no trace in the final representation.

This is not a minor implementation detail: it cuts to the heart of language understanding. Language is not a bag of words. The sentence "Dog bites man" and the sentence "Man bites dog" share exactly the same vocabulary, the same word co-occurrences, and the same pairwise relationships between tokens. Yet they describe opposite events. If a model cannot distinguish word order, it cannot distinguish these sentences. And if it cannot distinguish these sentences, it cannot reason about who is doing what to whom, which is arguably the most basic task in natural language understanding.

The position problem pervades nearly every aspect of language. Subject-verb agreement, pronoun resolution, negation scope, modifier attachment, syntactic parsing, temporal ordering of events, causal relationships between clauses: all of these linguistic phenomena depend critically on where words appear in a sentence, not merely on which words appear. A language model that ignores position is like a reader who sees all the words on a page at once but has no concept of reading left to right. The words are all there, but the structure that gives them meaning is lost.

This chapter explores why position blindness emerges from the mathematics of attention, why word order matters so critically for language understanding across multiple dimensions, and what properties any solution must possess to restore positional awareness to the model. We will set the stage for the various positional encoding schemes that subsequent chapters cover in depth. By the end, you will understand that position matters, why the attention mechanism discards it, and what the design criteria are for getting it back.

Where We Are in the Book

This chapter opens Part 4, which covers positional encoding. We assume familiarity with token embeddings (Part 1), the attention mechanism (Part 2), and the transformer architecture overview (Part 3). The chapters that follow this one present specific positional encoding solutions: sinusoidal encoding, learned embeddings, relative encodings, and RoPE. This chapter focuses on diagnosing the problem rather than solving it.

Attention Is Permutation Equivariant

To understand why self-attention ignores position, we need to examine its mathematical structure with care. The conclusion that self-attention is position-blind is not intuitive at first, because the computation does produce one output vector per position. It looks like position is being tracked. But appearances are deceiving. Let's build up the key insight step by step, starting with precise definitions and culminating in a demonstration that reveals where position information is missing.

Permutation Invariance vs. Equivariance

Two related properties appear throughout machine learning, and distinguishing them clarifies exactly what kind of symmetry self-attention has.

A function is permutation invariant if reordering its inputs leaves the output completely unchanged. Computing the mean of a set of numbers is permutation invariant: the mean of {3,7,2,5}\{3, 7, 2, 5\} is the same as the mean of {5,2,7,3}\{5, 2, 7, 3\}. Set membership and operations such as summation or taking a maximum have the same property. Permutation invariant functions discard order entirely.

A function is permutation equivariant if reordering the inputs causes the outputs to be reordered in exactly the same way. If you feed a permuted input, you get back the same outputs you would have gotten from the original input, just rearranged by the same permutation. Formally, a function ff is permutation equivariant if for any permutation π\pi:

f(π(X))=π(f(X))f(\pi(\mathbf{X})) = \pi(f(\mathbf{X}))

where:

  • X\mathbf{X}: the input sequence (a matrix where each row is a token's embedding)
  • π\pi: any permutation of the row indices
  • f(X)f(\mathbf{X}): the output sequence produced by the function ff
  • π(f(X))\pi(f(\mathbf{X})): that same output, with its rows rearranged by π\pi

Self-attention falls squarely into the equivariant category, not the invariant category.

Permutation Equivariance

A function ff is permutation equivariant if for any permutation π\pi applied to the input sequence, the output is permuted in the same way: f(π(X))=π(f(X))f(\pi(\mathbf{X})) = \pi(f(\mathbf{X})). Self-attention has this property because it processes all positions using the same computation with no position-specific parameters or inputs. Every position is treated identically by the machinery of attention.

Why does equivariance matter more than invariance here? Invariance would be even more extreme: the output itself would not change when we reorder inputs. Self-attention at least tracks which output corresponds to which input, which is why we describe it as equivariant rather than invariant. But equivariance is still problematic for language, because it means the model cannot tell from the outputs alone whether the input was "dog bites man" or "man bites dog." The relative ordering between positions carries no information. Two sequences that are permutations of each other produce outputs that are the same permutation of each other.

The key insight is this: permutation equivariance means that reordering the input has exactly the same effect as reordering the output. The attention mechanism does not encode any relationship between input positions that survives reordering. It is as if the attention computation is computed in a "coordinate-free" space where positions have no identity.

Tracing the Mathematics

To see why permutation equivariance emerges, we need to trace through the self-attention computation and identify every place where position could enter, but does not.

Self-attention produces an output vector for each position by aggregating information from all positions. For position ii, we compute a weighted sum where each position jj contributes its value vector, scaled by how relevant it is to position ii. This is the fundamental operation of attention:

yi=∑j=1nαijvj\mathbf{y}_i = \sum_{j=1}^{n} \alpha_{ij} \mathbf{v}_j

where:

  • yi∈Rdv\mathbf{y}_i \in \mathbb{R}^{d_v}: the output representation for position ii
  • αij\alpha_{ij}: the attention weight from position ii to position jj (determines how much jj contributes to ii's output)
  • vj∈Rdv\mathbf{v}_j \in \mathbb{R}^{d_v}: the value vector at position jj (the information that position jj contributes)
  • nn: the sequence length
  • ∑j=1n\sum_{j=1}^{n}: summation over all positions in the sequence

Notice that the formula is a weighted sum: a linear combination of value vectors. The weight αij\alpha_{ij} controls how much of position jj's value flows into position ii's output. If αij\alpha_{ij} is large, position jj has a strong influence on position ii's representation. If αij\alpha_{ij} is near zero, position jj contributes almost nothing.

This formula shows that the output at position ii depends on all value vectors vj\mathbf{v}_j and all attention weights αij\alpha_{ij}. But notice: the values come from the token content, and the weights come from content comparisons. Where does position enter?

The attention weights themselves come from applying softmax to scaled dot products between query and key vectors. The softmax function converts raw similarity scores into a probability distribution. This keeps all weights are positive and sum to 1:

αij=exp⁡ ⁣(qi⋅kj/dk)∑l=1nexp⁡ ⁣(qi⋅kl/dk)\alpha_{ij} = \frac{\exp\!\left(\mathbf{q}_i \cdot \mathbf{k}_j / \sqrt{d_k}\right)}{\sum_{l=1}^{n} \exp\!\left(\mathbf{q}_i \cdot \mathbf{k}_l / \sqrt{d_k}\right)}

where:

  • αij\alpha_{ij}: the attention weight from position ii to position jj (how much position ii attends to position jj)
  • qi∈Rdk\mathbf{q}_i \in \mathbb{R}^{d_k}: the query vector for position ii
  • kj∈Rdk\mathbf{k}_j \in \mathbb{R}^{d_k}: the key vector for position jj
  • dkd_k: the dimension of query and key vectors (used for scaling)
  • exp⁡(⋅)\exp(\cdot): the exponential function. This keeps all values are positive
  • ∑l=1n\sum_{l=1}^{n}: summation over all nn positions, serving as the normalizing constant

The queries and keys are computed by multiplying the input embedding by learned weight matrices. For position ii with input embedding xi\mathbf{x}_i:

qi=xiWQ,ki=xiWK,vi=xiWV\mathbf{q}_i = \mathbf{x}_i \mathbf{W}_Q, \quad \mathbf{k}_i = \mathbf{x}_i \mathbf{W}_K, \quad \mathbf{v}_i = \mathbf{x}_i \mathbf{W}_V

where WQ\mathbf{W}_Q, WK\mathbf{W}_K, and WV\mathbf{W}_V are the learned projection matrices shared across all positions. The same matrices are applied to every token, regardless of where it appears in the sequence.

The Missing Position Signal

Here is the critical observation: examine both equations carefully and notice what is absent. The indices ii and jj appear only as subscripts identifying which vectors to retrieve. They never appear as values in the computation itself. The attention weight αij\alpha_{ij} depends entirely on:

  1. The query vector qi\mathbf{q}_i, which is derived from the content at position ii (the token embedding xi\mathbf{x}_i)
  2. The key vector kj\mathbf{k}_j, which is derived from the content at position jj (the token embedding xj\mathbf{x}_j)
  3. The normalization over all keys, which also depends only on content

Position indices serve as labels that tell us which output corresponds to which input. They are bookkeeping, not input features. The number ii or jj itself never enters any arithmetic operation. If we swap the tokens at positions 2 and 5, the query and key vectors simply swap, and the computation produces swapped outputs. The absolute values 2 and 5 played no role.

Think of it as the difference between a library catalog (where shelf number matters for retrieval) and a set of books on a table (where the arrangement of books has no permanent identity). Self-attention is the table: books can be rearranged and the pairwise relationships remain identical, just reordered.

This contrasts sharply with recurrent networks. In an RNN or LSTM, position is implicitly encoded through sequential processing. The hidden state at step 5 depends on the hidden state at step 4, which depends on step 3, and so on. The temporal chain of computation bakes position into the representation as an inductive bias: earlier tokens shape later representations in ways that later tokens cannot shape earlier ones. Self-attention discards this sequential chain entirely, processing all positions in parallel and achieving massive efficiency gains, but at the cost of positional awareness.

Demonstrating Permutation Equivariance

The mathematical argument is clear, but seeing the effect in code makes it concrete. Let's implement self-attention and verify that permuting the input produces identically permuted output.

In[3]:
Code
import numpy as np

np.random.seed(42)


def softmax(x, axis=-1):
    """Compute softmax along the specified axis."""
    exp_x = np.exp(x - np.max(x, axis=axis, keepdims=True))
    return exp_x / np.sum(exp_x, axis=axis, keepdims=True)


def self_attention(X, W_Q, W_K, W_V):
    """
    Compute self-attention output.

    Args:
        X: Input embeddings of shape (seq_len, embed_dim)
        W_Q, W_K, W_V: Projection matrices

    Returns:
        Output representations of shape (seq_len, d_v)
    """
    Q = X @ W_Q
    K = X @ W_K
    V = X @ W_V

    d_k = K.shape[1]
    scores = Q @ K.T / np.sqrt(d_k)
    weights = softmax(scores, axis=-1)
    output = weights @ V

    return output, weights

We will create a simple 4-token sequence, apply a permutation that swaps two positions, and compare the outputs.

In[4]:
Code
# Create embeddings for a 4-token sequence
seq_len = 4
embed_dim = 8
d_k = d_v = 6

# Random embeddings representing different tokens
X = np.random.randn(seq_len, embed_dim)

# Initialize projection matrices
W_Q = np.random.randn(embed_dim, d_k) * 0.1
W_K = np.random.randn(embed_dim, d_k) * 0.1
W_V = np.random.randn(embed_dim, d_v) * 0.1

# Compute attention on original sequence
output_original, weights_original = self_attention(X, W_Q, W_K, W_V)

# Create a permutation: swap positions 1 and 2
permutation = [0, 2, 1, 3]
X_permuted = X[permutation]

# Compute attention on permuted sequence
output_permuted, weights_permuted = self_attention(X_permuted, W_Q, W_K, W_V)

# Apply the same permutation to the original output for comparison
output_original_reordered = output_original[permutation]
Out[5]:
Console
Permutation applied: [0, 2, 1, 3] (swap positions 1 and 2)

Original output shape: (4, 6)
Permuted output shape: (4, 6)

Max difference between pi(output) and output(pi(input)): 2.78e-17
Outputs are identical (within numerical precision)

The outputs match exactly, up to floating-point precision. This confirms permutation equivariance: permuting the input permutes the output in the same way. Self-attention does not "know" that we reordered the tokens. It computes the same pairwise relationships in a different order and produces the same outputs in a different order. No information about the original positions has been preserved beyond the assignment of outputs to input slots.

The Root Cause: Content-Only Dependencies

This experiment confirms what the mathematics predicted. Self-attention treats the input as an unordered set of tokens, using only their content to determine relationships. Each token's output depends on three things, none of which involve position:

  1. Its own content (through its query vector, which asks "what am I looking for?")
  2. The content of all other tokens (through their key vectors, which answer "what do I offer?")
  3. The pairwise content relationships (through dot products that measure semantic compatibility)

This is a deliberate architectural choice, not an oversight. The original Transformer designers wanted to avoid the sequential bottleneck of RNNs, where position is encoded implicitly by the order of computation. Parallelizing across all positions means forgoing positional inductive bias. The model is given maximum flexibility to learn relationships from content alone, but that flexibility comes with the responsibility of explicitly injecting position information from outside the attention mechanism.

Notice that this is quite different from convolutional neural networks (CNNs), which have a different kind of positional inductive bias. A CNN's filters are applied at specific spatial locations, and nearby positions share features through local connectivity. Position is implicitly encoded through the spatial structure of the convolution. Self-attention has no analogous spatial structure. It is maximally flexible with respect to distance.

Why Position Matters for Language

This position blindness might seem like a minor technical detail, but it is catastrophic for language understanding. Human language is deeply structured around word order, and this structure operates at multiple levels simultaneously. Let's examine why position is essential across several dimensions of linguistic meaning.

Grammatical Roles and Argument Structure

In English and many other languages, word order is the primary mechanism for encoding grammatical roles. Who is the agent (the one doing) and who is the patient (the one being done to) depends critically on position relative to the verb.

Consider these sentences:

  • "The dog chased the cat."
  • "The cat chased the dog."

The same words appear in both sentences, but they play opposite grammatical roles. In the first sentence, "dog" is the subject (the chaser) and "cat" is the object (the chased). In the second, these roles reverse entirely. Pure self-attention, seeing only that "dog," "chased," and "cat" co-occur, cannot distinguish which is the agent and which is the patient. It knows these three words appear together. It does not know who is biting whom.

This matters enormously for any downstream task. Sentiment analysis, information extraction, question answering, and reading comprehension all require knowing who does what to whom. A model that cannot distinguish subject from object cannot extract the event structure of a sentence. And event structure, it turns out, is most of what sentences express.

The technical term for this positional encoding of grammatical roles is argument structure. Verbs take arguments, and those arguments have specific semantic roles called thematic roles or theta roles. These include agent and patient. Other examples include beneficiary, instrument, or location. In English, the mapping from syntactic position to thematic role is largely positional. The entity before the verb tends to be the agent; the entity after the verb tends to be the patient. Without positional information, the model has no way to infer these roles.

Negation Scope

Word order determines what negation applies to, which can completely reverse the meaning of a sentence:

  • "I did not say he stole the money." (I said nothing; the accusation is false or came from someone else)
  • "I did say he did not steal the money." (I spoke up; I claimed his innocence)

The positioning of "not" relative to "say" and "steal" completely changes the assertion. Moving a single word maps an accusation of silence into a defense of character. A bag-of-words model sees "not," "say," and "steal" together in both sentences and cannot determine which action is negated.

This extends to other scope-sensitive phenomena. Quantifiers and modal verbs interact with scope in ways that depend on linear order; adverbs and conditionals do too. "Every student read some book" has a different meaning from "Some book was read by every student," and this meaning difference is entirely carried by word order.

Modifier Attachment and Structural Ambiguity

Consider the well-known ambiguity in:

  • "I saw the man with the telescope."

Does "with the telescope" modify "saw" (I used a telescope to see) or "man" (the man was carrying a telescope)? Both interpretations are syntactically and semantically plausible. The resolution depends on which positional interpretation the reader adopts. More precisely, it depends on which element in the sentence the prepositional phrase is attached to structurally, and that attachment is partially determined by proximity and order.

A second classic example:

  • "The horse raced past the barn fell."

This garden-path sentence initially appears to have "raced past the barn" as the main clause, but the final word "fell" reveals a different structure: "The horse [that was] raced past the barn fell." Readers process this sentence positionally, building a parse tree as they read left to right. Self-attention without positional information cannot capture this left-to-right progressive parsing that human language processing depends on.

Temporal and Causal Ordering

Sequence often implies temporal or causal relationships:

  • "She got her degree, then found a job."
  • "She found a job, then got her degree."

The order of clauses suggests different life trajectories. Both sentences contain the same events, but their relative positions encode temporal precedence. In narrative text, the order in which events are mentioned generally corresponds to their order in time (the "narrative sequence" principle). A model that cannot track positional order cannot extract the timeline of a story.

Causality also frequently depends on sequential position. "The bridge collapsed because the engineers cut corners" has a different causal structure from "The engineers cut corners because the bridge collapsed," even though both sentences contain the same propositions. Without position, the model cannot determine which event caused which.

Semantic Composition and Compound Formation

Languages compose meaning hierarchically, and order often determines the composition structure:

  • "Hot dog" (a type of food, very different from a warm canine)
  • "Dog hot" (not standard English, but in languages with post-nominal adjectives, this would mean "a dog that is warm")

English compounds typically have the modifier before the head noun. Without positional information, "hot" and "dog" are just two tokens that appear together, with no way to determine which modifies which, or even whether this is a compound or a two-word phrase.

At the phrasal level, the same words in different orders can produce entirely different semantic compositions:

  • "Only she spoke to him." (no one else spoke to him)
  • "She only spoke to him." (she spoke but did nothing else)
  • "She spoke only to him." (she excluded everyone else as a listener)

The adverb "only" is a classic focus particle that must be positioned immediately before the element it focuses on. Moving it changes the scope of exclusion and therefore the meaning of the sentence. Self-attention without positional information cannot recover which element "only" is focusing on.

In[6]:
Code
# Simulate the attention pattern for "dog bites man" vs "man bites dog"
# We'll use random embeddings but give them distinct identities

np.random.seed(123)

# Create distinct embeddings for each word
word_embeddings = {
    "dog": np.random.randn(embed_dim),
    "bites": np.random.randn(embed_dim),
    "man": np.random.randn(embed_dim),
}

# Construct the two sequences
sentence1 = np.stack(
    [word_embeddings["dog"], word_embeddings["bites"], word_embeddings["man"]]
)
sentence2 = np.stack(
    [word_embeddings["man"], word_embeddings["bites"], word_embeddings["dog"]]
)

# Compute attention weights for both
_, weights1 = self_attention(sentence1, W_Q, W_K, W_V)
_, weights2 = self_attention(sentence2, W_Q, W_K, W_V)
Out[7]:
Visualization
3x3 heatmap showing attention weights for dog bites man sentence.
Attention weights for 'dog bites man'. Each cell shows how much the row token attends to the column token. Note that the weight from 'dog' to 'bites' and from 'man' to 'bites' reflect only content similarity, not positional role.
3x3 heatmap showing attention weights for man bites dog sentence.
Attention weights for 'man bites dog'. The pairwise weights between corresponding word pairs are identical to the left panel. This shows that position information is absent from the computation despite the opposite sentence meaning.
Out[8]:
Console
--- Key Observation ---
Attention from 'dog' to 'bites': 0.3439 (sentence 1, dog at position 0)
Attention from 'dog' to 'bites': 0.3439 (sentence 2, dog at position 2)

These are identical because attention only depends on content, not position!

The heatmaps reveal the core problem vividly. Despite "dog bites man" and "man bites dog" having opposite meanings (opposite agent-patient roles), the attention weight between "dog" and "bites" is identical in both sentences. Self-attention cannot tell that in one sentence the dog is the biter, while in the other the dog is the bitten. The pairwise content relationships are the same; only the positions differ. The attention mechanism computes rich contextual representations, but all of that context is semantically based, with no positional signal anywhere in the computation.

A Worked Example: Positional Ambiguity at Scale

To make the problem even more concrete, consider a longer passage:

"The committee voted to ban the organization, and the organization appealed the decision."

A model with no positional encoding sees: {committee, voted, ban, organization, organization, appealed, decision}. It cannot tell that the first "organization" is the object of "ban" and the second is the subject of "appealed." In a bag-of-words view, these two occurrences of "organization" are identical tokens. Only their positions differentiate their grammatical roles and their referential relationship to the surrounding verbs.

Now consider coreference resolution, which is the task of identifying which noun phrases refer to the same entity. In this sentence, does "organization" in the second clause refer to the same organization as in the first? Almost certainly yes, but arriving at that conclusion requires tracking that "the organization" appears first as a patient (being banned) and then as an agent (doing the appealing). Positional information is essential for this tracking.

This example is not exotic. Most sentences of any complexity require positional reasoning. The longer and more syntactically complex a sentence becomes, the more critical it is for the model to track where each token appears relative to the governing verbs and the surrounding clause structure and connectives.

What We Need from Positional Information

Understanding why position matters helps us identify what properties any positional encoding scheme must have. The ideal solution must satisfy several requirements, and each requirement rules out simple approaches you might first consider.

Unique Position Identification

Each position in a sequence needs a distinct representation. If positions 3 and 7 have identical positional signals, the model cannot distinguish tokens at those positions based on their locations. This requirement seems obvious, but it rules out the simplest possible approaches. Using a constant vector (the same positional signal everywhere) provides no information. Using a random vector assigned at initialization but not varied across positions gives each position a signal, but that signal is arbitrary and meaningless to the model.

The uniqueness requirement means the positional representation must vary systematically with position. The question is how.

Bounded and Stable Values

Whatever numerical representation we use for position, it must have bounded, well-behaved values throughout. If position 1 is encoded as 1, position 1000 as 1000, and position 1 million as 1 million, the scale differences will overwhelm the semantic content of token embeddings. Neural networks are sensitive to the magnitudes of their inputs. Unbounded positional representations would create training instabilities and force the model to learn very different weight scales to handle early versus late positions in a sequence.

The key insight is that the positional representation should live in the same numerical space as the token embeddings it is being combined with. Token embeddings are typically distributed with mean near zero and standard deviation near one, thanks to embedding initialization schemes and layer normalization. Positional representations should match this scale.

Consistency Across Sequence Lengths

The encoding for position 5 should be the same whether it appears in a 10-token sequence or a 10,000-token sequence. If positional representations change based on sequence length, the model must learn different positional patterns for every possible length it might encounter. This would make it impossible to transfer knowledge from short sequences to long ones, and vice versa.

This requirement rules out a common naive approach: normalizing the position by the sequence length. Position pp in a sequence of length nn might be represented as p/np/n, giving values between 0 and 1. This is bounded, which is good. But it violates consistency: position 5 in a 10-token sequence gets value 0.5, while position 5 in a 100-token sequence gets value 0.05. The model learns that "0.5" means "halfway through," but the absolute meaning of position 5 changes with sequence length.

Relative Position Accessibility

Many linguistic relationships depend on relative rather than absolute position. A verb typically appears near its subject and object, regardless of where they fall in the sentence. An adjective typically immediately precedes the noun it modifies. A relative clause immediately follows the noun it modifies. These are all relative positional relationships. The model needs to compute "token A is 3 positions before token B," rather than only "token A is at position 7 and token B is at position 10."

Ideally, the positional encoding should make it easy to determine that position 7 is "2 steps after" position 5, rather than only showing that both are somewhere in the sequence. This is a stronger requirement: each position needs a unique signal, and the relationship between signals needs to encode relative distance in a way the model can exploit.

Generalization Beyond Training

Models trained on sequences of length 512 may encounter sequences of length 1024 at inference time. The positional encoding should degrade gracefully (or not at all) when extrapolating to positions not seen during training. This requirement is especially important for applications where input length is unpredictable, such as processing documents of arbitrary length.

Fixed mathematical functions (like sinusoids) can extrapolate naturally, because the formula applies to any position, including ones never seen during training. Learned embeddings, by contrast, have no entry in their lookup table for positions beyond the training maximum. Handling this gracefully requires interpolation, extrapolation heuristics, or other architectural choices.

Compatibility with Attention

The positional information must integrate with the attention mechanism in a way that allows position to influence attention patterns. If a query cannot distinguish nearby keys from distant keys, positional information is not flowing into the core computation. This requirement means we cannot just append position indices somewhere; the position information must be in a form that the dot-product attention can use.

The standard approach is to combine positional representations with token embeddings before computing QKV representations. This way, when the model computes dot products between queries and keys, those dot products depend on position as well as content.

In[9]:
Code
# Illustrate the problem with naive positional representations

positions_naive = np.arange(1, 513)

# Naive approach 1: Raw position index
naive_index = positions_naive

# Naive approach 2: Normalized position (0 to 1)
naive_normalized = positions_naive / 512

# Naive approach 3: Log-scaled position
naive_log = np.log(positions_naive)
Out[10]:
Visualization
Line plot showing raw position indices growing linearly from 1 to 512.
Raw position indices grow linearly and without bound. At position 512 the value is 512, which is hundreds of times larger than typical token embedding values near zero, causing the positional signal to overwhelm semantic content.
Line plot showing normalized positions from near 0 to 1.
Normalized positions change meaning with sequence length. Position 256 maps to 0.5 in a 512-token sequence but to 0.25 in a 1024-token sequence, so the model cannot generalize positional patterns across different sequence lengths.
Line plot showing log-scaled positions with diminishing growth rate.
Log scaling compresses distant positions, causing poor discrimination between later positions. Positions 400 and 500 differ by only about 0.22 in log scale, while positions 1 and 2 differ by 0.69, creating highly unequal positional resolution.

None of these naive approaches satisfy our requirements. Raw indices are unbounded, growing to hundreds or thousands while token embeddings remain in a range around zero. Normalized positions change meaning with sequence length: position 0.5 is location 256 in a 512-token sequence but location 512 in a 1024-token sequence, making the learned positional patterns non-transferable. Log scaling compresses distant positions so tightly that the model can barely distinguish position 400 from position 500, creating highly unequal positional resolution across the sequence.

The requirements we have identified describe a multi-dimensional representation with bounded values, stable meaning across sequence lengths, and enough variation to distinguish positions. This is the problem that sinusoidal encodings and learned embeddings were designed to solve, and we will explore both in detail in subsequent chapters.

Position Encoding vs. Position Embedding

Before diving into specific solutions, let us clarify an important terminological distinction that causes persistent confusion in the literature.

Encoding vs. Embedding

Positional encoding refers to a fixed, deterministic function that maps position indices to vectors. The function is defined by a mathematical formula, not by training. Positional embedding refers to learned vectors stored in a lookup table, one per position. Both inject positional information into the model, but through fundamentally different mechanisms with different trade-offs.

Positional encoding uses a predetermined formula to compute a position vector. Given position pp, the encoding function f(p)f(p) returns a fixed vector that never changes during training. No gradient flows through the positional encoding itself; it is a fixed input to the model. The classic example is the sinusoidal encoding from the original Transformer paper, which uses sine and cosine functions at different frequencies to construct a unique vector for each position. The primary advantage is generalization: because the formula is defined for all integers, it applies to any position, including those longer than anything seen during training. The model can, in principle, handle sequences of arbitrary length. The drawback is that the encoding is fixed and may not capture the positional patterns most useful for a specific task.

Positional embedding treats positions like vocabulary items. We create an embedding matrix P∈RL×d\mathbf{P} \in \mathbb{R}^{L \times d}, where LL is the maximum sequence length and dd is the embedding dimension. Position pp looks up row pp of this matrix, retrieving a dd-dimensional learned vector pp\mathbf{p}_p. These embeddings are initialized randomly and then updated through training by gradient descent, just like word embeddings. The advantage is flexibility: the model can discover whatever positional representations are useful given the training data and task. The disadvantages are that positions beyond LL have no representation, and the embeddings must be trained from scratch, requiring sufficient examples of each position to learn useful representations.

Most modern language models use positional embeddings (learned lookup tables) for short-to-medium contexts. Models designed for longer contexts or better generalization tend to use more sophisticated positional schemes like RoPE or ALiBi. The terminology gets mixed in the literature, with both approaches sometimes loosely called "positional encoding," so it is worth being precise about the distinction.

How Position Information Enters the Model

Now that we understand the two mechanisms for generating positional vectors, the remaining question is: how do we inject this information into the attention computation?

Recall from our earlier analysis that attention weights depend only on query and key vectors, which are projections of the input embeddings. To make attention position-aware, we need to modify what goes into those projections. There are several conceptual approaches:

  1. Input modification: Change the token embeddings before they enter the model, so that the Q/K/V projections see both content and position
  2. Attention score modification: Directly add a positional bias to the attention scores after the Q/K dot product but before softmax
  3. Relative encoding in projections: Modify the key vectors to include relative position terms

The standard approach in the original Transformer (and in most BERT-style models) is option 1: add the positional vector directly to the token embedding, creating a single combined representation that carries both semantic meaning and positional information:

xi′=xi+pi\mathbf{x}'_i = \mathbf{x}_i + \mathbf{p}_i

where:

  • xi∈Rd\mathbf{x}_i \in \mathbb{R}^d: the token embedding at position ii (carries semantic content)
  • pi∈Rd\mathbf{p}_i \in \mathbb{R}^d: the positional vector for position ii (carries positional information)
  • xi′∈Rd\mathbf{x}'_i \in \mathbb{R}^d: the combined representation fed to the self-attention layers
  • dd: the embedding dimension (must match for both token and position vectors, enabling element-wise addition)

Why addition rather than concatenation or multiplication? Addition is the simplest operation that preserves dimensionality. If we concatenated positional vectors, the combined representation would have dimension 2d2d, requiring all downstream weight matrices to be twice as wide. Addition keeps everything at dimension dd. More subtly, addition allows the model to learn how much weight to give positional versus semantic information through the QKV projection matrices. If the model finds certain dimensions of the positional encoding uninformative for a task, the learned projection matrices can simply down-weight those dimensions. The addition operation is also differentiable, so gradients flow through it cleanly during training.

In practice, the positional vectors are often scaled to have a similar magnitude to the token embeddings, to prevent either one from dominating the combined representation. This is particularly important with sinusoidal encodings, which we will examine in the next chapter.

In[11]:
Code
def add_positional_info(token_embeddings, position_vectors):
    """
    Combine token embeddings with positional information.

    Args:
        token_embeddings: Shape (seq_len, embed_dim)
        position_vectors: Shape (seq_len, embed_dim)

    Returns:
        Combined representations of shape (seq_len, embed_dim)
    """
    return token_embeddings + position_vectors


# Example: simple positional vectors (not a real encoding scheme)
seq_len_ex = 4
embed_dim_ex = 8

# Token embeddings
tokens_ex = np.random.randn(seq_len_ex, embed_dim_ex)

# Placeholder positional vectors (small scale to illustrate the combination)
positions_ex = np.random.randn(seq_len_ex, embed_dim_ex) * 0.1

# Combined representation
combined_ex = add_positional_info(tokens_ex, positions_ex)
Out[12]:
Console
Shapes:
  Token embeddings:     (4, 8)
  Position vectors:     (4, 8)
  Combined:             (4, 8)

Example: Position 0
  Token:    [[-1.254 -0.638  0.907 -1.429]...]
  Position: [[0.089 0.175 0.15  0.107]...]
  Combined: [[-1.165 -0.462  1.057 -1.322]...]

The shapes confirm that adding positional vectors preserves the embedding dimension. Looking at position 0, you can see how each dimension of the combined vector is simply the sum of the corresponding token and position dimensions. The positional vectors are scaled by 0.1 here to keep them relatively small compared to the token embeddings. This is a common practice that prevents positional information from dominating semantic content at the start of training. Over time, the model can adjust the influence of positional information through its learned projections.

The combined representation now carries both semantic content (from the token embedding) and positional context (from the positional vector). When this combined representation is projected into QKV representations, both types of information influence the attention pattern. A token's query vector asks "what am I looking for, given my content and position?" A token's key vector offers "what do I provide, given my content and position?" The dot product between these position-aware queries and keys produces attention weights that reflect both semantic relevance and positional relationships.

The Alternative: Modifying Attention Scores Directly

Some approaches, including ALiBi (which we will cover in a later chapter), bypass input modification entirely and instead add position-dependent biases directly to the attention scores:

αij=softmax ⁣(qi⋅kjdk+b(i,j))\alpha_{ij} = \text{softmax}\!\left(\frac{\mathbf{q}_i \cdot \mathbf{k}_j}{\sqrt{d_k}} + b(i, j)\right)

where b(i,j)b(i, j) is a bias term that depends on the positions ii and jj (typically just their distance ∣i−j∣|i - j|). This approach has the advantage that the position signal never gets mixed with the semantic content in the embeddings. Semantic relationships are computed purely from content, and then positional biases are applied afterward to adjust attention patterns. This cleaner separation can improve extrapolation: the semantic attention mechanism is unchanged, and positional biases are applied on top.

Understanding that there are multiple points of entry for positional information will be important as we survey positional encoding methods in subsequent chapters. Different methods inject position at different stages of the computation, leading to different trade-offs.

Absolute vs. Relative Position

A final conceptual distinction shapes the entire design space for positional representations: should we encode where each token is in absolute terms, or should we encode how tokens relate to each other in relative terms? This choice is not merely technical; it reflects a hypothesis about what kinds of positional patterns matter most for language.

Absolute positional encoding assigns each position a fixed vector. Position 0 always gets the same encoding, position 1 always gets the same encoding, and so on, regardless of the sequence. The attention mechanism receives absolute location signals and must learn to compute relative positional information by comparing absolute positions.

Relative positional encoding directly encodes the distance or relationship between pairs of positions. Instead of saying "this token is at position 5," relative encoding says "this token is 3 positions before the query token." The attention mechanism directly incorporates these relative distances, without needing to infer them from absolute locations.

Out[13]:
Visualization
Diagram showing five word boxes labeled The cat sat on mat with position indices 0 through 4.
Absolute position assigns a fixed index to each token regardless of context. The numbers 0 through 4 label where each word appears in the sequence. Any downstream computation that needs to know token distances must derive that information by comparing absolute indices.
Diagram showing relative distances from the query token sat to other tokens, with arrows and signed integer labels.
Relative position encodes the signed distance from a chosen query token (here 'sat' at position 2). Tokens to the left have negative offsets, tokens to the right have positive offsets. This representation directly supports distance-based reasoning without requiring absolute index comparison.

Each approach has significant trade-offs. Absolute encoding is simpler to implement: just add a position vector to each token embedding before feeding it to the model. The attention mechanism then sees combined content-position vectors and can, with sufficient capacity, learn to extract relative positional information by computing differences between absolute positions. The original Transformer and later models such as BERT and GPT all use absolute positional encoding.

The limitation of absolute encoding is that it requires the model to learn the mapping from absolute to relative positions implicitly. If the model sees "position 3 attends to position 1" during training, it must generalize to "position 103 attends to position 101" at inference. These are the same relative offset (+2) expressed as different absolute positions. With enough data and capacity, models can learn this generalization, but it is not guaranteed, and it may require more training examples than relative encoding.

Relative encoding directly captures these distance-based relationships. Instead of embedding absolute positions, the model embeds the distances i−ji - j between query position ii and key position jj. This representation is translation-invariant by construction: a distance of +2 means the same thing regardless of whether we are at positions 1 and 3 or positions 101 and 103. The model does not need to generalize across absolute positions.

The drawback of relative encoding is architectural complexity. Absolute positional vectors can be added to token embeddings once at the input layer, and the rest of the model is unchanged. Relative positions depend on pairs of positions, not individual positions, so they cannot simply be added to the input. The attention score between positions ii and jj must explicitly incorporate information about the distance i−ji - j. This requires modifying the attention computation itself, rather than changing only the input preprocessing. The result is a more expressive model that is also harder to implement efficiently.

Modern architectures increasingly favor relative or hybrid approaches. RoPE (Rotary Position Embedding) cleverly encodes relative positions through rotation operations that integrate naturally with the standard attention dot product. ALiBi (Attention with Linear Biases) adds a simple learned bias to attention scores based on distance, achieving relative encoding with minimal architectural change. We will explore both approaches in dedicated chapters.

Why Relative Positions Are Linguistically Natural

The preference for relative encoding in recent architectures reflects a linguistic observation about local relationships. Most syntactic and semantic relationships in language are local and relative, not global and absolute. A subject typically appears within a few words of its verb, and this proximity holds whether the subject-verb pair appears at the beginning of a sentence or in the middle of a long document. A determiner ("the," "a") always immediately precedes the noun phrase it introduces, and this is a relative relationship independent of absolute position. Anaphoric pronouns refer back to noun phrases that appeared somewhere earlier in the discourse, a backward-pointing relative relationship.

If positional representations encode relative distances well, the model can more easily learn these universal syntactic generalizations. If they encode only absolute positions, the model must learn to apply each generalization at every possible absolute location, multiplying the number of patterns it needs to internalize.

This does not mean absolute encoding is useless. Document-level reasoning, reference tracking across very long spans, and certain ordering tasks may benefit from knowing absolute position. A model that knows it is processing the first sentence of a document versus the twentieth sentence might behave differently. Hybrid approaches that encode both absolute and relative information are an active area of research.

A Formal Look at What Breaks Without Position

It is worth being precise about the consequences of permutation equivariance for specific NLP tasks. Different tasks suffer from position blindness in different ways.

Classification tasks are affected but sometimes less severely. If a task asks "does this review express positive or negative sentiment?", the answer often depends more on the presence of positive/negative words than on their order. "Terrible, not at all good" and "good, not at all terrible" have similar bag-of-words representations but different meanings. However, for most documents, sentiment is carried reasonably well by vocabulary, and classification models can sometimes compensate. The failure mode is subtle and appears when negation or ordering changes sentiment.

Sequence labeling tasks are severely affected. Named entity recognition, part-of-speech tagging, and syntactic parsing all require assigning labels to specific positions in the sequence. If the model cannot determine that "bank" is at position 3 and modifies "river" at position 4, it cannot correctly determine that "bank" is a noun functioning as part of a compound rather than as a financial institution. Positional relationships among adjacent tokens are essential for sequence labeling.

Generation tasks are catastrophically affected. Language modeling requires predicting the next token given the previous ones, and causal masking means only earlier positions are visible. Without positional encoding, the model cannot determine which tokens came earlier and which came later, making causal prediction undefined. The model cannot enforce that the output at position tt conditions only on positions 1,…,t−11, \ldots, t-1 without knowing what those position indices mean.

Question answering and reading comprehension sit somewhere in between. Finding the answer span in a passage requires understanding that a specific sequence of tokens at a specific location in the passage is the answer. This is heavily positional. But recognizing that the passage contains the relevant information is often more content-based.

The breadth of tasks affected makes positional encoding not an optional enhancement but a fundamental requirement for any transformer-based system doing real natural language processing.

Limitations of Any Positional Encoding Approach

No positional encoding scheme is perfect. Every approach makes trade-offs that are worth understanding before choosing a scheme for a given application.

Fixed encodings may not capture learned patterns. Sinusoidal encoding uses mathematically determined frequencies. These frequencies produce a particular kind of positional representation that may or may not align with what is most useful for a specific task. A model with learned embeddings can discover whatever positional representations the training data requires. The sinusoidal encoding makes an implicit assumption that periodic patterns across dimensions are the right way to represent position, but this assumption has no linguistic grounding.

Learned embeddings do not generalize to longer sequences. If you train with a maximum sequence length of 512, positions 513 and beyond have no learned representation. Some models handle this with interpolation (linearly interpolating between nearby learned embeddings) or by fine-tuning on longer sequences before deployment. But performance often degrades for positions not seen during training, sometimes dramatically. This is a fundamental limitation that requires either architectural solutions (relative encoding) or specific fine-tuning protocols.

Additive combination limits expressiveness. Adding positional vectors to token embeddings means the combined representation must encode both content and position in a shared dd-dimensional space. The model may struggle when positional and semantic information conflict. For example, if a word's semantic embedding happens to be similar to the positional encoding of position 5, the model may confuse that word's content with a positional signal. More fundamentally, addition forces the model to represent semantics and position in the same space, rather than in separate subspaces that might be more efficiently organized.

Relative positions require architectural changes. Directly encoding relative positions requires modifying the attention computation rather than changing only the input. While libraries like Hugging Face Transformers handle this transparently, implementing relative encoding from scratch is substantially more complex than the absolute approach. This complexity can introduce bugs and makes the model harder to reason about.

Very long sequences remain challenging. Even with sophisticated positional encoding, transformers face fundamental challenges with extremely long sequences. The attention mechanism has quadratic memory and computation complexity with respect to sequence length. Positional patterns learned on sequences of a few thousand tokens may not transfer to sequences of hundreds of thousands of tokens. Recent work on sparse attention, linear attention, and state-space models addresses the quadratic complexity issue, but the positional encoding problem for very long documents is still an active area of research.

Position alone is insufficient for some phenomena. Even with perfect positional encoding, some linguistic phenomena require representations that go beyond simple sequential position. Hierarchical structure (syntactic parse trees), coreference chains, discourse structure, and rhetorical relations are not captured by linear sequential position alone. Position encoding solves the "which slot in the sequence" problem, but does not solve the "what is the syntactic role" problem directly. Downstream learning from the encoded positions is still required.

Historical Context: From RNNs to Transformers

The position problem is unique to the transformer architecture. Recurrent neural networks (RNNs) and their variants (LSTMs, GRUs) have no analogous problem because sequential processing is their core inductive bias. The hidden state at each step carries information about everything that came before, building up positional awareness implicitly. The Transformer's "Attention Is All You Need" paper (Vaswani et al., 2017) made the deliberate decision to abandon sequential processing entirely in favor of fully parallel computation. This improved training speed and scalability dramatically, but it required a new solution for position. The sinusoidal encoding proposed in that paper was the first widely adopted answer, and its strengths and limitations have led to the wide range of positional encoding approaches used today.

In Practice: What Modern Models Use

Comparing these options helps interpret the design choices in state-of-the-art models. Here is a brief survey of what modern systems use; later chapters examine these methods in detail.

BERT (2018) uses absolute learned positional embeddings with a maximum sequence length of 512. These are trained from random initialization along with the word embeddings. BERT's success demonstrated that learned absolute embeddings work well for most NLP tasks, though they cannot generalize beyond the 512-position limit.

GPT-2 and GPT-3 also use absolute learned positional embeddings, with GPT-2 supporting 1024 tokens and GPT-3 supporting 2048. These are pure lookup tables, and positions beyond the training maximum are unsupported without fine-tuning.

RoBERTa also uses learned absolute positional embeddings (same architecture as BERT). This shows that the training recipe matters more than the positional encoding scheme for many tasks.

T5 uses relative positional encoding through learned scalar biases added to attention scores. Each attention head learns a separate bias for each relative position bucket. This provides relative position awareness without changing the input embedding process.

GPT-NeoX and LLaMA use Rotary Position Embedding (RoPE), which encodes relative positions through rotation matrices applied to query and key vectors. RoPE has become the dominant choice in recent open-source language models due to its strong generalization properties.

Mistral and Mixtral also use RoPE, with additional tricks like sliding window attention to extend effective context length.

The trend in the field is clearly moving toward relative encoding schemes, particularly RoPE, for their superior generalization and extrapolation properties. However, understanding why absolute encoding was the starting point, and what its limitations are, is essential context for understanding why these newer approaches were developed.

Summary

Self-attention is blind to position because its computation depends only on content, not on where tokens appear in the sequence. This permutation equivariance is a fundamental property of the attention mechanism, arising from the fact that position indices serve only as bookkeeping labels and never appear as values in any arithmetic operation. It is not a bug in any particular implementation; it is built into the mathematical structure of dot-product attention.

Key takeaways from this chapter:

  • Permutation equivariance: Self-attention produces the same outputs (reordered) regardless of input order. Shuffling the input shuffles the output identically, because attention weights depend only on query-key compatibility, not on position indices. The equation f(π(X))=π(f(X))f(\pi(\mathbf{X})) = \pi(f(\mathbf{X})) captures this precisely.
  • Language requires order: Grammatical roles, negation scope, modifier attachment, temporal relationships, and semantic composition all depend on word position. "Dog bites man" and "man bites dog" are opposite events expressed by the same vocabulary, and only position distinguishes them.
  • Requirements for positional information: Any solution must provide unique position identification, bounded values, consistency across sequence lengths, accessibility of relative positions, generalization beyond training, and compatibility with attention.
  • Encoding vs. embedding: Positional encoding uses fixed mathematical formulas (like sinusoids) that generalize to any position. Positional embedding uses learned lookup tables that adapt to training data but have a fixed maximum length. Both have significant trade-offs.
  • Absolute vs. relative: Absolute encoding assigns fixed vectors to positions, requiring the model to infer relative distances. Relative encoding directly captures distances between positions, supporting better generalization but requiring architectural modifications.
  • Injection mechanism: The standard approach adds positional vectors to token embeddings before computing queries and keys, combining semantic and positional information in a shared representation. Alternative approaches modify attention scores directly.
  • Modern trends: The field has moved from absolute learned embeddings (BERT, GPT) toward relative encoding schemes (T5, RoPE). This reflects the linguistic insight that relative distances matter more than absolute positions for most language tasks.

In the next chapter, we will examine the sinusoidal positional encoding introduced in the original Transformer paper. This elegant mathematical construction uses sine and cosine functions at different frequencies to create unique position vectors that support relative position computation through simple linear operations. You will see exactly how the requirements identified in this chapter shaped the design of that solution.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the position problem in self-attention.

The Position Problem

Question 1 of 80 of 8 completed
What property does self-attention have that makes it blind to word order?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025positionproblem, author = {Michael Brenndoerfer}, title = {The Position Problem}, year = {2025}, url = {https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). The Position Problem. Retrieved from https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order
MLAAcademic
Michael Brenndoerfer. "The Position Problem." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order>.
CHICAGOAcademic
Michael Brenndoerfer. "The Position Problem." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order.
HARVARDAcademic
Michael Brenndoerfer (2025) 'The Position Problem'. Available at: https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). The Position Problem. https://mbrenndoerfer.com/writing/position-problem-self-attention-word-order

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.