Part of Language AI Handbook
Explains why self-attention is blind to word order and what properties positional encodings need.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
The Position Problem: Why Self-Attention Is Blind to Word Order
In the previous chapters, we explored self-attention and the query-key-value mechanism in detail. We saw how each token can attend to every other token in a sequence, computing rich contextual representations that capture relationships regardless of distance. A token at position 1 can attend to a token at position 100 just as easily as to its immediate neighbor, which is one of self-attention's most powerful properties. But there is a fundamental limitation hiding in plain sight: self-attention has no notion of order. It treats the input sequence as an unordered bag of tokens, completely ignoring the positions where those tokens appear.
Think of it this way. Imagine you receive a bag of Scrabble tiles spelling out a message. You can see every letter and every pair of letters, but the bag has no concept of left-to-right sequence. You know which letters co-occur with which other letters, but you cannot reconstruct the original order from that information alone. Self-attention faces the same problem. It sees every pair of tokens and computes their compatibility, but the positions those tokens occupy leave no trace in the final representation.
This is not a minor implementation detail: it cuts to the heart of language understanding. Language is not a bag of words. The sentence "Dog bites man" and the sentence "Man bites dog" share exactly the same vocabulary, the same word co-occurrences, and the same pairwise relationships between tokens. Yet they describe opposite events. If a model cannot distinguish word order, it cannot distinguish these sentences. And if it cannot distinguish these sentences, it cannot reason about who is doing what to whom, which is arguably the most basic task in natural language understanding.
The position problem pervades nearly every aspect of language. Subject-verb agreement, pronoun resolution, negation scope, modifier attachment, syntactic parsing, temporal ordering of events, causal relationships between clauses: all of these linguistic phenomena depend critically on where words appear in a sentence, not merely on which words appear. A language model that ignores position is like a reader who sees all the words on a page at once but has no concept of reading left to right. The words are all there, but the structure that gives them meaning is lost.
This chapter explores why position blindness emerges from the mathematics of attention, why word order matters so critically for language understanding across multiple dimensions, and what properties any solution must possess to restore positional awareness to the model. We will set the stage for the various positional encoding schemes that subsequent chapters cover in depth. By the end, you will understand that position matters, why the attention mechanism discards it, and what the design criteria are for getting it back.
This chapter opens Part 4, which covers positional encoding. We assume familiarity with token embeddings (Part 1), the attention mechanism (Part 2), and the transformer architecture overview (Part 3). The chapters that follow this one present specific positional encoding solutions: sinusoidal encoding, learned embeddings, relative encodings, and RoPE. This chapter focuses on diagnosing the problem rather than solving it.
Attention Is Permutation Equivariant
To understand why self-attention ignores position, we need to examine its mathematical structure with care. The conclusion that self-attention is position-blind is not intuitive at first, because the computation does produce one output vector per position. It looks like position is being tracked. But appearances are deceiving. Let's build up the key insight step by step, starting with precise definitions and culminating in a demonstration that reveals where position information is missing.
Permutation Invariance vs. Equivariance
Two related properties appear throughout machine learning, and distinguishing them clarifies exactly what kind of symmetry self-attention has.
A function is permutation invariant if reordering its inputs leaves the output completely unchanged. Computing the mean of a set of numbers is permutation invariant: the mean of is the same as the mean of . Set membership and operations such as summation or taking a maximum have the same property. Permutation invariant functions discard order entirely.
A function is permutation equivariant if reordering the inputs causes the outputs to be reordered in exactly the same way. If you feed a permuted input, you get back the same outputs you would have gotten from the original input, just rearranged by the same permutation. Formally, a function is permutation equivariant if for any permutation :
where:
- : the input sequence (a matrix where each row is a token's embedding)
- : any permutation of the row indices
- : the output sequence produced by the function
- : that same output, with its rows rearranged by
Self-attention falls squarely into the equivariant category, not the invariant category.
A function is permutation equivariant if for any permutation applied to the input sequence, the output is permuted in the same way: . Self-attention has this property because it processes all positions using the same computation with no position-specific parameters or inputs. Every position is treated identically by the machinery of attention.
Why does equivariance matter more than invariance here? Invariance would be even more extreme: the output itself would not change when we reorder inputs. Self-attention at least tracks which output corresponds to which input, which is why we describe it as equivariant rather than invariant. But equivariance is still problematic for language, because it means the model cannot tell from the outputs alone whether the input was "dog bites man" or "man bites dog." The relative ordering between positions carries no information. Two sequences that are permutations of each other produce outputs that are the same permutation of each other.
The key insight is this: permutation equivariance means that reordering the input has exactly the same effect as reordering the output. The attention mechanism does not encode any relationship between input positions that survives reordering. It is as if the attention computation is computed in a "coordinate-free" space where positions have no identity.
Tracing the Mathematics
To see why permutation equivariance emerges, we need to trace through the self-attention computation and identify every place where position could enter, but does not.
Self-attention produces an output vector for each position by aggregating information from all positions. For position , we compute a weighted sum where each position contributes its value vector, scaled by how relevant it is to position . This is the fundamental operation of attention:
where:
- : the output representation for position
- : the attention weight from position to position (determines how much contributes to 's output)
- : the value vector at position (the information that position contributes)
- : the sequence length
- : summation over all positions in the sequence
Notice that the formula is a weighted sum: a linear combination of value vectors. The weight controls how much of position 's value flows into position 's output. If is large, position has a strong influence on position 's representation. If is near zero, position contributes almost nothing.
This formula shows that the output at position depends on all value vectors and all attention weights . But notice: the values come from the token content, and the weights come from content comparisons. Where does position enter?
The attention weights themselves come from applying softmax to scaled dot products between query and key vectors. The softmax function converts raw similarity scores into a probability distribution. This keeps all weights are positive and sum to 1:
where:
- : the attention weight from position to position (how much position attends to position )
- : the query vector for position
- : the key vector for position
- : the dimension of query and key vectors (used for scaling)
- : the exponential function. This keeps all values are positive
- : summation over all positions, serving as the normalizing constant
The queries and keys are computed by multiplying the input embedding by learned weight matrices. For position with input embedding :
where , , and are the learned projection matrices shared across all positions. The same matrices are applied to every token, regardless of where it appears in the sequence.
The Missing Position Signal
Here is the critical observation: examine both equations carefully and notice what is absent. The indices and appear only as subscripts identifying which vectors to retrieve. They never appear as values in the computation itself. The attention weight depends entirely on:
- The query vector , which is derived from the content at position (the token embedding )
- The key vector , which is derived from the content at position (the token embedding )
- The normalization over all keys, which also depends only on content
Position indices serve as labels that tell us which output corresponds to which input. They are bookkeeping, not input features. The number or itself never enters any arithmetic operation. If we swap the tokens at positions 2 and 5, the query and key vectors simply swap, and the computation produces swapped outputs. The absolute values 2 and 5 played no role.
Think of it as the difference between a library catalog (where shelf number matters for retrieval) and a set of books on a table (where the arrangement of books has no permanent identity). Self-attention is the table: books can be rearranged and the pairwise relationships remain identical, just reordered.
This contrasts sharply with recurrent networks. In an RNN or LSTM, position is implicitly encoded through sequential processing. The hidden state at step 5 depends on the hidden state at step 4, which depends on step 3, and so on. The temporal chain of computation bakes position into the representation as an inductive bias: earlier tokens shape later representations in ways that later tokens cannot shape earlier ones. Self-attention discards this sequential chain entirely, processing all positions in parallel and achieving massive efficiency gains, but at the cost of positional awareness.
Demonstrating Permutation Equivariance
The mathematical argument is clear, but seeing the effect in code makes it concrete. Let's implement self-attention and verify that permuting the input produces identically permuted output.
import numpy as np
np.random.seed(42)
def softmax(x, axis=-1):
"""Compute softmax along the specified axis."""
exp_x = np.exp(x - np.max(x, axis=axis, keepdims=True))
return exp_x / np.sum(exp_x, axis=axis, keepdims=True)
def self_attention(X, W_Q, W_K, W_V):
"""
Compute self-attention output.
Args:
X: Input embeddings of shape (seq_len, embed_dim)
W_Q, W_K, W_V: Projection matrices
Returns:
Output representations of shape (seq_len, d_v)
"""
Q = X @ W_Q
K = X @ W_K
V = X @ W_V
d_k = K.shape[1]
scores = Q @ K.T / np.sqrt(d_k)
weights = softmax(scores, axis=-1)
output = weights @ V
return output, weightsWe will create a simple 4-token sequence, apply a permutation that swaps two positions, and compare the outputs.
# Create embeddings for a 4-token sequence
seq_len = 4
embed_dim = 8
d_k = d_v = 6
# Random embeddings representing different tokens
X = np.random.randn(seq_len, embed_dim)
# Initialize projection matrices
W_Q = np.random.randn(embed_dim, d_k) * 0.1
W_K = np.random.randn(embed_dim, d_k) * 0.1
W_V = np.random.randn(embed_dim, d_v) * 0.1
# Compute attention on original sequence
output_original, weights_original = self_attention(X, W_Q, W_K, W_V)
# Create a permutation: swap positions 1 and 2
permutation = [0, 2, 1, 3]
X_permuted = X[permutation]
# Compute attention on permuted sequence
output_permuted, weights_permuted = self_attention(X_permuted, W_Q, W_K, W_V)
# Apply the same permutation to the original output for comparison
output_original_reordered = output_original[permutation]Permutation applied: [0, 2, 1, 3] (swap positions 1 and 2) Original output shape: (4, 6) Permuted output shape: (4, 6) Max difference between pi(output) and output(pi(input)): 2.78e-17 Outputs are identical (within numerical precision)
The outputs match exactly, up to floating-point precision. This confirms permutation equivariance: permuting the input permutes the output in the same way. Self-attention does not "know" that we reordered the tokens. It computes the same pairwise relationships in a different order and produces the same outputs in a different order. No information about the original positions has been preserved beyond the assignment of outputs to input slots.
The Root Cause: Content-Only Dependencies
This experiment confirms what the mathematics predicted. Self-attention treats the input as an unordered set of tokens, using only their content to determine relationships. Each token's output depends on three things, none of which involve position:
- Its own content (through its query vector, which asks "what am I looking for?")
- The content of all other tokens (through their key vectors, which answer "what do I offer?")
- The pairwise content relationships (through dot products that measure semantic compatibility)
This is a deliberate architectural choice, not an oversight. The original Transformer designers wanted to avoid the sequential bottleneck of RNNs, where position is encoded implicitly by the order of computation. Parallelizing across all positions means forgoing positional inductive bias. The model is given maximum flexibility to learn relationships from content alone, but that flexibility comes with the responsibility of explicitly injecting position information from outside the attention mechanism.
Notice that this is quite different from convolutional neural networks (CNNs), which have a different kind of positional inductive bias. A CNN's filters are applied at specific spatial locations, and nearby positions share features through local connectivity. Position is implicitly encoded through the spatial structure of the convolution. Self-attention has no analogous spatial structure. It is maximally flexible with respect to distance.
Why Position Matters for Language
This position blindness might seem like a minor technical detail, but it is catastrophic for language understanding. Human language is deeply structured around word order, and this structure operates at multiple levels simultaneously. Let's examine why position is essential across several dimensions of linguistic meaning.
Grammatical Roles and Argument Structure
In English and many other languages, word order is the primary mechanism for encoding grammatical roles. Who is the agent (the one doing) and who is the patient (the one being done to) depends critically on position relative to the verb.
Consider these sentences:
- "The dog chased the cat."
- "The cat chased the dog."
The same words appear in both sentences, but they play opposite grammatical roles. In the first sentence, "dog" is the subject (the chaser) and "cat" is the object (the chased). In the second, these roles reverse entirely. Pure self-attention, seeing only that "dog," "chased," and "cat" co-occur, cannot distinguish which is the agent and which is the patient. It knows these three words appear together. It does not know who is biting whom.
This matters enormously for any downstream task. Sentiment analysis, information extraction, question answering, and reading comprehension all require knowing who does what to whom. A model that cannot distinguish subject from object cannot extract the event structure of a sentence. And event structure, it turns out, is most of what sentences express.
The technical term for this positional encoding of grammatical roles is argument structure. Verbs take arguments, and those arguments have specific semantic roles called thematic roles or theta roles. These include agent and patient. Other examples include beneficiary, instrument, or location. In English, the mapping from syntactic position to thematic role is largely positional. The entity before the verb tends to be the agent; the entity after the verb tends to be the patient. Without positional information, the model has no way to infer these roles.
Negation Scope
Word order determines what negation applies to, which can completely reverse the meaning of a sentence:
- "I did not say he stole the money." (I said nothing; the accusation is false or came from someone else)
- "I did say he did not steal the money." (I spoke up; I claimed his innocence)
The positioning of "not" relative to "say" and "steal" completely changes the assertion. Moving a single word maps an accusation of silence into a defense of character. A bag-of-words model sees "not," "say," and "steal" together in both sentences and cannot determine which action is negated.
This extends to other scope-sensitive phenomena. Quantifiers and modal verbs interact with scope in ways that depend on linear order; adverbs and conditionals do too. "Every student read some book" has a different meaning from "Some book was read by every student," and this meaning difference is entirely carried by word order.
Modifier Attachment and Structural Ambiguity
Consider the well-known ambiguity in:
- "I saw the man with the telescope."
Does "with the telescope" modify "saw" (I used a telescope to see) or "man" (the man was carrying a telescope)? Both interpretations are syntactically and semantically plausible. The resolution depends on which positional interpretation the reader adopts. More precisely, it depends on which element in the sentence the prepositional phrase is attached to structurally, and that attachment is partially determined by proximity and order.
A second classic example:
- "The horse raced past the barn fell."
This garden-path sentence initially appears to have "raced past the barn" as the main clause, but the final word "fell" reveals a different structure: "The horse [that was] raced past the barn fell." Readers process this sentence positionally, building a parse tree as they read left to right. Self-attention without positional information cannot capture this left-to-right progressive parsing that human language processing depends on.
Temporal and Causal Ordering
Sequence often implies temporal or causal relationships:
- "She got her degree, then found a job."
- "She found a job, then got her degree."
The order of clauses suggests different life trajectories. Both sentences contain the same events, but their relative positions encode temporal precedence. In narrative text, the order in which events are mentioned generally corresponds to their order in time (the "narrative sequence" principle). A model that cannot track positional order cannot extract the timeline of a story.
Causality also frequently depends on sequential position. "The bridge collapsed because the engineers cut corners" has a different causal structure from "The engineers cut corners because the bridge collapsed," even though both sentences contain the same propositions. Without position, the model cannot determine which event caused which.
Semantic Composition and Compound Formation
Languages compose meaning hierarchically, and order often determines the composition structure:
- "Hot dog" (a type of food, very different from a warm canine)
- "Dog hot" (not standard English, but in languages with post-nominal adjectives, this would mean "a dog that is warm")
English compounds typically have the modifier before the head noun. Without positional information, "hot" and "dog" are just two tokens that appear together, with no way to determine which modifies which, or even whether this is a compound or a two-word phrase.
At the phrasal level, the same words in different orders can produce entirely different semantic compositions:
- "Only she spoke to him." (no one else spoke to him)
- "She only spoke to him." (she spoke but did nothing else)
- "She spoke only to him." (she excluded everyone else as a listener)
The adverb "only" is a classic focus particle that must be positioned immediately before the element it focuses on. Moving it changes the scope of exclusion and therefore the meaning of the sentence. Self-attention without positional information cannot recover which element "only" is focusing on.
# Simulate the attention pattern for "dog bites man" vs "man bites dog"
# We'll use random embeddings but give them distinct identities
np.random.seed(123)
# Create distinct embeddings for each word
word_embeddings = {
"dog": np.random.randn(embed_dim),
"bites": np.random.randn(embed_dim),
"man": np.random.randn(embed_dim),
}
# Construct the two sequences
sentence1 = np.stack(
[word_embeddings["dog"], word_embeddings["bites"], word_embeddings["man"]]
)
sentence2 = np.stack(
[word_embeddings["man"], word_embeddings["bites"], word_embeddings["dog"]]
)
# Compute attention weights for both
_, weights1 = self_attention(sentence1, W_Q, W_K, W_V)
_, weights2 = self_attention(sentence2, W_Q, W_K, W_V)

--- Key Observation --- Attention from 'dog' to 'bites': 0.3439 (sentence 1, dog at position 0) Attention from 'dog' to 'bites': 0.3439 (sentence 2, dog at position 2) These are identical because attention only depends on content, not position!
The heatmaps reveal the core problem vividly. Despite "dog bites man" and "man bites dog" having opposite meanings (opposite agent-patient roles), the attention weight between "dog" and "bites" is identical in both sentences. Self-attention cannot tell that in one sentence the dog is the biter, while in the other the dog is the bitten. The pairwise content relationships are the same; only the positions differ. The attention mechanism computes rich contextual representations, but all of that context is semantically based, with no positional signal anywhere in the computation.
A Worked Example: Positional Ambiguity at Scale
To make the problem even more concrete, consider a longer passage:
"The committee voted to ban the organization, and the organization appealed the decision."
A model with no positional encoding sees: {committee, voted, ban, organization, organization, appealed, decision}. It cannot tell that the first "organization" is the object of "ban" and the second is the subject of "appealed." In a bag-of-words view, these two occurrences of "organization" are identical tokens. Only their positions differentiate their grammatical roles and their referential relationship to the surrounding verbs.
Now consider coreference resolution, which is the task of identifying which noun phrases refer to the same entity. In this sentence, does "organization" in the second clause refer to the same organization as in the first? Almost certainly yes, but arriving at that conclusion requires tracking that "the organization" appears first as a patient (being banned) and then as an agent (doing the appealing). Positional information is essential for this tracking.
This example is not exotic. Most sentences of any complexity require positional reasoning. The longer and more syntactically complex a sentence becomes, the more critical it is for the model to track where each token appears relative to the governing verbs and the surrounding clause structure and connectives.
What We Need from Positional Information
Understanding why position matters helps us identify what properties any positional encoding scheme must have. The ideal solution must satisfy several requirements, and each requirement rules out simple approaches you might first consider.
Unique Position Identification
Each position in a sequence needs a distinct representation. If positions 3 and 7 have identical positional signals, the model cannot distinguish tokens at those positions based on their locations. This requirement seems obvious, but it rules out the simplest possible approaches. Using a constant vector (the same positional signal everywhere) provides no information. Using a random vector assigned at initialization but not varied across positions gives each position a signal, but that signal is arbitrary and meaningless to the model.
The uniqueness requirement means the positional representation must vary systematically with position. The question is how.
Bounded and Stable Values
Whatever numerical representation we use for position, it must have bounded, well-behaved values throughout. If position 1 is encoded as 1, position 1000 as 1000, and position 1 million as 1 million, the scale differences will overwhelm the semantic content of token embeddings. Neural networks are sensitive to the magnitudes of their inputs. Unbounded positional representations would create training instabilities and force the model to learn very different weight scales to handle early versus late positions in a sequence.
The key insight is that the positional representation should live in the same numerical space as the token embeddings it is being combined with. Token embeddings are typically distributed with mean near zero and standard deviation near one, thanks to embedding initialization schemes and layer normalization. Positional representations should match this scale.
Consistency Across Sequence Lengths
The encoding for position 5 should be the same whether it appears in a 10-token sequence or a 10,000-token sequence. If positional representations change based on sequence length, the model must learn different positional patterns for every possible length it might encounter. This would make it impossible to transfer knowledge from short sequences to long ones, and vice versa.
This requirement rules out a common naive approach: normalizing the position by the sequence length. Position in a sequence of length might be represented as , giving values between 0 and 1. This is bounded, which is good. But it violates consistency: position 5 in a 10-token sequence gets value 0.5, while position 5 in a 100-token sequence gets value 0.05. The model learns that "0.5" means "halfway through," but the absolute meaning of position 5 changes with sequence length.
Relative Position Accessibility
Many linguistic relationships depend on relative rather than absolute position. A verb typically appears near its subject and object, regardless of where they fall in the sentence. An adjective typically immediately precedes the noun it modifies. A relative clause immediately follows the noun it modifies. These are all relative positional relationships. The model needs to compute "token A is 3 positions before token B," rather than only "token A is at position 7 and token B is at position 10."
Ideally, the positional encoding should make it easy to determine that position 7 is "2 steps after" position 5, rather than only showing that both are somewhere in the sequence. This is a stronger requirement: each position needs a unique signal, and the relationship between signals needs to encode relative distance in a way the model can exploit.
Generalization Beyond Training
Models trained on sequences of length 512 may encounter sequences of length 1024 at inference time. The positional encoding should degrade gracefully (or not at all) when extrapolating to positions not seen during training. This requirement is especially important for applications where input length is unpredictable, such as processing documents of arbitrary length.
Fixed mathematical functions (like sinusoids) can extrapolate naturally, because the formula applies to any position, including ones never seen during training. Learned embeddings, by contrast, have no entry in their lookup table for positions beyond the training maximum. Handling this gracefully requires interpolation, extrapolation heuristics, or other architectural choices.
Compatibility with Attention
The positional information must integrate with the attention mechanism in a way that allows position to influence attention patterns. If a query cannot distinguish nearby keys from distant keys, positional information is not flowing into the core computation. This requirement means we cannot just append position indices somewhere; the position information must be in a form that the dot-product attention can use.
The standard approach is to combine positional representations with token embeddings before computing QKV representations. This way, when the model computes dot products between queries and keys, those dot products depend on position as well as content.
# Illustrate the problem with naive positional representations
positions_naive = np.arange(1, 513)
# Naive approach 1: Raw position index
naive_index = positions_naive
# Naive approach 2: Normalized position (0 to 1)
naive_normalized = positions_naive / 512
# Naive approach 3: Log-scaled position
naive_log = np.log(positions_naive)


None of these naive approaches satisfy our requirements. Raw indices are unbounded, growing to hundreds or thousands while token embeddings remain in a range around zero. Normalized positions change meaning with sequence length: position 0.5 is location 256 in a 512-token sequence but location 512 in a 1024-token sequence, making the learned positional patterns non-transferable. Log scaling compresses distant positions so tightly that the model can barely distinguish position 400 from position 500, creating highly unequal positional resolution across the sequence.
The requirements we have identified describe a multi-dimensional representation with bounded values, stable meaning across sequence lengths, and enough variation to distinguish positions. This is the problem that sinusoidal encodings and learned embeddings were designed to solve, and we will explore both in detail in subsequent chapters.
Position Encoding vs. Position Embedding
Before diving into specific solutions, let us clarify an important terminological distinction that causes persistent confusion in the literature.
Positional encoding refers to a fixed, deterministic function that maps position indices to vectors. The function is defined by a mathematical formula, not by training. Positional embedding refers to learned vectors stored in a lookup table, one per position. Both inject positional information into the model, but through fundamentally different mechanisms with different trade-offs.
Positional encoding uses a predetermined formula to compute a position vector. Given position , the encoding function returns a fixed vector that never changes during training. No gradient flows through the positional encoding itself; it is a fixed input to the model. The classic example is the sinusoidal encoding from the original Transformer paper, which uses sine and cosine functions at different frequencies to construct a unique vector for each position. The primary advantage is generalization: because the formula is defined for all integers, it applies to any position, including those longer than anything seen during training. The model can, in principle, handle sequences of arbitrary length. The drawback is that the encoding is fixed and may not capture the positional patterns most useful for a specific task.
Positional embedding treats positions like vocabulary items. We create an embedding matrix , where is the maximum sequence length and is the embedding dimension. Position looks up row of this matrix, retrieving a -dimensional learned vector . These embeddings are initialized randomly and then updated through training by gradient descent, just like word embeddings. The advantage is flexibility: the model can discover whatever positional representations are useful given the training data and task. The disadvantages are that positions beyond have no representation, and the embeddings must be trained from scratch, requiring sufficient examples of each position to learn useful representations.
Most modern language models use positional embeddings (learned lookup tables) for short-to-medium contexts. Models designed for longer contexts or better generalization tend to use more sophisticated positional schemes like RoPE or ALiBi. The terminology gets mixed in the literature, with both approaches sometimes loosely called "positional encoding," so it is worth being precise about the distinction.
How Position Information Enters the Model
Now that we understand the two mechanisms for generating positional vectors, the remaining question is: how do we inject this information into the attention computation?
Recall from our earlier analysis that attention weights depend only on query and key vectors, which are projections of the input embeddings. To make attention position-aware, we need to modify what goes into those projections. There are several conceptual approaches:
- Input modification: Change the token embeddings before they enter the model, so that the Q/K/V projections see both content and position
- Attention score modification: Directly add a positional bias to the attention scores after the Q/K dot product but before softmax
- Relative encoding in projections: Modify the key vectors to include relative position terms
The standard approach in the original Transformer (and in most BERT-style models) is option 1: add the positional vector directly to the token embedding, creating a single combined representation that carries both semantic meaning and positional information:
where:
- : the token embedding at position (carries semantic content)
- : the positional vector for position (carries positional information)
- : the combined representation fed to the self-attention layers
- : the embedding dimension (must match for both token and position vectors, enabling element-wise addition)
Why addition rather than concatenation or multiplication? Addition is the simplest operation that preserves dimensionality. If we concatenated positional vectors, the combined representation would have dimension , requiring all downstream weight matrices to be twice as wide. Addition keeps everything at dimension . More subtly, addition allows the model to learn how much weight to give positional versus semantic information through the QKV projection matrices. If the model finds certain dimensions of the positional encoding uninformative for a task, the learned projection matrices can simply down-weight those dimensions. The addition operation is also differentiable, so gradients flow through it cleanly during training.
In practice, the positional vectors are often scaled to have a similar magnitude to the token embeddings, to prevent either one from dominating the combined representation. This is particularly important with sinusoidal encodings, which we will examine in the next chapter.
def add_positional_info(token_embeddings, position_vectors):
"""
Combine token embeddings with positional information.
Args:
token_embeddings: Shape (seq_len, embed_dim)
position_vectors: Shape (seq_len, embed_dim)
Returns:
Combined representations of shape (seq_len, embed_dim)
"""
return token_embeddings + position_vectors
# Example: simple positional vectors (not a real encoding scheme)
seq_len_ex = 4
embed_dim_ex = 8
# Token embeddings
tokens_ex = np.random.randn(seq_len_ex, embed_dim_ex)
# Placeholder positional vectors (small scale to illustrate the combination)
positions_ex = np.random.randn(seq_len_ex, embed_dim_ex) * 0.1
# Combined representation
combined_ex = add_positional_info(tokens_ex, positions_ex)Shapes: Token embeddings: (4, 8) Position vectors: (4, 8) Combined: (4, 8) Example: Position 0 Token: [[-1.254 -0.638 0.907 -1.429]...] Position: [[0.089 0.175 0.15 0.107]...] Combined: [[-1.165 -0.462 1.057 -1.322]...]
The shapes confirm that adding positional vectors preserves the embedding dimension. Looking at position 0, you can see how each dimension of the combined vector is simply the sum of the corresponding token and position dimensions. The positional vectors are scaled by 0.1 here to keep them relatively small compared to the token embeddings. This is a common practice that prevents positional information from dominating semantic content at the start of training. Over time, the model can adjust the influence of positional information through its learned projections.
The combined representation now carries both semantic content (from the token embedding) and positional context (from the positional vector). When this combined representation is projected into QKV representations, both types of information influence the attention pattern. A token's query vector asks "what am I looking for, given my content and position?" A token's key vector offers "what do I provide, given my content and position?" The dot product between these position-aware queries and keys produces attention weights that reflect both semantic relevance and positional relationships.
The Alternative: Modifying Attention Scores Directly
Some approaches, including ALiBi (which we will cover in a later chapter), bypass input modification entirely and instead add position-dependent biases directly to the attention scores:
where is a bias term that depends on the positions and (typically just their distance ). This approach has the advantage that the position signal never gets mixed with the semantic content in the embeddings. Semantic relationships are computed purely from content, and then positional biases are applied afterward to adjust attention patterns. This cleaner separation can improve extrapolation: the semantic attention mechanism is unchanged, and positional biases are applied on top.
Understanding that there are multiple points of entry for positional information will be important as we survey positional encoding methods in subsequent chapters. Different methods inject position at different stages of the computation, leading to different trade-offs.
Absolute vs. Relative Position
A final conceptual distinction shapes the entire design space for positional representations: should we encode where each token is in absolute terms, or should we encode how tokens relate to each other in relative terms? This choice is not merely technical; it reflects a hypothesis about what kinds of positional patterns matter most for language.
Absolute positional encoding assigns each position a fixed vector. Position 0 always gets the same encoding, position 1 always gets the same encoding, and so on, regardless of the sequence. The attention mechanism receives absolute location signals and must learn to compute relative positional information by comparing absolute positions.
Relative positional encoding directly encodes the distance or relationship between pairs of positions. Instead of saying "this token is at position 5," relative encoding says "this token is 3 positions before the query token." The attention mechanism directly incorporates these relative distances, without needing to infer them from absolute locations.


Each approach has significant trade-offs. Absolute encoding is simpler to implement: just add a position vector to each token embedding before feeding it to the model. The attention mechanism then sees combined content-position vectors and can, with sufficient capacity, learn to extract relative positional information by computing differences between absolute positions. The original Transformer and later models such as BERT and GPT all use absolute positional encoding.
The limitation of absolute encoding is that it requires the model to learn the mapping from absolute to relative positions implicitly. If the model sees "position 3 attends to position 1" during training, it must generalize to "position 103 attends to position 101" at inference. These are the same relative offset (+2) expressed as different absolute positions. With enough data and capacity, models can learn this generalization, but it is not guaranteed, and it may require more training examples than relative encoding.
Relative encoding directly captures these distance-based relationships. Instead of embedding absolute positions, the model embeds the distances between query position and key position . This representation is translation-invariant by construction: a distance of +2 means the same thing regardless of whether we are at positions 1 and 3 or positions 101 and 103. The model does not need to generalize across absolute positions.
The drawback of relative encoding is architectural complexity. Absolute positional vectors can be added to token embeddings once at the input layer, and the rest of the model is unchanged. Relative positions depend on pairs of positions, not individual positions, so they cannot simply be added to the input. The attention score between positions and must explicitly incorporate information about the distance . This requires modifying the attention computation itself, rather than changing only the input preprocessing. The result is a more expressive model that is also harder to implement efficiently.
Modern architectures increasingly favor relative or hybrid approaches. RoPE (Rotary Position Embedding) cleverly encodes relative positions through rotation operations that integrate naturally with the standard attention dot product. ALiBi (Attention with Linear Biases) adds a simple learned bias to attention scores based on distance, achieving relative encoding with minimal architectural change. We will explore both approaches in dedicated chapters.
Why Relative Positions Are Linguistically Natural
The preference for relative encoding in recent architectures reflects a linguistic observation about local relationships. Most syntactic and semantic relationships in language are local and relative, not global and absolute. A subject typically appears within a few words of its verb, and this proximity holds whether the subject-verb pair appears at the beginning of a sentence or in the middle of a long document. A determiner ("the," "a") always immediately precedes the noun phrase it introduces, and this is a relative relationship independent of absolute position. Anaphoric pronouns refer back to noun phrases that appeared somewhere earlier in the discourse, a backward-pointing relative relationship.
If positional representations encode relative distances well, the model can more easily learn these universal syntactic generalizations. If they encode only absolute positions, the model must learn to apply each generalization at every possible absolute location, multiplying the number of patterns it needs to internalize.
This does not mean absolute encoding is useless. Document-level reasoning, reference tracking across very long spans, and certain ordering tasks may benefit from knowing absolute position. A model that knows it is processing the first sentence of a document versus the twentieth sentence might behave differently. Hybrid approaches that encode both absolute and relative information are an active area of research.
A Formal Look at What Breaks Without Position
It is worth being precise about the consequences of permutation equivariance for specific NLP tasks. Different tasks suffer from position blindness in different ways.
Classification tasks are affected but sometimes less severely. If a task asks "does this review express positive or negative sentiment?", the answer often depends more on the presence of positive/negative words than on their order. "Terrible, not at all good" and "good, not at all terrible" have similar bag-of-words representations but different meanings. However, for most documents, sentiment is carried reasonably well by vocabulary, and classification models can sometimes compensate. The failure mode is subtle and appears when negation or ordering changes sentiment.
Sequence labeling tasks are severely affected. Named entity recognition, part-of-speech tagging, and syntactic parsing all require assigning labels to specific positions in the sequence. If the model cannot determine that "bank" is at position 3 and modifies "river" at position 4, it cannot correctly determine that "bank" is a noun functioning as part of a compound rather than as a financial institution. Positional relationships among adjacent tokens are essential for sequence labeling.
Generation tasks are catastrophically affected. Language modeling requires predicting the next token given the previous ones, and causal masking means only earlier positions are visible. Without positional encoding, the model cannot determine which tokens came earlier and which came later, making causal prediction undefined. The model cannot enforce that the output at position conditions only on positions without knowing what those position indices mean.
Question answering and reading comprehension sit somewhere in between. Finding the answer span in a passage requires understanding that a specific sequence of tokens at a specific location in the passage is the answer. This is heavily positional. But recognizing that the passage contains the relevant information is often more content-based.
The breadth of tasks affected makes positional encoding not an optional enhancement but a fundamental requirement for any transformer-based system doing real natural language processing.
Limitations of Any Positional Encoding Approach
No positional encoding scheme is perfect. Every approach makes trade-offs that are worth understanding before choosing a scheme for a given application.
Fixed encodings may not capture learned patterns. Sinusoidal encoding uses mathematically determined frequencies. These frequencies produce a particular kind of positional representation that may or may not align with what is most useful for a specific task. A model with learned embeddings can discover whatever positional representations the training data requires. The sinusoidal encoding makes an implicit assumption that periodic patterns across dimensions are the right way to represent position, but this assumption has no linguistic grounding.
Learned embeddings do not generalize to longer sequences. If you train with a maximum sequence length of 512, positions 513 and beyond have no learned representation. Some models handle this with interpolation (linearly interpolating between nearby learned embeddings) or by fine-tuning on longer sequences before deployment. But performance often degrades for positions not seen during training, sometimes dramatically. This is a fundamental limitation that requires either architectural solutions (relative encoding) or specific fine-tuning protocols.
Additive combination limits expressiveness. Adding positional vectors to token embeddings means the combined representation must encode both content and position in a shared -dimensional space. The model may struggle when positional and semantic information conflict. For example, if a word's semantic embedding happens to be similar to the positional encoding of position 5, the model may confuse that word's content with a positional signal. More fundamentally, addition forces the model to represent semantics and position in the same space, rather than in separate subspaces that might be more efficiently organized.
Relative positions require architectural changes. Directly encoding relative positions requires modifying the attention computation rather than changing only the input. While libraries like Hugging Face Transformers handle this transparently, implementing relative encoding from scratch is substantially more complex than the absolute approach. This complexity can introduce bugs and makes the model harder to reason about.
Very long sequences remain challenging. Even with sophisticated positional encoding, transformers face fundamental challenges with extremely long sequences. The attention mechanism has quadratic memory and computation complexity with respect to sequence length. Positional patterns learned on sequences of a few thousand tokens may not transfer to sequences of hundreds of thousands of tokens. Recent work on sparse attention, linear attention, and state-space models addresses the quadratic complexity issue, but the positional encoding problem for very long documents is still an active area of research.
Position alone is insufficient for some phenomena. Even with perfect positional encoding, some linguistic phenomena require representations that go beyond simple sequential position. Hierarchical structure (syntactic parse trees), coreference chains, discourse structure, and rhetorical relations are not captured by linear sequential position alone. Position encoding solves the "which slot in the sequence" problem, but does not solve the "what is the syntactic role" problem directly. Downstream learning from the encoded positions is still required.
The position problem is unique to the transformer architecture. Recurrent neural networks (RNNs) and their variants (LSTMs, GRUs) have no analogous problem because sequential processing is their core inductive bias. The hidden state at each step carries information about everything that came before, building up positional awareness implicitly. The Transformer's "Attention Is All You Need" paper (Vaswani et al., 2017) made the deliberate decision to abandon sequential processing entirely in favor of fully parallel computation. This improved training speed and scalability dramatically, but it required a new solution for position. The sinusoidal encoding proposed in that paper was the first widely adopted answer, and its strengths and limitations have led to the wide range of positional encoding approaches used today.
In Practice: What Modern Models Use
Comparing these options helps interpret the design choices in state-of-the-art models. Here is a brief survey of what modern systems use; later chapters examine these methods in detail.
BERT (2018) uses absolute learned positional embeddings with a maximum sequence length of 512. These are trained from random initialization along with the word embeddings. BERT's success demonstrated that learned absolute embeddings work well for most NLP tasks, though they cannot generalize beyond the 512-position limit.
GPT-2 and GPT-3 also use absolute learned positional embeddings, with GPT-2 supporting 1024 tokens and GPT-3 supporting 2048. These are pure lookup tables, and positions beyond the training maximum are unsupported without fine-tuning.
RoBERTa also uses learned absolute positional embeddings (same architecture as BERT). This shows that the training recipe matters more than the positional encoding scheme for many tasks.
T5 uses relative positional encoding through learned scalar biases added to attention scores. Each attention head learns a separate bias for each relative position bucket. This provides relative position awareness without changing the input embedding process.
GPT-NeoX and LLaMA use Rotary Position Embedding (RoPE), which encodes relative positions through rotation matrices applied to query and key vectors. RoPE has become the dominant choice in recent open-source language models due to its strong generalization properties.
Mistral and Mixtral also use RoPE, with additional tricks like sliding window attention to extend effective context length.
The trend in the field is clearly moving toward relative encoding schemes, particularly RoPE, for their superior generalization and extrapolation properties. However, understanding why absolute encoding was the starting point, and what its limitations are, is essential context for understanding why these newer approaches were developed.
Summary
Self-attention is blind to position because its computation depends only on content, not on where tokens appear in the sequence. This permutation equivariance is a fundamental property of the attention mechanism, arising from the fact that position indices serve only as bookkeeping labels and never appear as values in any arithmetic operation. It is not a bug in any particular implementation; it is built into the mathematical structure of dot-product attention.
Key takeaways from this chapter:
- Permutation equivariance: Self-attention produces the same outputs (reordered) regardless of input order. Shuffling the input shuffles the output identically, because attention weights depend only on query-key compatibility, not on position indices. The equation captures this precisely.
- Language requires order: Grammatical roles, negation scope, modifier attachment, temporal relationships, and semantic composition all depend on word position. "Dog bites man" and "man bites dog" are opposite events expressed by the same vocabulary, and only position distinguishes them.
- Requirements for positional information: Any solution must provide unique position identification, bounded values, consistency across sequence lengths, accessibility of relative positions, generalization beyond training, and compatibility with attention.
- Encoding vs. embedding: Positional encoding uses fixed mathematical formulas (like sinusoids) that generalize to any position. Positional embedding uses learned lookup tables that adapt to training data but have a fixed maximum length. Both have significant trade-offs.
- Absolute vs. relative: Absolute encoding assigns fixed vectors to positions, requiring the model to infer relative distances. Relative encoding directly captures distances between positions, supporting better generalization but requiring architectural modifications.
- Injection mechanism: The standard approach adds positional vectors to token embeddings before computing queries and keys, combining semantic and positional information in a shared representation. Alternative approaches modify attention scores directly.
- Modern trends: The field has moved from absolute learned embeddings (BERT, GPT) toward relative encoding schemes (T5, RoPE). This reflects the linguistic insight that relative distances matter more than absolute positions for most language tasks.
In the next chapter, we will examine the sinusoidal positional encoding introduced in the original Transformer paper. This elegant mathematical construction uses sine and cosine functions at different frequencies to create unique position vectors that support relative position computation through simple linear operations. You will see exactly how the requirements identified in this chapter shaped the design of that solution.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about the position problem in self-attention.
The Position Problem
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!