Part of Language AI Handbook
Explains how RoPE encodes position through vector rotation, making attention scores depend on relative position.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Rotary Position Embedding (RoPE)
The transformer attention mechanism, as we've seen, is inherently position-blind. Sinusoidal encodings and learned position embeddings address this by adding position information to token embeddings before attention. But these approaches encode absolute position. Token at position 5 always receives the same positional signal, regardless of context. What if the relationship between positions 5 and 7 matters more than the absolute locations? Relative position encoding tackles this, but earlier methods required modifying the attention architecture or adding explicit bias terms.
Rotary Position Embedding, or RoPE, takes an elegant geometric approach. Instead of adding position to embeddings, it rotates them. Each position corresponds to a rotation angle, and the rotation is applied directly to query and key vectors. The ingenious part: when you compute the dot product between a rotated query at position and a rotated key at position , the result depends only on their relative distance . Absolute positions vanish, leaving only the relationship between tokens.
Think of it like the hands of a clock. The hour hand and minute hand each occupy some absolute position on the dial, but the angle between them tells you how much time has passed. RoPE works similarly for token positions. It encodes absolute position through rotation, but the inner product that drives attention extracts only the angular difference, the relative separation between tokens.
This design sidesteps a long-standing tension in position encoding. Absolute encodings are simple and cheap, but the model must learn implicitly that positions 3 and 5 are "two apart" the same way positions 97 and 99 are. Learned relative encodings solve that problem directly but require explicit bias matrices that grow with sequence length. RoPE achieves relative awareness as a free consequence of geometry, requiring no extra parameters and no changes to the attention formula itself.
The earliest transformers (Vaswani et al., 2017) used fixed sinusoidal encodings added to token embeddings before the first layer. These encodings were absolute: each position received a unique vector independent of surrounding context. Researchers quickly noticed this limited relative-position generalization, prompting work on relative encodings such as those by Shaw et al. (2018) and the Transformer-XL approach (Dai et al., 2019). Each of these required attention-formula modifications. Su et al. introduced RoPE in 2021, showing that rotation could deliver relative-position awareness without any architectural change. LLaMA and PaLM 2, along with Mistral and Falcon, were among the open-weight models that adopted RoPE soon afterward.
This chapter develops RoPE from first principles. We'll start with rotations in 2D, extend to higher dimensions through paired rotations, derive why the mechanism captures relative position, and implement it in code. By the end, you'll understand why RoPE has become the dominant position encoding in modern language models such as LLaMA and PaLM.
Why Rotation?
Consider what we want from a position encoding. When a query at position attends to a key at position , the attention score should somehow reflect their relative distance . If token 3 attends to token 1, the model should "know" they're 2 positions apart, exactly as if token 8 attends to token 6.
Rotations have a beautiful property that accomplishes this. If you rotate vector by angle and vector by angle , then compute their dot product, the result depends on the angle difference . The absolute angles cancel out.
For two vectors and , rotating both by the same angle preserves their dot product: . This is because rotations preserve lengths and angles between vectors.
The key insight is that this invariance extends to the difference of rotations. If you rotate by and by , the dot product reflects the angular difference , not the individual angles. This is why rotation is a natural fit for relative position encoding: position becomes an angle, and the angle difference between query and key positions is exactly the relative separation we want the model to observe.
Now imagine we associate each position with an angle: position gets angle for some base angle . If we rotate the query vector at position by and the key vector at position by , their dot product will involve the angle . That's exactly the relative position information we want.
Notice that no part of this requires changing the attention formula itself. Standard scaled dot-product attention simply computes . RoPE just ensures that the vectors being dotted already carry the right geometric relationship. The attention mechanism does its job as usual, and position awareness comes for free.
This is the core insight of RoPE: encode position through rotation, and let the geometry of dot products naturally extract relative position.
Rotation Matrices in 2D
To understand why RoPE works, you first need a firm grip on how rotation changes a vector and why the dot product between rotated vectors has a simple form. Let's build up the mechanics from scratch.
In two dimensions, rotating a vector by angle counterclockwise uses the rotation matrix:
where:
- : the 2D rotation matrix that rotates vectors by angle
- : the rotation angle in radians (counterclockwise is positive)
- , : trigonometric functions evaluated at angle
Applying this rotation matrix to a 2D vector transforms its coordinates:
where:
- : the original vector coordinates
- : the rotated vector coordinates
- The first row computes the new -coordinate by combining the original coordinates with and
- The second row computes the new -coordinate using and
The rotated vector has the same length as but points in a direction shifted by . This is because rotation matrices are orthogonal, meaning they preserve vector lengths (norms) and angles between vectors. Orthogonality is the key property that makes everything else work: when you rotate both a query and a key, their magnitudes stay the same, and the only thing that changes in their dot product is the angular relationship between them.
In practice, two properties of orthogonal matrices matter most for RoPE. First, the transpose of a rotation matrix equals its inverse: . This tells you that "undoing" a rotation by angle is the same as transposing the matrix, which will become important when we simplify the dot product expression. Second, multiplying two rotation matrices combines their angles: . This composition rule is what allows absolute angles to cancel and leave only their difference.
def rotate_2d(vector, theta):
"""Rotate a 2D vector by angle theta (in radians)."""
cos_t, sin_t = np.cos(theta), np.sin(theta)
rotation_matrix = np.array([[cos_t, -sin_t], [sin_t, cos_t]])
return rotation_matrix @ vectorLet's visualize how rotation transforms a vector:
# Original vector
original = np.array([1.0, 0.5])
# Rotate by several angles
angles = [0, np.pi / 6, np.pi / 3, np.pi / 2]
rotated_vectors = [rotate_2d(original, theta) for theta in angles]
The dashed circle shows the path traced by the vector tip as it rotates. The length (magnitude) never changes. Rotation is an isometry, a transformation that preserves distances. This matters for attention: the magnitude of a query or key vector carries information about how strongly that token "wants" to attend or "wants" to be attended to. Rotation leaves that strength unchanged while shifting the directional relationship between the two vectors.
From Rotation to Relative Position
Now let's see how rotation makes dot products depend on relative position. Take two 2D vectors (query) and (key). Rotate by angle (for position ) and by angle (for position ).
The dot product of the rotated vectors is:
where:
- : the query vector at position
- : the key vector at position
- : rotation matrix that rotates by angle (position times base angle)
- : rotation matrix that rotates by angle
Using properties of rotation matrices, we can simplify this expression step by step:
The derivation proceeds as follows:
- Rewrite the dot product as matrix multiplication:
- Apply the transpose-inverse property: The transpose of a rotation matrix equals its inverse, so . This gives us .
- Apply the composition property: Multiplying two rotation matrices adds their angles, so . Therefore .
The final result depends only on the difference , not on the absolute values of and separately.
When we rotate query by and key by , their dot product depends only on , the relative position. The absolute positions and vanish, replaced by their difference.
This is remarkable. We encode absolute position through rotation angle, but the attention mechanism, which uses dot products, automatically extracts relative position. No architectural changes needed. No explicit bias terms. Just geometry.
The derivation also reveals something that's easy to overlook: the original vectors and are still present in the final expression . Their content still matters, it's just that the rotation wraps that content in a position-dependent frame. The model can still learn that "verb tokens should attend to their subject nouns"; it just does so in a coordinate system where position is already factored out. This is a cleaner separation of concerns than additive encodings, where position and content signals are entangled from the start.
Let's verify this numerically:
def verify_relative_position():
"""Verify that rotated dot products depend only on relative position."""
# Random query and key vectors
q = np.random.randn(2)
k = np.random.randn(2)
theta = 0.5 # Base rotation angle
# Different absolute positions, same relative distance
results = []
for m, n in [(1, 3), (5, 7), (10, 12), (100, 102)]:
q_rotated = rotate_2d(q, m * theta)
k_rotated = rotate_2d(k, n * theta)
dot_product = np.dot(q_rotated, k_rotated)
results.append((m, n, n - m, dot_product))
return results
position_results = verify_relative_position()Verifying relative position property: Position m Position n n - m Dot Product -------------------------------------------------- 1 3 2 -0.651890 5 7 2 -0.651890 10 12 2 -0.651890 100 102 2 -0.651890
All four pairs have the same relative distance (2 positions apart), and their dot products are identical despite wildly different absolute positions. This confirms the relative position property holds numerically. Notice that positions 1 and 3, 5 and 7, 10 and 12, and 100 and 102 all produce the same score, exactly as the geometric derivation predicted.
Extending to Higher Dimensions
We've established that 2D rotation elegantly encodes relative position. But real transformer embeddings have hundreds or thousands of dimensions, not just 2. How do we extend this geometric insight to high-dimensional space?
The challenge is that rotations in high dimensions are more complex than in 2D. In three dimensions, a rotation requires specifying an axis and an angle. In four dimensions, there are six independent planes of rotation. In general, a -dimensional rotation matrix has degrees of freedom, and computing such rotations efficiently for large is expensive. More importantly, constructing a single rotation that somehow encodes both absolute position and relative-position semantics across all dimensions simultaneously is not straightforward.
A naive approach might try to define a single rotation that affects all dimensions simultaneously, but this would be computationally expensive and wouldn't preserve the relative position property we just derived. The key observation is that we don't need a full -dimensional rotation. We need something weaker: a transformation that maps position information into the dot product while keeping pairs of dimensions independent so the relative-position algebra still works.
RoPE's solution is both clever and efficient: treat the -dimensional embedding as independent pairs. A -dimensional embedding is split into pairs: , , ..., . Each pair is rotated independently as a 2D vector, and since the pairs don't interact, the relative position property holds for each pair separately. When we sum up the contributions from all pairs in a dot product, the overall score still depends only on relative position. The total dot product is the sum of independent 2D dot products, each of which depends only on the relative angle. Summing them preserves that property.
RoPE assigns each pair a different rotation frequency. The first pair might rotate by per position, the second by , the third by , and so on. Think of it like the hour and minute hands, plus the second hand, of a clock: each moves at a different rate, and together they can represent any time uniquely. Using multiple frequencies, RoPE creates an encoding where different dimension pairs capture position information at different scales.
The multi-frequency design solves a representational problem that single-frequency encodings cannot. If all pairs used the same rotation speed, two positions separated by exactly one full revolution would look identical, since . By spreading rotation speeds across many orders of magnitude, RoPE ensures that no two positions within the practical context window produce the same overall signature across all pairs simultaneously. It is the same principle behind mixed-radix number systems: using different place values lets you represent large numbers uniquely with a compact set of digits.
The rotation angle for dimension pair at position is:
where:
- : the rotation angle (in radians) for dimension pair at sequence position
- : the position in the sequence (0, 1, 2, ..., for a sequence of length )
- : the dimension pair index (0, 1, 2, ..., )
- : the total embedding dimension (must be even)
- : the base frequency for dimension pair , which decreases exponentially as increases
- 10000: the base constant (same as in sinusoidal position encodings), chosen empirically for good performance
To understand why this formula creates a multi-scale representation, consider the exponent :
- When : (fastest rotation, one radian per position)
- When : (slower rotation)
- When : (slowest rotation)
This exponential decay means early dimension pairs (small ) rotate quickly, capturing fine-grained position differences, while later dimension pairs (large ) rotate slowly, capturing longer-range relationships.
def compute_rope_frequencies(d_model, base=10000):
"""Compute rotation frequencies for each dimension pair."""
# Number of dimension pairs
num_pairs = d_model // 2
# Frequency for each pair: 1 / base^(2i/d)
i = np.arange(num_pairs)
frequencies = 1.0 / (base ** (2 * i / d_model))
return frequencies
# Example with 8 dimensions (4 pairs)
d_model = 8
freqs = compute_rope_frequencies(d_model)RoPE frequencies for 8-dimensional embeddings: Pair Frequency Wavelength (positions) --------------------------------------------- 0 1.000000 6.3 1 0.100000 62.8 2 0.010000 628.3 3 0.001000 6283.2
The table shows the exponential decay: pair 0 completes a full cycle in about 6 positions (high frequency), while pair 3 takes over 600 positions (low frequency). This 100-times difference in wavelength is what allows RoPE to encode positions at multiple scales simultaneously. In a full-size model with head dimensions and 64 pairs, the wavelengths span from roughly 6 positions all the way to 628,000 positions, far beyond any practical context window. That range guarantees that every position within the window has a unique multi-frequency fingerprint.
Let's visualize this frequency spectrum to see the exponential decay more clearly:
# Visualize frequency decay across dimension pairs
d_model_viz = 64 # Typical small model dimension
freqs_viz = compute_rope_frequencies(d_model_viz)
wavelengths = 2 * np.pi / freqs_viz

The wavelengths tell us how many positions before a dimension pair completes a full rotation (360°). Pair 0 completes a cycle in about 6 positions, while pair 3 takes over 600 positions. This exponential spread ensures RoPE can distinguish positions both locally and globally.
The Complete RoPE Formula
Now that we understand the individual components, let's bring everything together into the complete RoPE transformation. We've established three key ideas:
- Rotation encodes position: Each position corresponds to a rotation angle
- Dot products extract relative position: When query and key are rotated by different amounts, their dot product depends only on the angle difference
- Multiple frequencies create richness: Different dimension pairs rotate at different rates, capturing both local and global position information
The complete RoPE formula combines these insights into a single elegant operation. Given a query or key vector at position , we apply RoPE as follows:
where:
- : the rotated vector, a function of both the input vector and position
- : the input query or key vector with dimensions
- : the sequence position (integer index)
- : the 22 rotation matrix for angle , applied to dimension pair
- : the base frequency for dimension pair (decreases exponentially with )
- The block-diagonal structure means each 22 rotation block operates independently on its corresponding dimension pair
- Empty off-diagonal blocks are zeros, so dimensions in different pairs don't interact during rotation
The large block-diagonal matrix applies different rotations to each dimension pair simultaneously. This is efficient because each 2D rotation is independent of the others, allowing for parallel computation.
Expanding the rotation for a single dimension pair :
where:
- , : the -th and -th components of the input vector (using 1-based indexing)
- , : the corresponding components after rotation
- : the rotation angle, which increases linearly with position at a rate determined by frequency
Written out element-wise, the transformation is:
Complex Number Perspective
The matrix formulation we've developed is mathematically complete, but complex numbers provide another way to express RoPE. This notation reveals the connection between rotations and exponentials and leads to more efficient implementations.
The key insight is that a 2D rotation is equivalent to multiplication by a complex exponential. Every 2D vector can be viewed as a complex number , where is the imaginary unit. In this representation, rotating the vector by angle is simply multiplying by .
Think of it this way: the complex plane is just the 2D plane with an algebraic structure layered on top. The real axis is the -axis and the imaginary axis is the -axis. A complex number is just the point , and multiplying two complex numbers combines their magnitudes and adds their angles. This multiplicative structure is exactly what we need: composing rotations corresponds to multiplying complex exponentials, and the angle-addition law is the algebraic counterpart of the rotation composition rule .
For the dimension pair , interpret them as the real and imaginary parts of a complex number:
where:
- : a complex number representing dimension pair
- : the real part (first element of the pair)
- : the imaginary part (second element of the pair)
- : the imaginary unit (satisfying )
Rotation by angle in the complex plane is achieved by multiplication:
where:
- : the rotated complex number
- : the complex exponential, a point on the unit circle at angle
- : the expanded form via Euler's formula
Euler's formula states that . Geometrically, represents a point on the unit circle at angle from the positive real axis. Multiplying any complex number by rotates it by angle counterclockwise in the complex plane, preserving its magnitude.
For RoPE at position , we apply this rotation with the position-dependent angle:
where:
- : the sequence position
- : the base frequency for dimension pair
- : the total rotation angle (increases linearly with position)
This formulation is mathematically equivalent to the rotation matrix approach. The complex perspective leads to more concise code and can be more efficient on hardware with optimized complex number operations. PyTorch, for instance, provides native complex tensor support with torch.view_as_complex, which allows the rotation to be expressed as a single element-wise multiplication rather than explicit matrix-vector products. This is why most production RoPE implementations use the complex formulation even when the underlying mathematics is easier to explain via rotation matrices.
def rope_complex(x, position, frequencies):
"""Apply RoPE using complex number formulation.
Args:
x: Input vector of shape (d,) where d is even
position: Position index (integer)
frequencies: Pre-computed frequencies of shape (d/2,)
Returns:
Rotated vector of shape (d,)
"""
# Reshape to pairs and view as complex numbers
d = len(x)
x_pairs = x.reshape(-1, 2)
x_complex = x_pairs[:, 0] + 1j * x_pairs[:, 1]
# Rotation angles for this position
angles = position * frequencies
# Apply rotation via complex multiplication
rotated_complex = x_complex * np.exp(1j * angles)
# Convert back to real pairs
rotated_pairs = np.stack(
[rotated_complex.real, rotated_complex.imag], axis=1
)
return rotated_pairs.flatten()Let's verify that the complex formulation gives the same result as explicit rotation matrices:
def rope_matrix(x, position, frequencies):
"""Apply RoPE using explicit rotation matrices."""
d = len(x)
result = np.zeros_like(x)
for i in range(d // 2):
theta = position * frequencies[i]
cos_t, sin_t = np.cos(theta), np.sin(theta)
# Apply 2D rotation to dimension pair (2i, 2i+1)
x1, x2 = x[2 * i], x[2 * i + 1]
result[2 * i] = x1 * cos_t - x2 * sin_t
result[2 * i + 1] = x1 * sin_t + x2 * cos_t
return result
# Compare both implementations
test_vector = np.random.randn(8)
test_freqs = compute_rope_frequencies(8)
position = 5
result_complex = rope_complex(test_vector, position, test_freqs)
result_matrix = rope_matrix(test_vector, position, test_freqs)Comparing RoPE implementations: Original vector: [-0.2342 -0.2341 1.5792 0.7674 -0.4695 0.5426 -0.4634 -0.4657] Complex formulation: [-0.2909 0.1581 1.018 1.4306 -0.496 0.5184 -0.4611 -0.468 ] Matrix formulation: [-0.2909 0.1581 1.018 1.4306 -0.496 0.5184 -0.4611 -0.468 ] Max difference: 5.55e-17
The two implementations produce identical results (up to floating-point precision). The complex formulation is often preferred in practice because it is more concise and can use optimized complex number operations.
Worked Example: Step-by-Step RoPE for a 4-Dimensional Vector
Before we look at batched implementations, it helps to trace through a complete example by hand. This makes the abstract rotation formula concrete and shows exactly what changes when you apply RoPE to a real vector.
Suppose we have a 4-dimensional query vector at position :
With , we have two dimension pairs and two frequencies:
At position , the rotation angles for each pair are:
- Pair 0: radians
- Pair 1: radians
Now apply the 2D rotation formula to each pair separately. For pair 0, with and :
For pair 1, with and (small angle, so and ):
The rotated query is . Notice that the first pair changed dramatically (2 radians is a large rotation), while the second pair barely moved (0.02 radians is about 1 degree). This illustrates the multi-scale nature of RoPE: high-frequency pairs carry a strong position signal that varies quickly from token to token, while low-frequency pairs change slowly and provide a coarser sense of position.
To confirm the key property, you would perform the same calculation for a key vector at position , compute the dot product , and verify that it equals the dot product you would get at any other pair of positions with relative distance . The numbers work out exactly because the composition of rotations always produces the relative angle regardless of where in the sequence you started.
Visualizing RoPE Patterns
With both the matrix and complex formulations implemented, let's visualize how RoPE transforms embeddings. These patterns help explain why RoPE is so effective at encoding position information.
We'll plot the rotation patterns for each dimension pair, tracking how a unit vector moves as position increases:
# Generate RoPE patterns across positions
d_model = 8
max_position = 50
frequencies = compute_rope_frequencies(d_model)
# Track how a unit vector in each pair rotates with position
positions = np.arange(max_position)
rotation_patterns = np.zeros((max_position, d_model // 2, 2))
for pos in positions:
for i in range(d_model // 2):
angle = pos * frequencies[i]
rotation_patterns[pos, i, 0] = np.cos(angle)
rotation_patterns[pos, i, 1] = np.sin(angle)



The color gradient (dark to light) shows increasing position. Pair 0 makes multiple full rotations within 50 positions, while Pair 3 barely completes an arc. This multi-frequency structure is what gives RoPE its expressiveness.
Relative Position Through Dot Products
We've derived mathematically that RoPE should make attention scores depend only on relative position. Now let's verify this core property empirically and see what it looks like in practice.
The test is straightforward: create identical query and key vectors at different positions, apply RoPE, and compute attention scores. If RoPE works as intended, the scores should form a Toeplitz matrix, where each diagonal contains identical values. This structure proves that scores depend only on relative position (the difference between query and key positions), not on absolute positions.
Think of a Toeplitz matrix as a kind of shift-invariant fingerprint. In a standard matrix, the value at row , column can be anything. In a Toeplitz matrix, the value depends only on the offset . If you shift both row and column indices by the same amount, you get the same value. That is exactly what relative-position encoding produces: shifting the entire sequence by a constant does not change any attention score, because all relative distances remain the same.
def compute_rope_attention_scores(queries, keys, frequencies):
"""Compute attention scores with RoPE applied.
Args:
queries: Query vectors, shape (seq_len, d)
keys: Key vectors, shape (seq_len, d)
frequencies: RoPE frequencies, shape (d/2,)
Returns:
Attention scores, shape (seq_len, seq_len)
"""
seq_len, d = queries.shape
scores = np.zeros((seq_len, seq_len))
for m in range(seq_len):
# Rotate query at position m
q_rotated = rope_complex(queries[m], m, frequencies)
for n in range(seq_len):
# Rotate key at position n
k_rotated = rope_complex(keys[n], n, frequencies)
# Compute dot product
scores[m, n] = np.dot(q_rotated, k_rotated)
return scores
# Create test queries and keys
seq_len = 6
d_model = 8
# All queries are identical, all keys are identical
# This isolates the effect of position
q_template = np.random.randn(d_model)
k_template = np.random.randn(d_model)
queries = np.tile(q_template, (seq_len, 1))
keys = np.tile(k_template, (seq_len, 1))
frequencies = compute_rope_frequencies(d_model)
scores = compute_rope_attention_scores(queries, keys, frequencies)
The Toeplitz structure is clear: all entries along each diagonal are identical. Position (0,0), (1,1), (2,2) all have the same score (relative distance 0). Position (0,1), (1,2), (2,3) all match (relative distance 1). This is the relative position property in action.
Let's verify numerically by extracting scores for each relative distance:
# Extract scores by relative distance
relative_scores = {}
for m in range(seq_len):
for n in range(seq_len):
rel_dist = n - m
if rel_dist not in relative_scores:
relative_scores[rel_dist] = []
relative_scores[rel_dist].append(scores[m, n])Scores grouped by relative position (n - m): Relative Position Scores Std Dev --------------------------------------------------------------------------- -5 0.4769 0.00e+00 -4 0.1021, 0.1021 2.78e-16 -3 2.0972, 2.0972, 2.0972 0.00e+00 -2 4.4375, 4.4375, 4.4375, 4.4375 0.00e+00 -1 4.7684, 4.7684, 4.7684, 4.7684, ... 6.88e-16 0 2.5720, 2.5720, 2.5720, 2.5720, ... 4.80e-16 1 -0.3545, -0.3545, -0.3545, -0.3545, ... 3.10e-16 2 -1.5489, -1.5489, -1.5489, -1.5489 3.51e-16 3 -0.1455, -0.1455, -0.1455 1.89e-16 4 2.3313, 2.3313 3.14e-16 5 3.3711 0.00e+00
All standard deviations are effectively zero (within floating-point precision), confirming that scores at each relative distance are identical.
How Dot Products Vary with Relative Distance
The Toeplitz structure tells us scores depend only on relative position, but how do they vary? Let's trace how the dot product changes as we increase the relative distance between query and key:
# Analyze how dot product varies with relative distance
d_model = 32
frequencies = compute_rope_frequencies(d_model)
# Create random query and key vectors
q = np.random.randn(d_model)
k = np.random.randn(d_model)
# Compute dot products at various relative distances
max_rel_dist = 50
relative_distances = np.arange(-max_rel_dist, max_rel_dist + 1)
dot_products = []
# Fix query at position 50 (middle of range)
query_pos = 50
q_rotated = rope_complex(q, query_pos, frequencies)
for rel_dist in relative_distances:
key_pos = query_pos + rel_dist
k_rotated = rope_complex(k, key_pos, frequencies)
dot_products.append(np.dot(q_rotated, k_rotated))
dot_products = np.array(dot_products)
The oscillating pattern is characteristic of RoPE. The multiple frequencies create a complex interference pattern where some relative distances produce higher scores than others. This structure allows the model to learn position-dependent attention patterns during training.
Efficient Implementation
The implementations we've shown so far process one token at a time, which is clear for understanding but inefficient in practice. Modern deep learning frameworks excel at vectorized operations, so we want to apply RoPE to all tokens in a sequence simultaneously.
The key insight is that we can precompute all rotation angles as a matrix and apply them through broadcasting. Instead of looping over positions and dimension pairs, we compute everything in parallel. The rotation angles form an outer product: positions multiplied by frequencies gives a matrix of shape (seq_len, d/2) where entry [m, i] is the rotation angle for position m and dimension pair i. Computing cos and sin of this matrix gives two tensors of the same shape, and the rotation of all tokens can then be expressed as a single vectorized multiply-and-add.
def apply_rope_batch(x, frequencies):
"""Apply RoPE to a batch of vectors at consecutive positions.
Args:
x: Input tensor of shape (seq_len, d)
frequencies: Pre-computed frequencies of shape (d/2,)
Returns:
Rotated tensor of shape (seq_len, d)
"""
seq_len, d = x.shape
positions = np.arange(seq_len)
# Compute all rotation angles: (seq_len, d/2)
angles = np.outer(positions, frequencies)
# Compute cos and sin for all positions and frequencies
cos_angles = np.cos(angles)
sin_angles = np.sin(angles)
# Reshape input to pairs: (seq_len, d/2, 2)
x_pairs = x.reshape(seq_len, -1, 2)
# Apply rotation to each pair
# x_new[0] = x[0] * cos - x[1] * sin
# x_new[1] = x[0] * sin + x[1] * cos
x_rotated = np.stack(
[
x_pairs[:, :, 0] * cos_angles - x_pairs[:, :, 1] * sin_angles,
x_pairs[:, :, 0] * sin_angles + x_pairs[:, :, 1] * cos_angles,
],
axis=-1,
)
return x_rotated.reshape(seq_len, d)Let's verify this batch implementation matches the per-token version:
# Test batch vs individual application
seq_len = 10
d_model = 16
test_embeddings = np.random.randn(seq_len, d_model)
frequencies = compute_rope_frequencies(d_model)
# Batch application
batch_result = apply_rope_batch(test_embeddings, frequencies)
# Individual application
individual_result = np.zeros_like(test_embeddings)
for pos in range(seq_len):
individual_result[pos] = rope_complex(
test_embeddings[pos], pos, frequencies
)Batch vs individual implementation: Maximum difference: 2.22e-16 Implementations match: True
The maximum difference between implementations is on the order of , which is essentially machine epsilon for 64-bit floating point. This confirms that the batch implementation produces numerically identical results while being much more efficient through vectorization.
In a production PyTorch implementation, the efficiency gains go even further. The frequency table is computed once and stored as a buffer (not a parameter), so it is never updated during training and does not consume gradient memory. During inference with key-value caching, the frequencies for already-processed positions can be looked up directly from this precomputed table rather than recomputed. The rotation itself becomes a fused operation in CUDA kernels used by FlashAttention-compatible implementations, reducing memory reads and writes to the minimum possible. The overhead of RoPE in a well-optimized implementation is negligible compared to the QKV projections and the attention matrix itself.
Integration with Self-Attention
With efficient RoPE implementation in hand, let's see how it fits into a complete self-attention layer. The integration is remarkably clean: RoPE slots in between the QKV projections and the attention computation, requiring no architectural changes to the transformer.
Here's the complete flow:
- Project input embeddings to Q, K, V using learned weight matrices
- Apply RoPE to Q and K (but not to V)
- Compute scaled dot-product attention as usual
- Return the attention output
The critical detail is step 2: we rotate queries and keys but leave values untouched. Let's implement this:
class RoPEAttention:
"""Self-attention with Rotary Position Embedding."""
def __init__(self, d_model, d_k, base=10000, seed=None):
"""Initialize RoPE attention layer.
Args:
d_model: Input embedding dimension
d_k: Query/Key/Value dimension (must be even for RoPE)
base: Base for frequency computation
seed: Random seed for weight initialization
"""
if d_k % 2 != 0:
raise ValueError("d_k must be even for RoPE")
if seed is not None:
np.random.seed(seed)
# Projection matrices
scale = np.sqrt(2.0 / (d_model + d_k))
self.W_q = np.random.randn(d_model, d_k) * scale
self.W_k = np.random.randn(d_model, d_k) * scale
self.W_v = np.random.randn(d_model, d_k) * scale
# RoPE frequencies
self.frequencies = compute_rope_frequencies(d_k, base)
self.d_k = d_k
def forward(self, x):
"""Compute RoPE attention.
Args:
x: Input embeddings of shape (seq_len, d_model)
Returns:
output: Attention output of shape (seq_len, d_k)
attention_weights: Weights of shape (seq_len, seq_len)
"""
# Project to Q, K, V
Q = x @ self.W_q
K = x @ self.W_k
V = x @ self.W_v
# Apply RoPE to Q and K
Q_rope = apply_rope_batch(Q, self.frequencies)
K_rope = apply_rope_batch(K, self.frequencies)
# Scaled dot-product attention
scores = Q_rope @ K_rope.T / np.sqrt(self.d_k)
# Softmax
scores_stable = scores - scores.max(axis=1, keepdims=True)
attention_weights = np.exp(scores_stable) / np.exp(scores_stable).sum(
axis=1, keepdims=True
)
# Aggregate values (V is NOT rotated)
output = attention_weights @ V
return output, attention_weightsNote that RoPE is applied only to queries and keys, not to values. This is because:
- Queries and keys determine attention patterns: The dot product between them computes compatibility. RoPE makes this compatibility position-aware.
- Values carry content: They should not be position-encoded because the content itself doesn't depend on position, only how much weight it receives.
The asymmetry matters. Attention works in two stages: first decide how much to attend to each token (queries and keys), then aggregate what content to retrieve (values). Position belongs only to the first stage. A value vector encodes what a token means, and that meaning should not be rotated away just because the token appears at a different position. If you were to rotate values as well, the final attention output would carry a positional artifact mixed into the content representation, complicating what the next layer needs to do. Keeping values position-free separates the two functions: the attention distribution is position-aware, but the content it aggregates is not.
# Test the RoPE attention layer
seq_len = 8
d_model = 16
d_k = 8
x = np.random.randn(seq_len, d_model)
attention = RoPEAttention(d_model, d_k, seed=123)
output, weights = attention.forward(x)RoPE Attention Layer Test: Input shape: (8, 16) Output shape: (8, 8) Weights shape: (8, 8) Row sums (should be 1.0): [1. 1. 1. 1. 1. 1. 1. 1.]
The output confirms correct behavior: input of shape (8, 16) produces output of shape (8, 8) after projection to the query/key/value dimension. The attention weights form an 8×8 matrix where each row sums to exactly 1.0, confirming proper softmax normalization. The RoPE transformations are applied internally to queries and keys, making attention position-aware without changing the external interface.
The integration is truly a drop-in modification. An existing transformer implementation that computes Q, K, V = linear(x) followed by scores = Q @ K.T / sqrt(d_k) needs only one additional step: apply RoPE to Q and K before the matrix multiply. No changes to the attention score formula, no changes to the softmax, no changes to the value aggregation. This is why RoPE was adopted so quickly: it required almost no refactoring of existing transformer codebases.
RoPE Frequency Patterns
The frequency choices determine which relative offsets each dimension pair can distinguish. Let's visualize how the multi-frequency structure creates unique position signatures:
# Visualize how position affects embedding components
d_model = 64
max_pos = 100
frequencies = compute_rope_frequencies(d_model)
# For each position, compute the rotation for each dimension pair
# We'll track cos(pos * freq) and sin(pos * freq)
cos_components = np.zeros((max_pos, d_model // 2))
sin_components = np.zeros((max_pos, d_model // 2))
for pos in range(max_pos):
cos_components[pos] = np.cos(pos * frequencies)
sin_components[pos] = np.sin(pos * frequencies)

The heatmaps reveal the multi-scale nature of RoPE. Low-index dimension pairs (top rows) cycle rapidly, distinguishing nearby positions. High-index pairs (bottom rows) change slowly. This provides a coarse position signal. This structure resembles sinusoidal position encodings since both use similar frequency patterns. The key difference is that sinusoidal encodings add position information to embeddings, while RoPE rotates the embeddings themselves.
Every position in the sequence maps to a unique column pattern across all the dimension pairs. No two positions within the context window produce the same combination of cosine and sine values (unless positions are so far apart that even the slowest frequency has repeated, which requires distances well beyond any practical context). This uniqueness is what allows the model to discriminate between any two positions in the sequence.
We can also visualize the effective rotation that accumulates over a long sequence by tracking a single unit vector in each pair as position increases. The resulting picture shows clearly why high-frequency pairs lose discriminative power at large distances while low-frequency pairs preserve it:

Why RoPE Works So Well
Several properties make RoPE particularly effective, and understanding them helps explain why it displaced earlier relative encoding approaches so quickly.
Relative position by design. Unlike additive position encodings that must learn to extract relative position, RoPE provides it automatically through the geometry of rotations. The model doesn't need to learn that positions 5 and 7 are "2 apart"; the attention scores inherently reflect this. With sinusoidal encodings, the model sees a mixture of position and content signals and must disentangle them during training. With RoPE, the signals are separated from the start: content lives in the unrotated representation, and the rotation frame encodes how far apart two tokens are. This cleaner separation means the model can allocate its capacity to learning semantic patterns rather than arithmetic.
Length generalization. Because RoPE encodes relative rather than absolute position, models can often generalize to longer sequences than seen during training. Position 1000 rotating relative to position 1002 works the same as position 0 rotating relative to position 2. With learned absolute embeddings, position 1000 would simply be out of vocabulary if the model was trained only up to position 512. With RoPE, the underlying rotation frequencies are defined for any integer, so the mechanism itself never encounters an out-of-distribution position. In practice, other model components (feed-forward networks, layer normalizations) can still degrade at extreme lengths because they were optimized on short contexts, but RoPE itself is not the bottleneck.
Computational efficiency. RoPE requires no additional parameters beyond the pre-computed frequencies. The rotation can be implemented as element-wise operations, making it very fast. Compare this to relative position approaches like the one in Transformer-XL, which require computing a separate attention bias matrix of size and injecting it into the attention scores. RoPE has no such cost: compute and once per head per position, store them in a cache, and reuse across all layers. The per-token overhead is a handful of multiplications and additions, negligible compared to the matrix multiplications that dominate transformer compute.
Compatibility with linear attention. Some efficient attention approximations rely on the inner product structure of attention. RoPE preserves this structure (rotation is a linear transformation), making it compatible with these methods. Additive approaches that inject position bias after the softmax would not work with kernelized attention variants, but RoPE's pre-softmax rotation fits naturally into any formulation that starts from the dot product of query and key.
# Demonstrate length generalization
# Train on short sequences (conceptually)
train_seq_len = 20
# Test on longer sequence
test_seq_len = 100
d_model = 16
frequencies = compute_rope_frequencies(d_model)
# Check that relative distances produce consistent scores
# even at positions far beyond "training"
q = np.random.randn(d_model)
k = np.random.randn(d_model)
# Near the start (like training)
q_rot_5 = rope_complex(q, 5, frequencies)
k_rot_7 = rope_complex(k, 7, frequencies)
score_near = np.dot(q_rot_5, k_rot_7)
# Far beyond (test generalization)
q_rot_85 = rope_complex(q, 85, frequencies)
k_rot_87 = rope_complex(k, 87, frequencies)
score_far = np.dot(q_rot_85, k_rot_87)Length generalization test: Score at positions 5, 7 (relative distance 2): -5.592714 Score at positions 85, 87 (relative distance 2): -5.592714 Difference: 1.78e-15
Identical scores at identical relative distances, regardless of absolute position. This is why RoPE-based models can extrapolate to longer contexts more gracefully than models with absolute position encodings.
Let's test this length-generalization property across many absolute position pairs:
# Length generalization test
d_model = 32
frequencies = compute_rope_frequencies(d_model)
q = np.random.randn(d_model)
k = np.random.randn(d_model)
# Test relative distance of 2 at many different absolute positions
relative_dist = 2
absolute_positions = [0, 10, 50, 100, 500, 1000, 5000, 10000]
scores_at_positions = []
for pos in absolute_positions:
q_rot = rope_complex(q, pos, frequencies)
k_rot = rope_complex(k, pos + relative_dist, frequencies)
scores_at_positions.append(np.dot(q_rot, k_rot))
scores_at_positions = np.array(scores_at_positions)
The nearly zero standard deviation confirms that RoPE perfectly preserves relative position information regardless of where in the sequence we look. This is the mathematical foundation for length generalization in RoPE-based models.
Comparing RoPE to Other Position Encodings
Let's compare RoPE with other position encodings:
| Property | Sinusoidal | Learned | Relative (Shaw) | RoPE |
|---|---|---|---|---|
| Parameters | 0 | 0 | ||
| Position type | Absolute | Absolute | Relative | Relative |
| Attention modified | No | No | Yes | No (uses rotation) |
| Length extrapolation | Moderate | Poor | Moderate | Good |
| Computational cost | Low | Low | Higher | Low |
RoPE combines the parameter efficiency of sinusoidal encodings with the relative position benefits of learned relative encodings, without the architectural complexity. This balance explains its widespread adoption. The zero-parameter property is especially valuable at large scale: a 70-billion-parameter model trained with learned positional embeddings would need to store and train a separate matrix just for positions, which adds memory overhead and limits the maximum sequence length to the number of positions for which embeddings were learned. RoPE has neither of these constraints.
To make this comparison concrete, let's visualize how position similarity decays with distance for different encoding schemes:
# Compare position similarity decay across encoding methods
d_model_cmp = 64
max_distance = 50
# RoPE: dot product between rotated vectors at different distances
frequencies_cmp = compute_rope_frequencies(d_model_cmp)
rope_similarities = []
base_vec = np.ones(d_model_cmp) / np.sqrt(d_model_cmp) # Unit vector
base_rotated = rope_complex(base_vec, 0, frequencies_cmp)
for dist in range(max_distance):
other_rotated = rope_complex(base_vec, dist, frequencies_cmp)
rope_similarities.append(np.dot(base_rotated, other_rotated))
# Sinusoidal: dot product between position encodings
def sinusoidal_encoding(position, d_model):
"""Generate sinusoidal position encoding."""
pe = np.zeros(d_model)
for i in range(0, d_model, 2):
div_term = 10000 ** (i / d_model)
pe[i] = np.sin(position / div_term)
if i + 1 < d_model:
pe[i + 1] = np.cos(position / div_term)
return pe
sinusoidal_similarities = []
base_sin = sinusoidal_encoding(0, d_model_cmp)
for dist in range(max_distance):
other_sin = sinusoidal_encoding(dist, d_model_cmp)
sinusoidal_similarities.append(np.dot(base_sin, other_sin))
# Convert to arrays
rope_similarities = np.array(rope_similarities)
sinusoidal_similarities = np.array(sinusoidal_similarities)
Both methods show oscillating similarity patterns due to their multi-frequency structure. The key difference: sinusoidal encodings add this pattern to the input, while RoPE modulates the attention computation directly through rotation. For sinusoidal encodings, position similarity measures how similar two position vectors look in embedding space, which is a property of the encoding itself. For RoPE, what we measured is how a fixed content vector's score with itself changes across positions, a more direct measure of how position affects attention. The two are related but not the same, which is why the curves have different shapes even though both use the same base frequency formula.
Limitations and Considerations
Despite its elegance, RoPE has limitations worth understanding. Knowing where the mechanism can fail helps you make better architecture and training choices.
Frequency base sensitivity. The base (typically 10000) determines the frequency range and therefore which context lengths RoPE can distinguish reliably. Models trained with one base may not transfer well to contexts requiring different frequency patterns. If you train with base 10000 and then want to serve contexts of 100,000 tokens, many high-frequency pairs will have wrapped around so many times that distant positions become confused with nearby ones. Recent work like YaRN and NTK-aware scaling addresses this by adjusting frequencies for longer contexts, either by interpolating positions into the original range or by modifying the base value before the model is fine-tuned on longer data. These techniques are practical and widely adopted, but they do require additional training or fine-tuning steps.
High-frequency aliasing. At very long positions, high-frequency dimension pairs may "wrap around" multiple times, potentially creating aliasing where distant positions appear similar. In practice, this is rarely problematic within reasonable context lengths, but it's a theoretical limitation. The fastest-rotating pair completes a full cycle roughly every 6 positions, so by position 6,000 it has rotated 1,000 times. The cosine value at that point is the same as it was at position 0, 6, 12, and so on. The model can still distinguish positions globally because the slower pairs have not yet repeated, but if you rely only on the fast pairs or truncate the head dimension, aliasing becomes a real concern.
Dimension divisibility. RoPE requires even-dimensional queries and keys since it operates on pairs. This is a minor constraint but must be considered in architecture design. If a model uses a head dimension that is odd (unusual but not impossible), RoPE cannot be applied directly without padding. Standard models with have no issue, but custom architectures should verify this before adopting RoPE.
Training distribution effects. While RoPE theoretically supports any position, the model's other components (feed-forward networks, layer norms) are trained on a specific position distribution. Significant extrapolation may still degrade performance due to these other components, not RoPE itself. A model trained on sequences up to 2,048 tokens has feed-forward layers that have seen only hidden states arising from those contexts. At 32,768 tokens the activations may fall outside the distribution those layers were optimized for, causing unexpected behavior even though RoPE's geometric properties remain perfect.
These limitations are generally manageable. The community has developed extensions like Position Interpolation and NTK-aware RoPE that modify the frequency computation for better long-context performance. The core rotation mechanism remains unchanged.
In Practice: RoPE in Modern Language Models
RoPE has become the near-universal choice for position encoding in open-weight large language models released since 2023. LLaMA 1 and 2 adopted it directly from the original paper, as did Mistral 7B, Falcon, and Qwen. PaLM 2 from Google uses RoPE internally, and the pattern has continued into LLaMA 3 and its derivatives.
The implementation details vary slightly across these models. Most use the complex-number formulation with precomputed frequency caches that are generated once at model initialization and reused across all attention layers. Some implementations share the same frequency cache across all heads and layers, while others compute per-layer or per-head variants for ablation studies. The base frequency is sometimes increased (LLaMA 3 uses a base of 500,000 rather than 10,000) to support longer context windows without additional fine-tuning.
During inference with key-value caching, RoPE requires careful handling. The position indices passed to the rotation must reflect absolute sequence positions rather than the length of the current input chunk. When generating token by token, position 0 is the first token of the prompt, and each new token receives the next integer in the sequence. This bookkeeping is handled by the inference framework, but understanding it matters when you implement custom generation loops or want to correctly continue generation after a prefix.
Key Parameters
When implementing RoPE in your models, a small number of parameters control the behavior of the rotation. Understanding what each one does makes it easier to adapt RoPE for non-standard architectures or longer contexts.
When implementing RoPE in your models, these parameters control its behavior:
-
d_model(embedding dimension): The total dimension of query and key vectors. Must be even since RoPE operates on dimension pairs. Common values range from 64 to 4096, typically matching the model's hidden dimension divided by the number of attention heads. -
base(frequency base): Controls the range of rotation frequencies. The default value of 10000 provides a good balance between local and global position sensitivity. Larger values (e.g., 100000) extend the effective context length by slowing all rotations; smaller values make the encoding more sensitive to nearby positions. -
theta_i(per-dimension frequency): Computed as for dimension pair . Not typically set directly, but understanding this helps diagnose behavior: the first pair rotates once per radian, while the last pair completes a full rotation over approximately positions. -
Position offset: Some implementations support a starting position offset for key-value caching during inference. This allows continuing generation from a specific position without recomputing RoPE for all previous positions.
Choosing the right base value is the most consequential decision when deploying RoPE. The original value of 10,000 works well for sequences up to about 4,000 tokens. For 32,000-token contexts, values around 100,000 to 500,000 are common. For contexts of a million tokens or more, researchers have explored dynamic base scaling, where the effective base is adjusted at inference time based on the actual input length. These extensions preserve the core rotation mechanism while stretching the frequency range to cover longer distances without aliasing.
Summary
Rotary Position Embedding encodes position through geometric rotation rather than additive signals. This approach elegantly captures relative position through the natural properties of dot products between rotated vectors.
The core geometric argument is clean and complete. A 2D rotation preserves vector lengths and changes dot products only through the angular difference between the two rotated vectors. Applying a different rotation angle to each token position, with angle proportional to position, causes the attention score between any query and key to depend only on their positional separation. Extending this to high-dimensional embeddings is straightforward: split dimensions into pairs, rotate each pair independently at a different frequency, and sum the per-pair dot products. The relative-position property holds for each pair and therefore for the total.
Key takeaways:
-
Rotation as position encoding. Each position corresponds to a rotation angle. Rotating query and key vectors embeds position information directly into their geometric relationship.
-
Relative position emerges. When a rotated query at position attends to a rotated key at position , the dot product depends only on . Absolute positions cancel out through rotation mathematics.
-
Multi-frequency structure. Different dimension pairs rotate at different frequencies, creating a rich position representation. High frequencies capture local position differences; low frequencies capture global structure.
-
No additional parameters. Like sinusoidal encodings, RoPE uses deterministic frequencies based on dimension index. The only computation is the rotation itself.
-
Applied to Q and K only. Values are not rotated because they carry content, not position information. Rotation affects attention patterns, not the content that flows through them.
-
Good extrapolation. Because relative position is baked into the mechanism, models can often generalize to longer sequences than seen during training, though other model components may still limit this.
-
Widely adopted. LLaMA and Mistral, along with Falcon and Qwen, are among the recent open-weight models that use RoPE, often with a larger base frequency to support longer contexts.
RoPE has become the dominant position encoding in modern large language models. Its combination of theoretical elegance, computational efficiency, and practical effectiveness makes it a foundational technique for transformer architectures. The design is also notable for what it avoids: no extra parameters, no attention-formula modification, no explicit relative-position matrices. Position awareness comes from geometry itself, woven into the dot product by the mathematics of rotation. In the next chapter, we'll explore ALiBi, an alternative approach that adds relative position bias directly to attention scores.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about Rotary Position Embedding (RoPE).
Rotary Position Embedding (RoPE) Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!