Learned Position Embeddings

Michael BrenndoerferUpdated June 1, 202548 min read

Part of Language AI Handbook

How GPT and BERT encode position through learnable parameters. Topics include embedding tables, position similarity, interpolation techniques.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Learned Position Embeddings

The previous chapter introduced sinusoidal position encoding, a fixed mathematical function that assigns each position a unique pattern of oscillating values. The approach is elegant and principled: its wavelengths are designed according to a specific formula so that relative position can be detected through simple dot products. But there is an alternative philosophy worth examining carefully. Instead of designing position representations by hand, why not learn them from data?

This question sits at the heart of a broader tension in machine learning: should we encode our prior knowledge about a problem into the model's architecture, or should we give the model the capacity to discover what matters on its own? Sinusoidal encodings represent the "encode prior knowledge" camp. They bake in a specific mathematical structure based on our understanding of how position should behave. Learned position embeddings represent the "let the data decide" camp. They provide the model with a general-purpose table of parameters and let gradient descent shape those parameters into whatever positional representations best serve the task.

The distinction matters in practice. When you design a position encoding by hand, you make assumptions: that relative position matters more than absolute position, that frequency-based representations are informative, that the same encoding works across all tasks. Learned embeddings make no such assumptions. They can represent position in any way the training data supports, including patterns that a human designer would never think to impose. If a particular task benefits from having position 1 be very different from position 2 but position 50 look very similar to position 51, the model can learn exactly that. No hand-designed formula would produce such a pattern, but gradient descent will find it if it reduces loss.

Learned position embeddings take a different stance on what "representing position" means. Rather than encoding position through a predetermined formula, they treat position representations as trainable parameters. Just as word embeddings are learned vectors that capture semantic meaning through exposure to large text corpora, position embeddings can be learned vectors that capture positional meaning through training. The model discovers what aspects of position matter for the task at hand.

This approach powers many of the most influential language models in history, including GPT-2, GPT-3, and BERT. Understanding learned position embeddings is essential for working with these architectures and for making informed choices about position encoding when designing new models. The concept is also a gateway to understanding more advanced schemes like relative position encodings and rotary embeddings, which build on the intuitions developed here.

Historical Context: A Pragmatic Innovation

Learned position embeddings appeared in the very first transformer paper, "Attention Is All You Need" (Vaswani et al., 2017), as one of two options for position encoding. The authors noted that both sinusoidal and learned approaches performed similarly on the machine translation benchmarks they studied, suggesting that the exact form of position encoding matters less than the presence of any positional signal. Subsequent influential models like GPT and BERT adopted learned embeddings, establishing them as the default choice for language modeling. The decision was partly pragmatic: learned embeddings are conceptually simpler and require no hyperparameter choices about frequency ranges.

The Position Embedding Table

The core idea is simple enough to state in one sentence: maintain a lookup table of position vectors, one for each position in the sequence, and when processing a sequence, retrieve the embedding for each position and add it to the corresponding token embedding. This sentence contains the entire mechanism. Everything else is elaboration on why it works, what the learned vectors look like, and what constraints the approach imposes.

Think of the position embedding table as an analog to the token embedding table. When a transformer processes text, it first converts each token into a dense vector by looking up that token in the vocabulary embedding matrix. This matrix has one row per token in the vocabulary, and each row is a learnable parameter vector. The position embedding table works identically: it has one row per position up to the maximum sequence length, and each row is a learnable parameter vector. The only difference is what gets indexed. Token embeddings are indexed by token identity (which word is this?), while position embeddings are indexed by position (where is this word in the sequence?).

Position Embedding Table

A position embedding table is a trainable matrix of shape (Lmax⁡,d)(L_{\max}, d), where Lmax⁡L_{\max} is the maximum sequence length and dd is the embedding dimension. Position ii is represented by the ii-th row of this matrix. These vectors are learned during training, not computed from a formula.

Mathematically, if P∈RLmax⁡×d\mathbf{P} \in \mathbb{R}^{L_{\max} \times d} is the position embedding table, then the position embedding for position ii is:

pi=P[i,:]\mathbf{p}_i = \mathbf{P}[i, :]

where:

  • pi∈Rd\mathbf{p}_i \in \mathbb{R}^d: the position embedding vector for position ii
  • P∈RLmax⁡×d\mathbf{P} \in \mathbb{R}^{L_{\max} \times d}: the position embedding table (a learnable parameter matrix)
  • Lmax⁡L_{\max}: the maximum sequence length the model can handle
  • dd: the embedding dimension (same as token embeddings)
  • ii: the position index (0-indexed, so i∈{0,1,…,Lmax⁡−1}i \in \{0, 1, \ldots, L_{\max} - 1\})

This is identical to how word embeddings work: just as we look up a word's embedding from a vocabulary table, we look up a position's embedding from a position table. The key insight is that the lookup operation is differentiable with respect to the table entries, so gradients can flow through the lookup and update the position vectors during training.

Notice that this formula contains no mathematical structure about position. The vector p5\mathbf{p}_5 has no a priori relationship to p6\mathbf{p}_6. The vectors start as random noise and only acquire structure through training. This is both the strength and the weakness of the approach: the model is free to discover any positional structure, but it must discover it from scratch using training data.

In[3]:
Code
import matplotlib.pyplot as plt  # noqa: F401
import numpy as np


class LearnedPositionEmbedding:
    """
    Learned position embedding table.

    This is a simplified version that stores the embedding table as a NumPy array.
    In practice, this would be a PyTorch nn.Embedding layer.
    """

    def __init__(self, max_seq_len, embed_dim, seed=None):
        """
        Initialize the position embedding table.

        Args:
            max_seq_len: Maximum sequence length (L_max)
            embed_dim: Embedding dimension (d)
            seed: Random seed for reproducibility
        """
        if seed is not None:
            np.random.seed(seed)

        # Initialize with small random values (like word embeddings)
        # Using normal initialization scaled by 1/sqrt(embed_dim)
        self.embeddings = np.random.randn(max_seq_len, embed_dim) * (
            1.0 / np.sqrt(embed_dim)
        )
        self.max_seq_len = max_seq_len
        self.embed_dim = embed_dim

    def __call__(self, positions):
        """
        Look up position embeddings.

        Args:
            positions: Array of position indices

        Returns:
            Position embeddings of shape (len(positions), embed_dim)
        """
        return self.embeddings[positions]

The implementation mirrors how word embedding layers work in deep learning frameworks. In PyTorch, you would use nn.Embedding(max_seq_len, embed_dim), which handles the lookup and gradient computation automatically.

Let's create a position embedding table and examine what the initial (random) embeddings look like:

In[4]:
Code
# Create a position embedding table
max_seq_len = 128
embed_dim = 64

pos_embed = LearnedPositionEmbedding(max_seq_len, embed_dim, seed=42)

# Look up embeddings for positions 0 through 9
positions = np.arange(10)
pos_vectors = pos_embed(positions)
Out[5]:
Console
Position embedding table shape: (128, 64)
  - 128 positions
  - 64 dimensions per position

Sample position embeddings (first 5 dimensions):
  Position 0: [ 0.062 -0.017  0.081  0.19  -0.029] ...
  Position 1: [ 0.102  0.17  -0.009  0.125  0.045] ...
  Position 2: [ 0.012 -0.063 -0.194  0.009 -0.133] ...
  Position 3: [ 0.027 -0.156  0.022  0.048 -0.11 ] ...
  Position 4: [ 0.158 -0.088  0.055  0.097 -0.116] ...

At initialization, the position embeddings are random vectors with no meaningful structure. The magic happens during training, when gradients flow back through these embeddings and reshape them to capture positional patterns useful for the task. The initialization values are small (scaled by 1/d1/\sqrt{d}) to keep the initial representations in a reasonable range that avoids saturating activations early in training.

Out[6]:
Visualization
Heatmap of random position embeddings showing no coherent structure or patterns.
Random position embeddings at initialization. Each row represents a dimension and each column a position. The lack of any vertical structure shows that adjacent positions have completely unrelated embeddings at the start of training. Gradient descent will transform this noise into smooth, meaningful positional patterns over the course of training.

The heatmap reveals the chaotic structure of random initialization. Each column (position) has values that appear unrelated to neighboring columns. There is no smooth gradient, no pattern that would help the model understand that position 5 is closer to position 6 than to position 50. This randomness is the starting point. Training will transform this noise into meaningful structure through the pressure of minimizing prediction loss.

Combining Token and Position Embeddings

Once you have position embeddings, the question is how to incorporate them into the model. The standard approach, used by virtually all transformer architectures, is simple addition: take the token embedding for each word and add the position embedding for that word's location. The result is a single vector that encodes both the identity of the token and its position in the sequence.

Just as with sinusoidal encoding, learned position embeddings are added to token embeddings:

hi=ewi+pi\mathbf{h}_i = \mathbf{e}_{w_i} + \mathbf{p}_i

where:

  • hi∈Rd\mathbf{h}_i \in \mathbb{R}^d: the input representation for position ii, combining token and position information
  • ewi∈Rd\mathbf{e}_{w_i} \in \mathbb{R}^d: the token embedding for the word at position ii
  • pi∈Rd\mathbf{p}_i \in \mathbb{R}^d: the position embedding for position ii

This additive combination means position information is blended into the token representation from the very first layer. The model sees each token as existing at a particular position, not as a position-agnostic entity. When the attention mechanism computes its QKV projections from hi\mathbf{h}_i, those projections carry information about both what the token is and where it sits.

You might wonder whether addition is the right operation here. Why not concatenation, which would keep the two types of information separate? The answer is partly practical and partly principled. Concatenation would double the dimensionality of the representation, increasing the cost of all subsequent computations. Addition keeps the dimension fixed. More importantly, addition forces the model to learn token and position representations that work together harmoniously in the same vector space. The two sources of information must share the same dimensions, which creates pressure for them to be complementary rather than redundant.

In[7]:
Code
def combine_embeddings(token_embeddings, pos_embed):
    """
    Combine token embeddings with position embeddings.

    Args:
        token_embeddings: Token embeddings of shape (seq_len, embed_dim)
        pos_embed: LearnedPositionEmbedding instance

    Returns:
        Combined embeddings of shape (seq_len, embed_dim)
    """
    seq_len = token_embeddings.shape[0]
    positions = np.arange(seq_len)
    position_embeddings = pos_embed(positions)
    return token_embeddings + position_embeddings


# Example: combine with some token embeddings
np.random.seed(123)
token_embeds = np.random.randn(10, embed_dim) * 0.5  # 10 tokens

combined = combine_embeddings(token_embeds, pos_embed)
Out[8]:
Console
Combining token and position embeddings:
  Token embeddings shape:    (10, 64)
  Position embeddings shape: (10, 64)
  Combined shape:            (10, 64)

Example vector norms (showing position contribution):
  Position 0: token=4.709, pos=0.910, combined=4.554
  Position 5: token=3.886, pos=0.881, combined=3.958
  Position 9: token=4.019, pos=1.000, combined=4.333

The combined embeddings carry information about both what token appears at each position and where that position is in the sequence. This is the representation that attention layers will operate on. The norms show that position embeddings contribute a meaningful fraction of the total vector magnitude, confirming that positional information is not drowned out by token information.

In practice, a layer normalization step often follows the embedding addition in models like BERT, which further mixes the token and position signals before passing them to the attention layers. This normalization ensures the combined representation has a consistent scale regardless of individual embedding magnitudes.

How Position Embeddings Learn

Position embeddings learn through the same backpropagation process as all other parameters. When the model makes a prediction, gradients flow backward through the network, and the position embeddings receive updates that push them toward configurations that reduce the loss. The position embedding table is just a matrix of floats, and each float gets its own gradient computed via the chain rule.

The key question is: what signal drives these updates? Consider a language model predicting the next word. The prediction depends on the current token and its context. Position information affects what constitutes good context: a verb near the beginning of a sentence plays a different role than a verb near the end. When the model consistently makes better predictions by recognizing these positional patterns, the gradients will reinforce the position embeddings that support those patterns. Over many training examples, the embeddings learn to encode whichever aspects of position matter for reducing loss.

But what patterns do position embeddings learn? Research has revealed several consistent findings that appear across different models and architectures.

Position embeddings typically learn to encode absolute position information directly. Nearby positions tend to have similar embeddings, creating a smooth gradient across the sequence. This makes intuitive sense: positions 5 and 6 are more similar in terms of their role in the sequence than positions 5 and 50. A token at position 6 can attend to a token at position 5 with roughly the same relationship as any two adjacent tokens, while a token at position 5 and a token at position 50 are separated by a much wider context gap.

The key insight is that the model does not need to be told that nearby positions should have similar embeddings. This property emerges automatically because the training signal rewards it. When the model uses position embeddings to distinguish nearby versus distant context, it naturally learns that fine-grained positional distinctions (position 5 versus 6) look different from coarse distinctions (position 5 versus 50), and the embeddings reflect this.

Let's simulate what trained position embeddings might look like by creating a toy example with smoothly varying patterns:

In[9]:
Code
def create_trained_position_embeddings(max_len, embed_dim, seed=42):
    """
    Create position embeddings that simulate trained patterns.

    Real trained embeddings show smooth variation across positions.
    We simulate this with a combination of:
    - Low-frequency sinusoids (capturing global position)
    - Medium-frequency patterns (capturing local structure)
    - Some learned noise (capturing task-specific patterns)
    """
    np.random.seed(seed)
    embeddings = np.zeros((max_len, embed_dim))
    positions = np.arange(max_len)

    for dim in range(embed_dim):
        # Mix of frequencies, simulating what training discovers
        freq = 0.1 + 0.5 * (dim / embed_dim)  # Lower dims = lower freq
        phase = np.random.rand() * 2 * np.pi

        # Smooth sinusoidal component
        embeddings[:, dim] = 0.5 * np.sin(freq * positions + phase)

        # Add some learned variation
        embeddings[:, dim] += (
            0.2 * np.random.randn(max_len) * np.exp(-0.01 * positions)
        )

    # Normalize to have reasonable magnitude
    embeddings = embeddings / np.sqrt(embed_dim)

    return embeddings


trained_embeddings = create_trained_position_embeddings(128, 64)
Out[10]:
Visualization
Heatmap of position embeddings showing smooth vertical patterns that represent learned positional information.
Simulated trained position embeddings showing smooth patterns across positions. Different dimensions capture different scales of positional information. Low-index dimensions vary slowly (encoding global position), while higher-index dimensions oscillate more rapidly (capturing local structure). This multi-scale representation lets the model detect both coarse and fine-grained positional distinctions.

The visualization shows the characteristic structure of learned position embeddings. Different dimensions capture different aspects of position: some vary slowly across the entire sequence (low-frequency components), while others vary more rapidly (high-frequency components). This multi-scale representation allows the model to detect both global position ("near the start versus near the end") and local structure ("two positions apart"). Notice the similarity to sinusoidal encodings, but with irregular phase offsets and amplitude variations that reflect task-specific patterns rather than a fixed formula.

In practice, the learned embeddings from models like GPT-2 look remarkably similar to what our simulation shows: smooth variations across positions, with different dimensions operating at different frequencies. This convergence suggests that the training objective itself, minimizing next-word prediction loss over large corpora, naturally induces a frequency-based positional structure even without it being designed in.

Worked Example: Tracing Position Through a Transformer

To make the mechanism concrete, let's trace exactly what happens to position information when a transformer processes a simple sentence. Consider the sentence "The cat sat." with tokens at positions 0, 1, and 2.

At the input layer, each token is converted to its embedding and the corresponding position embedding is added. The token embedding for "cat" encodes something like "feline animal, common noun." The position embedding for position 1 encodes something like "second position in the sequence, often subject or modifier." Their sum is a single vector that carries both pieces of information.

When the attention layer processes this combined representation, it computes attention scores between all pairs of positions. The attention score between position 0 ("The") and position 1 ("cat") depends partly on their position embeddings. If those embeddings are similar (adjacent positions), the attention mechanism can easily learn to attend to nearby tokens. If they are dissimilar (distant positions), the mechanism must work harder to bridge the distance.

The query and key vectors are linear projections of the combined representation, so they carry positional information mixed with token identity information. A query at position 2 ("sat") asking "what is the subject of this verb?" will produce high attention scores for position 1 ("cat") partly because the position embeddings have been trained to make subject-verb relationships detectable by attention. The gradient signals that drive this learning come from the model's attempts to predict which tokens likely follow which patterns, and position is one of the key distinguishing features.

After attention, the position information propagates forward through the feed-forward layers, becoming increasingly mixed with the token's semantic content. By the final layer, the representation at each position has been enriched by context from other positions, with the original position embedding serving as one component of the total positional signal.

This concrete tracing illustrates a key property: position embeddings don't encode position in isolation. They work together with the attention mechanism to enable position-sensitive pattern matching. The embeddings provide raw positional coordinates, and the attention weights learn to use those coordinates to implement meaningful linguistic operations like subject-verb agreement, modifier-head relationships, and sentence boundary detection.

Position Similarity Analysis

One way to understand what position embeddings have learned is to compute similarities between positions. If the model uses position for word order and syntax, nearby positions should be more similar than distant ones.

The cosine similarity between two position embeddings pi\mathbf{p}_i and pj\mathbf{p}_j measures how aligned their directions are in the embedding space:

sim(i,j)=pi⋅pj∥pi∥⋅∥pj∥\text{sim}(i, j) = \frac{\mathbf{p}_i \cdot \mathbf{p}_j}{\|\mathbf{p}_i\| \cdot \|\mathbf{p}_j\|}

where:

  • pi⋅pj\mathbf{p}_i \cdot \mathbf{p}_j: the dot product of position embeddings ii and jj
  • ∥pi∥\|\mathbf{p}_i\|: the Euclidean norm of position embedding ii
  • sim(i,j)∈[−1,1]\text{sim}(i, j) \in [-1, 1]: a value of 1 means identical direction, 0 means orthogonal, -1 means opposite direction

When we compute this similarity for all pairs of positions and visualize the result as a matrix, we get a window into the geometric structure that training has imposed on the position embeddings. If positions are organized smoothly, we expect a band of high similarity along the diagonal, gradually fading to lower values as the distance between positions increases.

In[11]:
Code
def cosine_similarity_matrix(embeddings):
    """Compute pairwise cosine similarities between embeddings."""
    # Normalize each embedding to unit length
    norms = np.linalg.norm(embeddings, axis=1, keepdims=True)
    normalized = embeddings / (norms + 1e-8)

    # Cosine similarity = dot product of normalized vectors
    return normalized @ normalized.T


# Compute similarity for trained embeddings
sim_matrix = cosine_similarity_matrix(trained_embeddings)
Out[12]:
Visualization
Heatmap showing cosine similarity between positions with a bright diagonal and gradually decreasing values off-diagonal.
Cosine similarity between position embeddings. The bright diagonal confirms each position is most similar to itself. The gradual fade off the diagonal shows that nearby positions have more similar embeddings than distant ones, a property that emerges from training and enables the model to distinguish local from distant context.

The similarity matrix reveals the structure learned by position embeddings. The bright diagonal indicates that each position is most similar to itself (similarity = 1). The gradual darkening as we move away from the diagonal shows that nearby positions are more similar than distant ones. This band structure is characteristic of well-trained position embeddings. The pattern differs from sinusoidal encodings, which produce a more regular, wave-like similarity structure. Learned embeddings develop their own characteristic shape based on how positions are used in the training data.

Let's quantify how similarity decreases with distance:

In[13]:
Code
def similarity_by_distance(sim_matrix, max_distance=30):
    """Compute average similarity as a function of position distance."""
    n = sim_matrix.shape[0]
    distances = []
    avg_similarities = []

    for d in range(max_distance + 1):
        sims = []
        for i in range(n - d):
            sims.append(sim_matrix[i, i + d])
        distances.append(d)
        avg_similarities.append(np.mean(sims))

    return np.array(distances), np.array(avg_similarities)


distances, avg_sims = similarity_by_distance(sim_matrix, max_distance=40)
Out[14]:
Visualization
Line plot showing cosine similarity decreasing from 1.0 at distance 0 to near 0 at distance 40.
Average cosine similarity between position embeddings as a function of positional distance. Similarity starts at 1.0 for distance 0 (each position compared to itself) and decays smoothly toward 0 as distance increases. This decay profile is the geometric signature of a well-trained position embedding: the model has learned to map nearby positions to similar vectors and distant positions to unrelated vectors.

The decay curve shows that similarity drops off smoothly with distance. Positions 1 apart have high similarity (around 0.9), while positions 20 or more apart have near-zero similarity. This pattern allows the model to detect relative position through embedding similarity: if two position embeddings are very similar, the positions are likely close together. Notice that this relative position information emerges implicitly from absolute position embeddings. The model does not explicitly store "distance 3" as a concept; it learns absolute positions whose geometry happens to encode relative distance as a byproduct.

The Maximum Sequence Length Constraint

Unlike sinusoidal encodings, learned position embeddings have a hard constraint: the model can only handle sequences up to length Lmax⁡L_{\max}. If you train with Lmax⁡=512L_{\max} = 512, you have exactly 512 position embeddings. Position 513 simply does not exist in the table.

This constraint is fundamental to the design, not a bug or an oversight. The position embedding table has exactly Lmax⁡L_{\max} rows, and there is no formula to generate additional rows. A sinusoidal encoding can compute the representation for position 10,000 even if the model was trained on sequences of length 512, because the values come from a mathematical formula that can be evaluated at any integer. Learned embeddings cannot do this: position 10,000 is simply absent from the table.

In practice, this creates several design challenges.

During training, you choose Lmax⁡L_{\max} based on your computational budget and data characteristics. Larger Lmax⁡L_{\max} means more parameters (the position table has Lmax⁡×dL_{\max} \times d values) and higher memory usage during training. But the more significant constraint is attention complexity: self-attention scales quadratically with sequence length, so doubling Lmax⁡L_{\max} quadruples the attention cost. For this reason, models typically choose Lmax⁡L_{\max} as the longest sequence they expect to encounter in deployment, plus a small buffer, rather than setting it to the maximum they could technically accommodate.

During inference, sequences longer than Lmax⁡L_{\max} cannot be processed directly. You must either truncate the sequence to the first Lmax⁡L_{\max} tokens (discarding the rest), use a sliding window approach that processes overlapping chunks, or extend the model through fine-tuning on longer sequences.

In[15]:
Code
# Demonstrate the sequence length constraint
max_len = 128
pos_embed_small = LearnedPositionEmbedding(max_len, 64, seed=42)


def check_position_access(pos_embed, position):
    """Check if a position can be accessed."""
    if position < pos_embed.max_seq_len:
        return True, pos_embed(np.array([position]))
    else:
        return False, None


# Test various positions
test_positions = [0, 50, 127, 128, 200]
Out[16]:
Console
Position embedding table size: 128

Position access test:
  Position   0: Valid
  Position  50: Valid
  Position 127: Valid
  Position 128: Out of range
  Position 200: Out of range

The constraint is fundamental to the approach. Sinusoidal encodings can generate a representation for any position using the formula, but learned embeddings must be stored explicitly. This is a key trade-off between the two approaches, and it motivates the extrapolation techniques discussed in the next section.

In practice, most language modeling applications work within the context window. BERT was trained with Lmax⁡=512L_{\max} = 512 because most documents and question-answer pairs fit comfortably within 512 tokens. GPT-2 used Lmax⁡=1024L_{\max} = 1024 to accommodate longer generations. Modern models have pushed these limits to 2048, 4096, 8192, and beyond, accepting higher training costs in exchange for the ability to process longer contexts without truncation.

Extrapolation: Beyond Training Length

What happens if you need to process sequences longer than Lmax⁡L_{\max}? With learned embeddings, you have several options, none of which are ideal.

Option 1: Extend and fine-tune. Add new position embeddings for positions beyond Lmax⁡L_{\max}, initialize them (perhaps by interpolating from existing embeddings), and fine-tune on longer sequences. This works but requires additional training. The existing embeddings for positions 0 through Lmax⁡−1L_{\max} - 1 remain intact, but the new embeddings need training data that includes long sequences to learn meaningful representations for the extended positions. This approach is used in practice when a model trained on short sequences needs to handle longer documents: extend the table, fine-tune on long-document data, and the model adapts.

Option 2: Position interpolation. If you trained with Lmax⁡=512L_{\max} = 512 but need to handle 1024 tokens, you can interpolate the position indices. Position 512 in the long sequence uses the embedding for position 256 in the original table. This trick, called position interpolation, works surprisingly well for moderate extensions.

In[17]:
Code
def position_interpolation(pos_embed, target_len, original_max_len):
    """
    Interpolate position embeddings for longer sequences.

    Maps positions [0, target_len) to [0, original_max_len) and uses
    linear interpolation between the two nearest positions.
    """
    # Scale factor: how much to compress positions
    scale = original_max_len / target_len

    interpolated = np.zeros((target_len, pos_embed.embed_dim))

    for i in range(target_len):
        # Map to position in original table
        orig_pos = i * scale

        # Linear interpolation between floor and ceil positions
        low_pos = int(np.floor(orig_pos))
        high_pos = min(int(np.ceil(orig_pos)), original_max_len - 1)

        if low_pos == high_pos:
            interpolated[i] = pos_embed.embeddings[low_pos]
        else:
            # Weight based on fractional position
            weight = orig_pos - low_pos
            interpolated[i] = (1 - weight) * pos_embed.embeddings[
                low_pos
            ] + weight * pos_embed.embeddings[high_pos]

    return interpolated


# Extend from 128 to 256 positions using interpolation
original_max = 128
target_len = 256
interpolated_embeds = position_interpolation(
    pos_embed, target_len, original_max
)
Out[18]:
Console
Position Interpolation:
  Original max length: 128
  Target length:       256
  Scale factor:        0.50

Example mappings:
  New position   0 -> Original position 0.0
  New position  64 -> Original position 32.0
  New position 128 -> Original position 64.0
  New position 192 -> Original position 96.0
  New position 255 -> Original position 127.5

Position interpolation effectively "stretches" the position embedding table. The model sees positions at half the resolution, but the relative ordering is preserved. This works because the model primarily cares about relative positions, and those relationships are maintained under scaling. Research on extending GPT-2's context from 1024 to 2048 using interpolation found surprisingly small degradation in language modeling performance, suggesting that the absolute scale of position representations matters less than their relative structure.

The mathematical intuition is straightforward. If position embeddings encode a smooth manifold in the embedding space (which they do after training), then interpolated points along that manifold are reasonable approximations of what trained embeddings at those positions would look like. The interpolation maintains the smooth variation property that makes position embeddings effective.

Let's visualize how position interpolation affects the embedding structure:

In[19]:
Code
# Compare similarity structure before and after interpolation
# First, create a simulated trained embedding for the original table
original_trained = create_trained_position_embeddings(128, 64, seed=42)


# Interpolate to 256 positions
class SimpleEmbed:
    def __init__(self, embeddings):
        self.embeddings = embeddings
        self.embed_dim = embeddings.shape[1]


interpolated_256 = position_interpolation(
    SimpleEmbed(original_trained), 256, 128
)

# Compute similarity matrices for both
sim_original = cosine_similarity_matrix(original_trained)
sim_interpolated = cosine_similarity_matrix(interpolated_256)
Out[20]:
Visualization
Heatmap of 64x64 position similarities with bright diagonal band.
Original position similarity matrix for 128 positions. The diagonal band of high similarity, gradually fading with distance, shows the smooth positional geometry learned during training.
Heatmap of 128x128 interpolated position similarities with similar but stretched pattern.
Interpolated position similarity matrix for 256 positions. The same structural pattern is preserved but stretched across twice as many positions. Each original position maps to two new positions, and the relative relationships between all pairs are maintained.

The side-by-side comparison shows that position interpolation preserves the essential structure: nearby positions remain similar, and similarity decays with distance. The interpolated version has twice as many positions, but the relative relationships are maintained. This explains why interpolation works reasonably well for moderate length extensions.

Option 3: Sliding window. Process long sequences in overlapping chunks of length Lmax⁡L_{\max}. Each chunk gets proper position embeddings within the window (positions 0 through Lmax⁡−1L_{\max} - 1), but the model cannot attend across window boundaries. Predictions near the boundary of each chunk may be less accurate because they lack full context. This is common for very long documents where sliding window processing is acceptable.

The extrapolation problem is a significant limitation of learned position embeddings. Models trained with shorter contexts may struggle when forced to process longer sequences, even with interpolation tricks. This limitation has motivated a body of research into position encodings that generalize better to unseen lengths, including relative position encodings (which encode the distance between positions rather than their absolute locations) and rotary position embeddings (RoPE, which encodes position by rotating query and key vectors). We'll explore these approaches in later chapters.

Initialization and Training Dynamics

The initialization of position embeddings deserves careful attention because it affects how quickly they learn meaningful structure. Several initialization strategies have been used in practice.

The most common approach, used in GPT-2, initializes position embeddings from a normal distribution with standard deviation 0.02. This small value keeps the initial representations in a range where the network activations are well-behaved: not so small that position information vanishes in the noise, not so large that it dominates the token embeddings and destabilizes early training.

Another approach, sometimes used in models that inherit architecture from Word2Vec-style pretraining, initializes position embeddings with the same distribution as token embeddings. This treats position as just another feature to embed, with no special consideration for its unique properties. In practice, both approaches work similarly because training quickly overrides the initialization.

The training dynamics of position embeddings differ from other parameters in one subtle way: positions are sampled non-uniformly during training. In most text corpora, short sequences are more common than long ones. A model trained on Wikipedia articles will see many examples of positions 0 through 50 but relatively few examples of positions 400 through 512. This means later positions receive weaker gradient signals and may learn less reliable representations. This sampling bias can cause models to perform slightly worse on text that uses the full context window compared to text that fits comfortably in a shorter window.

In practice, the gradient signal is strong enough for the positions that appear most frequently in the training data, and those are typically the positions that matter most in deployment. Positions near the end of the context window are also positions where the model has the most context to work with, so their representations may need to distinguish among more possible contexts. The potential weakness in their training is partially offset by the fact that gradient signals from attention over long contexts provide rich learning signal when long sequences do occur.

GPT-Style Position Embeddings

GPT-2 and GPT-3 use learned position embeddings in a straightforward way. The architecture adds token and position embeddings at the input layer, then processes the combined representation through transformer blocks. There are no segment embeddings, no special handling of sentence boundaries, and no modifications to the basic additive scheme.

The simplicity is part of the design philosophy of GPT-style models. By using a single contiguous sequence as input and learned position embeddings for position, the model places the full burden of learning useful representations on the training process. The architecture provides no inductive biases beyond "tokens at different positions have different representations." Everything else, including how position interacts with token identity, how attention uses position information, and what positional patterns are linguistically meaningful, emerges from training on large text corpora.

In[21]:
Code
class GPTStyleEmbedding:
    """
    GPT-style embedding layer with learned token and position embeddings.

    This combines:
    - Token embeddings: lookup table for vocabulary
    - Position embeddings: lookup table for positions
    """

    def __init__(self, vocab_size, max_seq_len, embed_dim, seed=None):
        if seed is not None:
            np.random.seed(seed)

        # Token embedding table
        self.token_embeddings = np.random.randn(vocab_size, embed_dim) * 0.02

        # Position embedding table
        self.position_embeddings = (
            np.random.randn(max_seq_len, embed_dim) * 0.02
        )

        self.vocab_size = vocab_size
        self.max_seq_len = max_seq_len
        self.embed_dim = embed_dim

    def forward(self, token_ids):
        """
        Compute embeddings for a sequence of token IDs.

        Args:
            token_ids: Array of token indices, shape (seq_len,)

        Returns:
            Embeddings of shape (seq_len, embed_dim)
        """
        seq_len = len(token_ids)

        if seq_len > self.max_seq_len:
            raise ValueError(
                f"Sequence length {seq_len} exceeds maximum {self.max_seq_len}"
            )

        # Look up token embeddings
        token_embeds = self.token_embeddings[token_ids]

        # Look up position embeddings (positions 0 through seq_len-1)
        positions = np.arange(seq_len)
        pos_embeds = self.position_embeddings[positions]

        # Add them together
        return token_embeds + pos_embeds
In[22]:
Code
# Demonstrate GPT-style embedding
vocab_size = 50257  # GPT-2 vocabulary size
max_seq_len = 1024  # GPT-2 context length
embed_dim = 768  # GPT-2 embedding dimension

gpt_embed = GPTStyleEmbedding(vocab_size, max_seq_len, embed_dim, seed=42)

# Simulate processing a sequence
# (in practice, token_ids come from a tokenizer)
sample_token_ids = np.array(
    [464, 3290, 318, 845, 1310]
)  # "The cat is very small"
embeddings = gpt_embed.forward(sample_token_ids)
Out[23]:
Console
GPT-Style Embedding Layer
==================================================
Vocabulary size:       50,257
Max sequence length:   1,024
Embedding dimension:   768

Parameter counts:
  Token embeddings:    38,597,376 (38.6M)
  Position embeddings: 786,432 (0.79M)
  Total:               39,383,808 (39.4M)

Input sequence:  [ 464 3290  318  845 1310]
Output shape:    (5, 768)

The parameter count reveals an important insight: position embeddings are relatively cheap. In GPT-2 with its 50K vocabulary and 1024 positions, the position embedding table has only 0.79M parameters versus 38.6M for token embeddings. The position table represents about 2% of the embedding layer. This means increasing Lmax⁡L_{\max} is computationally inexpensive from a parameter standpoint, though it increases the quadratic attention cost.

This asymmetry matters for understanding why modern models keep pushing context length. Adding another 1024 positions doubles the position table from 0.79M to 1.58M parameters, a trivial increase relative to the billions of parameters in the rest of the model. The real cost of longer contexts is the self-attention computation, which requires O(L2)O(L^2) time and memory. This is why extending context length is an engineering challenge, not a parameter count challenge.

BERT-Style Position Embeddings

BERT also uses learned position embeddings, but with a few differences. BERT includes segment embeddings to distinguish between sentence pairs and uses a masked language modeling objective that trains the model to fill in masked tokens rather than predict the next token. These differences change how position embeddings are used but not how they are defined: BERT's position embedding table works identically to GPT's.

The segment embeddings in BERT represent a third type of input information beyond token identity and position. BERT processes pairs of sentences concatenated together (for tasks like question answering and natural language inference), and the segment embedding distinguishes which sentence each token belongs to. The three embeddings are summed together:

hi=ewi+pi+sgi\mathbf{h}_i = \mathbf{e}_{w_i} + \mathbf{p}_i + \mathbf{s}_{g_i}

where:

  • ewi∈Rd\mathbf{e}_{w_i} \in \mathbb{R}^d: the token embedding for word wiw_i
  • pi∈Rd\mathbf{p}_i \in \mathbb{R}^d: the position embedding for position ii
  • sgi∈Rd\mathbf{s}_{g_i} \in \mathbb{R}^d: the segment embedding for segment gi∈{0,1}g_i \in \{0, 1\}
  • hi∈Rd\mathbf{h}_i \in \mathbb{R}^d: the combined input representation
In[24]:
Code
class BERTStyleEmbedding:
    """
    BERT-style embedding layer with token, position, and segment embeddings.
    """

    def __init__(
        self, vocab_size, max_seq_len, embed_dim, num_segments=2, seed=None
    ):
        if seed is not None:
            np.random.seed(seed)

        # Token embeddings
        self.token_embeddings = np.random.randn(vocab_size, embed_dim) * 0.02

        # Position embeddings
        self.position_embeddings = (
            np.random.randn(max_seq_len, embed_dim) * 0.02
        )

        # Segment embeddings (for distinguishing sentence A vs sentence B)
        self.segment_embeddings = (
            np.random.randn(num_segments, embed_dim) * 0.02
        )

        self.vocab_size = vocab_size
        self.max_seq_len = max_seq_len
        self.embed_dim = embed_dim

    def forward(self, token_ids, segment_ids=None):
        """
        Compute embeddings for a sequence.

        Args:
            token_ids: Array of token indices, shape (seq_len,)
            segment_ids: Array of segment indices (0 or 1), shape (seq_len,)

        Returns:
            Embeddings of shape (seq_len, embed_dim)
        """
        seq_len = len(token_ids)

        if segment_ids is None:
            segment_ids = np.zeros(seq_len, dtype=int)

        # Look up all three embedding types
        token_embeds = self.token_embeddings[token_ids]
        pos_embeds = self.position_embeddings[np.arange(seq_len)]
        seg_embeds = self.segment_embeddings[segment_ids]

        # Sum all three
        return token_embeds + pos_embeds + seg_embeds

The addition of segment embeddings allows BERT to understand sentence structure in tasks like next sentence prediction and question answering. The position embeddings work identically to GPT-style: a simple lookup and addition. The key difference is that BERT's position embeddings must encode position within the full 512-token input, which may span two distinct sentences, while GPT's position embeddings encode position within a single continuous text stream.

This structural difference means BERT's position embeddings potentially encode slightly different information than GPT's. BERT's embeddings must help the model recognize when a position is near the boundary between two sentences, since the [SEP] token that separates sentences provides explicit signal but position provides context about how far into each sentence each token sits.

Analyzing Real Position Embeddings

When researchers analyze trained position embeddings from models like GPT-2 and BERT, several patterns emerge consistently. Understanding these empirical properties helps build intuition about what learned embeddings represent.

Low-rank structure. The 768-dimensional position embeddings can often be approximated well by a much lower-dimensional subspace. The first 50-100 principal components typically capture most of the variance. This suggests that despite living in a high-dimensional space, the effective information content of position embeddings is much lower. The extra dimensions provide capacity for task-specific adjustments during training, but the core positional structure is low-dimensional.

Smooth interpolation. Adjacent positions have similar embeddings, and this similarity decreases smoothly with distance. The embeddings form a continuous manifold in the embedding space, which is the geometric property that makes position interpolation possible.

Boundary effects. The first few positions (0, 1, 2) and positions near Lmax⁡L_{\max} sometimes show different patterns, possibly because they are encountered in distinct contexts during training. Position 0 always corresponds to special tokens like [CLS] in BERT or the beginning of a document in GPT, giving it atypical statistics. Positions near Lmax⁡L_{\max} are rare in training data if most documents are shorter than the maximum length.

Similarity to sinusoidal encodings. When you project trained position embeddings onto their principal components, the resulting patterns often look qualitatively similar to sinusoidal encodings at different frequencies. This convergence happens because sinusoidal patterns are efficient at encoding position in a way that supports the training objectives, and gradient descent discovers this even without the patterns being designed in.

Let's visualize the low-rank structure:

In[25]:
Code
from numpy.linalg import svd


def analyze_position_embedding_rank(embeddings):
    """Analyze the effective rank of position embeddings via SVD."""
    U, S, Vt = svd(embeddings, full_matrices=False)

    # Compute cumulative explained variance
    total_var = np.sum(S**2)
    cumulative_var = np.cumsum(S**2) / total_var

    return S, cumulative_var


# Analyze our simulated trained embeddings
singular_values, cumulative_var = analyze_position_embedding_rank(
    trained_embeddings
)
Out[26]:
Visualization
Line plot of singular values on log scale showing rapid exponential decay.
Singular value spectrum of position embeddings on a log scale. The rapid decay indicates that a small number of components captures most of the structural information, confirming the low effective dimensionality of learned positional representations.
Cumulative variance curve rising steeply then plateauing near 100%.
Cumulative explained variance as a function of the number of principal components. The curve reaches 90% with relatively few components, showing that the 64-dimensional position embedding space is spanned by a much smaller effective subspace.

The rapid decay of singular values confirms the low-rank structure. Most of the information in position embeddings can be captured by a handful of principal components. This suggests that positions are fundamentally simple, even though we represent them in high-dimensional space. The extra dimensions provide capacity for task-specific adjustments during training, but the core positional geometry is low-dimensional.

Visualizing Position Embedding Geometry

The similarity analysis and singular value decomposition give us statistical summaries of position embedding structure. But we can also visualize the geometry directly by projecting position embeddings into two dimensions using principal component analysis.

In[27]:
Code
from numpy.linalg import svd as np_svd


def pca_projection(embeddings, n_components=2):
    """Project embeddings to 2D using PCA."""
    # Center the data
    centered = embeddings - embeddings.mean(axis=0)

    # SVD for PCA
    U, S, Vt = np_svd(centered, full_matrices=False)

    # Project onto top 2 components
    projected = centered @ Vt[:n_components].T
    return projected, S[:n_components] / S.sum() * 100


# Project trained embeddings to 2D
proj_2d, explained = pca_projection(trained_embeddings)
Out[28]:
Visualization
2D scatter plot of position embeddings colored by position index showing a smooth ordered curve.
PCA projection of trained position embeddings into 2D. Positions are colored from early (dark blue) to late (yellow). The smooth color gradient reveals that the embeddings form an ordered curve through the embedding space, where adjacent positions lie close together and the overall trajectory traces from beginning to end of the sequence.

The PCA visualization reveals the geometric structure of position embeddings in a way that statistics cannot. The embeddings trace a smooth curve through the two-dimensional subspace, with early positions at one end and late positions at the other. This is the geometric signature of a well-trained position embedding: positions are ordered, and the ordering is smooth. The model has learned to map the discrete sequence of positions onto a continuous manifold.

Notice that this manifold structure is exactly what makes position interpolation work. If you need to insert a new position between positions 50 and 51, you can linearly interpolate between their embeddings and land on the manifold at the right location. The interpolated point is geometrically consistent with its neighbors in the same way that any point on a smooth curve is consistent with the points around it.

Trade-offs: Learned vs. Sinusoidal

The choice between learned and sinusoidal position encodings involves several trade-offs that depend on the application's requirements. Neither approach is universally better; each has situations where it is the more appropriate choice.

Flexibility. Learned embeddings can capture any pattern the data requires, including task-specific positional biases that no designer would anticipate. Sinusoidal encodings impose a fixed mathematical structure that may not match the task's needs. For most NLP tasks, learned embeddings perform as well as or better than sinusoidal encodings when the model is trained on sufficient data. The flexibility advantage becomes most pronounced in specialized domains where text structure differs from general language, such as programming languages, structured documents, or domain-specific formats.

Generalization to longer sequences. Sinusoidal encodings can generate representations for any position, including those never seen during training. A model trained with sinusoidal encodings on sequences of length 512 can immediately process sequences of length 1000 without any modification, because the formula simply evaluates at the new position indices. Learned embeddings are limited to Lmax⁡L_{\max} and may degrade for positions near the boundary where less training signal exists. For applications requiring length generalization, sinusoidal or other fixed encodings have a meaningful advantage.

Parameter count. Learned embeddings add Lmax⁡×dL_{\max} \times d parameters. For typical transformer sizes, this is a small fraction of total parameters. Sinusoidal encodings add zero parameters since they are computed from a formula. The parameter difference is usually negligible in practice.

Interpretability. Sinusoidal encodings have clear mathematical properties: the dot product of two sinusoidal position encodings has a specific relationship to the relative distance between positions, and different frequency components capture different scales of positional structure. These properties can be verified analytically. Learned embeddings are opaque; their properties must be discovered empirically through analysis like the similarity matrices and PCA plots we computed above.

Training efficiency. Learned embeddings must be trained, which requires gradients to flow through positions encountered during training. Rare positions (near Lmax⁡L_{\max} when most training sequences are short) may receive insufficient updates and develop weaker representations. Sinusoidal encodings work correctly immediately, with no training needed for the position component. This makes sinusoidal encodings particularly attractive when training data is limited or when position coverage is uneven.

Empirical performance. The original transformer paper found negligible performance difference between sinusoidal and learned position encodings on machine translation. Subsequent work has largely confirmed this: for standard language modeling and classification tasks, both approaches achieve similar performance when other architectural choices are equal. The community's drift toward learned embeddings reflects their simplicity and flexibility more than any performance advantage.

In practice, most modern language models use learned position embeddings. The flexibility to adapt to task-specific positional patterns outweighs the generalization advantages of fixed encodings for most applications. The sequence length limit is addressed by choosing Lmax⁡L_{\max} large enough for the target use case or by using techniques like position interpolation for moderate extensions.

Implementation in PyTorch

In practice, you would implement learned position embeddings using PyTorch's nn.Embedding layer. Here's what a real implementation looks like:

In[44]:
Code
import torch
import torch.nn as nn


class TransformerEmbedding(nn.Module):
    """
    Transformer embedding layer with learned token and position embeddings.
    """

    def __init__(self, vocab_size, max_seq_len, embed_dim, dropout=0.1):
        super().__init__()

        # Token embedding table
        self.token_embed = nn.Embedding(vocab_size, embed_dim)

        # Position embedding table
        self.pos_embed = nn.Embedding(max_seq_len, embed_dim)

        # Dropout for regularization
        self.dropout = nn.Dropout(dropout)

        # Store for position indexing
        self.max_seq_len = max_seq_len

    def forward(self, token_ids):
        """
        Args:
            token_ids: LongTensor of shape (batch_size, seq_len)

        Returns:
            Embeddings of shape (batch_size, seq_len, embed_dim)
        """
        batch_size, seq_len = token_ids.shape

        # Create position indices: [0, 1, 2, ..., seq_len-1]
        positions = torch.arange(seq_len, device=token_ids.device)
        positions = positions.unsqueeze(0).expand(batch_size, -1)

        # Look up embeddings and add
        token_embeds = self.token_embed(token_ids)
        pos_embeds = self.pos_embed(positions)

        return self.dropout(token_embeds + pos_embeds)

The PyTorch implementation is clean and efficient. The nn.Embedding layer handles the lookup table and gradient computation automatically. Position indices are created on-the-fly based on sequence length, and the same position embeddings are shared across all examples in a batch.

Notice the positions.unsqueeze(0).expand(batch_size, -1) idiom. This creates a position index tensor of shape (batch_size, seq_len) where every row is [0, 1, 2, ..., seq_len-1]. The expansion is memory-efficient because it shares the underlying storage rather than creating copies. When the embedding lookup runs, every example in the batch uses identical position indices (positions 0 through seq_len-1), so they all get the same position embeddings. This is correct: position 0 in sequence A should use the same position embedding as position 0 in sequence B, because position is an absolute index that does not depend on the content of the sequence.

The dropout applied after combining embeddings is regularization. By randomly zeroing out dimensions of the combined embedding during training, dropout prevents the model from relying too heavily on any single dimension of the positional or token representation. This tends to improve generalization, especially for smaller datasets.

Limitations and Impact

Learned position embeddings represent a pragmatic approach to position encoding. By treating positions as learnable parameters rather than fixed functions, they allow models to discover optimal position representations for their training data. This flexibility has proven valuable across many tasks, from language modeling to machine translation.

The primary limitation is the fixed sequence length. Models cannot process sequences longer than their training length without additional techniques like position interpolation or fine-tuning with extended context. This constraint has driven research into position encodings that generalize better to unseen lengths. Relative position encodings, which we will examine in the next chapter, address this by encoding the distance between positions rather than their absolute locations. Rotary position embeddings (RoPE), used in LLaMA and many recent models, go further by encoding position as a rotation of the query and key vectors, which produces excellent length generalization.

Another limitation is the lack of theoretical guarantees about what the embeddings learn. Unlike sinusoidal encodings, which have clear mathematical properties ensuring that the dot product of two position encodings encodes relative distance, learned embeddings are empirical objects whose properties must be discovered through analysis. When debugging a model that seems to struggle with positional reasoning, you cannot reason mathematically about what the position embeddings can and cannot represent. You must measure empirically, which is time-consuming and may not yield actionable insights.

A subtler limitation relates to the distribution of positions during training. When the training data contains many short sequences and few long ones, early positions receive much stronger gradient signals than late positions. The model develops better representations for common positions and potentially weaker representations for rare ones. This can cause unexpected behavior at inference time when processing sequences that are longer than typical training examples, even if they are within Lmax⁡L_{\max}.

Despite these limitations, learned position embeddings remain a dominant choice for transformer architectures. They are simple to implement and can adapt to task-specific patterns. Their effectiveness has made them the default in GPT-2, GPT-3, BERT, and many other influential models. The technique demonstrates a broader principle in deep learning: when you have enough data, learned representations often match or exceed the performance of hand-designed ones, and the engineering simplicity of learning everything from data tends to pay dividends at scale.

The impact of learned position embeddings extends beyond their direct use in models. The realization that position can be represented as a learnable parameter rather than a fixed function opened up research into other aspects of model architecture that were previously treated as fixed design choices. If position can be learned, what else can be learned? This perspective influenced the development of fully learned tokenizers, learned layer norms, and ultimately architectures like the Universal Transformer that learn structural biases from data.

Key Parameters

When implementing learned position embeddings, these parameters determine the capacity and behavior of the position encoding:

  • max_seq_len (Lmax⁡L_{\max}): Maximum sequence length the model can handle. This determines the number of rows in the position embedding table. Common values range from 512 (BERT) to 2048 and beyond in modern GPT variants. Larger values increase memory usage linearly but allow processing longer documents without truncation. The key tradeoff is that attention cost scales quadratically with sequence length, so doubling Lmax⁡L_{\max} quadruples the attention computation.

  • embed_dim (dd): Dimension of each position embedding vector. Must match the token embedding dimension for additive combination. Typical values range from 256 to 4096 depending on model size. Higher dimensions provide more capacity but increase parameter count proportionally. In practice, the effective dimensionality of position information is much lower than dd, so there is ample capacity even in moderate-dimension models.

  • Initialization scale: Position embeddings are typically initialized with small random values, often scaled by 1/d1/\sqrt{d} or using a fixed standard deviation (e.g., 0.02 in GPT-2). Smaller initialization helps training stability by keeping initial representations in a reasonable range. Very large initializations can cause position information to overwhelm token information early in training, slowing convergence.

  • dropout: Dropout rate applied after combining token and position embeddings. Values of 0.1 are common. This regularizes the model by randomly zeroing embedding dimensions during training, preventing over-reliance on specific features and improving generalization.

The ratio of position parameters to total model parameters is typically small (around 2% for GPT-2), making Lmax⁡L_{\max} relatively cheap to increase from a parameter perspective. The practical bottleneck is attention complexity, which constrains how long a context the model can attend over during training and inference.

Summary

Learned position embeddings treat position representations as trainable parameters rather than fixed formulas. This simple idea has proven remarkably effective across many language modeling tasks and underpins some of the most influential transformer architectures ever built.

Key takeaways from this chapter:

  • Position embedding table: A learnable matrix of shape (Lmax⁡,d)(L_{\max}, d) stores one embedding vector per position. Position ii is represented by the ii-th row of this table, analogous to how token embeddings work.

  • Additive combination: Position embeddings are added to token embeddings at the input layer, creating representations that encode both what token appears and where it appears. The sum is passed to all subsequent attention and feed-forward layers.

  • Training discovers structure: Through backpropagation, position embeddings learn to encode positional information useful for the task. Nearby positions typically develop similar embeddings, and the embeddings form a smooth low-rank manifold in the embedding space.

  • Maximum sequence length constraint: Unlike sinusoidal encodings, learned embeddings cannot represent positions beyond Lmax⁡L_{\max}. Processing longer sequences requires techniques like position interpolation (which stretches the existing embedding table) or fine-tuning with extended context.

  • Low effective dimensionality: Despite being stored in high-dimensional space, position embeddings often lie in a low-rank subspace. A few principal components capture most of the positional information, suggesting that the fundamental structure of position is simple even when represented in high dimensions.

  • GPT and BERT style: Major models like GPT-2 and BERT use learned position embeddings with simple additive combination. BERT adds segment embeddings on top to support sentence-pair inputs. The approach is straightforward to implement and works well in practice.

  • Trade-offs vs. sinusoidal: Learned embeddings offer more flexibility and task-adaptability but are limited in length generalization. Sinusoidal encodings generalize to any length but impose a fixed mathematical structure. For most large-scale language modeling applications, learned embeddings are the practical default.

  • PCA geometry: The PCA projection of trained position embeddings reveals an ordered curve through the low-dimensional subspace, confirming that training imposes a smooth, ordered geometry consistent with the sequence structure of language.

In the next chapter, we'll explore relative position encodings, which address the extrapolation problem by encoding the distance between positions rather than their absolute locations. This design helps models generalize to different sequence lengths without interpolation tricks.

Quiz

Test your understanding of learned position embeddings and how they differ from fixed encoding schemes.

Learned Position Embeddings

Question 1 of 100 of 10 completed
What is the shape of a learned position embedding table?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2025learnedposition, author = {Michael Brenndoerfer}, title = {Learned Position Embeddings}, year = {2025}, url = {https://mbrenndoerfer.com/writing/learned-position-embeddings}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2025). Learned Position Embeddings. Retrieved from https://mbrenndoerfer.com/writing/learned-position-embeddings
MLAAcademic
Michael Brenndoerfer. "Learned Position Embeddings." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/learned-position-embeddings>.
CHICAGOAcademic
Michael Brenndoerfer. "Learned Position Embeddings." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/learned-position-embeddings.
HARVARDAcademic
Michael Brenndoerfer (2025) 'Learned Position Embeddings'. Available at: https://mbrenndoerfer.com/writing/learned-position-embeddings (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2025). Learned Position Embeddings. https://mbrenndoerfer.com/writing/learned-position-embeddings

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.