Embedding Models: Architecture, Pooling & Selection

Michael BrenndoerferJanuary 24, 202665 min read

Part of Language AI Handbook

Explains how embedding models convert text to vectors for RAG. Topics include bi-encoder architecture, pooling strategies, dimensionality trade-offs.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Embedding Model Architecture and Selection

In the previous chapters on dense retrieval and contrastive learning, we established why we need dense vector representations and how models learn to place semantically similar texts near each other in vector space. We explored the mechanics of contrastive loss functions and saw how training pairs can sculpt a continuous geometry of meaning. However, we did not address a necessary question: how does a transformer, which produces a sequence of token-level hidden states, yield a single vector that represents an entire sentence or paragraph? The answer lies in the architecture and design choices of embedding models, the specialized systems that convert variable-length text into fixed-dimensional vectors suitable for similarity search.

Embedding models sit at the heart of every RAG pipeline. The quality of your retrieval, and therefore the quality of your generated answers, depends directly on how well your embedding model captures semantic meaning. A poorly chosen or misconfigured embedding model will return irrelevant chunks no matter how sophisticated the rest of your pipeline is. You could have the most advanced reranker, the most carefully designed prompts, and the most capable generative LLM in the world, but if your retriever returns the wrong passages, none of that sophistication can compensate. The embedding model is the foundation on which everything else stands.

Think of an embedding model as a very sophisticated filing system. When you file a document, you decide what folder it belongs in, and later you can retrieve it by looking in the right folder. An embedding model does something far more fine-grained: it places each piece of text at a precise coordinate in a high-dimensional space, where the coordinates encode the text's meaning. Documents about similar topics land near each other; documents about unrelated topics land far apart. When a query arrives, you convert it to coordinates in the same space and look for nearby documents. The quality of this "filing" determines everything about your retrieval system's accuracy.

This chapter examines embedding models from the inside out. We'll start with the architectural foundations and the non-trivial problem of distilling a sequence into a single vector. Then we'll explore the pooling strategies that perform this distillation, examining each one's mathematical properties and practical implications. We'll discuss the role of embedding dimensionality, including the elegant Matryoshka technique that allows a single model to serve multiple dimensionality requirements. We'll work through a concrete numerical example that makes these concepts tangible. Finally, we'll cover how to select and evaluate a model before deploying it for your use case, including limitations you need to be aware of.

Historical Context

The challenge of creating sentence-level representations from word-level models predates transformers. Early approaches averaged word2vec or GloVe vectors, which worked reasonably well for short texts but failed to capture word order, negation, and compositional meaning. The InferSent model (2017) used BiLSTMs trained on natural language inference data to produce better sentence embeddings. The Universal Sentence Encoder (2018) used transformer architectures but relatively shallow ones. Sentence-BERT (2019) marked a turning point by showing that deep transformer encoders, fine-tuned with siamese networks and contrastive objectives, could produce dramatically better semantic embeddings than any prior approach. Since then, the field has advanced rapidly, with modern models like E5-Mistral and GTE-Qwen2 using billion-parameter LLMs as embedding backbones.

From Token Embeddings to Sentence Embeddings

As covered in previous chapters, encoder models like BERT produce contextual token embeddings: one vector per input token, where each vector is influenced by the surrounding context through self-attention. A sentence of 12 tokens passed through BERT yields 12 vectors of dimension 768. But for retrieval, we need one vector per document chunk. The bridge between token-level representations and a single sentence-level vector is what defines an embedding model, and this bridge is anything but trivial.

To appreciate why this distillation is non-trivial, consider the nature of the problem more carefully. A transformer encoder processes text through multiple layers of self-attention and feed-forward transformations, creating at each layer a set of contextualized representations, one per token. By the final layer, each token's hidden state is richly informed by all the other tokens in the input. The vector for "bank" in "the river bank was muddy" encodes the word "bank" and the fact that it is a physical location near water rather than a financial institution, because self-attention has allowed the model to blend contextual information from all surrounding tokens into each position's representation. These representations are extraordinarily rich, but they remain anchored to individual token positions. There is no built-in mechanism in a standard transformer encoder that automatically distills an entire sequence's meaning into a single, compact vector. That distillation is precisely what an embedding model must accomplish, and the method it uses to do so has a significant impact on retrieval quality.

The key insight is that collapsing 12 token vectors into 1 sentence vector is fundamentally a lossy compression problem. You cannot preserve every nuance of every token's contextual representation in a single fixed-length vector. The question is: which information should you keep? Different pooling strategies make different choices about this compression, and the right choice depends on both the model architecture and the training objective. A pooling strategy is a design choice that the model must be trained to exploit, rather than a mathematical operation applied only after training. The model's weights, shaped by the training loss, determine what information ends up in each position's hidden state. If the model was trained assuming mean pooling would be used, the hidden states will encode information in a form optimized for averaging. If you then apply CLS pooling, you're reading information out of a representation that was never designed to concentrate meaning at the CLS position, and quality degrades.

Embedding Model

An embedding model is a neural network, typically built on a transformer encoder, that maps variable-length text to a fixed-dimensional vector. Unlike general-purpose language models, embedding models are specifically trained so that the geometric relationships between output vectors reflect semantic relationships between input texts. Two texts that mean the same thing should produce vectors that point in nearly the same direction; two texts about completely different topics should produce vectors that are roughly orthogonal.

The naive approach of simply using a pre-trained BERT model to generate sentence embeddings performs surprisingly poorly. Reimers and Gurevych (2019) demonstrated that averaging BERT's token embeddings without further training produces representations that are often worse than GloVe averages for semantic similarity tasks. This finding was surprising at the time because BERT represented a massive leap forward in NLP performance on almost every other benchmark. The problem is that BERT's pre-training objectives (masked language modeling and next sentence prediction) optimize for token-level predictions, not for creating globally meaningful sentence vectors. BERT learns to predict missing words and to decide whether two sentences follow each other in the original document, but neither of these tasks requires the model to compress the overall meaning of a sentence into a single point in vector space. The token-level hidden states carry abundant contextual information, but that information is distributed across positions in a way that does not naturally aggregate into a coherent whole without explicit guidance.

This insight motivated the development of dedicated embedding models: transformer encoders that are fine-tuned with objectives that explicitly shape the geometry of the output space. The contrastive learning techniques from the previous chapter are precisely how this fine-tuning is done. By training the model to push embeddings of semantically similar texts closer together while pulling embeddings of dissimilar texts apart, the fine-tuning process turns a collection of token-level representations into a structured embedding space where distances encode semantic similarity. The pre-trained transformer provides a strong foundation of linguistic knowledge; the contrastive fine-tuning teaches the model how to organize that knowledge into a space that supports retrieval.

Embedding Model Architectures

Modern embedding models share a common blueprint, but they differ in their choice of base encoder, training procedure, and output processing. Understanding these architectural families helps clarify what happens under the hood when you call a model's encode method and receive a vector back. The architectural family determines which pooling strategy is appropriate, how the model should be prompted, and what kinds of texts it can handle effectively.

Before we examine specific families, it is worth understanding the general structure that all embedding models share. Every embedding model takes a text string as input and produces a fixed-dimensional vector as output. Between these two endpoints, there are three main components. The tokenizer converts the raw text into a sequence of integer token IDs, handling subword splitting, special token insertion, and sequence padding or truncation. The transformer encoder then processes these token IDs through multiple layers of self-attention and feed-forward networks, creating a matrix of hidden states, one row per token. Finally, the pooling and projection layer collapses this matrix into a single vector and optionally applies normalization or dimensionality reduction. Each of these components has design choices with significant consequences for the quality and behavior of the resulting embeddings.

Sentence-BERT (SBERT) and the Bi-Encoder Pattern

Sentence-BERT, introduced by Reimers and Gurevych (2019), established the dominant paradigm for embedding models. The core idea is elegant: take a pre-trained transformer encoder, add a pooling layer on top, and fine-tune the entire stack using siamese or triplet networks with a contrastive objective. What makes this approach so powerful is the combination of a strong pre-trained foundation, which provides rich contextual understanding of language built from billions of tokens of web text, with a training objective that explicitly organizes the output space for semantic comparison. The pre-training instills deep linguistic knowledge; the fine-tuning teaches the model how to express that knowledge as a geometry of meaning.

The architecture uses what's called a bi-encoder (or dual-encoder) pattern, which we encountered in the dense retrieval chapter. The query and the document are encoded independently by the same model, without any cross-attention between them. This independence is what makes bi-encoders fast at retrieval time: document embeddings can be precomputed and indexed once, and at query time only the query needs to be encoded and compared against the pre-indexed vectors. The key insight is that because the same encoder processes both queries and documents, both types of text are mapped into the same shared vector space. Proximity in this space reflects semantic similarity, regardless of whether the two texts being compared are both queries, both documents, or one of each. The shared space is what enables the model to match query intent to document content even when the specific words differ.

The forward pass for a single input proceeds through four key steps. First, tokenization splits the input text into subword tokens using the encoder's vocabulary, inserting special tokens like [CLS] at the beginning and [SEP] at the end. Second, encoding passes the token IDs through the transformer layers, creating contextualized token embeddings where each position's representation reflects the surrounding context. Third, pooling reduces the token embeddings to a single vector using a strategy like mean pooling. Fourth and optionally, the vector is L2-normalized to unit length, which simplifies similarity computation. Each of these steps involves design choices that affect the quality of the resulting embedding, and the pooling step in particular, which we will examine in depth shortly, is where much of the embedding model's character is determined.

Training SBERT-style models uses a siamese architecture during the training phase. Two copies of the same encoder (sharing all weights) process a pair of texts in parallel, creating two embeddings. A contrastive loss then trains these embeddings to be similar for semantically related pairs and dissimilar for unrelated pairs. The weight-sharing is needed: it ensures that the same text always produces the same embedding regardless of which "branch" of the siamese network processes it, which is a prerequisite for consistent similarity scoring. After training, only one copy of the encoder is needed for inference.

Instructor and Task-Aware Models

A limitation of standard embedding models is that they produce the same vector for a piece of text regardless of the intended task. The text "Python is a programming language" gets the same embedding whether you're doing topic classification, semantic search over technical documentation, or clustering discussion forum posts. This is a meaningful limitation because the aspects of a text that matter most depend entirely on how the embedding will be used. For a classification task, you might want the embedding to emphasize the category or topic. For a semantic search task, you might want it to emphasize the specific technical facts or claims that a query could ask about. For a clustering task, you might want it to emphasize the linguistic style and register of the text. A single, task-agnostic embedding cannot optimally serve all of these purposes simultaneously.

Task-aware models like Instructor (Su et al., 2023) address this limitation by prepending a natural language task instruction to each input. Instead of encoding just the text, you encode something like: "Represent the science document for retrieval: Python is a programming language." The instruction conditions the model to produce embeddings optimized for the specified task. By seeing the instruction as part of the input sequence, the model can attend to the instruction when building the contextual representations of the main text, effectively shifting what aspects of meaning it emphasizes.

This approach does not require architectural changes compared to standard embedding models: it uses the transformer's ability to attend across the full input, including the instruction prefix. The model is trained on diverse tasks with diverse instructions, learning to shift its embedding space based on the instruction context. The result is a single model that behaves like many specialized models, adapting its output representations to the task at hand simply by reading a natural language instruction. For a RAG system serving multiple query types, for example factual lookup, comparison, and definition retrieval, you can use different instructions to bias the embeddings appropriately for each type.

Late Interaction Models

While not strictly single-vector embedding models, late interaction architectures like ColBERT deserve mention because they represent a middle ground between bi-encoders and cross-encoders. The basic trade-off in retrieval model design is between expressiveness and efficiency. Bi-encoders are efficient because they compress each text to a single vector, but this compression sacrifices fine-grained token-level matching. Two passages that contain the same key phrase but in very different contexts might receive similar embeddings from a bi-encoder, even though their relevance to a specific query differs substantially. Cross-encoders are expressive because they perform full token-level cross-attention between a query and a document, but this means the document cannot be pre-encoded independently of the query, making them too slow for first-stage retrieval over large corpora.

Late interaction models handle this trade-off by retaining more information than bi-encoders while remaining far more efficient than cross-encoders. Instead of compressing each text into a single vector, ColBERT retains all token-level embeddings and computes similarity using a MaxSim operation: for each query token, find its maximum cosine similarity to any document token, then sum these maximums across all query tokens. This mechanism allows ColBERT to capture fine-grained lexical and semantic matches between individual query terms and document terms, something that a single-vector comparison cannot achieve.

The MaxSim operation is defined as:

Score(q,d)=∑i∈qmax⁡j∈dcos⁡(hiq,hjd)\text{Score}(q, d) = \sum_{i \in q} \max_{j \in d} \cos(\mathbf{h}_i^q, \mathbf{h}_j^d)

where:

  • qq: the query text with token positions ii
  • dd: the document text with token positions jj
  • hiq\mathbf{h}_i^q: the embedding of query token ii
  • hjd\mathbf{h}_j^d: the embedding of document token jj
  • cos⁡(⋅,⋅)\cos(\cdot, \cdot): cosine similarity between two token embeddings

The key insight here is that every query token gets to "find" its best match in the document, so even rare or domain-specific terms that might be washed out in a mean-pooled embedding can contribute to the overall relevance score. This preserves more fine-grained information than single-vector approaches while remaining more efficient than full cross-attention, since document token embeddings can still be precomputed and indexed. We'll see how ColBERT fits into retrieval pipelines when we discuss reranking later in this part.

Modern Embedding Architectures

Recent embedding models have pushed performance significantly through several complementary innovations. These advances reflect a broader trend in the field: rather than relying on a single breakthrough technique, the best embedding models combine multiple strategies to achieve state-of-the-art results. Understanding these innovations helps you evaluate available embedding models and anticipate where the field is heading.

The first major innovation is the use of larger and better base models. Early embedding models used BERT-base (110M parameters) or BERT-large (340M parameters) as their backbone. More recent models have moved to much larger encoders or even decoder-based models. Models like E5-Mistral use a decoder-only LLM (Mistral 7B, with 7 billion parameters) as the backbone, applying special pooling to extract embeddings. The intuition here is that larger language models develop richer internal representations of language during pre-training on vastly larger datasets. These richer representations provide a better starting point for embedding fine-tuning, and they can encode more fine-grained semantic distinctions than smaller models.

The second major innovation is multi-stage training. The GTE, BGE, and E5 model families use a progressive training pipeline with distinct stages. The first stage uses large-scale weakly-supervised data: for example, title-body pairs from web pages, question-answer pairs from forums, or paraphrase pairs from multilingual corpora. This teaches the model broad notions of textual relevance from massive, automatically collected data. The second stage fine-tunes on curated, human-annotated data where quality is much higher but quantity is lower. This refines the model's understanding using high-quality human judgments about relevance. An optional third stage performs knowledge distillation from a larger teacher model into a smaller student, compressing performance from a model too large for practical inference into one that can be deployed efficiently.

Third is Matryoshka representation learning, which trains embeddings so that any prefix of the vector is itself a useful embedding. This technique, which we will examine in detail later in this chapter, allows a single model to serve multiple dimensionality requirements by simply truncating its output vector at different lengths. Fourth is unified models for multiple tasks: models like GTE-Qwen2 handle embedding, retrieval, reranking, and classification within a single architecture, using natural language instructions to switch between modes. This unification simplifies deployment because a single model can serve multiple roles in the pipeline, reducing the number of models you need to manage.

Pooling Strategies

The pooling layer is arguably the most necessary design choice in an embedding model. It determines how the rich, token-level information from the transformer encoder is compressed into a single vector. To understand why this choice matters so much, consider what happens at the output of a transformer encoder. You have a matrix of hidden states, one row per token, each row a high-dimensional vector capturing that token's meaning in context. The pooling layer must somehow condense this entire matrix, which can have hundreds of rows for longer inputs, into a single row. Different strategies for performing this condensation emphasize different aspects of the input, and the choice of strategy interacts deeply with how the model was trained.

Think of the pooling operation as deciding how to summarize a meeting transcript. You could take the notes from the meeting chair (CLS pooling), who was designated as the official recorder but might have focused on their own agenda items. You could average everyone's notes together (mean pooling), which captures the collective sense of the meeting but might dilute points that were only raised by one person. You could take the most emphatic point from each person (max pooling), which highlights the strongest sentiments but misses the fine-grained discussion. Or you could take the notes from the last person to speak (last token pooling), who had heard the entire discussion and could synthesize it, but might have overweighted the final points.

CLS Token Pooling

BERT and its variants include a special [CLS] token at the beginning of every input sequence. During BERT's pre-training, the hidden state corresponding to [CLS] is used as the "aggregate sequence representation" for the next sentence prediction task: a linear classifier applied to this single vector must decide whether the second sentence follows the first. CLS pooling simply takes this one vector as the sentence embedding, designating the first position as the dedicated aggregation point.

Formally, the CLS pooling operation extracts the hidden state at the very first position of the output:

e=h[CLS]\mathbf{e} = \mathbf{h}_{\text{[CLS]}}

where:

  • e\mathbf{e}: the resulting sentence embedding vector of dimension dd
  • h[CLS]\mathbf{h}_{\text{[CLS]}}: the final hidden state vector at the [CLS] token position, which is always the first token in BERT's tokenization scheme

The appeal of CLS pooling is that the model has a dedicated token whose job is to aggregate information from the entire sequence. Through bidirectional self-attention, the [CLS] token can attend to every other token at every layer, and the model can learn to pack a summary of the input into this position. Because self-attention is fully connected (every token can interact with every other token), the [CLS] token has, in principle, access to the full content of the input by the time it reaches the final layer. If the training objective rewards the [CLS] position for capturing the overall meaning of the input, the model has both the mechanism (self-attention) and the incentive (the loss function) to make this work.

However, CLS pooling has a significant weakness: the [CLS] token's representation is heavily shaped by its pre-training objective (next sentence prediction in BERT), which does not align well with semantic similarity. The next sentence prediction task is a binary classification problem. It asks a coarse question: does sentence B follow sentence A in the original document? This coarse signal does not require, and does not reward, the model for encoding fine-grained semantic meaning into the [CLS] representation. As a result, without explicit fine-tuning for embedding quality, [CLS] representations can be poorly organized for semantic search, placing dissimilar texts near each other and similar texts far apart. After proper contrastive fine-tuning that specifically rewards the [CLS] representation for semantic discrimination, CLS pooling works well and is used by several strong models including DeBERTa-based models and many decoder-based embedding models.

Mean Pooling

Mean pooling takes a fundamentally different approach from CLS pooling. Rather than relying on a single designated token to carry all the information, mean pooling distributes the responsibility across every token in the sequence. It computes the sentence embedding by averaging the hidden states of all input tokens, excluding padding tokens, to produce a single representative vector.

To understand the motivation, consider what happens at the final layer of a transformer encoder. Each token's hidden state is a rich, contextualized representation that has been refined through many layers of self-attention and feed-forward processing. The hidden state at the word "photosynthesis" in a biology passage does not just encode the word itself: it encodes the word in the context of the entire passage, informed by every other word through self-attention. By averaging all of these context-rich representations, mean pooling creates a vector that reflects the collective semantic content of the entire sequence, weighted equally across all positions. The formal expression is:

e=1∣T∣∑i∈Thi\mathbf{e} = \frac{1}{|\mathcal{T}|} \sum_{i \in \mathcal{T}} \mathbf{h}_i

where:

  • e\mathbf{e}: the resulting sentence embedding vector
  • T\mathcal{T}: the set of indices corresponding to real (non-padding) tokens
  • hi\mathbf{h}_i: the hidden state vector output by the transformer at position ii
  • ∣T∣|\mathcal{T}|: the count of real tokens in the sequence, which is the normalizing denominator

In practice, this is implemented using the attention mask to zero out padding positions before averaging. Padding tokens are added to make all sequences in a batch have the same length, but they carry no meaningful content and should not influence the embedding. The attention mask, a binary vector with 1 for real tokens and 0 for padding, provides exactly the information needed to exclude padding positions:

e=∑i=1nmi⋅hi∑i=1nmi\mathbf{e} = \frac{\sum_{i=1}^{n} m_i \cdot \mathbf{h}_i}{\sum_{i=1}^{n} m_i}

where:

  • e\mathbf{e}: the resulting sentence embedding vector
  • nn: the total length of the tokenized sequence, including padding
  • mim_i: the attention mask value at position ii (1 for real tokens, 0 for padding)
  • hi\mathbf{h}_i: the hidden state vector at position ii

The numerator sums only the hidden states of real tokens, since multiplying by the mask zeroes out padding positions. The denominator counts the number of real tokens. This ensures that the average is taken only over meaningful positions, regardless of how much padding was added.

Mean pooling is the most widely used strategy in modern embedding models, and for good reason. By averaging over all tokens, it captures information distributed across the entire sequence rather than relying on a single position that might or might not have attended equally to all parts of the input. It is more reliable to the arbitrary characteristics of specific token positions, and the training signal can flow to every token equally, giving the model more flexibility in how it distributes semantic information across the sequence. For longer inputs, where useful information may appear anywhere from the opening sentence to the final clause, this distributed representation is especially valuable.

Sentence-BERT showed that mean pooling consistently outperforms CLS pooling when fine-tuning from a pre-trained BERT checkpoint. Most state-of-the-art models, including the E5, GTE, and BGE families, use mean pooling as their default strategy. The reason mean pooling often outperforms CLS pooling is not that CLS is inherently inferior, but that mean pooling is more forgiving: even a model that has not perfectly learned to concentrate information at the CLS position will produce reasonable mean-pooled embeddings, because the average captures information wherever it happens to reside in the hidden states.

Weighted Mean Pooling

A refinement of mean pooling introduces position-dependent weights, letting the model to emphasize certain token positions over others. Typically, later tokens in the sequence receive higher weight, based on the intuition that they have "seen" more context through the attention mechanism and may carry more refined semantic information:

e=∑i=1nwi⋅mi⋅hi∑i=1nwi⋅mi\mathbf{e} = \frac{\sum_{i=1}^{n} w_i \cdot m_i \cdot \mathbf{h}_i}{\sum_{i=1}^{n} w_i \cdot m_i}

where:

  • e\mathbf{e}: the resulting sentence embedding vector
  • nn: the total length of the tokenized sequence
  • wiw_i: the weight assigned to position ii (common choices include wi=iw_i = i for linear weighting or wi=i2w_i = i^2 for quadratic weighting)
  • mim_i: the attention mask value at position ii
  • hi\mathbf{h}_i: the hidden state vector at position ii

The intuition for this weighting is that in some models or some types of text, information is not uniformly distributed across positions. For instance, the predicate of a sentence (often appearing in the middle or near the end) might be more semantically informative than the subject (often appearing near the beginning). In a bidirectional encoder where every token attends to every other token, this intuition is weaker than it might seem, because the model can freely route information to any position. However, in practice, positional weighting can provide a small boost in some settings.

In production models, weighted mean pooling offers marginal improvements and is not widely adopted. The added complexity of choosing and justifying a weighting scheme rarely justifies the modest gains, and standard mean pooling remains the preferred default. We mention it here because you may encounter it in research papers or as a hyperparameter option in some libraries.

Max Pooling

Max pooling takes the element-wise maximum across all token positions, selecting the strongest activation for each dimension independently:

ej=max⁡i∈Thi,je_j = \max_{i \in \mathcal{T}} h_{i,j}

where:

  • eje_j: the value of the jj-th dimension in the final embedding vector
  • T\mathcal{T}: the set of non-padding token positions
  • hi,jh_{i,j}: the scalar value of the jj-th dimension of the hidden state at token position ii

To understand what this formula does, consider a single dimension jj of the embedding. Across all token positions in the input, each token's hidden state contributes a scalar value for this dimension. Max pooling selects the largest of these values. The resulting embedding vector is assembled dimension by dimension, where each dimension's value comes from whichever token produced the strongest activation along that particular feature axis. Different dimensions can be satisfied by different tokens: dimension 47 might take its value from token position 3, while dimension 312 takes its value from token position 11.

Max pooling captures the presence of a feature anywhere in the sequence, regardless of how infrequently it appears. If a particular dimension activates strongly when the model detects a mention of a geographic location, max pooling will preserve that signal even if the location is mentioned only once in a long passage. Mean pooling would dilute this signal by averaging with all the other tokens where that dimension has lower activation. In image processing, max pooling over spatial regions is a standard technique for detecting the presence of local features regardless of position, and the intuition is similar for text.

However, max pooling is sensitive to outlier activations: a single anomalous token can dominate the embedding along many dimensions simultaneously, creating a vector that reflects the most extreme aspects of the input rather than its typical character. It is also more sensitive to padding-related artifacts, which is why the mask-based exclusion of padding tokens is particularly important. Max pooling is rarely used as the sole pooling strategy in modern production models, but it is sometimes combined with mean pooling in ensemble approaches.

Last Token Pooling

For decoder-only models like GPT or LLaMA that use causal attention, the information flow through the transformer has a fundamentally different directional character than in bidirectional encoders. In causal attention, each token can only attend to itself and to tokens that precede it in the sequence. The first token sees only itself; the second token sees the first and itself; and each successive token has access to a progressively larger window of context. By the time we reach the last token, it has attended to every previous token in the sequence, which makes it the only position that has incorporated information from the complete input:

e=hn\mathbf{e} = \mathbf{h}_n

where:

  • e\mathbf{e}: the resulting sentence embedding vector
  • hn\mathbf{h}_n: the hidden state vector at the last non-padding token position nn

This is the analog of CLS pooling for causal models, but the reasoning for using the last position rather than the first is almost exactly reversed. In a bidirectional encoder, the [CLS] token at the beginning can attend to all other tokens through bidirectional self-attention, which makes it a well-informed aggregation point. In a causal decoder, the same logic applies to the last token: it is the last position, not the first, that has access to the full sequence. Using the first token for a causal model would be deeply counterproductive, because the first token has attended only to itself and has no information about what follows.

Models like E5-Mistral, SFR-Embedding, and GTE-Qwen2 use this strategy. Some models append a special end-of-sequence token or a dedicated aggregation token (such as </s> or [EOS]) to each input and use its representation. This special token acts as a dedicated aggregation point, analogous to BERT's [CLS], but positioned at the end where it can attend to the full prefix. Adding a specific token for this purpose gives the model a consistent, well-defined position to learn to concentrate embedding information, rather than relying on the last token of the input text, which varies in identity across different inputs.

The practical consequence of last token pooling is that decoder-based embedding models often benefit from appending a short instruction or query prefix to the input text, because the final position then reflects the task context as well as the text content. This is why models like E5-Mistral expect inputs formatted as "Instruct: Retrieve relevant passages. Query: {query text}" or similar patterns.

Comparing Pooling Strategies

The choice of pooling strategy interacts strongly with the model architecture and training procedure, and this interaction is one of the most important things to understand about embedding models. A pooling strategy cannot be evaluated in isolation, because its effectiveness depends entirely on how the model was designed and trained to use it. A CLS token from a model trained with mean pooling is not the same kind of object as a CLS token from a model trained to use CLS pooling: in the former case, the training never rewarded the CLS position for putting semantic information, so it may contain arbitrary or residual information rather than a meaningful summary.

For encoder models (BERT, RoBERTa, DeBERTa), mean pooling generally wins over CLS pooling when fine-tuned with contrastive objectives, because the training signal distributes more evenly across token positions. CLS pooling is competitive after proper fine-tuning where the loss specifically rewards the CLS representation for semantic discrimination. For decoder-only models (Mistral, LLaMA, Qwen), last token pooling is the natural and appropriate choice due to causal attention masking: it is the only position with sufficient context to serve as a standalone summary. For encoder-decoder models (T5, BART), the encoder's output is typically mean-pooled, using the same reasoning that favors mean pooling in encoder-only models, since the encoder in these architectures uses bidirectional attention.

The practical guideline is simple: always use the pooling strategy specified in the model's documentation or that was used during its training. Mixing pooling strategies produces degraded embeddings not because one strategy is universally better, but because the model's weights encode expectations about how the pooling layer will read out its hidden states. Violating those expectations corrupts the information being read out.

Let's implement the main pooling strategies and see how they differ in practice. We'll encode three sentences through the same model and apply each pooling method to the raw hidden states, then compare the resulting cosine similarity matrices.

In[3]:
Code
!uv pip install transformers sentence-transformers matplotlib numpy torch einops

import torch
import torch.nn.functional as F
import matplotlib.pyplot as plt
import numpy as np
from transformers import AutoModel, AutoTokenizer

model_name = "sentence-transformers/all-MiniLM-L6-v2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
model.eval()

sentences = [
    "The cat sat on the mat.",
    "A feline rested on the rug.",
    "Stock prices surged on Monday."
]

We encode these sentences and extract the raw hidden states before applying different pooling strategies. By working with the raw outputs rather than the model's built-in encoding pipeline, we can apply each pooling strategy independently to the same hidden states. This keeps a fair comparison.

In[4]:
Code
encoded = tokenizer(
    sentences, padding=True, truncation=True, return_tensors="pt"
)

with torch.no_grad():
    outputs = model(**encoded)

token_embeddings = outputs.last_hidden_state
attention_mask = encoded["attention_mask"]
In[5]:
Code
def cls_pooling(token_embeddings):
    return token_embeddings[:, 0, :]


def mean_pooling(token_embeddings, attention_mask):
    mask_expanded = attention_mask.unsqueeze(-1).float()
    sum_embeddings = (token_embeddings * mask_expanded).sum(dim=1)
    sum_mask = mask_expanded.sum(dim=1).clamp(min=1e-9)
    return sum_embeddings / sum_mask


def max_pooling(token_embeddings, attention_mask):
    mask_expanded = attention_mask.unsqueeze(-1).float()
    token_embeddings = token_embeddings.masked_fill(mask_expanded == 0, -1e9)
    return token_embeddings.max(dim=1).values


cls_emb = cls_pooling(token_embeddings)
mean_emb = mean_pooling(token_embeddings, attention_mask)
max_emb = max_pooling(token_embeddings.clone(), attention_mask)

Now let's compare the cosine similarity matrices produced by each strategy. Cosine similarity measures the angle between two vectors, with values near 1.0 showing that the vectors point in nearly the same direction (high semantic similarity) and values near 0 showing that the vectors are roughly orthogonal (low semantic similarity). A good pooling strategy should produce high similarity between the two animal-on-surface sentences and lower similarity between either of those and the stock market sentence.

In[6]:
Code
def cosine_sim_matrix(embeddings):
    normed = F.normalize(embeddings, p=2, dim=1)
    return (normed @ normed.T).numpy()


cls_sims = cosine_sim_matrix(cls_emb)
mean_sims = cosine_sim_matrix(mean_emb)
max_sims = cosine_sim_matrix(max_emb)
Out[7]:
Visualization
Heatmap of cosine similarity matrix using CLS pooling for three sentences.
CLS pooling cosine similarity matrix for three sentences. The model was trained with mean pooling, so CLS representations are less well-organized for semantic comparison, creating less clear separation between the similar pair and the dissimilar sentence.
Heatmap of cosine similarity matrix using mean pooling for three sentences.
Mean pooling cosine similarity matrix for the same three sentences. Because this model was trained with mean pooling, it correctly shows high similarity between the cat/mat and feline/rug sentences and lower similarity with the stock prices sentence.
Heatmap of cosine similarity matrix using max pooling for three sentences.
Max pooling cosine similarity matrix for the same three sentences. Max pooling captures the strongest per-dimension activations but produces less calibrated semantic distances than mean pooling for a model trained with mean pooling.

The heatmaps reveal the differences between pooling strategies. This particular model (all-MiniLM-L6-v2) was trained with mean pooling, so that strategy produces the most semantically meaningful similarities: the two animal-on-surface sentences are highly similar, while the stock market sentence is more distant. CLS and max pooling, applied to a model trained for mean pooling, produce less differentiated similarity scores. The practical point is that the pooling strategy must match what the model was trained with. Using CLS pooling on a model fine-tuned with mean pooling, or vice versa, will degrade quality in ways that may not be immediately obvious but will show up as poorer retrieval performance in your RAG pipeline.

Key Parameters

The key parameters for the tokenizer used in this example are:

  • padding: Adds special padding tokens to ensure all sequences in a batch have the same length, letting batch matrix operations.
  • truncation: Truncates sequences that exceed the model's maximum context length to prevent out-of-range errors.
  • return_tensors: Specifies the return type ("pt" for PyTorch) for compatibility with the model.

Worked Example: Mean Pooling Step by Step

To make mean pooling concrete, let's walk through a numerical example with a toy model. Suppose we have a toy transformer with hidden dimension d=4d = 4, and we encode the three-token sequence ["The", "cat", "meows"] with one padding token added to match a batch length of 4.

After the final transformer layer, suppose the hidden states are as follows (each row is one token's hidden state, each column is one of the four dimensions):

H=(0.8−0.20.50.1−0.30.90.4−0.60.60.1−0.70.30.00.00.00.0)\mathbf{H} = \begin{pmatrix} 0.8 & -0.2 & 0.5 & 0.1 \\ -0.3 & 0.9 & 0.4 & -0.6 \\ 0.6 & 0.1 & -0.7 & 0.3 \\ 0.0 & 0.0 & 0.0 & 0.0 \end{pmatrix}

where the first row is the hidden state for "The", the second for "cat", the third for "meows", and the fourth for the padding token. The attention mask is m=[1,1,1,0]\mathbf{m} = [1, 1, 1, 0], showing that the first three tokens are real and the fourth is padding.

Step 1: Apply the mask. Multiply each hidden state row by its mask value:

m1⋅h1=1⋅(0.8,−0.2,0.5,0.1)=(0.8,−0.2,0.5,0.1)m_1 \cdot \mathbf{h}_1 = 1 \cdot (0.8, -0.2, 0.5, 0.1) = (0.8, -0.2, 0.5, 0.1) m2⋅h2=1⋅(−0.3,0.9,0.4,−0.6)=(−0.3,0.9,0.4,−0.6)m_2 \cdot \mathbf{h}_2 = 1 \cdot (-0.3, 0.9, 0.4, -0.6) = (-0.3, 0.9, 0.4, -0.6) m3⋅h3=1⋅(0.6,0.1,−0.7,0.3)=(0.6,0.1,−0.7,0.3)m_3 \cdot \mathbf{h}_3 = 1 \cdot (0.6, 0.1, -0.7, 0.3) = (0.6, 0.1, -0.7, 0.3) m4⋅h4=0⋅(0.0,0.0,0.0,0.0)=(0.0,0.0,0.0,0.0)m_4 \cdot \mathbf{h}_4 = 0 \cdot (0.0, 0.0, 0.0, 0.0) = (0.0, 0.0, 0.0, 0.0)

Step 2: Sum the masked hidden states across all token positions, dimension by dimension:

∑i=14mi⋅hi=(0.8+(−0.3)+0.6+0.0,  −0.2+0.9+0.1+0.0,  0.5+0.4+(−0.7)+0.0,  0.1+(−0.6)+0.3+0.0)\sum_{i=1}^{4} m_i \cdot \mathbf{h}_i = (0.8 + (-0.3) + 0.6 + 0.0,\; -0.2 + 0.9 + 0.1 + 0.0,\; 0.5 + 0.4 + (-0.7) + 0.0,\; 0.1 + (-0.6) + 0.3 + 0.0) =(1.1,  0.8,  0.2,  −0.2)= (1.1,\; 0.8,\; 0.2,\; -0.2)

Step 3: Divide by the number of real tokens. The sum of the mask values is ∑i=14mi=1+1+1+0=3\sum_{i=1}^{4} m_i = 1 + 1 + 1 + 0 = 3, so we divide each dimension by 3:

e=13(1.1,  0.8,  0.2,  −0.2)=(0.367,  0.267,  0.067,  −0.067)\mathbf{e} = \frac{1}{3}(1.1,\; 0.8,\; 0.2,\; -0.2) = (0.367,\; 0.267,\; 0.067,\; -0.067)

Step 4: L2-normalize. Compute the Euclidean norm of this vector:

∥e∥2=0.3672+0.2672+0.0672+(−0.067)2\|\mathbf{e}\|_2 = \sqrt{0.367^2 + 0.267^2 + 0.067^2 + (-0.067)^2} =0.1347+0.0713+0.0045+0.0045=0.2150≈0.464\begin{aligned} &= \sqrt{0.1347 + 0.0713 + 0.0045 + 0.0045} \\ &= \sqrt{0.2150} \\ &\approx 0.464 \end{aligned}

Then normalize:

e^=e∥e∥2=(0.367,0.267,0.067,−0.067)0.464≈(0.791,0.575,0.144,−0.144)\hat{\mathbf{e}} = \frac{\mathbf{e}}{\|\mathbf{e}\|_2} = \frac{(0.367, 0.267, 0.067, -0.067)}{0.464} \approx (0.791, 0.575, 0.144, -0.144)

The key insight from this example is that mean pooling is a simple weighted average where the weights come from the attention mask, and the normalization step ensures that the output vector lies on the unit sphere. The entire computation is differentiable with respect to the hidden states H\mathbf{H}, which means gradients can flow back through the pooling operation during training, teaching the transformer to produce hidden states that aggregate well when averaged.

Embedding Dimensions

The dimensionality of the output embedding, the length of the vector returned by the model, is a important parameter that affects retrieval quality, storage costs, and computation speed. Every embedding model produces vectors of a fixed size, and this size determines how much information the model can pack into each representation. To understand why dimensionality matters, think of each dimension as an independent axis along which the model can distinguish between texts. More dimensions mean more axes of distinction, giving the model a richer vocabulary of features to describe semantic content. But more dimensions also mean larger vectors to store, more computation to compare them, and more memory to load them into RAM.

The challenge, then, is to choose a dimensionality that is large enough to capture the semantic distinctions your application requires, without incurring unnecessary computational and storage costs. There is a well-established empirical relationship between dimensionality and quality: below a certain threshold, quality degrades sharply because the model lacks the representational capacity to encode fine-grained distinctions. Above a certain threshold, quality plateaus and the marginal benefit of additional dimensions is small. The location of this threshold depends on the complexity of your domain, the diversity of your corpus, and the model's architecture and training.

Common Dimensionalities

Embedding dimensions across popular models span a wide range, which reflects the diversity of model architectures and intended use cases:

Common embedding model dimensions and parameter counts.
ModelDimensionsParameters
all-MiniLM-L6-v238422M
bge-base-en-v1.5768109M
e5-large-v21024335M
gte-Qwen2-1.5B-instruct15361.5B
E5-Mistral-7B-instruct40967B

The dimension is typically determined by the hidden size of the underlying transformer. BERT-base has a hidden size of 768, so BERT-based embedding models naturally produce 768-dimensional vectors. Smaller models like MiniLM distill into a 384-dimensional space, while larger models based on LLMs produce vectors with thousands of dimensions. Notice the general trend: as models grow in parameter count, their embedding dimensionality tends to increase. This correlation is not coincidental. Larger models have wider hidden layers, which means more dimensions are available at the output, and training objectives can exploit this additional capacity to encode finer-grained distinctions.

The Dimension Trade-Off

Higher dimensions offer more representational capacity. Each dimension is a feature axis that the model can use to encode some aspect of meaning. With more dimensions, the model can represent finer-grained distinctions between texts that differ in subtle ways, such as two scientific passages about related but distinct phenomena. The extra dimensions give the model room to encode features that smaller spaces cannot accommodate without interference between distinct features.

But higher dimensions come with real costs. Storage costs scale linearly: each vector requires d×4d \times 4 bytes in float32 (or d×2d \times 2 bytes in float16). For 10 million documents at 1024 dimensions in float32, that's roughly 40 GB of vector storage, before accounting for any index overhead. At 4096 dimensions (E5-Mistral), that grows to 160 GB. Computation costs also scale linearly: cosine similarity between two vectors requires O(d)O(d) operations, so doubling the dimension doubles the similarity computation time. For brute-force search over nn vectors, the total cost is O(nd)O(nd), so both nn and dd directly multiply the search time. Approximate nearest neighbor indices (which the upcoming chapters on HNSW and IVF will cover) also scale with dimensionality, requiring more memory and longer index-build times for higher-dimensional vectors.

The practical sweet spot for most applications is 384 to 1024 dimensions. Below 256, quality degrades noticeably for complex semantic tasks because the model simply does not have enough dimensions to express the nuances that differentiate closely related but distinct meanings. Above 1024, the gains are marginal for most general-purpose retrieval scenarios, though specific domains or multilingual settings may benefit from higher dimensions. Multilingual models must simultaneously represent the semantic spaces of many different languages in the same vector space, and this additional complexity can benefit from the extra capacity that higher dimensionality provides.

Matryoshka Representation Learning

A clever technique called Matryoshka Representation Learning (MRL), introduced by Kusupati et al. (2022), trains embeddings so that any prefix of the vector is a valid, useful embedding. The name comes from Russian nesting dolls (matryoshka): just as each doll contains a smaller doll inside, a 1024-dimensional Matryoshka embedding contains a useful 512-dimensional embedding in its first 512 components, a useful 256-dimensional embedding in its first 256 components, a useful 128-dimensional embedding in its first 128 components, and so on. This is a remarkable property because it decouples the choice of dimensionality from the choice of model. Without Matryoshka training, if you want embeddings of a different size, you need a different model entirely. With Matryoshka training, a single model can serve multiple dimensionality requirements simply by truncating its output vectors at different lengths.

The training procedure is elegant. During training, the loss is computed on the full-dimensional embedding and on truncated versions at multiple target sizes. The model is simultaneously asked to produce good embeddings at every target dimensionality, which means the learning signal flows back through the network from multiple truncation points at once:

LMRL=∑d∈DLcontrastive(e[:d])\mathcal{L}_{\text{MRL}} = \sum_{d \in \mathcal{D}} \mathcal{L}_{\text{contrastive}}(\mathbf{e}_{[:d]})

where:

  • LMRL\mathcal{L}_{\text{MRL}}: the total Matryoshka Representation Learning loss
  • D\mathcal{D}: a set of target output dimensions (e.g., {64,128,256,512,1024}\{64, 128, 256, 512, 1024\})
  • dd: a specific dimension size from the set D\mathcal{D}
  • e[:d]\mathbf{e}_{[:d]}: the embedding vector truncated to keep only the first dd dimensions
  • Lcontrastive\mathcal{L}_{\text{contrastive}}: the contrastive loss function applied to the truncated vectors

To understand what this loss function achieves, consider what happens at each truncation point. For a given target dimension dd, the model takes only the first dd components of its output vector and computes the contrastive loss using those components alone. This loss penalizes the model if the truncated vectors do not preserve the correct ranking of similar and dissimilar pairs. By summing these penalties across all target dimensions in D\mathcal{D}, the total loss forces the model to ensure that every prefix, from the shortest to the longest, produces embeddings that respect semantic relationships.

This loss has a fascinating effect on what the model learns: it forces the model to pack the most important information into the earlier dimensions. The first 64 dimensions must capture the broadest semantic content, the coarsest features that distinguish major topics and categories. The next 64 dimensions refine this picture, adding details that separate more closely related texts. Each subsequent block of dimensions adds progressively finer-grained distinctions, so that the full-dimensional vector captures the most fine-grained semantic information the model can represent. At inference time, you can truncate embeddings to any prefix length, trading some retrieval quality for storage and computation savings without retraining.

Think of Matryoshka embeddings as a compressed image in a hierarchical format. The first few dimensions give you a coarse, low-resolution picture of the document's meaning. Each additional block of dimensions adds detail, sharpening the picture. You can stop decoding at any resolution level depending on how much precision you need for your current task.

Many modern embedding models support Matryoshka embeddings, including nomic-embed-text-v1.5, models in the gte family, and OpenAI's text-embedding-3 series. Let's see this in action by encoding a query and three documents at multiple dimensionalities and observing how retrieval quality changes as we increase the number of dimensions.

In[8]:
Code
from sentence_transformers import SentenceTransformer

# nomic-embed-text-v1.5 supports Matryoshka dimensions
matryoshka_model = SentenceTransformer(
    "nomic-ai/nomic-embed-text-v1.5", trust_remote_code=True
)

query = "search_query: What causes the northern lights?"
docs = [
    "search_document: The aurora borealis is caused by charged particles from the sun interacting with Earth's magnetic field.",
    "search_document: The northern lights are visible in high-latitude regions near the Arctic.",
    "search_document: LED lights are more energy efficient than incandescent bulbs.",
]

full_dim = 768
dims_to_test = [64, 128, 256, 512, full_dim]
In[9]:
Code
all_texts = [query] + docs
full_embeddings = matryoshka_model.encode(all_texts, convert_to_tensor=True)

results = {}
for d in dims_to_test:
    truncated = full_embeddings[:, :d]
    truncated_norm = F.normalize(truncated, p=2, dim=1)
    query_emb = truncated_norm[0:1]
    doc_embs = truncated_norm[1:]
    sims = (query_emb @ doc_embs.T).squeeze().cpu().numpy()
    results[d] = sims
Out[10]:
Visualization
Grouped bar chart showing similarity scores for three documents across five embedding dimensions from 64 to 768.
Cosine similarity scores for three documents against a query about the northern lights, shown across five embedding dimensionalities from 64 to 768 dimensions. The Matryoshka-trained model maintains the correct ranking order even at 64 dimensions, with the aurora/magnetic field document consistently scoring highest. The margin of separation between the relevant and irrelevant documents increases with dimensionality, illustrating the quality-efficiency trade-off that Matryoshka training enables.

The results demonstrate the Matryoshka property nicely. Even at just 64 dimensions, the model correctly identifies the aurora/magnetic field document as most relevant. As we increase dimensions, the separation between relevant and irrelevant documents improves, and the absolute similarity scores become more calibrated. For a production system, this means you could use 128 or 256 dimensions for a fast initial retrieval pass over a large corpus, then use the full 768 dimensions for a more precise re-scoring of the top candidates, all from the same embedding model and without any retraining.

Key Parameters

The key parameters used in this implementation are:

  • trust_remote_code: Required for models with custom architectures (like Nomic's) that use code not yet merged into the standard Transformers library. Always review the model's repository before letting this option.
  • convert_to_tensor: Instructs the model to return PyTorch tensors instead of NumPy arrays, letting efficient GPU-accelerated operations.

Embedding Model Selection

Choosing the right embedding model for your RAG pipeline is one of the highest-impact decisions you will make. A good model makes retrieval accurate; a poor choice creates a performance ceiling that no amount of downstream engineering can overcome. The decision involves multiple competing considerations: retrieval quality, inference cost, context length, language coverage, and operational simplicity.

The embedding model selection problem is analogous to choosing a foundation for a building. The foundation determines what structures are possible above it. A weak foundation limits everything built on top, no matter how carefully the upper floors are constructed. An overly expensive foundation might provide unnecessary strength at a cost that makes the whole project impractical. You want the foundation that is exactly right for the structure you intend to build, evaluated in the context of your actual requirements.

The MTEB Benchmark

The Massive Text Embedding Benchmark (MTEB) by Muennighoff et al. (2023) is the standard benchmark for evaluating embedding models. It covers a wide range of tasks across multiple domains and languages. This provides a complete picture of model capabilities. The tasks it evaluates include the following:

  • Retrieval: Finding relevant documents given a query, which is the most directly relevant task for RAG applications.
  • Semantic Textual Similarity (STS): Scoring how similar two sentences are, which tests whether the model's similarity scores are well-calibrated.
  • Classification: Using embeddings as features for text classification, which tests whether the embedding captures task-relevant semantic content.
  • Clustering: Grouping similar documents together, which tests the global structure of the embedding space.
  • Pair Classification: Determining if two texts are paraphrases, entailments, or contradictions.
  • Reranking: Reordering candidate documents by relevance to a query.
  • Summarization: Evaluating summary quality via embedding similarity.

The MTEB leaderboard (available at huggingface.co/spaces/mteb/leaderboard) ranks models across these tasks and provides aggregated scores. For RAG applications, the retrieval score matters most because it most directly reflects how well the model will perform at finding relevant passages. Good performance on STS and clustering suggests the model has learned a well-organized semantic space overall, which is a positive indicator even for retrieval tasks not directly evaluated in MTEB.

However, you should treat MTEB scores as a starting point, not a definitive answer. MTEB is computed on general-purpose benchmarks in a handful of domains. If your corpus contains specialized content such as legal contracts, medical literature, scientific papers, or software documentation, the benchmark scores may not accurately predict performance on your data. The model that ranks first on MTEB might rank second or third on your specific domain if the benchmark does not cover similar content.

Selection Criteria

When choosing an embedding model for a RAG system, consider these factors in roughly this order of importance, adjusting based on your specific constraints and requirements.

Retrieval quality on your domain is the primary criterion. MTEB scores are computed on general-purpose benchmarks. If you're working in a specialized domain such as legal, medical, or scientific content, benchmark scores may not reflect actual performance. Always evaluate candidate models on a sample of your own data before committing to a model in production. Even a small evaluation set of 50 to 100 query-document pairs representative of your use case will give you far more reliable signal than public benchmarks alone.

Language support is non-negotiable if your corpus includes non-English content. You need a multilingual model such as multilingual-e5-large, bge-m3, or paraphrase-multilingual-mpnet-base-v2. English-only models will fail silently on non-English text: they will produce embeddings, but these embeddings will not be comparable across languages in a useful way, and retrieval quality will degrade substantially. If your corpus mixes languages, even a small fraction of non-English content can create confusing retrieval behavior if the model was not trained multilingually.

Context length determines the maximum size of a document chunk the model can process. Standard BERT-based models handle 512 tokens. Many modern models support 8,192 tokens (like nomic-embed-text-v1.5 and gte-large-en-v1.5) or even longer. Your chunk sizes from the document chunking stage should fit comfortably within the model's context window. If your chunks exceed the model's limit, the text will be silently truncated, losing information without any warning or error.

Embedding dimension and inference cost matter enormously at scale. For a corpus of 10 million documents, the difference between a 22M-parameter model and a 7B-parameter model is enormous in compute cost. Consider whether the quality improvement justifies the cost for your application, and whether you have the infrastructure to support larger models at your required throughput.

Matryoshka support gives you flexibility to trade quality for speed at serving time. If storage costs are a concern, or if you want to use a two-stage retrieval approach, models with Matryoshka embeddings give you this option without retraining.

Instruction support allows task-aware models to improve retrieval quality by conditioning the embedding on the retrieval intent. This is especially useful when the same corpus is used for different types of queries, or when the gap between query style and document style is large.

Practical Model Tiers

Based on the trade-offs above, embedding models fall into three practical tiers for RAG applications.

Lightweight models (under 50M parameters, 384 dimensions) include models like all-MiniLM-L6-v2. They offer fast inference and small vector sizes, making them suitable for prototyping, low-resource deployments, or applications where latency matters more than peak accuracy. Encoding millions of documents is fast and cheap, and you can iterate quickly on other parts of your pipeline. The trade-off is lower retrieval quality, particularly on specialized domains or subtle semantic distinctions.

Mid-range models (100M to 350M parameters, 768 to 1024 dimensions) include models like bge-base-en-v1.5, gte-large-en-v1.5, and e5-large-v2. These represent the best quality-cost trade-off for most production systems. They offer strong retrieval quality on general-purpose and many specialized domains, while remaining affordable to run on GPUs or even CPUs for moderate-scale corpora.

Heavy models (1B or more parameters, 1536 to 4096 dimensions) include models like gte-Qwen2-1.5B-instruct and E5-Mistral-7B-instruct. These push the quality frontier but require GPU inference and produce large vectors. They are appropriate when retrieval quality takes priority over compute cost, or when serving a smaller, high-value corpus.

Out[11]:
Visualization
Scatter plot on a logarithmic x-axis showing approximate MTEB retrieval scores versus model parameter counts for five embedding models, with colored regions marking the Lightweight, Mid-range, and Heavy tiers.
Trade-off between model parameter count and approximate retrieval quality (MTEB retrieval score) for five representative embedding models. Larger models in the Heavy tier achieve higher retrieval scores but incur significantly greater computational costs, while Lightweight models offer lower latency with reduced accuracy. The logarithmic x-axis highlights the orders-of-magnitude difference in parameter count between tiers.

Evaluating on Your Own Data

Let's walk through a practical evaluation workflow that you can adapt for your own use case. We'll compare two models on a small retrieval task using Mean Reciprocal Rank (MRR) and Recall at kk as our metrics. These two metrics capture complementary aspects of retrieval quality and together give a complete picture of model performance.

In[12]:
Code
from sentence_transformers import SentenceTransformer

models_to_compare = {
    "MiniLM-L6 (22M, d=384)": SentenceTransformer(
        "sentence-transformers/all-MiniLM-L6-v2"
    ),
    "BGE-base (109M, d=768)": SentenceTransformer("BAAI/bge-base-en-v1.5"),
}
In[13]:
Code
# Simple evaluation dataset: queries with relevant document indices
queries = [
    "How does photosynthesis work?",
    "What is the capital of Japan?",
    "Explain the theory of relativity",
]

corpus = [
    "Photosynthesis converts sunlight into chemical energy in plants using chlorophyll.",
    "Tokyo is the capital city of Japan, located on the eastern coast of Honshu.",
    "Einstein's theory of relativity describes the relationship between space, time, and gravity.",
    "The stock market experienced significant volatility last quarter.",
    "Machine learning algorithms learn patterns from training data.",
    "Chloroplasts contain chlorophyll, which absorbs light during photosynthesis.",
    "Japan's largest city by population is Tokyo, which serves as its capital.",
    "Special relativity shows that the speed of light is constant for all observers.",
]

# Ground truth: indices of relevant documents for each query
relevant_docs = [
    [0, 5],  # photosynthesis docs
    [1, 6],  # Japan capital docs
    [2, 7],  # relativity docs
]
In[14]:
Code
def evaluate_retrieval(model, queries, corpus, relevant_docs, k=3):
    query_embs = model.encode(queries, normalize_embeddings=True)
    corpus_embs = model.encode(corpus, normalize_embeddings=True)

    sim_matrix = query_embs @ corpus_embs.T

    mrr_scores = []
    recall_at_k = []

    for i, (sims, rels) in enumerate(zip(sim_matrix, relevant_docs)):
        ranked_indices = np.argsort(-sims)

        # MRR: reciprocal rank of first relevant doc
        for rank, idx in enumerate(ranked_indices, 1):
            if idx in rels:
                mrr_scores.append(1.0 / rank)
                break

        # Recall@k: fraction of relevant docs in top-k
        top_k = set(ranked_indices[:k].tolist())
        recall = len(top_k.intersection(set(rels))) / len(rels)
        recall_at_k.append(recall)

    return {
        "MRR": np.mean(mrr_scores),
        f"Recall@{k}": np.mean(recall_at_k),
    }


eval_results = {}
for name, m in models_to_compare.items():
    eval_results[name] = evaluate_retrieval(
        m, queries, corpus, relevant_docs, k=3
    )
Out[15]:
Console

MiniLM-L6 (22M, d=384):
  MRR: 1.0000
  Recall@3: 1.0000

BGE-base (109M, d=768):
  MRR: 1.0000
  Recall@3: 1.0000

The high scores in this example confirm that both models successfully retrieve the relevant documents for these clear, well-formed queries. In practice, you want at least 50 to 100 query-document pairs representative of your actual use case, including difficult queries where similar but incorrect documents might be retrieved, queries that use different vocabulary than the documents, and queries that require understanding of domain-specific terminology.

MRR (Mean Reciprocal Rank) measures how high the first relevant document appears in the ranked list. An MRR of 1.0 means the correct document is always ranked first; an MRR of 0.5 means it appears on average at position 2. MRR is most informative when you retrieve only a small number of candidates (for example, if you pass only the top-3 results to the LLM). Recall@k measures what fraction of all relevant documents appear in the top-kk results. If a query has two relevant documents and both appear in the top 3, Recall@3 is 1.0 for that query. Recall@k is most informative when you want to ensure complete coverage: if you miss any relevant document, the LLM cannot use that information to answer the question.

Key Parameters

The key parameters for SentenceTransformer models are:

  • model_name_or_path: The Hugging Face model ID (e.g., "BAAI/bge-base-en-v1.5") or local path to the model directory.
  • normalize_embeddings: If True during encoding, normalizes output vectors to unit length, letting cosine similarity computation via a simple dot product.

API-Based Embedding Models

Not all embedding models need to run locally. Several providers offer embedding models as API services, which can be convenient for prototyping or for small-scale applications where managing GPU infrastructure is not practical.

OpenAI offers text-embedding-3-small (1536 dimensions) and text-embedding-3-large (3072 dimensions), both supporting Matryoshka-style dimension reduction via the dimensions parameter. Cohere offers embed-english-v3.0 and multilingual variants, with distinct input type parameters (search_query vs. search_document) that function similarly to instruction-conditioned models. Google provides Vertex AI text embeddings optimized for Google Cloud deployments. Voyage AI offers domain-specific models specialized for code, legal text, and financial documents.

API models trade control and cost predictability for operational convenience. For production RAG systems with large corpora, the per-token cost of API embeddings adds up quickly: at typical API pricing, embedding a 10-million-document corpus can cost hundreds of dollars or more, and re-embedding when the model changes costs the same again. Open-source models you run locally avoid these risks, though they introduce operational complexity in exchange.

A practical concern with API embedding models is vendor lock-in. Your entire vector index is tied to a specific model version from a specific provider. If the provider deprecates the model, changes pricing substantially, or experiences an outage, your entire retrieval system is affected. This has happened before with older OpenAI embedding model versions, requiring teams to re-embed their entire corpora. Open-source models you run locally avoid these risks entirely.

Normalization and Post-Processing

After pooling, most embedding models apply one final transformation before returning the output vector: L2 normalization. This step projects the embedding onto the surface of a unit hypersphere. This keeps every vector has exactly the same Euclidean length of 1, regardless of the input text. The formula for this normalization divides the raw embedding by its Euclidean length:

e^=e∥e∥2\hat{\mathbf{e}} = \frac{\mathbf{e}}{\|\mathbf{e}\|_2}

where:

  • e^\hat{\mathbf{e}}: the L2-normalized embedding vector, which has unit length
  • e\mathbf{e}: the raw embedding vector output by the pooling layer
  • ∥e∥2\|\mathbf{e}\|_2: the L2 norm (Euclidean length) of vector e\mathbf{e}, computed as ∑jej2\sqrt{\sum_j e_j^2}

To see why this matters, recall that cosine similarity between two vectors measures the cosine of the angle between them. The formula for cosine similarity is:

cos⁡(a,b)=a⋅b∥a∥2∥b∥2\cos(\mathbf{a}, \mathbf{b}) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\|_2 \|\mathbf{b}\|_2}

where the denominator normalizes for the lengths of both vectors. Without normalization, two embedding vectors might differ in length for reasons unrelated to semantic content. Two texts of very different lengths might produce embeddings of very different magnitudes simply because the mean over more tokens amplifies or dilutes activations in certain ways. These length differences would contaminate similarity computations by conflating directional similarity (which reflects semantic relatedness) with magnitude differences (which do not).

By normalizing all vectors to unit length, we eliminate magnitude as a factor. When ∥e^∥2=1\|\hat{\mathbf{e}}\|_2 = 1 for all embeddings, the cosine similarity formula simplifies to a plain dot product:

cos⁡(e^a,e^b)=e^a⋅e^b\cos(\hat{\mathbf{e}}_a, \hat{\mathbf{e}}_b) = \hat{\mathbf{e}}_a \cdot \hat{\mathbf{e}}_b

where:

  • e^a,e^b\hat{\mathbf{e}}_a, \hat{\mathbf{e}}_b: the normalized embedding vectors for two texts aa and bb
  • ⋅\cdot: the dot product operation

This equivalence holds because when both norms are exactly 1, the denominator of the cosine similarity formula becomes 1×1=11 \times 1 = 1, leaving just the numerator (the dot product). This simplification is computationally valuable because dot products are faster to compute than cosine similarity (no normalization step at search time), and are well-supported by vector index implementations like FAISS. Once you normalize embeddings at encoding time, you can use the simpler and faster dot product throughout your retrieval pipeline without any loss of accuracy.

Out[16]:
Visualization
Diagram of a unit circle with two raw embedding vectors and their L2-normalized projections, illustrating that normalization equalizes vector lengths while preserving direction.
Geometric interpretation of L2 normalization projecting two-dimensional vectors onto the unit circle. Raw embedding vectors (faded arrows) have different magnitudes but the same directions as their normalized counterparts (solid arrows), which lie exactly on the unit circle. This transformation ensures that cosine similarity depends solely on the angle between vectors, removing the influence of vector magnitude on similarity scores.

Some models include a learned linear projection layer after pooling, mapping the hidden dimension to a different output dimension. For example, some models might project from 768 hidden dimensions down to 256 output dimensions to produce more compact embeddings suitable for memory-constrained deployments. These projection layers are learned during training, so the model learns which linear combinations of hidden dimensions are most informative for the target embedding size. Unlike Matryoshka training, a fixed projection layer commits you to a single output size at inference time: you cannot truncate the projected vector at an intermediate length and expect the result to be useful.

Encoding Long Documents

A practical challenge for embedding models is handling text that exceeds the model's context window. Transformers have a fixed maximum sequence length, determined by the positional encoding scheme used during training and the memory allocated to the attention matrices. Any input that exceeds this limit must be handled explicitly.

As we discussed in the document chunking chapter, the standard approach is to split documents into chunks that fit within the model's limit. Each chunk is encoded independently, creating one embedding per chunk. At retrieval time, the query is compared against all chunk embeddings, and the most relevant chunks are retrieved. This approach has several advantages: it is simple, it is compatible with any embedding model regardless of context length, and it produces fine-grained retrieval (the specific relevant section is retrieved, not the entire document). The disadvantage is that information spanning chunk boundaries is split across two embeddings, and inter-chunk context is lost.

There are alternative strategies worth understanding, each with distinct trade-offs. Chunked mean pooling encodes each chunk separately and then averages the resulting embeddings into a single document embedding. This is simple and captures some sense of the document's overall content, but it loses the fine-grained resolution needed for question-specific retrieval. If you're looking for the answer to a specific question, averaging embeddings from all sections of a 50-page document will bury the relevant section's signal in the noise from all the irrelevant sections.

Hierarchical encoding produces chunk-level embeddings, then uses a separate aggregation model to combine them into a document embedding. This can potentially capture inter-chunk relationships if the aggregation model has sufficient expressive power, but it adds a second model to the pipeline, increases complexity, and the aggregation model itself requires careful design and training.

Long-context embedding models offer a third path. Newer models like jina-embeddings-v2-base-en (8,192 tokens) and nomic-embed-text-v1.5 (8,192 tokens) natively handle sequences many times longer than the original BERT limit of 512 tokens. These models use modified positional encoding schemes, such as ALiBi or extended RoPE contexts, that generalize to longer sequences than those seen during training.

For RAG applications, per-chunk embeddings are almost always preferable to document-level embeddings, regardless of whether the embedding model supports long contexts. When you ask a specific question, you want to retrieve the specific chunk that answers it, not an entire document that might bury the relevant passage among thousands of irrelevant tokens. The granularity of retrieval directly determines the quality of the context you provide to the LLM: a 500-token chunk containing exactly the relevant information is far more useful than a 10,000-token document where that information appears on page 8.

Limitations and Practical Considerations

Embedding models, despite their impressive capabilities, carry important limitations that every practitioner building RAG systems must understand. Recognizing these limitations allows you to design systems that compensate for them, to interpret retrieval failures correctly, and to set realistic expectations for system performance.

The most basic limitation is the information bottleneck. Compressing a passage of potentially hundreds of tokens into a single fixed-length vector inevitably loses information. A 768-dimensional vector, rich as it is, cannot encode every nuance of a 500-token passage. Important details that are not prominent in the aggregate representation, such as a specific date mentioned once in the middle of a paragraph, may not survive the compression process strongly enough to drive retrieval when a query asks specifically for that information. This is why cross-encoder rerankers, which perform full token-level cross-attention between the query and each candidate document, can substantially improve accuracy after first-stage embedding retrieval: they can recover fine-grained signals that were compressed away by the pooling operation. The two-stage architecture of coarse embedding retrieval followed by cross-encoder reranking is specifically designed to work around this bottleneck, and we will examine it in detail when we cover reranking later in this part.

Another significant limitation is domain mismatch. General-purpose embedding models are trained primarily on web text, Wikipedia, news articles, and common NLP benchmarks. When deployed on specialized domains such as medical literature, legal contracts, financial reports, or software codebases, their performance can drop substantially. The embedding space may not have learned the fine-grained distinctions that matter in your domain. A general model might place "myocardial infarction" and "cardiac arrest" close together in the embedding space because they are both heart-related medical conditions with overlapping vocabulary, even though they are clinically distinct conditions requiring different treatments and showing different pathologies. A medical embedding model trained on clinical text would learn to distinguish these concepts because the training data contains many examples where their differences matter. Domain-specific fine-tuning with contrastive learning on in-domain pairs can address this limitation, but it requires curated labeled data and training infrastructure.

There is also the issue of embedding drift over time. As better models become available or as your domain evolves, you may want to upgrade your embedding model. But all embeddings in your vector index must come from the same model: you cannot mix embeddings from different models because they lie in fundamentally different vector spaces where proximity has no consistent meaning across models. Upgrading your embedding model requires re-embedding the entire corpus and rebuilding the vector index from scratch. For large corpora, this can be a multi-day operation that consumes significant compute and storage and must be coordinated carefully to avoid downtime. This operational cost creates a form of technical debt: once you build a vector index with a particular model, switching models is expensive, which can deter you from adopting newer, better models.

Finally, embedding models inherit the biases present in their training data. If certain topics, perspectives, writing styles, or demographic groups are underrepresented in training, the embedding space will be poorly calibrated for inputs from those underrepresented categories. Queries from users asking about underrepresented topics, or queries phrased in language styles not well-represented in training data, may receive systematically worse retrieval results. This creates both fairness and reliability problems for applications serving diverse users or topics. Evaluating your embedding model across diverse query types and user populations, not just on the most common or representative cases, is important for understanding where these blind spots lie.

Summary

This chapter covered the architecture and design decisions that make embedding models work for retrieval in RAG systems:

  • Embedding models are transformer encoders fine-tuned with contrastive objectives to produce semantically meaningful sentence and passage vectors. The key challenge they solve is distilling a sequence of token-level hidden states into a single vector that captures the overall meaning of the text. Standard pre-trained models like BERT do not do this well without further fine-tuning.

  • Pooling strategies determine how token embeddings are compressed into a single vector. Mean pooling is the most common and effective strategy for encoder models, computing a mask-weighted average of all token hidden states. CLS pooling uses only the first token's hidden state, which is effective after proper contrastive fine-tuning. Last token pooling is the natural choice for decoder-based embedding models where causal attention makes the last position the best-informed one. The pooling strategy must match what the model was trained with.

  • The worked example showed that mean pooling is a simple masked average followed by L2 normalization. The entire operation is differentiable, allowing training to shape the hidden states so they aggregate into a useful representation. The normalization step places all embeddings on the unit hypersphere, simplifying cosine similarity to a dot product at retrieval time.

  • Embedding dimensionality involves a trade-off between representational capacity and computational cost. Most production systems use 384 to 1024 dimensions. Matryoshka representation learning enables flexible dimensionality from a single model by training the loss jointly at multiple truncation lengths, forcing the model to pack the most important semantic content into the earliest dimensions.

  • Model selection should be guided by retrieval performance on your specific data, language requirements, context length needs, and inference cost constraints. MTEB benchmarks provide a useful starting point but should always be supplemented by evaluation on representative samples of your actual domain.

  • Normalization to unit length simplifies cosine similarity to a dot product, letting faster search and compatibility with standard vector index implementations. The geometric intuition is that normalization projects all embeddings onto the surface of a unit hypersphere, so similarity is purely a function of the angle between vectors.

  • Embedding models have important limitations: the information bottleneck from pooling, domain mismatch for specialized content, embedding drift when upgrading models, and inherited biases from training data. Designing RAG systems that account for these limitations, particularly through reranking and domain-specific evaluation, is needed for production deployments.

With embeddings in hand, the next step in a RAG pipeline is storing them in a way that enables fast similarity search. The upcoming chapters on vector similarity search, HNSW, and IVF indices will show how to search over millions of vectors efficiently without resorting to brute-force comparison.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about embedding models.

Embedding Models Knowledge Check

Question 1 of 70 of 7 completed
Why are bi-encoder architectures preferred over cross-encoders for the initial retrieval step in large-scale RAG pipelines?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026embeddingmodels, author = {Michael Brenndoerfer}, title = {Embedding Models: Architecture, Pooling & Selection}, year = {2026}, url = {https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Embedding Models: Architecture, Pooling & Selection. Retrieved from https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection
MLAAcademic
Michael Brenndoerfer. "Embedding Models: Architecture, Pooling & Selection." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection>.
CHICAGOAcademic
Michael Brenndoerfer. "Embedding Models: Architecture, Pooling & Selection." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Embedding Models: Architecture, Pooling & Selection'. Available at: https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Embedding Models: Architecture, Pooling & Selection. https://mbrenndoerfer.com/writing/embedding-models-architecture-pooling-selection

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.