Dense Retrieval: Semantic Search & Bi-Encoder Implementation

Michael BrenndoerferJanuary 22, 202662 min read

Part of Language AI Handbook

Covers dense retrieval for semantic search. Examines bi-encoder architectures, embedding metrics, and contrastive learning to overcome keyword limitations.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Dense Retrieval: Semantic Search and Bi-Encoders

In the previous chapter, we examined the RAG architecture and how it combines retrieval with generation. At the heart of any RAG system lies a necessary question: how do we find the most relevant documents for a given query? The answer to that question determines whether the LLM receives the right context to produce an accurate, grounded response, or whether it receives noise that misleads it. Traditional search engines rely on lexical matching, counting how often query terms appear in documents and scoring results by term frequency statistics. This approach has served us well for decades and remains the backbone of most search infrastructure today.

But lexical matching has a basic blind spot: it operates entirely on surface form. If the document uses different words to express the same concept, a lexical retriever will miss it entirely. What happens when a patient asks "heart attack symptoms" but the medical textbook discusses "myocardial infarction warning signs"? What happens when a software engineer searches for "how to speed up my code" but the documentation describes "performance optimization techniques"? The words are completely different, yet the meaning is identical. A retriever that counts term overlap would score these relevant documents at zero, returning results that match the words but not the intent.

Dense retrieval addresses this basic limitation by representing both queries and documents as continuous vectors in a shared semantic space. Rather than matching exact words, dense retrieval measures the similarity between the meanings of queries and documents. A query about "climate change impacts" can match documents discussing "global warming effects" because both map to nearby points in the embedding space, even though they share no common terms. Think of this as a map of meaning: semantically similar concepts are placed close together, and retrieval becomes a nearest-neighbor search in that space.

This shift from discrete term matching to continuous semantic similarity is one of the most significant advances in information retrieval in the past decade. It moves retrieval from pattern matching toward language understanding. Building on the transformer architectures and embedding techniques we have explored throughout this book, dense retrieval enables systems that compare meaning rather than merely counting words. The bi-encoder architecture at the heart of dense retrieval allows us to encode entire corpora offline, then retrieve relevant passages at query time with a single vector similarity computation.

Understanding dense retrieval deeply matters because it forms the foundation of modern RAG pipelines, semantic search products, question answering systems, and knowledge-intensive language tasks. Every time you receive an accurate, knowledge-grounded response from a production LLM assistant, there is almost certainly a dense retrieval system working behind the scenes. In this chapter, we build that system from the ground up: starting from the vocabulary mismatch problem that motivates it, through the bi-encoder architecture that makes it practical, into the similarity metrics and training objectives that make it accurate, and finally through a complete implementation you can adapt for your own applications.

Historical Context

The origins of dense retrieval trace back to the neural information retrieval research of the 2010s, when researchers began exploring whether distributed word representations like Word2Vec could improve document ranking. Early efforts, including the Paragraph Vector model (2014) and various latent semantic analysis approaches, showed that dense representations could capture some semantic relationships. However, these early models were shallow and struggled to match the quality of BM25 on benchmark retrieval tasks.

The breakthrough came in 2019 when dense retrieval based on fine-tuned BERT encoders began outperforming classical methods on open-domain question answering benchmarks. The Dense Passage Retrieval (DPR) paper by Karpukhin et al. (2020) demonstrated that BERT-based bi-encoders, trained with contrastive learning on natural question-answer pairs, could substantially outperform BM25 on the Natural Questions and TriviaQA benchmarks. This result was surprising to many researchers because it contradicted the conventional wisdom that BM25 was difficult to beat with neural methods. The key insight was that simple embedding approaches based on generic pre-trained models were insufficient; models needed domain-specific fine-tuning with carefully constructed negatives to learn meaningful retrieval representations. Since DPR, the field has produced a rapid succession of improvements including improved negative mining strategies, stronger pre-training objectives, and more efficient architectures, all of which we examine in this chapter.

From Lexical to Semantic Matching

Recall from Part II that BM25 retrieves documents by computing term-frequency statistics: documents score higher when they contain more query terms, with diminishing returns for repeated terms and penalties for common words. This approach works remarkably well for navigational queries, entity lookups, and any search where the vocabulary of the query and the relevant documents naturally aligns. The design is elegant in its simplicity: store an inverted index mapping each term to the documents that contain it, and at query time, look up each query term and aggregate the scores.

The vocabulary mismatch problem becomes acute in several important scenarios that arise constantly in real applications. Consider searching a medical knowledge base with the query "heart attack symptoms." A relevant document might discuss "myocardial infarction warning signs" without ever using the words "heart" or "attack." BM25 would score this document at zero because it shares no terms with the query. Yet a physician would immediately recognize these as describing the same condition. The mismatch here is between everyday language and technical terminology, and it is systematic: patients speak in plain language while medical literature uses precise clinical terms.

The vocabulary gap manifests in many forms, and each form defeats lexical retrieval in a different way:

  • Synonyms and paraphrases: "car" vs "automobile," "purchase" vs "buy," "use" vs "utilize." Natural language has extraordinary breadth of expression for any given concept.
  • Technical versus casual language: "myocardial infarction" vs "heart attack," "hypertension" vs "high blood pressure." Users often prefer informal phrasing while documents use formal or specialized terms.
  • Abbreviations and expansions: "ML" vs "machine learning," "NLP" vs "natural language processing." Abbreviations are common in casual queries but documents may spell things out.
  • Conceptual similarity without lexical overlap: "renewable energy policy" might relate to documents about "solar panel subsidies" or "wind farm regulation," even though these share no query terms.
  • Different grammatical forms and derivations: "optimize" vs "optimization," "predict" vs "prediction." Stemming helps partially, but doesn't solve cross-form semantic relationships.

Dense retrieval sidesteps vocabulary mismatch entirely by operating in a different representational space. Instead of asking "which words match?", it asks "which meanings match?" By encoding both queries and documents into a continuous vector space where semantically similar texts cluster together, dense retrieval can find relevant documents regardless of the specific words they use.

The key insight is that transformer models, pre-trained on enormous amounts of text, develop internal representations that capture meaning at a level far beyond surface form. The model has seen "myocardial infarction" and "heart attack" used in similar contexts thousands of times, so their representations end up in similar regions of the embedding space. Dense retrieval uses this learned semantic geometry. When we ask the model to encode a query, it produces a vector that points in a direction in semantic space that represents the query's meaning. Document vectors that point in similar directions represent documents about related concepts.

This is a fundamentally different philosophy from lexical retrieval. In lexical retrieval, the representation is determined entirely by the words: the same words always produce the same representation, regardless of context. In dense retrieval, the representation is learned from data and captures distributional semantic relationships. Words that appear in similar contexts across millions of documents end up with similar representations, creating an implicit semantic knowledge graph embedded in the geometry of the vector space.

The Bi-Encoder Architecture

The dominant architecture for dense retrieval is the bi-encoder, which uses two separate encoder networks to produce embeddings for queries and documents independently. This separation is important for efficiency at scale, and understanding why requires us to think carefully about what happens during retrieval.

Think of a bi-encoder as two photographers: the query encoder takes a photograph of the question, and the document encoder takes a photograph of each document. The photographs are all printed in the same format, so you can compare them side by side using a simple similarity measure. Because the document photographs can be taken before any query arrives (offline), you only need to take the query photograph at search time and then compare it against the pre-existing document photographs. This is what makes the bi-encoder scalable: the expensive work of encoding documents happens once, not once per query.

Bi-Encoder

A neural architecture that encodes queries and documents using separate (but often identical) transformer encoders, creating fixed-dimensional vectors that can be compared using simple similarity metrics. The independence of query and document encoding enables pre-computation of document embeddings, making retrieval over large corpora computationally feasible.

Architecture Overview

The bi-encoder consists of two components that work in tandem to transform text into comparable numerical representations. The first component is the query encoder, denoted EqE_q, which takes a natural language query qq as input and produces a dense vector representation. The second component is the document encoder, denoted EdE_d, which performs the analogous transformation for documents.

Formally, these transformations are:

  • Query encoder EqE_q: Maps a query qq to a dense vector q=Eq(q)∈RD\mathbf{q} = E_q(q) \in \mathbb{R}^D
  • Document encoder EdE_d: Maps a document dd to a dense vector d=Ed(d)∈RD\mathbf{d} = E_d(d) \in \mathbb{R}^D

Both encoders are typically initialized from the same pre-trained transformer (such as BERT, as we covered in Part XXIV), and they may share weights or be fine-tuned independently. The encoders produce embeddings of the same dimensionality DD, letting direct comparison. This shared dimensionality is needed because it allows us to measure distances and angles between query and document vectors in the same geometric space. Without a shared dimensionality, there would be no principled way to compare a query vector to a document vector.

In practice, the query and document encoders are often initialized identically and may even share weights during training. Weight sharing reduces the total parameter count and can improve training stability. Some architectures allow the query and document encoders to diverge during fine-tuning, which can be beneficial when the distributional characteristics of queries and documents are very different. Short, keyword-like queries and long, paragraph-length documents are stylistically distinct, and separate encoders give the model flexibility to specialize. For many practical applications, however, a single shared encoder works well and is simpler to maintain.

A key challenge in building these encoders is converting the variable-length output of a transformer into a single fixed-size vector suitable for comparison. Recall from our BERT chapters that when a BERT model processes an input sequence, it produces a hidden state vector for each token in the sequence. A sentence of 20 words produces 20 hidden state vectors (plus special tokens), while a paragraph of 200 words produces 200 hidden state vectors. We need to compress this variable-length sequence into a single fixed-size vector, a process called pooling.

BERT-based encoders typically extract the representation of the special [CLS] token from the final layer:

q=BERT(q)[CLS]\mathbf{q} = \text{BERT}(q)_{[\text{CLS}]}

where:

  • q\mathbf{q} is the dense vector representation of the query
  • BERT(q)\text{BERT}(q) is the sequence of hidden states from the last layer of the BERT model
  • [CLS][\text{CLS}] indicates extraction of the vector corresponding to the special classification token, which is designed to aggregate sequence-level information during pre-training

The [CLS] token was specifically designed during BERT's pre-training to aggregate information from the entire sequence. During the next-sentence prediction pre-training task, the [CLS] representation is used to classify whether two sentences are consecutive, forcing it to encode sequence-level rather than token-level information. This makes it a natural candidate for creating a single vector that summarizes the entire input.

Alternatively, to capture information distributed across the entire sequence rather than relying on a single token, some models compute the average of all token representations. This approach, known as mean pooling, treats each token's contribution equally:

q=1n∑i=1nBERT(q)i\mathbf{q} = \frac{1}{n} \sum_{i=1}^{n} \text{BERT}(q)_i

where:

  • q\mathbf{q} is the dense vector representation of the query
  • nn is the number of tokens in the query
  • BERT(q)i\text{BERT}(q)_i is the vector representation of the ii-th token
  • ∑\sum computes the element-wise sum of all token vectors, and dividing by nn finds the geometric centroid

The intuition behind mean pooling is that important semantic information may be spread across multiple tokens rather than concentrated in the [CLS] token. By averaging, we create a representation that balances contributions from all parts of the input. Consider a sentence like "The quick brown fox jumped over the lazy dog." The [CLS] representation tries to summarize this whole sentence, but mean pooling creates a vector that represents the average meaning of all the individual tokens. For retrieval purposes, mean pooling often works slightly better because it ensures no important content is lost.

In practice, the choice between [CLS] pooling and mean pooling affects retrieval quality, and different models adopt different strategies based on their training objectives. Empirically, models trained with mean pooling as part of their objective tend to perform better when evaluated using mean pooling, and likewise for [CLS] pooling. The Sentence-BERT family of models (which we use in our implementation) were specifically trained with mean pooling and perform substantially better with it.

Why Separate Encoders?

The bi-encoder's separation of query and document encoding enables a necessary optimization: pre-computation of document embeddings. To appreciate why this matters, consider the computational demands of a retrieval system at real scale.

In a production retrieval system serving a corpus of 10 million documents, we need to compare every query against every potential document to find the most relevant ones. If we process each query-document pair jointly, a single query requires 10 million transformer forward passes. A transformer forward pass on a modern BERT-sized model takes approximately 10 milliseconds on a GPU. This means a single query would take 100,000 seconds, or about 27 hours. That is obviously not a retrieval system; it is a computational nightmare.

With the bi-encoder, we pre-compute document embeddings once. At query time, we run the transformer once for the query (taking 10 milliseconds), then compute dot products between the query vector and all 10 million pre-computed document vectors. Computing 10 million dot products in a 768-dimensional space takes microseconds using vectorized operations. The total query latency drops from 27 hours to roughly 50 milliseconds. This is the bi-encoder's core value proposition: it separates the expensive transformer computation (done once per document, offline) from the cheap similarity computation (done at query time, over pre-computed vectors).

This contrasts with cross-encoders, which concatenate the query and document and process them jointly through a single transformer. Cross-encoders can capture fine-grained interactions between query and document tokens. The attention mechanism allows every query token to attend to every document token, letting the model to reason about how specific parts of the document relate to specific parts of the query. This richer interaction produces more accurate relevance judgments in controlled experiments.

However, cross-encoders require running the transformer once for every query-document pair. For a corpus of 10 million documents, this means 10 million transformer forward passes per query, which is computationally infeasible for first-stage retrieval. Cross-encoders therefore serve a different role in the retrieval pipeline: they are used as rerankers that refine the top candidates returned by a faster first-stage retriever. We will explore reranking with cross-encoders in a later chapter.

The bi-encoder architecture trades some modeling power for massive efficiency gains:

Comparison of Bi-Encoder and Cross-Encoder architectures.
AspectBi-EncoderCross-Encoder
Query-time computationEncode query onceEncode query-doc pairs
Document pre-computationYes (offline)No
Query-document interactionNone (independent encoding)Full attention
Typical useFirst-stage retrievalReranking
ScalabilityMillions of documentsHundreds of candidates

We will explore reranking with cross-encoders in a later chapter, where they serve as a second-stage refinement over bi-encoder candidates. The combination of bi-encoder retrieval followed by cross-encoder reranking is a powerful and widely used two-stage pipeline.

Pooling Strategies in Practice

Beyond [CLS] and mean pooling, several other pooling strategies have been explored. Max pooling takes the maximum value in each dimension across all token vectors, which can emphasize strong signals from individual tokens. Weighted mean pooling assigns importance weights to tokens before averaging, potentially downweighting stopwords and upweighting content words. Some specialized architectures produce multiple vectors per document (one per sentence or paragraph) to support more granular retrieval.

The Sentence Transformers library, which we use in our code implementation, uses mean pooling combined with optional L2 normalization by default. This combination has proven reliable across many retrieval benchmarks. The normalization step ensures that embedding magnitudes are uniform, which simplifies the similarity computation and prevents magnitude biases from distorting retrieval results.

One important consideration is the maximum sequence length. Transformer models have a fixed context window, typically 512 tokens for BERT-based models. Documents longer than this limit must be truncated or split into multiple chunks. In RAG systems, documents are usually pre-chunked into passages of a few hundred words specifically to fit within these limits. Retrieval quality depends on both chunk size and overlap, including whether chunks respect sentence boundaries. Longer chunks preserve more context but may dilute relevance signals; shorter chunks are more focused but may miss context needed to answer the query.

Embedding Similarity Metrics

Once we have query and document embeddings, we need a similarity function to rank documents. The three most common metrics are dot product, cosine similarity, and Euclidean distance. Each of these metrics captures a different notion of what it means for two vectors to be "similar," and understanding their geometric interpretations helps us choose the right metric for a given application.

Think of the embedding space as a high-dimensional room where every piece of text has a specific location. Similarity metrics are different ways to measure how "close" two locations are. The dot product measures closeness while also caring about how far each vector is from the origin. Cosine similarity measures only the angle between vectors, ignoring how far they are from the origin. Euclidean distance measures the straight-line distance between the two locations in the room. Each notion of closeness has its uses, and the right choice depends on how the embeddings were trained and what properties you want retrieval to have.

Dot Product

The dot product quantifies the similarity between two vectors by aggregating their aligned components. To understand this intuitively, imagine two vectors as arrows pointing in some direction in high-dimensional space. The dot product measures how much one vector "projects" onto another, combining information about both their directions and their lengths. For vectors q\mathbf{q} and d\mathbf{d}, the dot product is defined as:

sim(q,d)=q⋅d=∑i=1Dqidi\text{sim}(\mathbf{q}, \mathbf{d}) = \mathbf{q} \cdot \mathbf{d} = \sum_{i=1}^{D} q_i d_i

where:

  • sim(q,d)\text{sim}(\mathbf{q}, \mathbf{d}) is the similarity score
  • q,d\mathbf{q}, \mathbf{d} are the dense vectors for the query and document
  • DD is the dimensionality of the embedding space (e.g., 768)
  • qi,diq_i, d_i are the ii-th scalar components of each vector
  • ∑\sum aggregates alignment across all DD dimensions

The geometric interpretation of this formula is instructive. Each dimension of the embedding space captures some aspect of meaning. When both qiq_i and did_i are large and positive, their product contributes positively to the similarity, showing that both the query and document exhibit that particular semantic feature strongly. Conversely, when one component is positive and the other is negative, their product is negative, reducing the overall similarity score. The dot product is therefore highest when the two vectors point in the same direction and have large magnitudes.

The dot product is computationally efficient and captures both the alignment of vectors (their angular similarity) and their magnitudes. Larger embeddings produce larger dot products, which can be useful when magnitude carries semantic meaning. For example, longer, more detailed documents might have larger embedding norms, and this property could be desirable if we want to favor complete, information-rich documents over brief ones.

Many dense retrieval models, including the original DPR (Dense Passage Retrieval), use dot product similarity because it can be computed extremely efficiently using matrix multiplication. Given a query vector and a matrix of document vectors, all similarity scores can be computed in a single matrix-vector multiplication using highly optimized BLAS libraries that run at near-peak hardware throughput.

Cosine Similarity

While the dot product captures both direction and magnitude, there are situations where we want to focus purely on semantic direction, ignoring how "long" the vectors are. Cosine similarity addresses this need by normalizing the dot product by the magnitudes of both vectors, measuring only the angular alignment:

cos⁡(q,d)=q⋅d∥q∥∥d∥=∑i=1Dqidi∑i=1Dqi2∑i=1Ddi2\begin{aligned} \cos(\mathbf{q}, \mathbf{d}) &= \frac{\mathbf{q} \cdot \mathbf{d}}{\|\mathbf{q}\| \|\mathbf{d}\|} \\ &= \frac{\sum_{i=1}^{D} q_i d_i}{\sqrt{\sum_{i=1}^{D} q_i^2} \sqrt{\sum_{i=1}^{D} d_i^2}} \end{aligned}

where:

  • cos⁡(q,d)\cos(\mathbf{q}, \mathbf{d}) is the cosine similarity score, ranging from −1-1 to +1+1
  • q⋅d\mathbf{q} \cdot \mathbf{d} is the unnormalized dot product (raw alignment)
  • ∥q∥,∥d∥\|\mathbf{q}\|, \|\mathbf{d}\| are the Euclidean norms (lengths) of the vectors
  • ∑qi2\sqrt{\sum q_i^2} computes the vector's magnitude as the square root of the sum of squared components

The cosine similarity ranges from −1-1 (vectors pointing in opposite directions) to +1+1 (vectors pointing in the same direction), with 00 showing orthogonality. By removing magnitude effects, cosine similarity focuses purely on semantic direction. This is particularly useful when we want to compare texts of different lengths on an equal footing, since a longer document might naturally produce a larger-magnitude embedding without being more relevant to any particular query.

The key insight is that by normalizing, we make the similarity score a function only of the angle between the vectors. Two documents that say exactly the same thing but one is much longer (perhaps with more examples and explanations) should have similar cosine similarity to a query, even if the longer document's raw embedding vector is larger in magnitude. This length-invariant behavior is often what we want in retrieval.

A useful equivalence emerges when embeddings are L2-normalized, meaning each vector is divided by its magnitude so that ∥q∥=∥d∥=1\|\mathbf{q}\| = \|\mathbf{d}\| = 1. In this case, cosine similarity equals the dot product:

cos⁡(q^,d^)=q^⋅d^\cos(\hat{\mathbf{q}}, \hat{\mathbf{d}}) = \hat{\mathbf{q}} \cdot \hat{\mathbf{d}}

Here, q^\hat{\mathbf{q}} and d^\hat{\mathbf{d}} are the unit-length vectors derived by dividing q\mathbf{q} and d\mathbf{d} by their respective norms. The dot product operation, applied to these normalized vectors, directly yields the cosine of the angle between the original vectors.

This equivalence is practically very useful: we can pre-normalize all document embeddings once during indexing, then use the faster dot product operation at query time while still computing cosine similarity. This is exactly what many production systems do, including most models in the Sentence Transformers library. Pre-normalizing once means we never need to compute norms at query time, keeping retrieval latency minimal.

Euclidean Distance

The third common metric takes a different perspective entirely. Rather than measuring how vectors align, Euclidean (L2) distance measures the straight-line distance between vectors in the embedding space:

dist(q,d)=∥q−d∥=∑i=1D(qi−di)2\text{dist}(\mathbf{q}, \mathbf{d}) = \|\mathbf{q} - \mathbf{d}\| = \sqrt{\sum_{i=1}^{D} (q_i - d_i)^2}

where:

  • dist(q,d)\text{dist}(\mathbf{q}, \mathbf{d}) is the Euclidean distance (lower means more similar)
  • q−d\mathbf{q} - \mathbf{d} is the difference vector, computed element-wise
  • (qi−di)2(q_i - d_i)^2 is the squared difference in dimension ii, which is always non-negative
  • ⋅\sqrt{\cdot} converts the sum of squared differences to linear distance units

Unlike the similarity metrics above, smaller Euclidean distances indicate greater similarity. This is an important distinction to keep in mind when implementing retrieval systems: with Euclidean distance, we search for the nearest neighbors rather than the highest-scoring documents.

Euclidean distance is related to dot product for normalized vectors. When vectors have unit length, minimizing Euclidean distance is mathematically equivalent to maximizing the dot product. To see why, expand the distance formula:

∥q−d∥2=∥q∥2−2q⋅d+∥d∥2=1−2q⋅d+1(for unit vectors)=2(1−q⋅d)\begin{aligned} \|\mathbf{q} - \mathbf{d}\|^2 &= \|\mathbf{q}\|^2 - 2\mathbf{q} \cdot \mathbf{d} + \|\mathbf{d}\|^2 \\ &= 1 - 2\mathbf{q} \cdot \mathbf{d} + 1 \quad \text{(for unit vectors)} \\ &= 2(1 - \mathbf{q} \cdot \mathbf{d}) \end{aligned}

Since 22 and 11 are constants, minimizing ∥q−d∥2\|\mathbf{q} - \mathbf{d}\|^2 is equivalent to maximizing q⋅d\mathbf{q} \cdot \mathbf{d}, which for normalized vectors equals cosine similarity. This relationship allows certain approximate nearest neighbor indexing algorithms designed for Euclidean distance to be adapted for dot product or cosine similarity retrieval, which is important for efficient large-scale search.

Out[3]:
Visualization
Using Python 3.11.14 environment at: /private/tmp/mb-language-ai-modern-plots/books/_quarto_language-ai-handbook/.venv
Checked 2 packages in 3ms
2D vector space diagram showing three arrows from origin: neutral Query A, green Match B forming small angle theta-1 with A, red Mismatch C forming larger angle theta-2 with A. C is longer but points away from A.
Geometric representation of query and document embeddings in 2D vector space. The query vector (A) forms a smaller angle with the semantically similar document (B) than with the dissimilar document (C), illustrating how angular alignment captures semantic relevance. Document C has a larger magnitude than B, showing why unnormalized dot products can assign higher scores to longer but less relevant documents.
Out[4]:
Visualization
Grouped bar chart comparing cosine similarity and dot product for Match B and Mismatch C documents. Under cosine sim, B scores higher. Under dot product, C scores higher due to length bias.
Comparison of cosine similarity and dot product scores for the Match (B) and Mismatch (C) documents from the previous figure. Cosine similarity correctly ranks Match (B) higher by ignoring vector magnitude. The unnormalized dot product incorrectly favors Mismatch (C) because its larger magnitude inflates the raw score, showing why magnitude normalization is often needed for fair retrieval.

Choosing a Metric

The choice of similarity metric should match the training objective used when the embedding model was created. This is not a trivial constraint: using the wrong metric can substantially degrade retrieval quality even for an otherwise excellent model.

  • Models trained with dot product loss perform best with dot product similarity at inference time
  • Models trained with cosine loss perform best with cosine similarity
  • Many embedding models normalize their outputs, making the choice less necessary in practice

In practice, cosine similarity (or equivalently, dot product with normalized embeddings) is the most common choice because it is reliable to variations in embedding magnitude and provides easily interpretable scores. When similarity ranges from −1-1 to +1+1, it becomes straightforward to set thresholds and compare scores across different queries and different corpora. A score of 0.8 means something consistent regardless of whether the query is long or short, whether the document is a sentence or a paragraph.

One subtle consideration is how the model was trained with respect to embedding norms. Some contrastive learning setups allow the model to encode confidence or importance in the embedding norm, making dot product the preferred metric. Others explicitly normalize embeddings during training, making cosine similarity the correct choice. Always check the documentation for any pre-trained embedding model to understand what metric it expects.

Dense vs Sparse Retrieval: A Detailed Comparison

Understanding when to use dense versus sparse retrieval requires examining their complementary strengths and weaknesses in detail. These are not competing technologies where one is simply better; they are different tools suited to different problems, and the best production systems often combine both.

Think of sparse retrieval as a librarian who finds books by searching the catalog for exact title and keyword matches, and dense retrieval as a librarian who understands the topic you are researching and can recommend relevant books even when they use completely different terminology. The first librarian works quickly and predictably without making creative associations. The second understands your intent but might occasionally recommend something that sounds related but misses the point.

Strengths of Dense Retrieval

Dense retrieval excels in scenarios where semantic understanding matters more than exact term matching:

Semantic matching: Dense retrieval captures meaning beyond surface forms. The query "best laptop for programming" matches documents about "developer-friendly notebooks with good keyboards" even without shared terms. This is the core value proposition: meaning-based matching rather than word-based matching.

Handling synonyms and paraphrases: Medical queries like "hypertension treatment" naturally match documents about "high blood pressure medication" because both concepts map to similar regions in the embedding space. The model has learned from large training data that these terms are used interchangeably in similar contexts.

Cross-lingual retrieval: With multilingual encoders trained on parallel corpora, the same query can retrieve relevant documents in different languages, as semantically equivalent text in different languages clusters together in the embedding space. This is particularly valuable for organizations with multilingual knowledge bases.

Tolerance of typos and minor variations: Minor spelling variations like "recieve" vs "receive" often produce similar embeddings because the transformer tokenizer breaks them into similar subword units, and the model has learned to handle such variations. Exact-match systems would fail entirely on misspelled queries.

Natural language and conversational queries: Queries that read like natural sentences ("What are the side effects of ibuprofen?") are well-handled by dense retrieval because the model understands intent rather than relying on specific keyword matches. This is especially important as voice search and conversational interfaces become more common.

Strengths of Sparse Retrieval

Sparse retrieval (BM25, TF-IDF) maintains important advantages in several areas that should not be dismissed:

Exact match requirements: When you search for specific identifiers like "CVE-2024-1234," "iPhone 15 Pro Max," or "Python 3.11.2," exact term matching is needed. Dense models may confuse similar but distinct identifiers, potentially returning results for "CVE-2024-1235" or "iPhone 15 Pro" when you need an exact match. This distinction can be necessary in security and legal work as well as technical documentation.

Rare terms and proper nouns: Uncommon terms carry high information value in BM25 through high IDF (inverse document frequency) weights. Dense models may underweight rare terms that were not well-represented in training data, potentially missing specialized topics. A query about an obscure historical figure or a niche technical protocol may be better served by a system that knows rare term occurrences are highly informative.

Interpretability and debugging: BM25 scores are explainable. You can tell a user: "This document ranked high because it contains the query terms 'machine' (3 times) and 'learning' (5 times)." Dense similarity scores are opaque black boxes. When retrieval goes wrong, debugging a BM25 system is straightforward; debugging a dense retrieval system requires understanding what the model learned and why it assigned certain similarity scores.

Zero-shot generalization: BM25 works immediately on any corpus without training or fine-tuning. Dense retrievers need training data that matches the target domain. For organizations with specialized text (legal contracts, scientific literature, source code, internal documentation), building training data and fine-tuning models requires significant investment.

Memory efficiency: BM25 uses inverted indices that scale efficiently to billions of documents with compact storage. Dense retrieval requires storing one full embedding vector per document, which at 768 dimensions in float32 requires approximately 3 KB per document. For a corpus of 100 million documents, that is 300 GB of embedding storage before any indexing overhead.

When Each Approach Fails

Dense retrieval struggles predictably with:

  • Keyword-heavy queries: Searching for "Python pandas DataFrame merge" requires exact library and function names. The user knows the specific API they need and does not want semantically related but incorrect functions.
  • Out-of-domain text: Models trained on Wikipedia may perform poorly on legal contracts, clinical notes, or source code because the semantic relationships in those domains differ substantially from general text.
  • Negation: "hotels without pools" might match documents about "hotels with pools" because both contain similar concepts. Dense models struggle with negation because their training rarely includes negative contrastive examples at this semantic level.
  • Numerical and logical constraints: "apartments under $2000/month in Seattle" requires understanding numerical constraints and geographic specificity, not just semantic similarity.

Sparse retrieval struggles predictably with:

  • Vocabulary mismatch: "affordable housing" vs "low-cost apartments" versus "subsidized residence" all mean similar things but share no terms.
  • Paraphrased queries: "what causes climate change" vs "global warming factors" uses different vocabulary for the same question.
  • Conceptual queries: "books like Harry Potter" requires understanding the genre and writing style, including narrative characteristics, rather than just term overlap.
  • Short queries: Single-word or two-word queries provide little context for term weighting. "Python" could mean the programming language, the snake, or the British comedy group.

The Case for Hybrid Approaches

Given these complementary strengths, many production systems combine both approaches. A typical hybrid retrieval strategy runs BM25 and dense retrieval in parallel, creating two ranked lists of candidates, then merges them using reciprocal rank fusion or learned score combination. The merged list captures both lexical precision and semantic recall, performing better than either approach alone on most benchmarks.

The simplest combination method is Reciprocal Rank Fusion (RRF), which scores each document by the sum of the reciprocals of its ranks in each list. If a document appears at rank 3 in the BM25 list and rank 7 in the dense list, its RRF score is 13+60+17+60\frac{1}{3+60} + \frac{1}{7+60} where 60 is a smoothing constant. Documents that appear highly ranked in both lists receive the highest combined scores, while documents unique to one list receive moderate scores based on their rank in that list.

We will explore hybrid search techniques in detail in a later chapter on combining retrieval signals, including how to learn optimal combination weights from data and how to handle score normalization across different retrieval systems.

Training Dense Retrievers

Creating effective dense retrievers requires specialized training to produce embeddings where similar queries and documents cluster together. This section covers the key components of the training process, from the mathematical objective that guides learning to the practical considerations of data collection and negative sampling. Understanding training is important for practitioners who fine-tune their own models and for anyone who wants to understand why off-the-shelf models have certain strengths and weaknesses.

The basic challenge in training a dense retriever is supervision: how do we tell the model what "similar" means? Unlike classification tasks where labels are clear-cut, relevance in retrieval is gradational and context-dependent. A document about photosynthesis is highly relevant to "how do plants make food" but somewhat relevant to "renewable energy sources." The training framework needs to capture this nuance.

The Training Objective

The goal of dense retriever training is to learn an embedding space where relevant query-document pairs have high similarity and irrelevant query-document pairs have low similarity. The key insight is that we can formulate this as a contrastive learning problem: teach the model to distinguish between documents that satisfy the information need and documents that do not.

Given a query qq, a relevant (positive) document d+d^+, and irrelevant (negative) documents {d1−,d2−,...,dk−}\{d^-_1, d^-_2, ..., d^-_k\}, we want:

sim(Eq(q),Ed(d+))>sim(Eq(q),Ed(di−))∀i\text{sim}(E_q(q), E_d(d^+)) > \text{sim}(E_q(q), E_d(d^-_i)) \quad \forall i

where:

  • sim(⋅)\text{sim}(\cdot) is the similarity function (e.g., dot product or cosine)
  • Eq(q),Ed(⋅)E_q(q), E_d(\cdot) are the encoder embeddings for queries and documents respectively
  • d+d^+ is the relevant (positive) document
  • di−d^-_i is the ii-th irrelevant (negative) document
  • ∀i\forall i indicates the condition should hold for all negative samples

The intuition behind this formulation is geometric: we want the query embedding to lie closer to the positive document embedding than to any negative document embedding in the vector space. The training loss operationalizes this geometric constraint.

The most common loss function is the contrastive loss (also called InfoNCE loss or NT-Xent loss). This function treats retrieval as a multi-class classification task over the set of candidates in a training batch. It first computes the probability that the positive document is the correct match among all candidates (positive plus negatives), then minimizes the negative log of that probability.

To build the loss step by step: we start with the query embedding q\mathbf{q} and compute its similarity to the positive document embedding d+\mathbf{d}^+ and each negative embedding di−\mathbf{d}^-_i. We then convert these similarities to a probability distribution using softmax with temperature τ\tau:

P(d+∣q)=exp⁡(sim(q,d+)/τ)exp⁡(sim(q,d+)/τ)+∑i=1kexp⁡(sim(q,di−)/τ)L=−log⁡P(d+∣q)\begin{aligned} P(d^+|q) &= \frac{\exp(\text{sim}(\mathbf{q}, \mathbf{d}^+) / \tau)}{\exp(\text{sim}(\mathbf{q}, \mathbf{d}^+) / \tau) + \sum_{i=1}^{k} \exp(\text{sim}(\mathbf{q}, \mathbf{d}^-_i) / \tau)} \\ \mathcal{L} &= -\log P(d^+|q) \end{aligned}

where:

  • P(d+∣q)P(d^+|q) is the probability that d+d^+ is the correct match for query qq
  • L\mathcal{L} is the contrastive loss we minimize during training
  • q,d+,di−\mathbf{q}, \mathbf{d}^+, \mathbf{d}^-_i are the embeddings for query, positive, and negative documents
  • τ\tau is the temperature parameter that controls the sharpness of the probability distribution
  • kk is the number of negative samples per positive

The numerator exp⁡(sim(q,d+)/τ)\exp(\text{sim}(\mathbf{q}, \mathbf{d}^+) / \tau) is the exponentiated similarity score for the positive pair. The denominator normalizes this by summing the exponentiated scores for the positive document and all kk negative documents. The resulting P(d+∣q)P(d^+|q) is the probability that the model assigns to the positive document being the correct match, given all the candidates.

The loss L=−log⁡P(d+∣q)\mathcal{L} = -\log P(d^+|q) is minimized when P(d+∣q)P(d^+|q) approaches 1, which happens when the positive document's similarity is much higher than all negative documents' similarities. Maximizing this probability is exactly what we want: make the positive pair the most similar, by a large margin.

The temperature parameter τ\tau deserves special attention. When τ\tau is small (close to 0), the exponential function amplifies differences between similarity scores. This makes the model more confident in its distinctions and giving a sharper training signal. When τ\tau is large, the probability distribution becomes more uniform. This provides a weaker gradient signal. Finding the right temperature, typically in the range of 0.05 to 0.1 for dense retrieval, is often done through experimentation, though some training frameworks learn it automatically.

Out[5]:
Visualization
Scatter plot showing a neutral query point at origin, teal positive document nearby, and three coral negative documents scattered further away. Arrows show pull toward positive and push away from negatives. Two dashed circles mark High Similarity Zone.
Visualization of the contrastive learning objective in 2D embedding space. The loss function forces the query (q) and positive document (d+) closer together in the vector space while pushing all negative documents (d-) further away. This push-pull dynamic creates a semantic region around each query where only relevant documents cluster, letting accurate nearest-neighbor retrieval at inference time.

Training Data Sources

Dense retrievers require training data consisting of query-document pairs with relevance labels. The quality and diversity of training data is one of the most important factors determining retrieval performance. Common sources include:

Natural Questions (NQ): Google's dataset of real questions from Google Search users, paired with Wikipedia passages containing answers. This provides a natural distribution of factual queries and their supporting evidence. The queries reflect how real users phrase questions, which is valuable for training a retriever that handles diverse natural language inputs. The limitation is that coverage is restricted to knowledge available on Wikipedia.

MS MARCO: Microsoft's large-scale reading comprehension dataset with approximately 500,000 queries drawn from Bing search logs, each paired with relevant passages from web documents. Its scale and web domain make it the most popular dataset for training general-purpose retrievers. Models trained on MS MARCO generalize reasonably well to many retrieval tasks, which makes it a common pre-training choice before domain-specific fine-tuning.

Synthetic data from LLMs: A recent and increasingly popular approach uses large language models to generate queries for existing documents. Given a passage, an LLM generates realistic questions that the passage would answer. This allows creating training data for any domain without requiring expensive human annotation. The quality of synthetic data has improved substantially as LLMs have become more capable, and recent work shows that models trained on synthetic data can approach the performance of models trained on human-annotated data.

Click logs from production systems: In deployed search or recommendation systems, user clicks provide implicit relevance signals. Documents that users click after submitting a query are treated as positive examples. This provides enormous amounts of training signal at low cost. This approach has been used by large-scale systems like Google, Bing, and Amazon to continuously improve their retrieval models. The challenge is handling noise: users sometimes click irrelevant results out of curiosity or mistake.

Domain-specific curated datasets: For specialized domains like legal, medical, or scientific retrieval, researchers and organizations have created curated datasets. Examples include BEIR (a benchmark with 18 diverse retrieval datasets), BioASQ for biomedical question answering, and MLDR for multilingual dense retrieval. Fine-tuning on these domain-specific datasets is often necessary for high-quality retrieval in specialized applications.

Negative Sampling Strategies

The choice of negative documents significantly impacts training quality, and this is one of the most studied aspects of dense retrieval training. The intuition is simple: the negatives define what "not relevant" means, and the model learns to distinguish positives from negatives. If negatives are too easy (obviously irrelevant), the model learns weak distinctions that do not help at inference time. If negatives are too hard (indistinguishable from positives), the training signal becomes noisy and inconsistent.

Random negatives: Sample random documents from the corpus. This is simple but often too easy. Random documents from a large corpus are typically obviously irrelevant to any given query. A retriever trained only on random negatives learns to distinguish relevant documents from random text, which is easy and does not prepare it for the harder cases where it needs to distinguish between two semantically similar documents, only one of which answers the query.

BM25 negatives: Use BM25 to find documents that contain query terms but are not relevant. These are called "hard negatives" because they share surface features with the query but differ semantically, forcing the model to learn representations that go beyond lexical overlap. For example, for the query "causes of inflation," a BM25 negative might be an article that mentions "inflation" and "causes" frequently but discusses historical inflation rather than economic causes. Training on such negatives teaches the model to distinguish semantic nuances.

In-batch negatives: Use positive documents from other queries in the same training batch as negatives for the current query. This approach is computationally efficient because it requires no additional encoding: the positive document embeddings already computed for other queries can be reused as negatives for the current query. With a batch size of 64, each query has 63 in-batch negatives at no extra encoding cost. The main limitation is that in-batch negatives may be easy (the other queries in the batch are typically about different topics), but with large batch sizes and diverse training data, in-batch negatives can be quite effective.

Self-mined hard negatives: Use the current model to find high-scoring documents that are not labeled as relevant. These are the hardest possible negatives because the current model finds them confusable with true positives. After each training epoch (or after a fixed number of steps), re-rank the corpus using the current model and mine new hard negatives. This iterative mining process progressively makes training harder, keeping the model in a regime where it is always learning something challenging.

Research has consistently shown that combining strategies works best. A typical approach: start training with in-batch negatives plus some BM25 negatives for initial training, then periodically re-mine hard negatives using the current model to refine the embedding space. The original DPR paper used BM25 negatives and in-batch negatives. More recent work like ANCE (Approximate Nearest Neighbor Negative Contrastive Estimation) demonstrates that self-mined hard negatives can substantially improve over BM25 negatives, especially later in training when the model has become good enough that BM25 negatives are no longer challenging.

The Training Pipeline

A typical dense retriever training pipeline proceeds as follows:

  1. Initialize query and document encoders from a pre-trained language model (BERT, RoBERTa, or a domain-specific model)
  2. Collect training data: query-positive document pairs plus negative documents using one of the strategies above
  3. For each training batch, encode all queries and documents using the two encoders
  4. Compute the contrastive loss using the similarity scores between query embeddings and all document embeddings in the batch
  5. Backpropagate gradients through both encoders and update their weights
  6. Periodically evaluate on a held-out validation set to monitor retrieval quality
  7. Optionally, re-mine hard negatives using the current model state after every few thousand steps

The training process is computationally intensive because we need to encode many negatives per query to provide a strong training signal, and BERT-based encoders are large models. Techniques like in-batch negatives help by reusing computation across the batch, but training a high-quality dense retriever still typically requires substantial GPU resources and hours to days of compute time. Smaller, more efficient encoder architectures like MiniLM or DistilBERT can reduce training time substantially while maintaining competitive retrieval quality.

An important practical consideration is gradient flow through both encoders. Both the query encoder and the document encoder receive gradients during training, so both are updated simultaneously. This means the document embeddings change at every training step, which invalidates any cached embeddings. In practice, this is handled by recomputing document embeddings periodically (not after every step) or by freezing one encoder and only training the other, depending on the application.

Worked Example: Semantic Similarity in Action

Let's trace through a concrete numerical example to build intuition for how dense retrieval works from query input to ranked output. This example makes the abstract concepts concrete and shows exactly what happens at each step.

Setup: We have a three-dimensional toy embedding space. Query qq = "How do plants make food?" has been encoded to the vector q=[0.6,0.5,0.6]\mathbf{q} = [0.6, 0.5, 0.6]. We have three candidate documents with pre-computed embeddings:

  • Document A: "Photosynthesis turns light energy into chemical energy in plants." Embedding: dA=[0.7,0.6,0.4]\mathbf{d}_A = [0.7, 0.6, 0.4]
  • Document B: "Plants produce glucose from carbon dioxide and sunlight." Embedding: dB=[0.5,0.7,0.5]\mathbf{d}_B = [0.5, 0.7, 0.5]
  • Document C: "My grandmother makes delicious food using fresh vegetables." Embedding: dC=[0.1,0.2,0.9]\mathbf{d}_C = [0.1, 0.2, 0.9]

Documents A and B are about plant biology and photosynthesis; Document C is about cooking. The embeddings reflect this: A and B have high values in dimensions 1 and 2 (which we can think of as capturing "biology" and "plant science"), while C has a high value in dimension 3 (capturing "food/cooking").

Step 1: Normalize all embeddings. Before computing similarities, we normalize each vector to unit length:

∥q∥=0.62+0.52+0.62=0.36+0.25+0.36=0.97≈0.985\|\mathbf{q}\| = \sqrt{0.6^2 + 0.5^2 + 0.6^2} = \sqrt{0.36 + 0.25 + 0.36} = \sqrt{0.97} \approx 0.985 q^=[0.6,0.5,0.6]0.985≈[0.609,0.508,0.609]\hat{\mathbf{q}} = \frac{[0.6, 0.5, 0.6]}{0.985} \approx [0.609, 0.508, 0.609]

Similarly normalizing: d^A≈[0.706,0.606,0.404]\hat{\mathbf{d}}_A \approx [0.706, 0.606, 0.404], d^B≈[0.488,0.683,0.488]\hat{\mathbf{d}}_B \approx [0.488, 0.683, 0.488], d^C≈[0.104,0.208,0.932]\hat{\mathbf{d}}_C \approx [0.104, 0.208, 0.932].

Step 2: Compute cosine similarities (dot products of normalized vectors).

sim(q^,d^A)=(0.609)(0.706)+(0.508)(0.606)+(0.609)(0.404)=0.430+0.308+0.246=0.984\begin{aligned} \text{sim}(\hat{\mathbf{q}}, \hat{\mathbf{d}}_A) &= (0.609)(0.706) + (0.508)(0.606) + (0.609)(0.404) \\ &= 0.430 + 0.308 + 0.246 = 0.984 \end{aligned} sim(q^,d^B)=(0.609)(0.488)+(0.508)(0.683)+(0.609)(0.488)=0.297+0.347+0.297=0.941\begin{aligned} \text{sim}(\hat{\mathbf{q}}, \hat{\mathbf{d}}_B) &= (0.609)(0.488) + (0.508)(0.683) + (0.609)(0.488) \\ &= 0.297 + 0.347 + 0.297 = 0.941 \end{aligned} sim(q^,d^C)=(0.609)(0.104)+(0.508)(0.208)+(0.609)(0.932)=0.063+0.106+0.568=0.737\begin{aligned} \text{sim}(\hat{\mathbf{q}}, \hat{\mathbf{d}}_C) &= (0.609)(0.104) + (0.508)(0.208) + (0.609)(0.932) \\ &= 0.063 + 0.106 + 0.568 = 0.737 \end{aligned}

Step 3: Rank by similarity. The ranking is: Document A (0.984) > Document B (0.941) > Document C (0.737).

Step 4: Interpret the result. Despite Document C containing the word "food" (which appears in the query "How do plants make food?"), it ranks last. Documents A and B, which describe photosynthesis without using the words "make" or "food," rank first and second. The embedding space has captured the semantic intent of the query: we are asking about plant biology, and the model correctly identifies plant biology documents as most relevant, while recognizing that "making food" in the cooking sense is semantically different.

Now consider what BM25 would do. The query contains terms: "how," "do," "plants," "make," "food." Document A contains "plants" once. Document B contains "Plants" once. Document C contains "food" once and is about cooking with vegetables from a garden (semantically related to "plants" in a different sense). BM25 might rank Document C highest due to the "food" term match, which appears to be a highly specific query term. Document A and B each only match "plants," giving them lower BM25 scores. This illustrates the vocabulary mismatch problem directly: the lexical method is misled by surface-level matches, while the semantic method captures true intent.

The key insight from this numerical example is that dense retrieval's power comes entirely from the quality of the embedding space. The toy embeddings above were constructed to make the example clear, but real embedding spaces are 768-dimensional and learned from billions of text examples. The geometric relationships in those learned spaces capture fine-grained semantic associations that no human could engineer by hand.

Out[6]:
Visualization
Grouped bar chart comparing BM25 (orange) and dense semantic (blue) retrieval scores. Doc A and B receive high dense scores but low BM25 scores. Doc C receives high BM25 but low dense score.
Comparison of lexical BM25 and dense retrieval scores for the query 'How do plants make food?'. The lexical model incorrectly favors the culinary document (C) due to surface-level keyword overlap with 'food'. The dense retriever correctly assigns the highest scores to the scientific documents (A and B) because it understands that the query asks about plant biology, not cooking, showing how semantic matching overcomes vocabulary mismatch.

Code Implementation

Let's implement dense retrieval using the sentence-transformers library, which provides pre-trained bi-encoder models optimized for semantic similarity. The library wraps BERT-based models with sensible defaults for pooling and normalization, which makes it straightforward to get a working retrieval system up and running.

In[7]:
Code
!uv pip install sentence-transformers scikit-learn

import numpy as np
import matplotlib.pyplot as plt
from collections import Counter
from sklearn.decomposition import PCA
from sentence_transformers import SentenceTransformer

# Load a pre-trained bi-encoder model
# This model was trained on over 1 billion sentence pairs
# using contrastive learning on diverse text collections
model = SentenceTransformer('all-MiniLM-L6-v2')

The all-MiniLM-L6-v2 model is a compact but effective retriever, creating 384-dimensional embeddings. It was distilled from a larger model and trained on over 1 billion sentence pairs using contrastive learning, which makes it a strong general-purpose retriever despite its small size. For production RAG systems requiring higher accuracy, larger models like msmarco-distilbert-base-v4 or all-mpnet-base-v2 provide better retrieval quality at higher computational cost.

In[8]:
Code
# Sample document corpus about machine learning
documents = [
    "Neural networks learn hierarchical representations of data through multiple layers.",
    "Gradient descent optimizes model parameters by iteratively updating weights.",
    "Transformers use self-attention to capture dependencies regardless of distance.",
    "Random forests combine multiple decision trees for reliable predictions.",
    "Support vector machines find optimal hyperplanes to separate classes.",
    "Convolutional neural networks excel at processing grid-structured data like images.",
    "Recurrent neural networks maintain hidden states to process sequential data.",
    "The backpropagation algorithm computes gradients through the chain rule.",
]

# Encode all documents (this would be done offline in production)
# normalize_embeddings=True divides each vector by its L2 norm
# letting dot product to equal cosine similarity at query time
document_embeddings = model.encode(
    documents, convert_to_numpy=True, normalize_embeddings=True
)
Out[9]:
Console
Encoded 8 documents
Embedding shape: (8, 384)
Embedding dimension: 384
First embedding norm (should be ~1.0): 1.0000

Each document is now represented as a 384-dimensional unit vector. The pre-normalization confirms that all embeddings have L2 norm of exactly 1.0, meaning we can use dot product as a drop-in replacement for cosine similarity at query time. In a production system, these embeddings would be stored in a vector index for efficient approximate nearest neighbor search over millions of documents.

Computing Similarity Scores

Let's implement a function to retrieve documents for a query:

In[10]:
Code
def dense_retrieve(query, doc_embeddings, docs, model, top_k=3):
    """
    Retrieve top-k documents using dense retrieval.

    Args:
        query: The search query string
        doc_embeddings: Pre-computed and normalized document embeddings
        docs: List of document strings
        model: The sentence transformer model
        top_k: Number of documents to retrieve

    Returns:
        List of (document, score) tuples sorted by descending relevance
    """
    # Encode the query (normalize to unit length for cosine similarity)
    query_embedding = model.encode(
        query, convert_to_numpy=True, normalize_embeddings=True
    )

    # Compute cosine similarity with all documents
    # For normalized embeddings, dot product equals cosine similarity
    # Shape of result: (num_docs,)
    similarities = np.dot(doc_embeddings, query_embedding)

    # Get indices of top-k highest scores (argsort returns ascending,
    # so we reverse with [::-1] and take first top_k)
    top_indices = np.argsort(similarities)[::-1][:top_k]

    # Return documents with their similarity scores
    return [(docs[i], similarities[i]) for i in top_indices]

Let's test the retriever with a query that demonstrates semantic matching across vocabulary:

In[11]:
Code
query = "How do neural networks learn?"
results = dense_retrieve(query, document_embeddings, documents, model, top_k=3)
Out[12]:
Console
Query: How do neural networks learn?

Top retrieved documents:

1. Score: 0.681
   Neural networks learn hierarchical representations of data through multiple layers.

2. Score: 0.532
   The backpropagation algorithm computes gradients through the chain rule.

3. Score: 0.437
   Recurrent neural networks maintain hidden states to process sequential data.

The retriever correctly identifies documents about neural network learning mechanisms: the first result explains hierarchical learning directly, the second covers gradient descent (the mechanism by which neural networks learn), and the third covers backpropagation (how gradients are computed to drive learning). None of these documents contain the exact phrase "neural networks learn" yet they are all semantically relevant. The model understands that learning in neural networks is achieved through gradient descent and backpropagation, not just through documents that use the word "learn."

Comparing Dense and Lexical Retrieval

Let's compare dense retrieval to a simple lexical approach on a query that demonstrates vocabulary mismatch:

In[13]:
Code
def bm25_score(query, doc, avg_doc_len, k1=1.5, b=0.75):
    """Compute simplified BM25 score for a single document."""
    query_terms = query.lower().split()
    doc_terms = doc.lower().split()
    doc_len = len(doc_terms)

    term_freq = Counter(doc_terms)

    score = 0.0
    for term in query_terms:
        if term in term_freq:
            tf = term_freq[term]
            # BM25 term frequency normalization
            numerator = tf * (k1 + 1)
            denominator = tf + k1 * (1 - b + b * (doc_len / avg_doc_len))
            score += numerator / denominator

    return score


def lexical_retrieve(query, docs, top_k=3):
    """Retrieve using simplified BM25."""
    avg_len = np.mean([len(d.split()) for d in docs])
    scores = [(doc, bm25_score(query, doc, avg_len)) for doc in docs]
    scores.sort(key=lambda x: x[1], reverse=True)
    return scores[:top_k]

Now let's test both approaches on a query that uses different terminology than the documents:

In[14]:
Code
# Query that uses different terminology than the documents
# "deep learning optimization methods" uses vocabulary not present in most docs
query_mismatch = "deep learning optimization methods"

dense_results = dense_retrieve(
    query_mismatch, document_embeddings, documents, model, top_k=3
)
lexical_results = lexical_retrieve(query_mismatch, documents, top_k=3)
Out[15]:
Console
Query: 'deep learning optimization methods'

Dense Retrieval Results:
  1. [0.513] Gradient descent optimizes model parameters by iteratively updating we...
  2. [0.379] Neural networks learn hierarchical representations of data through mul...
  3. [0.369] The backpropagation algorithm computes gradients through the chain rul...

Lexical (BM25) Results:
  1. [0.000] Neural networks learn hierarchical representations of data through mul...
  2. [0.000] Gradient descent optimizes model parameters by iteratively updating we...
  3. [0.000] Transformers use self-attention to capture dependencies regardless of ...

Dense retrieval correctly identifies gradient descent (an optimization method), backpropagation (how the optimization is computed), and neural network learning (the context of deep learning), all without those exact terms appearing in the query. The lexical system returns documents with BM25 scores of 0.0 for most candidates because the query terms "deep," "learning," "optimization," and "methods" rarely appear verbatim in the corpus. When results do appear, they are based on coincidental term matches rather than semantic relevance.

Visualizing the Embedding Space

Understanding the geometry of the embedding space helps build intuition for why dense retrieval works. Let's project the high-dimensional embeddings down to two dimensions using PCA to see how documents cluster:

Out[16]:
Visualization
2D scatter plot of eight labeled document points in PCA space. Neural nets, Gradient descent, and Backpropagation cluster near a coral star marking the query point. Transformers, CNNs, RNNs, SVMs, and Random forests occupy other regions.
PCA projection of 384-dimensional document embeddings to 2D space, showing how semantically related documents cluster together. The query 'How do neural networks learn?' (red star) falls near the neural network and optimization documents cluster, confirming that the bi-encoder places queries close to their relevant documents in embedding space. Documents about classical ML methods (SVMs, Random forests) appear in a separate region.

The PCA projection reveals the clustering structure in the embedding space. Documents about neural network training mechanisms (gradient descent, backpropagation, hierarchical representations) cluster together in one region, while classical machine learning methods (SVMs, Random forests) cluster elsewhere, and specialized architectures (CNNs, RNNs, Transformers) form their own neighborhood. The query "How do neural networks learn?" falls naturally near the learning-mechanism cluster, which is exactly where dense retrieval would look for relevant documents.

Batch Retrieval for Efficiency

In production systems, we typically process multiple queries simultaneously. Encoding queries one at a time misses the opportunity to parallelize computation on GPU hardware. Batch encoding is one of the key practical optimizations for real retrieval systems:

In[17]:
Code
def batch_retrieve(queries, doc_embeddings, docs, model, top_k=3):
    """
    Efficiently retrieve for multiple queries using batched encoding
    and matrix multiplication.

    The key efficiency: encoding all queries simultaneously on GPU
    is much faster than encoding them sequentially.
    """
    # Encode all queries at once (enables GPU parallelism)
    query_embeddings = model.encode(
        queries, convert_to_numpy=True, normalize_embeddings=True
    )

    # Compute all query-document similarities with a single matrix multiply
    # query_embeddings shape: (num_queries, embedding_dim)
    # doc_embeddings.T shape: (embedding_dim, num_docs)
    # Result shape: (num_queries, num_docs)
    all_similarities = np.dot(query_embeddings, doc_embeddings.T)

    results = []
    for i, query in enumerate(queries):
        similarities = all_similarities[i]
        top_indices = np.argsort(similarities)[::-1][:top_k]
        query_results = [(docs[j], similarities[j]) for j in top_indices]
        results.append((query, query_results))

    return results


# Test batch retrieval with three semantically diverse queries
test_queries = [
    "attention mechanisms in deep learning",
    "ensemble methods for classification",
    "processing sequential information",
]

batch_results = batch_retrieve(
    test_queries, document_embeddings, documents, model, top_k=2
)
Out[18]:
Console
Batch Retrieval Results:

Query: 'attention mechanisms in deep learning'
  [0.451] Neural networks learn hierarchical representations of data t...
  [0.376] Recurrent neural networks maintain hidden states to process ...

Query: 'ensemble methods for classification'
  [0.513] Support vector machines find optimal hyperplanes to separate...
  [0.491] Random forests combine multiple decision trees for reliable ...

Query: 'processing sequential information'
  [0.601] Recurrent neural networks maintain hidden states to process ...
  [0.265] Neural networks learn hierarchical representations of data t...

The matrix multiplication computes all query-document similarities simultaneously in a single operation. For a batch of QQ queries and DD documents with embedding dimension dd, this requires O(Q⋅D⋅d)O(Q \cdot D \cdot d) operations performed in parallel on GPU, which is orders of magnitude faster than computing Q⋅DQ \cdot D dot products sequentially. In practice, batching queries together is one of the simplest and most effective optimizations for production retrieval throughput.

Building a Simple Search Function

Let's wrap everything into a clean, reusable search interface that demonstrates the full retrieval pipeline:

In[19]:
Code
class DenseRetriever:
    """
    A simple dense retrieval system showing the full pipeline:
    offline indexing and online query serving.
    """

    def __init__(self, model_name="all-MiniLM-L6-v2"):
        self.model = SentenceTransformer(model_name)
        self.documents = []
        self.embeddings = None

    def index(self, documents):
        """
        Offline step: encode all documents and store embeddings.
        This is done once and can be cached to disk.
        """
        self.documents = documents
        self.embeddings = self.model.encode(
            documents,
            convert_to_numpy=True,
            normalize_embeddings=True,
            show_progress_bar=False,
        )
        return self

    def search(self, query, top_k=5):
        """
        Online step: encode query and find nearest neighbors.
        Only the query encoding happens at search time.
        """
        query_emb = self.model.encode(
            query, convert_to_numpy=True, normalize_embeddings=True
        )
        scores = np.dot(self.embeddings, query_emb)
        top_indices = np.argsort(scores)[::-1][:top_k]
        return [
            {
                "rank": i + 1,
                "score": float(scores[idx]),
                "text": self.documents[idx],
            }
            for i, idx in enumerate(top_indices)
        ]
In[20]:
Code
# Build the index (offline step)
retriever = DenseRetriever()
retriever.index(documents)
Out[21]:
Console
Query: 'sequence modeling with memory'

Rank 1 (score: 0.542)
  Recurrent neural networks maintain hidden states to process sequential data.

Rank 2 (score: 0.306)
  Gradient descent optimizes model parameters by iteratively updating weights.

Rank 3 (score: 0.247)
  Neural networks learn hierarchical representations of data through multiple layers.

The separation between index() (offline) and search() (online) mirrors the architecture of production retrieval systems. The index step can be precomputed and cached to disk; in real systems it is typically stored in a vector database like FAISS, Pinecone, or Weaviate. The search step is the only latency-necessary path, and it reduces to a single transformer forward pass plus a matrix-vector multiply.

Key Parameters

Understanding the key parameters of a dense retrieval system helps you tune it for your specific application:

  • model_name_or_path: The pre-trained model weights for SentenceTransformer. Different models offer different trade-offs between speed (smaller models like MiniLM), accuracy (larger models like MPNet), and domain coverage (domain-specific fine-tuned models). Always evaluate on your target domain before choosing a model.
  • normalize_embeddings: When set to True in model.encode(), produces unit-length vectors and enables dot product to equal cosine similarity. Most use cases should set this to True for consistent, interpretable similarity scores.
  • top_k: The number of results to retrieve. In a RAG pipeline, this controls how much context the language model receives. Retrieving more candidates provides richer context but increases prompt length and can introduce irrelevant content. Values between 3 and 10 are typical, with larger values used when a reranker will further filter results.
  • batch_size: The number of documents or queries to encode simultaneously. Larger batches use GPU memory more efficiently but require more memory. Typical values range from 32 to 256 depending on document length and GPU memory.
  • Temperature (τ\tau): During training, controls the sharpness of the contrastive loss. Smaller values produce harder gradients; larger values produce softer distributions. At inference time, the metric is computed without temperature.

Limitations and Impact

Dense retrieval has transformed information retrieval and forms the foundation of modern knowledge-intensive NLP systems. However, understanding its limitations is needed for effective deployment. No retrieval technology is universal, and dense retrieval has specific failure modes that practitioners must account for.

Key Limitations

Dense retrievers require substantial training data to perform well on any given domain. Unlike BM25, which works out-of-the-box on any text collection using only the document corpus itself, dense models need labeled query-document pairs that match the target domain. A retriever trained on Wikipedia passages may perform poorly on legal contracts, scientific papers, clinical notes, or source code because those domains use different vocabulary and writing styles to express their semantic relationships. This domain sensitivity creates a "cold start" problem: organizations adopting dense retrieval for specialized applications need to either find a pre-trained model that covers their domain, collect domain-specific training data, or accept degraded performance. Collecting labeled training data is expensive and time-consuming, and for highly specialized domains, it may require domain experts who are scarce and expensive.

The computational and memory requirements of dense retrieval are significant and should not be underestimated at scale. Each document must be processed by a transformer encoder, and the resulting embeddings must be stored and indexed. For a corpus of 100 million documents with 384-dimensional embeddings in float32 precision, storage alone requires approximately 150 gigabytes, before any indexing structure overhead. While approximate nearest neighbor indices like HNSW or IVF-PQ can reduce search latency to milliseconds, they require additional memory and careful tuning. Sparse indices based on inverted lists, as used by BM25, typically require far less memory and can be searched efficiently with simple term lookups. The infrastructure investment for production dense retrieval, including GPU servers for encoding, vector databases for indexing, and replication for high availability, is substantially larger than for sparse retrieval.

Dense retrievers can also fail silently in ways that sparse retrievers do not. When BM25 returns poor results, the explanation is often immediately clear: the query terms simply do not appear in relevant documents. When a dense retriever fails, diagnosing the problem is much harder. The model might not have learned good representations for certain concepts, might confuse similar-sounding but semantically distinct entities (e.g., "Python the language" vs "Python the snake"), might not handle negation correctly (retrieving documents that say "A does not cause B" when the query asks "what causes A"), or might have distribution shift between its training data and the deployment domain. These opacity challenges make debugging and systematically improving dense retrieval systems more difficult than improving lexical systems.

Negation and logical constraints present a well-known challenge. Dense retrieval operates on semantic similarity, so "hotels without swimming pools" may retrieve documents about "hotels with swimming pools" because both texts are semantically about hotels and pools. The model lacks a mechanism to understand that the user explicitly wants to exclude certain properties. Similarly, numerical constraints ("papers from 2020 to 2022") and logical combinations ("A but not B") are poorly handled by similarity-based retrieval and often require post-retrieval filtering or more specialized architectures.

Impact on NLP Systems

Despite these limitations, dense retrieval has had enormous impact on NLP and AI systems. Semantic matching at scale made several previously impractical capabilities feasible:

Open-domain question answering systems improved dramatically when they could retrieve passages by meaning rather than keywords. Systems like RAG, FiD (Fusion-in-Decoder), and REALM all rely on dense retrieval as their backbone, achieving human-competitive performance on open-domain QA benchmarks by grounding generation in retrieved evidence. This was a major milestone in moving NLP from narrow task-specific systems to general knowledge-intensive question answering.

Retrieval-Augmented Generation systems, which we introduced in the previous chapters, depend critically on dense retrieval. The ability to find semantically relevant passages allows LLMs to answer questions about events after their knowledge cutoff, provide responses grounded in specific private documents, and reason over large knowledge bases without memorizing everything in their parameters. Every major commercial LLM assistant now uses some form of retrieval augmentation, and dense retrieval is the workhorse of the retrieval component in most of these systems.

Semantic search products from Google, Microsoft Bing, and other major search providers now incorporate dense retrieval techniques to handle conversational queries, long-form questions, and semantic information needs that were previously poorly served by keyword search. The transition from keyword search to semantic search at commercial scale represents one of the most significant applied NLP developments of the 2020s.

Cross-lingual applications became far more practical because multilingual dense models can encode meaning across languages, letting search systems where queries and documents may be in different languages. A user querying in English can retrieve relevant passages from documents in Spanish, French, or Japanese, as long as the multilingual encoder has learned to map semantically equivalent content in different languages to similar regions of the embedding space.

The combination of transformer pre-training, contrastive learning for retrieval, and approximate nearest neighbor search established a new paradigm for information access that goes far beyond traditional search. Dense retrieval combines representation learning with information retrieval, an area that continues to drive active research and development. The next chapter examines contrastive learning for retrieval: the training objectives, loss functions, and hard negative mining techniques that determine the quality of the learned embedding space.

Summary

Dense retrieval represents a major change from lexical to semantic matching in information retrieval. Rather than counting term overlaps, dense retrievers encode queries and documents into continuous vector spaces where similarity reflects semantic relatedness. The core value proposition is overcoming vocabulary mismatch: finding relevant documents regardless of whether they use the same specific words as the query.

The bi-encoder architecture enables scalable dense retrieval by separating query and document encoding. Documents are pre-computed and indexed offline, requiring only a single query encoding at search time. This design reduces query-time computation from millions of transformer forward passes to just one, letting practical retrieval over corpora with millions or billions of documents. The trade-off is that independent encoding sacrifices some modeling power compared to cross-encoders that jointly process query-document pairs, which is why dense bi-encoders are used for first-stage retrieval while cross-encoders serve as second-stage rerankers.

Embedding similarity metrics, particularly cosine similarity and dot product on normalized embeddings, quantify semantic relatedness between query and document vectors. The equivalence of these metrics under L2 normalization allows production systems to pre-normalize embeddings once at index time and use fast dot products at query time. The choice of metric should match the model's training objective.

Dense and sparse retrieval have complementary strengths that are best exploited by hybrid systems. Dense retrieval handles semantic matches and synonyms well, including conversational queries. Sparse retrieval provides interpretable exact-term matching with strong rare-term handling and zero-shot domain transfer. Hybrid retrieval combining both approaches consistently outperforms either approach alone on diverse retrieval benchmarks.

Training dense retrievers requires query-document pairs and careful negative sampling. The contrastive learning objective pushes the model to maximize similarity between relevant pairs while minimizing similarity with negatives. The quality of negatives significantly impacts the learned embedding space: hard negatives mined from BM25 or from the model itself produce substantially better representations than random negatives.

The next chapter examines contrastive learning for retrieval, including the training objectives, loss functions, and techniques that produce effective dense retrieval models from raw question-answer data.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about dense retrieval and semantic matching.

Dense Retrieval Architecture & Concepts

Question 1 of 70 of 7 completed
What is the fundamental advantage of dense retrieval over lexical methods like BM25?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026denseretrieval, author = {Michael Brenndoerfer}, title = {Dense Retrieval: Semantic Search & Bi-Encoder Implementation}, year = {2026}, url = {https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Dense Retrieval: Semantic Search & Bi-Encoder Implementation. Retrieved from https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders
MLAAcademic
Michael Brenndoerfer. "Dense Retrieval: Semantic Search & Bi-Encoder Implementation." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders>.
CHICAGOAcademic
Michael Brenndoerfer. "Dense Retrieval: Semantic Search & Bi-Encoder Implementation." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Dense Retrieval: Semantic Search & Bi-Encoder Implementation'. Available at: https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Dense Retrieval: Semantic Search & Bi-Encoder Implementation. https://mbrenndoerfer.com/writing/dense-retrieval-semantic-search-bi-encoders

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.