Part of Language AI Handbook
Explains how reranking with cross-encoders solves bi-encoder limitations. Topics include two-stage retrieval, training strategies.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Reranking
After retrieving candidate documents using dense vectors or hybrid search, you face a basic limitation: embedding similarity is a coarse measure of relevance. As we discussed in Dense Retrieval, bi-encoders compress documents into fixed-length vectors to enable fast approximate search, but this compression discards fine-grained interactions between query and document tokens. A document might share vocabulary with your query yet answer a completely different question, or it might contain the answer buried in irrelevant context that the embedding averages away. This loss of fidelity creates an information bottleneck that limits the precision of first-stage retrieval systems. When you ask a specific question requiring a fine-grained answer, the bi-encoder's single-vector summary often lacks the discriminative power to distinguish between superficially similar documents with different informational content.
Think of bi-encoder retrieval as a librarian who quickly scans the titles and back covers of thousands of books to create a shortlist for you. The librarian is fast and generally right about which shelf to check, but they cannot read every page of every book. Reranking is the second step: you hand that shortlist to an expert who opens each book, reads the relevant passages, and tells you which one answers your question. The expert is slow, but once the librarian has narrowed the field to a manageable set, the expert's careful judgment turns an approximate result into a precise one.
Reranking solves this by introducing a second, more careful judgment phase. Instead of comparing precomputed vectors, a reranker processes the query and candidate document together through a deep cross-attention mechanism. This produces a precise relevance score. This two-stage architecture, retrieve-then-rerank, has become the standard pattern for high-accuracy information retrieval systems. You first cast a wide net with fast bi-encoders to recall potentially relevant documents, then apply a computationally expensive but accurate cross-encoder to the top candidates to rank them precisely. This architectural pattern recognizes an needed truth about information retrieval: efficiency and accuracy often demand different computational strategies. The initial retrieval stage optimizes for speed and recall. This keeps no relevant document is overlooked, while the reranking stage optimizes for precision. This keeps that the most relevant documents surface to the top.
The distinction matters in practice. Modern search engines, question-answering systems, and retrieval-augmented generation (RAG) pipelines all depend on this two-stage design. Without the first stage, the reranker would have to score millions of documents, which is computationally impossible at interactive latencies. Without the second stage, the system would return documents that look topically relevant but fail to answer the user's specific question. This chapter examines how the two stages work together and the engineering decisions required to balance their competing demands.
This chapter examines the cross-encoder architecture that powers modern reranking, the procedures for integrating reranking into retrieval pipelines, how to train rerankers effectively, and the necessary latency tradeoffs that determine when and how to deploy them. We will also explore intermediate architectures like ColBERT that occupy the efficiency-accuracy spectrum between pure bi-encoders and full cross-encoders. Understanding these elements enables you to balance accuracy against speed within a fixed computational budget. Along the way, we will pay close attention to the mathematical foundations that explain why each design choice works, and we will ground every concept in concrete examples that make the abstractions tangible.
The idea of two-stage retrieval predates neural networks. Classical information retrieval systems from the 1990s already used a coarse first-pass filter followed by a more expensive reranking step, often combining BM25 with learned feature extractors. The 2019 paper "Passage Re-ranking with BERT" by Nogueira and Cho demonstrated that fine-tuned BERT models as cross-encoders dramatically outperformed classical rerankers on MS MARCO, a large-scale question answering dataset collected from Bing search logs. This work catalyzed the neural reranking paradigm and triggered a wave of research into efficient two-stage architectures. The MS-MARCO Passage Ranking leaderboard became a central benchmark, and models trained on it (often released as cross-encoder/ms-marco-* on HuggingFace) remain standard starting points today. The subsequent development of ColBERT in 2020 by Khattab and Zaharia pushed the boundary further by showing that token-level late interactions could match cross-encoder accuracy while retaining precomputable document representations, opening a new research direction that continues to evolve.
The Limitations of Bi-Encoders
To understand why reranking is necessary, you must first appreciate the constraints of bi-encoder retrieval. In bi-encoder architectures, which we covered in Dense Retrieval and Embedding Models, the query and document are encoded independently. The scoring function takes the form:
where:
- : the query text
- : the document text
- : the query encoder (a transformer network with parameters )
- : the document encoder (a transformer network with parameters )
- : similarity function (typically cosine similarity or dot product)
This factorization enables precomputation: you encode your document corpus once and index the vectors for fast nearest-neighbor search. This architectural choice makes bi-encoders remarkably efficient at scale, letting millions or billions of documents to be searched in milliseconds. The document vectors are computed offline, and at query time you only need to encode the query and perform a fast approximate nearest-neighbor lookup. This scalability is what makes bi-encoders needed in production systems.
However, independence comes at a cost. When you encode a document, you must produce a single vector representation without knowing what query it will be compared against. This creates an asymmetric compression problem: the document encoder must anticipate all possible questions you might ask and pack the answer into a fixed-dimensional vector. For complex documents containing multiple facts or fine-grained arguments, this is inherently lossy. The encoder essentially performs lossy compression of the document's semantic content, averaging over potentially relevant information in a way that preserves general topical similarity but often obliterates the specific details that determine relevance to a particular query. From an information-theoretic perspective, the document encoder must construct a sufficient statistic for all possible queries, which is impossible for documents containing many distinct facts and arguments.
The key insight is that the independence assumption is not a bug but a deliberate architectural choice that trades representation fidelity for precomputability. The bi-encoder cannot be more accurate without sacrificing the ability to precompute document vectors, because any interaction between query and document tokens would require knowing the query before encoding the document. This tradeoff is basic, not accidental.
Consider a query: "What temperature causes chocolate to bloom?" A bi-encoder might retrieve a document about chocolate manufacturing that mentions blooming, but it cannot distinguish whether the document states the specific temperature threshold or discusses blooming as a general phenomenon. The embedding similarity captures semantic relatedness but not factual sufficiency. The vector representation might encode that the document discusses "chocolate," "bloom," and "temperature" in close semantic proximity, but it cannot verify whether the specific causal relationship or numerical value requested by the query is present in the text. This distinction between topical relevance and factual sufficiency lies at the heart of the bi-encoder's limitations.
Think of the bi-encoder as a fingerprint scanner. It can tell you that two things share the same general shape and pattern, but it cannot tell you whether a specific detail, a particular scar, a unique ridge, is present. The cross-encoder, by contrast, is a forensic analyst who examines every feature of both prints side by side, looking for specific matches and mismatches. The analyst is slower, but the analyst's judgment is far more reliable for cases where the difference between relevant and irrelevant comes down to a single specific detail.
Beyond factual sufficiency, bi-encoders also struggle with negation and contrast. If a query asks "What are the risks of using aspirin?" and a document discusses the benefits of aspirin, the embedding similarity may still be high because both texts share the topic of aspirin. A cross-encoder can detect the conceptual mismatch between "risks" in the query and "benefits" in the document because it processes them together and can attend to the semantic polarity of each claim. Similarly, queries involving temporal qualifiers ("What changed after 2020?"), causal relationships ("Why did X cause Y?"), or conditional constraints ("How do you do X without Y?") all present challenges for bi-encoders that cross-encoders handle more naturally through joint reasoning over the full text.
Cross-Encoder Architecture
Cross-encoders eliminate the independence assumption by processing query and document jointly. Rather than compressing each into a vector separately, they concatenate the texts and feed them through a transformer, letting every query token to attend directly to every document token. This joint processing enables the model to perform fine-grained comparisons that are impossible when representations are computed in isolation. When the query token "temperature" can directly attend to the document token "32°C," the model can verify the presence of specific information rather than merely estimating semantic overlap.
The architecture is conceptually simple but computationally powerful. You take a pre-trained transformer encoder, the same kind of BERT model we examined in BERT Architecture, and fine-tune it on relevance-labeled query-document pairs. The transformer's self-attention mechanism, applied across the concatenated input, creates a rich interaction matrix between all query and document tokens. Every query token can ask "Is this information present in the document?" and receive an answer informed by the full document context. Every document token can ask "Does this piece of information satisfy the query?" and receive an answer informed by the full query intent.
Input Representation
The standard cross-encoder format concatenates the query and document with special separator tokens:
where:
- : classification token marking the start of the sequence (following BERT conventions from BERT Architecture)
- : token sequence of the query
- : token sequence of the document
- : separator token marking boundaries between segments
- : token sequence concatenation operator
The entire sequence is processed through a transformer encoder, typically BERT or a similar masked language model. The use of special separator tokens is important because it allows the transformer to distinguish between query and document content. Without these explicit boundaries, the model would struggle to determine which part of the concatenated sequence represents your information need versus the potential source of that information. The [CLS] token, originally introduced for next-sentence prediction tasks in BERT, is an aggregate representation of the interaction between the two texts.
The segment embeddings that BERT uses for next-sentence prediction tasks serve a secondary role here. Segment A typically corresponds to the query, and segment B corresponds to the document. These segment embeddings provide the model with a structural prior about which text is the "question" and which is the "answer," even before the attention mechanism processes the content. In practice, models fine-tuned specifically for passage reranking often discard the segment embeddings or treat the entire concatenation as a single segment, finding that the [SEP] token alone is sufficient to signal the boundary.
Unlike the encoder-decoder cross-attention we explored in Cross-Attention, cross-encoders usually employ full bidirectional self-attention across the concatenated query-document sequence. Every token can attend to every other token, creating a dense interaction matrix of size , where and are the query and document lengths respectively. This allows the model to capture fine-grained alignments, such as matching specific numerical values or entity relationships between query and document. While encoder-decoder architectures use unidirectional cross-attention from encoder outputs to decoder states, cross-encoders maintain bidirectional flow throughout, letting richer contextualization of every token based on the full context of both texts. The term "cross-encoder" refers to the fact that information crosses between the query and document during encoding, not that the architecture uses a separate cross-attention module.
Scoring Head
After processing through the transformer layers, the final hidden state of the [CLS] token, , is the pooled representation of the query-document pair. This single vector condenses all the pairwise interactions between query and document tokens into a fixed-size summary. A scoring head, typically a simple feed-forward network or even a linear projection, maps this to a relevance score:
where:
- : the final hidden state of the [CLS] token, serving as the pooled representation
- : weight vector of the scoring head
- : bias term of the scoring head
- : the predicted relevance score (logit) for query-document pair
The [CLS] token is an interesting choice for pooling. During pre-training, BERT trained this token to predict whether two sentences were consecutive, which forced it to encode the relationship between the two input segments. Fine-tuning on relevance prediction adapts this representational capacity to encode query-document compatibility rather than sentence adjacency. Alternative pooling strategies, such as mean-pooling over all tokens or max-pooling over position-wise representations, are possible but empirically the [CLS] approach works well, especially when the model is initialized from a BERT checkpoint that already has strong [CLS]-based representations.
For classification-based training, the score is often passed through a sigmoid to produce a probability of relevance:
where:
- : the relevance score (logit) computed by the scoring head
- : the sigmoid function, mapping logits to probabilities in
- : probability that document is relevant to query
The sigmoid transformation is necessary when you need calibrated probability estimates rather than ordinal rankings. For a well-calibrated reranker, means that documents assigned this score are relevant 80% of the time. Calibration matters for downstream decisions like whether to include a document in the context window for generation, or whether to return "no relevant document found" rather than returning the top result by default.
Alternatively, some rerankers output logits for multiple relevance grades (e.g., not relevant, somewhat relevant, highly relevant) using a softmax over multiple output classes. This graded approach allows the model to express degrees of relevance rather than binary judgments, which can be particularly useful when training on data with multiple relevance levels, such as human judgments on a three-point or five-point scale. The choice between regression, binary classification, or ordinal classification depends on the nature of your training data and the specific requirements of your ranking task.
Computational Complexity
The richness of interaction comes at steep computational cost. While bi-encoders require transformer passes per query-document pair, reduced to if documents are pre-encoded, cross-encoders require attention computations per pair.
Specifically, for a sequence of length , the self-attention mechanism computes:
where:
- : query matrix representing projected input tokens
- : key matrix representing projected input tokens
- : value matrix representing projected input tokens
- : sequence length (total tokens in query + document + special tokens)
- : dimension of the key/query projections
- : normalization function converting attention scores to probabilities
The matrix multiplication alone costs . If your query is 10 tokens and your document is 512 tokens, a bi-encoder processes them separately (cost proportional to ), while a cross-encoder processes the concatenation (cost proportional to ). However, since bi-encoders pre-encode documents, the per-query cost is just for the query encoding plus for the dot product, making them thousands of times faster at query time.
This computational asymmetry fundamentally shapes system architecture. While the bi-encoder pays a high upfront cost to encode the corpus (which can be amortized over billions of queries), the cross-encoder must perform expensive computation for every query-document pair at query time. This means that if you attempt to rerank thousands of documents per query, latency becomes prohibitive, often reaching several seconds or even minutes. The quadratic scaling with sequence length also means that long documents dramatically increase computational cost, creating pressure to truncate texts aggressively, which risks losing relevant information located in later portions of the document.
Think of the computational cost as a restaurant analogy. A bi-encoder is like a restaurant that prepares all the dishes in advance and stores them in a display case. When a customer arrives, serving them is nearly instantaneous. A cross-encoder is like a restaurant that cooks every dish to order, personalized for each specific customer. The food is better, but every customer has to wait. The two-stage architecture is like a restaurant that prepares a large batch of semi-finished dishes in advance, then does a final personalized preparation for each customer from a shortlist. You get quality close to fully custom-cooked, at a speed close to the display case.

The Reranking Procedure
Reranking is not a replacement for retrieval but a refinement layer. Understanding the full pipeline, from query to final ranked list, requires understanding each stage's purpose, its failure modes, and the parameters that govern its behavior. The complete pipeline follows a retrieve-then-rerank pattern, and each stage contributes a distinct kind of intelligence to the final result.
The first stage is designed for speed and recall. It casts a wide net, accepting some irrelevant documents in exchange for the assurance that relevant documents are unlikely to be missed. The second stage is designed for precision. It carefully reads the shortlist produced by the first stage, identifies the documents that satisfy the query, and moves them to the top. Neither stage is sufficient alone: retrieval without reranking produces results that are topically relevant but often imprecise; reranking without retrieval would be prohibitively slow at corpus scale.
The complete pipeline:
-
Initial Retrieval: Use fast methods (dense retrieval with bi-encoders, BM25 from BM25, or hybrid search from Hybrid Search) to fetch a candidate set of documents. This stage prioritizes recall over precision; you want to ensure the relevant documents are somewhere in this set, even if ranked poorly. The goal is complete coverage: you are willing to accept some irrelevant documents in the initial set to avoid missing any relevant ones. This is often called the "recall-oriented" stage of the pipeline.
-
Reranking: Feed each query-candidate pair through the cross-encoder to obtain precise relevance scores. Sort by these scores to produce a reranked list. This stage prioritizes precision: you want to identify exactly which of the candidates best answers the query, even if this process is computationally expensive. The cross-encoder acts as a careful judge, examining each candidate in detail to determine its true relevance.
-
Truncation: Return the top- documents from the reranked list (where ) to the downstream generation system or user. Typically, you might retrieve documents but only return the top or . Because users rarely examine more than a handful of results, the top positions need the highest possible quality.
The parameter represents your reranking budget: how many candidates you can afford to score with the expensive cross-encoder. This is typically determined by latency constraints. If your cross-encoder takes 10ms per document and your latency budget is 100ms, you might set . This budget is often one of the most necessary tuning parameters in production systems, requiring careful balancing between accuracy (higher gives the reranker more options) and user experience (lower ensures faster response times). In practice, is often determined empirically through latency testing and quality evaluation, with typical values ranging from 10 to 1000 depending on the computational resources available and the latency requirements of the application.
The key insight here is that the value of reranking depends critically on the quality of the initial retrieval. If the first-stage retriever fails to include the relevant document in its top results, no amount of reranking can recover it. This is why we speak of retrieve-then-rerank as a recall-then-precision strategy. You must ensure that the first stage achieves high recall, then rely on the second stage to improve precision within the recalled set. Measuring recall@k for the initial retriever on a validation set is therefore an needed diagnostic step before deploying any reranking system.
Cascade Reranking
In large-scale systems, you might employ multiple reranking stages of increasing computational cost. Rather than applying the most expensive model to all candidates immediately, you use a sequence of progressively more sophisticated and costly models, each filtering the candidate set for the next stage. This creates a funnel where cheap models eliminate obviously irrelevant documents, leaving expensive models to focus computational effort only on the most promising candidates.
Think of cascade reranking as a talent competition with multiple rounds. In the first round, thousands of candidates perform a brief audition, and only the most promising advance. In the second round, those finalists perform a full set, and a smaller group moves to the finals. In the final round, a handful of top performers receive the most intensive scrutiny. Each round's judges spend proportionally more time per candidate, letting increasingly fine-grained discrimination. The key is that you never waste your most expensive judges on candidates who would clearly be eliminated in the first round.
For example, you might begin with a bi-encoder to retrieve one thousand documents, then apply a lightweight cross-encoder to select the best one hundred, followed by a heavy cross-encoder to identify the top ten, and finally employ a large language model as a judge to select the single best document or generate an answer based on the top three. Each stage in this cascade reduces the candidate set while increasing computational investment per document, letting you to apply your most expensive computational resources only where they matter most. This approach recognizes that most retrieved documents are clearly irrelevant and do not merit deep analysis, while a small subset requires careful scrutiny to distinguish subtle differences in relevance.
The engineering of cascade systems involves careful measurement at each stage. You need to know the precision-recall tradeoff of each filter, because overly aggressive filtering at an early stage can eliminate relevant documents that later stages can never recover. A common strategy is to run the full system offline on a validation set with varied cascade thresholds, then select thresholds that maintain a recall of at least 95% after each filtering stage.
Training Rerankers
Training a cross-encoder reranker requires supervised data consisting of query-document pairs labeled with relevance judgments. Unlike bi-encoders, which we trained using contrastive objectives in Contrastive Learning for Retrieval, cross-encoders can use standard classification or ranking losses because they process pairs jointly. This joint processing means that the model can learn directly from explicit relevance signals rather than needing to construct contrastive pairs that approximate the ranking task. The availability of labeled relevance judgments, typically obtained through human annotators or implicit feedback signals like clicks, allows cross-encoders to optimize directly for the probability of relevance or relative ranking preferences.
The choice of training objective has significant practical consequences. Binary classification losses train the model to output calibrated probabilities, which is useful when you want to apply a fixed threshold to determine whether any document is relevant. Pairwise ranking losses train the model to correctly order document pairs, which is more directly aligned with the ranking task but may produce poorly calibrated scores. Listwise losses optimize the entire ranked list simultaneously. This provides the strongest alignment with end-to-end ranking quality but requiring richer supervision signals. In practice, a combination of these objectives often works best, with pairwise or listwise losses used during primary training and calibration steps applied afterwards.
Binary Classification
The simplest approach treats relevance as a binary classification problem. Given a dataset where indicates relevance, you minimize cross-entropy loss. For each training example, the loss measures the discrepancy between the predicted probability and the true binary label. The cross-entropy formulation naturally handles the probabilistic nature of relevance judgments and provides smooth gradients that enable efficient training via gradient descent:
where:
- : number of training examples
- : the -th query in the training set
- : the -th document in the training set
- : binary relevance label (1 if relevant, 0 if irrelevant)
- : predicted relevance score for the -th pair
- : sigmoid function converting scores to probabilities
- : the cross-entropy loss measuring prediction error
When (relevant), the loss reduces to , which is minimized when the model assigns a high score to the pair. When (irrelevant), the loss becomes , which is minimized when the model assigns a low score. The average over all training examples pushes the model to simultaneously score all relevant documents highly and all irrelevant documents lowly. The key property of this loss is that the gradient does not vanish until the model is extremely confident in the correct direction. This keeps training continues to make meaningful progress even when the model is already doing reasonably well.
This works well when you have explicit relevance labels, such as click-through data or human judgments on a binary scale (relevant/irrelevant). However, collecting reliable binary labels presents significant challenges. Human annotators often disagree on relevance boundaries, and click-through data can be noisy due to position bias (users are more likely to click top results regardless of relevance) and presentation bias (attractive titles may generate clicks even for irrelevant documents). Active learning strategies can help prioritize which documents to label, focusing annotation effort on the boundary cases where the model is most uncertain. Additionally, class imbalance is common in retrieval datasets: for most queries, the large majority of documents are irrelevant, so you may need to use weighted loss functions or careful sampling strategies to prevent the model from learning to always predict "irrelevant."
Pairwise Ranking Loss
Often, you have relative judgments: for query , document is more relevant than . Rather than requiring absolute relevance labels, pairwise approaches only need to know the relative ordering between pairs. This is a more natural fit for how humans perceive relevance: it is often easier to say "Document A is better than Document B for this query" than to assign a specific numeric relevance score to each document in isolation.
The margin ranking loss encourages the model to score the positive higher than the negative by at least a margin :
where:
- : the query being processed
- : a document labeled as relevant (positive) to query
- : a document labeled as irrelevant (negative) to query
- : relevance score assigned to query-document pair
- : margin hyperparameter (typically set to 1.0) specifying the minimum desired score difference
- : the margin ranking loss penalizing violations of the ranking constraint
This is analogous to the hinge loss used in support vector machines. If , the loss is zero; otherwise, it increases linearly with the violation. The margin is a hyperparameter, typically set to 1.0. The intuition behind the margin is that you want the model to rank the positive above the negative with confidence. A small score difference might indicate uncertainty, while a large margin suggests the model has learned a reliable distinction between the two documents. The margin also prevents the loss from rewarding overly conservative models that barely distinguish positive from negative examples.
Building on our understanding of the Bradley-Terry model from Bradley-Terry Model, we can also interpret pairwise preferences probabilistically. The probability that is preferred over is:
where:
- : event showing document is preferred over document
- : relevance score for the positive document
- : relevance score for the negative document
- : sigmoid function converting score differences to probabilities
- : probability that is more relevant than given query
The Bradley-Terry formulation says that preferences between items can be modeled through a latent "strength" parameter for each item. The probability that one item is preferred over another depends only on the difference in their strengths, transformed through a sigmoid. In our case, the relevance score plays the role of strength, and we train the model to assign higher strength to relevant documents.
Maximizing the log-likelihood of observed preferences yields the logistic loss:
where:
- : relevance score for the positive document
- : relevance score for the negative document
- : sigmoid function
- : negative log-likelihood loss (logistic loss)
This formulation, used in RankNet and similar neural ranking models, naturally handles noisy preferences and provides well-calibrated probability estimates. Unlike the margin ranking loss, which treats preference violations with uniform penalty regardless of magnitude (beyond the margin), the logistic loss penalizes violations based on their probability, with larger mistakes receiving exponentially higher penalties. This probabilistic interpretation also allows you to combine evidence from multiple annotators or to model uncertainty in the preference labels themselves. A document pair that one annotator labels as but another labels as can be treated as giving partial, uncertain preference evidence rather than a hard constraint.


Listwise Approaches
While pairwise losses compare documents two at a time, listwise approaches consider the entire ranked list for a query. These methods often optimize target metrics more directly, such as Normalized Discounted Cumulative Gain (NDCG). The key insight behind listwise approaches is that ranking quality depends on individual document scores and on the relative ordering of all documents in the list. A pairwise approach might correctly order 99 pairs but still produce a suboptimal list if it makes a necessary error at the top position, while a listwise approach can directly optimize for the importance of top-ranked positions.
Think of listwise approaches as optimizing for the experience of someone who reads through the ranked list from top to bottom. The further down the list a reader goes, the less attention they typically pay to each result. Listwise losses incorporate this attention structure by weighting mistakes more heavily at higher positions. A relevant document ranked first is vastly more valuable than the same document ranked tenth, and listwise losses formalize this intuition mathematically.
One popular listwise loss is ListNet, which treats ranking as a probability distribution over permutations. For a list of documents , define the probability of document being at the top using the softmax over scores:
where:
- : the -th document in the list of candidates
- : relevance score assigned to document by the model
- : number of documents in the ranked list
- : exponential function so positive values
- : probability that document should be ranked at the top position
The softmax formulation ensures that the probabilities sum to one across all documents in the list, creating a proper probability distribution over which document deserves the top position. Similarly, define target probabilities based on relevance grades (e.g., exponential decay with position in the ideal ranking). The loss is the cross-entropy between these distributions:
where:
- : number of documents in the list
- : target probability for document (derived from relevance grades)
- : predicted probability from the softmax over scores
- : cross-entropy loss between target and predicted ranking distributions
The target distribution encodes the ideal ranking. If you have graded relevance labels (highly relevant, somewhat relevant, not relevant), you might set , so that highly relevant documents receive proportionally higher target probability. The cross-entropy loss then pushes the model's predicted distribution to match this target, training the model to assign high probability to the documents that should be at the top.
Listwise losses can better model the dependencies between documents and directly optimize for ranking quality, though they are computationally more intensive due to the need to process entire lists. They also allow you to incorporate position-dependent importance naturally, since the loss is sensitive to which documents appear at the top of the list. However, computing these losses over large lists can be expensive, and the softmax normalization can become numerically unstable with large score ranges. Variants like ListMLE simplify the computation by focusing on the probability of the observed permutation rather than the full distribution over all possible permutations.
Hard Negative Mining
A necessary aspect of training rerankers is negative sampling. If you train only on easy negatives (randomly sampled documents), the model learns little beyond what the bi-encoder already knows. The cross-encoder must learn to distinguish between documents that appear relevant to a fast retriever but are unsuitable for answering the query. Without hard negatives, the reranker fails to develop the discriminative capacity needed to improve upon the initial retrieval.
Think of easy negatives as training a chess player by having them practice against beginners. The beginner makes obvious blunders, and the student never needs to develop deep strategy. Hard negatives are like practice against near-equal opponents: the student must learn subtle principles and find the small advantages that separate winning from losing. A reranker trained only on easy negatives is like the chess player who can easily beat beginners but struggles against anyone who plays competently.
Instead, you should mine hard negatives: documents that the bi-encoder retrieves despite being irrelevant. These are documents that share semantic similarity with the query, perhaps containing overlapping vocabulary or related concepts, but fail to satisfy the specific information need. Hard negatives teach the model the subtle distinctions that separate plausible-looking documents from relevant ones.
The standard procedure is:
- Retrieve top- documents for each training query using your bi-encoder
- Filter out truly relevant documents (based on your labeled data)
- Use the remaining retrieved documents as negatives
This forces the cross-encoder to learn the subtle distinctions that separate plausible-looking but irrelevant documents from relevant ones. As training progresses, you should periodically refresh the hard negatives using the current best retrieval model, a technique known as iterative hard negative mining. This iterative process creates a curriculum where the model progressively learns to distinguish finer and finer distinctions. Early in training, the model benefits from easier negatives to learn basic semantic distinctions. As the model improves, harder negatives force it to focus on increasingly subtle cues. However, care must be taken not to select negatives that are too hard: if a negative document is relevant (a false negative), training on it as a negative can degrade model performance. Quality control of negative mining is needed, often requiring human validation of edge cases or the use of multiple retrieval systems to identify likely true negatives.
The key insight is that the distribution of training negatives should mirror the distribution of errors you expect to encounter in production. Since the bi-encoder is your first-stage retriever in deployment, the documents it retrieves but should rank lower are exactly the errors you need the reranker to correct. Training on these errors creates a reranker that is specifically calibrated to improve upon the bi-encoder's weaknesses, which is precisely the objective of the two-stage architecture.

Latency Considerations and Optimization
The primary barrier to deploying cross-encoders is latency. If your retrieval system returns 100 candidates and your cross-encoder requires 50ms per document, reranking adds 5 seconds of latency, unacceptable for most interactive applications. In web search or conversational AI systems, users expect sub-second response times, often targeting latencies below 200 milliseconds. When the cross-encoder becomes the bottleneck, you must either reduce the number of documents processed, reduce the time per document, or apply the cross-encoder selectively.
Understanding the sources of this latency is needed for effective optimization. The cross-encoder's latency has two main components: the compute cost of the transformer forward pass (which scales quadratically with sequence length) and the overhead of data loading and tokenization, followed by result extraction. For modern BERT-base models processing 512-token sequences, the transformer compute is the dominant cost. Optimization strategies must therefore either reduce the effective sequence length (document truncation, chunking), reduce the model's computational requirements (quantization, distillation, smaller models), or reduce the number of documents processed (selective reranking, cascade architectures).

Selective Reranking
You need not rerank all retrieved documents. Confidence-based filtering uses the bi-encoder scores to identify high-confidence candidates. If the top bi-encoder result has a score significantly higher than the rest, you might return it directly without reranking, applying the cross-encoder only to ambiguous cases where scores are close. This approach recognizes that for many queries, the bi-encoder is sufficiently confident in its top prediction that the expensive reranking provides minimal benefit.
By measuring the gap between the top score and the mean or second-highest score, you can estimate the uncertainty of the initial retrieval and allocate computational resources only where ambiguity exists. A large gap between the first and second retrieved document suggests that the bi-encoder is very confident in its top result, and the cross-encoder is unlikely to change the ranking. A small gap suggests that the relative ordering is uncertain, and the cross-encoder's fine-grained discrimination is most likely to add value.
This strategy requires careful calibration to ensure that high-confidence predictions are indeed accurate; otherwise, you risk bypassing the reranker precisely when it would have caught a retrieval error. Empirical validation on a held-out test set is needed: measure how often the bi-encoder's top result is also the cross-encoder's top result as a function of the score gap, and use this relationship to set the threshold for when to skip reranking.
Document Truncation
Cross-encoder complexity grows quadratically with sequence length. For long documents, you typically truncate to the first tokens (often 512 or 1024 tokens). However, relevant information might appear at the end of a document, a phenomenon known as the "lost in the middle" problem. Early portions of documents often contain boilerplate, introductions, or general context, while specific answers may appear later.
Sliding window approaches split long documents into overlapping chunks, rerank each chunk independently, and aggregate scores. For a document chunked into overlapping segments , the final document score is:
where:
- : the full document being scored
- : a specific chunk (text segment) of the document
- : set of overlapping text segments extracted from document
- : relevance score between query and chunk
- : final document score, computed as the maximum score across all chunks
Taking the maximum over chunks is motivated by the intuition that a document is relevant if any part of it answers the query. Alternative aggregation strategies include taking the average of the top few chunk scores, which is more reliable to chunks with spuriously high scores, or using a learned aggregation function that weights chunks based on their position or content. The choice of chunk size and overlap involves tradeoffs: smaller chunks reduce the quadratic cost per chunk but may lose context that spans chunk boundaries, while larger chunks preserve more context but increase computational cost.
The key insight here is that chunking shifts the granularity of the retrieval unit from the document to the passage. This is often a fundamentally better alignment with how queries are answered: most questions are answered in a specific paragraph or section, not by a document in its entirety. Many production systems chunk documents at indexing time rather than reranking time, treating chunks as first-class retrieval units and only reassembling them when necessary.
Late Interaction Models
Researchers have developed architectures that balance the efficiency of bi-encoders with the accuracy of cross-encoders. Late interaction models, exemplified by ColBERT (Contextualized Late Interaction over BERT), encode query and document separately but allow token-level interactions at scoring time. The core idea is to preserve one dense vector per token rather than collapsing the entire document into a single vector. This dramatically increases storage requirements compared to single-vector bi-encoders, but enables much richer matching at query time.
The ColBERT scoring function captures the degree to which each query token is "satisfied" by the best-matching document token:
where:
- : number of tokens in the query
- : number of tokens in the document
- : contextualized embedding of the -th query token
- : contextualized embedding of the -th document token
- : dot product similarity between query token and document token
- : final relevance score (sum of maximum similarities for each query token)
This operation is known as the MaxSim (maximum similarity) operator. For each query token, it finds the document token that most closely matches it, then sums these maximum similarities across all query tokens. The intuition is that a document is relevant when it contains a good match for every important concept in the query, not just when some average representation is similar.
The query and document are encoded independently (fast, precomputable), but the scoring involves a maximum similarity computation between all query-document token pairs, costing . This is more expensive than a dot product but vastly cheaper than full cross-attention per layer. Because documents can be pre-encoded and indexed, ColBERT achieves latency much closer to bi-encoders while retaining much of the cross-encoder's ability to model fine-grained interactions. Recent improvements in ColBERTv2 introduce residual compression and denoised supervision, further reducing storage requirements and improving retrieval accuracy. These models represent a promising middle ground for applications where full cross-encoders are too slow but bi-encoders are insufficiently accurate.
Approximate Reranking
For extremely large candidate sets, you can apply vector quantization techniques from Product Quantization to cross-encoder representations, or use knowledge distillation to train a smaller student model to mimic the large cross-encoder's scores. Distillation involves training a compact student model (perhaps a smaller transformer or even a bi-encoder) to reproduce the ranking decisions of the large cross-encoder teacher. The student learns to approximate the teacher's score function, achieving much faster inference while preserving much of the accuracy.
This approach is particularly effective when combined with the cascade architecture: the student model filters the bulk of candidates, and the teacher model verifies only the most promising ones. Distillation also allows you to compress a large cross-encoder into a faster bi-encoder that has been "taught" to produce embeddings that approximate cross-encoder scores, effectively transferring some of the cross-encoder's discriminative ability into a precomputable representation. Models like TAS-B and Augmented SBERT use variants of this distillation approach to produce bi-encoders that significantly outperform bi-encoders trained without a cross-encoder teacher.
Worked Example
Let us walk through a concrete reranking scenario in full numerical detail. Suppose you have built a technical support system for a software company. You ask:
"How do I reset my password if I forgot the email associated with my account?"
This query has a specific constraint that distinguishes it from a generic password reset question: the user has forgotten both the email address and the password. The forgotten email address is the crux of the relevance judgment, and it is exactly the kind of subtle distinction that bi-encoders tend to miss.
Step 1: Initial bi-encoder retrieval. Your bi-encoder encodes the query into a 384-dimensional vector and computes cosine similarity against your pre-indexed document embeddings. The system returns these three candidates:
Doc A: "To reset your password, click 'Forgot Password' on the login page. Enter your registered email address, and we will send you a reset link."
Doc B: "Account recovery options include password reset via email, security questions, or contacting support. If you no longer have access to your registered email, use the alternative verification methods available in your account settings."
Doc C: "Email notifications can be configured in your account settings. You can choose to receive updates about password changes, login attempts, and security alerts to any verified email address."
The bi-encoder ranking: Doc A (score 0.89), Doc C (score 0.85), Doc B (score 0.82). Doc A mentions "reset your password" and "email" prominently and matches the query's surface vocabulary well. Doc C mentions "email" and "password" but in the context of notifications. Doc B mentions "access to your registered email" but uses different vocabulary than the query, earning the lowest bi-encoder score despite being the most contextually relevant.
Step 2: Cross-encoder input construction. For each candidate, we construct the input sequence as [CLS] + query_tokens + [SEP] + document_tokens + [SEP]. For Doc A, the query tokens include "forgot," "email," and "account," while the document contains "registered email address" and "reset link." The [CLS] token will attend to all of these simultaneously.
Step 3: Joint attention and scoring. The cross-encoder processes each concatenated pair. For Doc A, the attention mechanism highlights the tension between "forgot the email" in the query and "Enter your registered email address" in the document. The document presupposes that the user knows their registered email, which directly contradicts the query's stated constraint. The [CLS] representation captures this semantic mismatch, and the scoring head assigns a low relevance score of 0.31.
For Doc B, the cross-encoder identifies the phrase "no longer have access to your registered email" as directly addressing the query's constraint. The attention heads align "forgot the email" in the query with "no longer have access to your registered email" in the document, recognizing semantic equivalence despite lexical difference. The phrase "alternative verification methods" is identified as the solution to the user's problem. The cross-encoder assigns a high score of 0.94.
For Doc C, the cross-encoder recognizes that while "email" and "password" appear, they are in the context of notification settings, not account recovery. The query's intent, recovering account access, finds no matching information in this document. The cross-encoder assigns a score of 0.12.
Step 4: Reranked results. The cross-encoder scores: Doc B (0.94), Doc A (0.31), Doc C (0.12). The reranker correctly identifies Doc B as the most relevant, despite it having the lowest bi-encoder score. This dramatic reversal illustrates how cross-encoders capture pragmatic constraints and presuppositions that bi-encoders miss. The necessary signal is not lexical overlap but the alignment of the specific condition stated in the query (forgotten email) with the specific scenario described in the document (no access to registered email).
Step 5: Interpreting the result. The ranking reversal between Doc A and Doc B illustrates the bi-encoder's weakness. Doc A scores high on the bi-encoder because it contains the exact words from the query: "reset," "password," "email." But it is a worse answer because it assumes the user knows their email. Doc B uses different words ("no longer have access") but addresses the exact situation the user is in. The cross-encoder, reading both texts together, can make this distinction. The bi-encoder, working with independent embeddings, cannot.

Code Implementation
Let us implement a complete retrieve-then-rerank pipeline using the sentence-transformers library, which provides convenient interfaces for both bi-encoders and cross-encoders. This implementation will make the concepts concrete and give you a working starting point for your own retrieval systems.
# Install required packages (run once locally, then comment out):
# import subprocess
# import sys
# subprocess.check_call([sys.executable, "-m", "pip", "install", "-q",
# "sentence-transformers", "scikit-learn", "numpy"])# Sample corpus of technical documentation
documents = [
"To reset your password, click 'Forgot Password' on the login page. Enter your registered email address, and we will send you a reset link.",
"Account recovery options include password reset via email, security questions, or contacting support. If you no longer have access to your registered email, use the alternative verification methods available in your account settings.",
"Email notifications can be configured in your account settings. You can choose to receive updates about password changes, login attempts, and security alerts to any verified email address.",
"Two-factor authentication adds an extra layer of security to your account. You can enable 2FA using an authenticator app or SMS verification.",
"If you suspect unauthorized access to your account, immediately change your password and review recent login activity in your security dashboard.",
"Premium users have access to priority support and advanced security features including hardware key authentication.",
"Mobile app users can enable biometric authentication for faster, secure login using fingerprint or facial recognition.",
"Account deletion is permanent and cannot be undone. All data associated with your account will be removed within 30 days.",
]
query = "How do I reset my password if I forgot the email associated with my account?"We begin by loading a bi-encoder for initial retrieval. We will use a small model suitable for demonstration, though production systems might use larger embeddings. The bi-encoder encodes both the query and documents into dense vectors that capture semantic meaning, letting us to quickly identify the subset of documents that are most likely to be relevant based on vector similarity.
## Load bi-encoder for initial retrieval
print("Loading bi-encoder...")
import numpy as np
from sentence_transformers import SentenceTransformer
from sklearn.metrics.pairwise import cosine_similarity
bi_encoder = SentenceTransformer("all-MiniLM-L6-v2")
## Encode corpus (this would be precomputed and indexed in production)
print("Encoding documents...")
doc_embeddings = bi_encoder.encode(documents, convert_to_tensor=True)
query_embedding = bi_encoder.encode([query], convert_to_tensor=True)
## Compute similarities
similarities = cosine_similarity(
query_embedding.cpu().numpy(), doc_embeddings.cpu().numpy()
)[0]
## Get top-k from bi-encoder
k = 4 # Retrieve top 4 for reranking
top_k_indices = np.argsort(similarities)[::-1][:k]
top_k_docs = [documents[i] for i in top_k_indices]
top_k_scores = similarities[top_k_indices]Top-4 documents from bi-encoder: 1. [Score: 0.750] To reset your password, click 'Forgot Password' on the login page. Enter your re... 2. [Score: 0.698] Account recovery options include password reset via email, security questions, o... 3. [Score: 0.473] If you suspect unauthorized access to your account, immediately change your pass... 4. [Score: 0.447] Email notifications can be configured in your account settings. You can choose t...
The bi-encoder returns candidates based on semantic similarity. Notice that documents mentioning "password" and "email" rank highly, but the specific condition about forgetting the email is not well captured by vector similarity alone. The bi-encoder identifies documents in the correct semantic neighborhood but lacks the discriminative power to verify whether the specific constraint mentioned in the query is satisfied by each document.
Now we load a cross-encoder to rerank these candidates. Cross-encoders are typically larger models that process query-document pairs jointly. Unlike the bi-encoder, which produced independent vectors, the cross-encoder will process the concatenated text of the query with each candidate document, letting fine-grained interaction between the query's specific constraints and the document's content.
## Load cross-encoder for reranking
print("Loading cross-encoder...")
import time
from sentence_transformers import CrossEncoder
cross_encoder = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
## Prepare pairs for cross-encoder
pairs = [[query, doc] for doc in top_k_docs]
## Time the reranking
start_time = time.time()
cross_scores = cross_encoder.predict(pairs)
rerank_time = time.time() - start_time
## Sort by cross-encoder scores (argsort gives indices of sorted elements, so reranked_indices[i] is the original position)
reranked_order = np.argsort(cross_scores)[::-1]
reranked_docs = [top_k_docs[i] for i in reranked_order]
reranked_scores = cross_scores[reranked_order]
## For display: reranked_indices[i] should give the original rank (1-indexed) of the i-th reranked doc
reranked_indices = reranked_order + 1 # Convert to 1-indexed original ranksReranking completed in 61.4ms Reranked results: 1. [Score: 7.156, originally #1] To reset your password, click 'Forgot Password' on the login page. Enter your re... 2. [Score: 4.691, originally #2] Account recovery options include password reset via email, security questions, o... 3. [Score: -2.313, originally #4] Email notifications can be configured in your account settings. You can choose t... 4. [Score: -2.403, originally #3] If you suspect unauthorized access to your account, immediately change your pass...
The cross-encoder reranks the documents, potentially changing the order based on fine-grained relevance signals that the bi-encoder missed. The document about alternative verification methods for lost email access should now rank higher than the generic password reset instructions. This reordering demonstrates how the cross-encoder identifies the specific relevance signals that distinguish a satisfactory answer from a superficially similar but ultimately unhelpful document.
Let us compare the latency characteristics of both approaches to understand the computational tradeoff. This benchmark illustrates why we reserve the cross-encoder for only a small set of candidates rather than applying it to the entire corpus.
# Benchmark latency comparison
print("Latency Benchmark")
print("=" * 50)
# Bi-encoder timing (query encoding only, assuming pre-indexed docs)
n_trials = 10
bi_times = []
for _ in range(n_trials):
start = time.time()
q_emb = bi_encoder.encode([query])
# Simulate dot product with precomputed index (very fast)
_ = cosine_similarity(q_emb, doc_embeddings.cpu().numpy()[:k])
bi_times.append(time.time() - start)
# Cross-encoder timing
cross_times = []
for _ in range(n_trials):
start = time.time()
_ = cross_encoder.predict(pairs)
cross_times.append(time.time() - start)Bi-encoder (query + top-4 retrieval): 10.34ms +/- 1.02ms Cross-encoder (score 4 pairs): 13.15ms +/- 1.70ms Slowdown factor: 1.3x
The cross-encoder is significantly slower, often by one to two orders of magnitude. This demonstrates why we use it only on a small candidate set rather than the entire corpus. In production systems, this latency difference necessitates careful architectural decisions about how many documents to rerank and whether to employ optimization techniques such as batching, quantization, or cascade filtering.
Finally, let us examine the attention patterns to understand how the cross-encoder makes its decisions. We will use the cross-encoder's ability to output attention weights. This visualization helps illustrate the token-level interactions that enable the cross-encoder to distinguish between documents that share vocabulary but differ in their relevance to the specific query.
# Use the worked example's relevant document so the simulated alignments remain
# stable even if an upstream model version ranks the small demo corpus differently.
attention_document = documents[1]
# Tokenize to show alignment
query_tokens = query.split()
doc_tokens = attention_document.split()
use_book_style()
plt.rcParams["figure.figsize"] = (6.0, 4.0)
# Create a simplified visualization of token-level interaction
# In practice, you would extract actual attention weights from the model
fig = plt.figure()
ax = plt.gca()
# Simulate attention pattern showing key alignments
# Real attention would be extracted from model layers
attention_matrix = np.random.rand(len(query_tokens), len(doc_tokens)) * 0.1
# Highlight key alignments
query_keywords = ["forgot", "email", "account"]
doc_keywords = ["access", "registered", "email", "alternative", "verification"]
for i, qt in enumerate(query_tokens):
for j, dt in enumerate(doc_tokens):
qt_lower = qt.lower().strip("?,.\"'")
dt_lower = dt.lower().strip(".,;\"'")
# Strong attention for keyword matches
if qt_lower in ["email", "password", "account"] and dt_lower in [
"email",
"password",
"account",
]:
attention_matrix[i, j] = 0.8
elif qt_lower == "forgot" and dt_lower in ["access", "longer"]:
attention_matrix[i, j] = 0.7
elif qt_lower == "reset" and dt_lower in ["recovery", "reset"]:
attention_matrix[i, j] = 0.6
image_artist = ax.imshow(
attention_matrix, cmap="Blues", aspect="auto", interpolation="nearest"
)
# Set ticks
ax.set_xticks(range(len(doc_tokens)), doc_tokens, rotation=45, ha="right")
ax.set_yticks(range(len(query_tokens)), query_tokens)
ax.set_xlabel("Document Tokens")
ax.set_ylabel("Query Tokens")
ax.set_title("Cross-Encoder Attention: Query-Document Alignment")
ax.grid(False)
fig.colorbar(image_artist, ax=ax, label="Attention Weight")
plt.show()
Query: How do I reset my password if I forgot the email associated with my account? Document visualized: Account recovery options include password reset via email, security questions, or contacting support...
The attention visualization reveals how the cross-encoder aligns specific query tokens with relevant document tokens. Unlike the bi-encoder, which produces a single similarity score, the cross-encoder can match "forgot" with "no longer have access" even though these phrases share no words. This shows the power of contextualized cross-attention. This ability to recognize semantic equivalence across different phrasings is what enables the cross-encoder to correctly identify Doc B as the most relevant, despite its lower lexical overlap with the query compared to Doc A.
Key Parameters
The key parameters for the retrieve-then-rerank implementation deserve careful consideration, as they fundamentally shape the behavior and performance of your system. Choosing these parameters correctly requires understanding both the mathematical properties of the models and the empirical constraints of your deployment environment.
The most important parameters are:
-
k (reranking budget): Number of candidates to retrieve for reranking. Higher values improve recall for the reranker but increase latency linearly. Selecting the appropriate value for requires balancing the probability that the relevant document appears in the top (recall) against the latency constraints of your application. In practice, is often determined empirically by measuring the recall@k of the initial retriever on a validation set and choosing the smallest that achieves acceptable recall. A common heuristic is to set ten to twenty times larger than the number of final results you want to return.
-
convert_to_tensor: Whether to return PyTorch tensors for the embeddings, letting GPU acceleration for similarity computations. Using GPU acceleration for both bi-encoder and cross-encoder operations is needed for achieving acceptable latency in production systems. The tensor format also facilitates batch processing, which can significantly improve throughput when processing multiple queries simultaneously.
-
model: The specific cross-encoder architecture (e.g.,
cross-encoder/ms-marco-MiniLM-L-6-v2). Larger models provide better accuracy at the cost of increased inference time. Model selection involves trading off between the improved discrimination of larger architectures (such as large BERT or RoBERTa variants) and the latency requirements of your use case. Models trained on MS MARCO, a large-scale passage ranking dataset, generally provide strong zero-shot performance for general-domain retrieval tasks. -
max_length: Maximum sequence length for the cross-encoder (typically 512 tokens). Documents longer than this are truncated, potentially losing relevant information at the end. When working with long documents, consider using the sliding window approach discussed earlier, or select models specifically designed for long sequences (such as those using Longformer or BigBird architectures) that can handle thousands of tokens.
Limitations and Impact
While cross-encoder reranking dramatically improves retrieval accuracy, it introduces significant operational constraints that shape system architecture. The quadratic complexity with respect to sequence length means that processing long documents becomes prohibitively expensive. You cannot easily rerank entire books or lengthy research papers without aggressive truncation or chunking, which risks losing relevant information buried in the middle or end of texts. This limitation has driven research into efficient transformers and hierarchical approaches that can process long documents without quadratic scaling. The practical consequence is that most reranking systems operate on chunks or passages rather than full documents, which shifts complexity from the reranking step to the document preprocessing and chunking strategy.
The lack of precomputation also requires running a forward pass through a large transformer for every query-document pair at query time, creating a throughput bottleneck that limits the number of candidates you can rerank within latency budgets. Unlike bi-encoders, which amortize document encoding over all queries, cross-encoders must reprocess documents for every new query. This architectural constraint means that cross-encoders scale with query volume in ways that can strain GPU resources during traffic spikes. A system that works well for ten queries per second may fail catastrophically at one hundred queries per second if the GPU cannot keep up with the reranking demands. Horizontal scaling requires careful load balancing and may require maintaining multiple replicas of the cross-encoder model, increasing infrastructure costs significantly.
Cross-encoders also inherit the context length limitations of the underlying transformer. Standard BERT-based models process at most 512 tokens, which is sufficient for short passages but inadequate for longer documents without chunking. Even extended-context models face practical limitations: processing a 4,096-token document requires roughly 64 times more compute than processing a 512-token document due to quadratic attention scaling. This creates asymmetric behavior where the reranker performs well on short, focused passages but degrades on longer, more complex documents that may require evidence from across the full text to judge relevance.
These limitations have spurred the development of intermediate architectures like ColBERT and SPLADE that attempt to preserve token-level interaction capabilities while maintaining indexable representations. They also motivate the cascade approaches mentioned earlier, where you progressively filter candidates through increasingly expensive models. Reranking established the two-stage retrieve-then-rerank paradigm as the standard for neural information retrieval, allowing systems to balance recall and precision within a latency budget. Without reranking, dense retrieval systems would struggle to achieve the accuracy required for production search engines and retrieval-augmented generation applications, where returning an irrelevant document in the top position significantly degrades user experience.
Looking ahead, the reranking paradigm is increasingly intersecting with large language models as judges. Rather than using a fine-tuned cross-encoder, some recent approaches use a general-purpose LLM to assess the relevance of each retrieved document by prompting it with the query and document and asking it to rate relevance on a numerical scale. This uses the LLM's language understanding without requiring task-specific fine-tuning data, though it comes with even higher latency costs than traditional cross-encoders. As we move toward more sophisticated RAG pipelines in upcoming chapters on RAG Prompt Engineering and Evaluation, the reranker serves as a quality filter so that only the most relevant context reaches the language generator. Selecting the right passage can mean the difference between a correct and incorrect answer, making the reranker one of the most consequential components in the RAG system.
Summary
Reranking addresses the basic limitation of bi-encoder retrieval systems: the inability to model fine-grained interactions between queries and documents due to independent encoding. By processing query-document pairs through cross-attention mechanisms, cross-encoders capture fine-grained relevance signals that vector similarity misses. The key mechanism is joint bidirectional self-attention over the concatenated query-document sequence, which allows every query token to directly interact with every document token, letting the model to verify the presence of specific information, detect semantic mismatches, and assess whether a document's content truly satisfies the query's constraints.
The standard retrieve-then-rerank pipeline uses fast bi-encoders to recall a candidate set, then applies computationally intensive cross-encoders to precisely rank the top candidates. Training these models involves binary classification, pairwise ranking losses, or listwise objectives, with hard negative mining being needed for learning discriminative representations. The curriculum of hard negatives, progressing from easy to difficult examples, enables the model to learn the subtle distinctions necessary to improve upon initial retrieval. Choosing the right training objective, sampling strategy, and negative mining procedure are as important to final model quality as the architecture choices.
The primary constraint is latency: cross-encoders are orders of magnitude slower than bi-encoders due to their inability to precompute document representations. This necessitates careful budgeting of the number of candidates to rerank, cascade architectures with multiple ranking stages, and the exploration of efficient alternatives like late interaction models. Late interaction models like ColBERT occupy an intermediate position in the efficiency-accuracy spectrum, using token-level representations that can be precomputed and indexed while still letting fine-grained matching at query time through the MaxSim operator.
When implementing retrieval systems, you should view reranking not as an optional enhancement but as a necessary component for achieving high precision, particularly in domains where subtle distinctions determine relevance. The choice of reranking budget involves trading accuracy for latency, a decision that depends on your specific application requirements and user expectations. The two-stage retrieve-then-rerank architecture has become the foundation of modern information retrieval and RAG systems precisely because it resolves the basic tension between the efficiency needed for corpus-scale search and the accuracy needed for high-quality answers. Understanding how its components interact and which tradeoffs shape optimization gives you the tools to build retrieval systems that perform well both in the laboratory and in production.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about reranking in information retrieval systems.
Reranking: Cross-Encoders and Retrieve-Then-Rerank Pipelines
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!