All articlesEngineering

A field guide to RAG reranking

Marcus Ibe10 min read
Engineering123

Cross-encoder rerankers score a query against each candidate document jointly, which makes them the most accurate option and the slowest — every candidate needs its own forward pass through the model. Understanding the tradeoffs across the main reranking approaches is worth doing properly, because the wrong choice for your data shows up as either wasted latency or missed accuracy, and the difference isn't always obvious until you've measured it directly.

We've now tried all three of the approaches below in production at some point, on different parts of our own retrieval pipeline, and this is the practical comparison we wish we'd had going in.

Cross-encoderspeedLLM-as-judgespeedHybridspeed

Cross-encoder rerankers

A cross-encoder takes the query and a single candidate document together as joint input, letting the model attend across both simultaneously rather than comparing two separately-computed representations. That joint attention is what makes cross-encoders the most accurate reranking approach we've tested — they can catch subtle relevance signals that a separately-encoded comparison misses entirely.

The cost is that every candidate needs its own full forward pass, which doesn't parallelize as cheaply as comparing pre-computed embeddings. For a shortlist of ten or twenty candidates from an initial retrieval pass, this is entirely manageable. For a much larger shortlist, the latency adds up quickly, which is why cross-encoders are almost always used as a second-stage reranker over a small candidate set, never as the first-stage retrieval mechanism itself.

LLM-as-judge reranking

LLM-as-judge reranking asks a general-purpose model to read the candidates and pick the best one directly, or score each one, using its general reasoning ability rather than a model trained specifically for relevance scoring. It's more flexible since there's no separate model to train or maintain, but noticeably slower and more expensive per query than a dedicated cross-encoder, because a general-purpose model is doing more computation than the narrow task strictly requires.

The flexibility is real, though, and matters most when the notion of "relevance" itself is unusual or hard to define cleanly with training data — a general-purpose model can be given nuanced instructions about what counts as a good match for a specific use case, in a way that's much harder to bake into a purpose-trained cross-encoder without collecting a large, carefully-labeled dataset first.

Hybrid scoring

Hybrid scoring blends the original vector similarity score with a lightweight secondary signal — recency, source authority, or a cheap keyword match — without adding a full second model pass. It's the fastest option and good enough when your documents are already well-curated, because a lot of the relevance signal that a heavier model would need to infer is already implicitly captured by simple metadata like how recently a document was updated or which source it came from.

The tradeoff is that hybrid scoring can't catch the kind of subtle semantic mismatch a cross-encoder would — a document that's recent, from an authoritative source, and shares vocabulary with the query, but still doesn't actually answer the specific question, will score well under hybrid scoring and poorly under a cross-encoder's joint judgment. Hybrid scoring is a bet that this specific failure mode is rare enough in your data not to matter much.

Which one to pick

We default new deployments to hybrid scoring and only reach for a cross-encoder once a team's document set is large and varied enough that similarity alone starts surfacing the wrong answer often enough to notice — which, in our experience, is a threshold most teams don't hit until their Knowledge Base has grown well past its initial size and started accumulating documents with genuinely overlapping vocabulary covering different specifics.

LLM-as-judge reranking, in practice, ends up being the right choice least often for us — not because it doesn't work, but because the specific flexibility it offers over a cross-encoder rarely justifies its added latency and cost for the fairly well-defined notion of relevance that customer support questions typically need. We've kept it available for specific cases with unusual relevance criteria, but it's not our default recommendation.

Common questions

Can you mix approaches — hybrid scoring for most queries, cross-encoder for a specific subset? Yes, and this is common in practice once a team has enough traffic to distinguish query types worth treating differently, though we'd recommend starting with a single consistent approach until you have the data to justify the added complexity of routing between two.

How much labeled data does training a custom cross-encoder actually require? Less than teams typically expect, if the labels come from real usage rather than being manufactured from scratch — a few thousand query-document pairs with a relevance judgment, ideally derived from actual customer interactions and reviewer feedback, is usually enough to meaningfully outperform an off-the-shelf reranker on your specific domain.

A direct comparison on the same data

To make the tradeoffs between these three approaches concrete rather than theoretical, we ran all three against the exact same evaluation set of customer support queries and candidate documents, specifically to see how the abstract tradeoffs described above played out on data that actually matters to a support use case rather than a generic benchmark.

Hybrid scoring, as expected, was the fastest by a wide margin and performed well on the majority of queries — anything where the correct document was both topically similar and reasonably recent or authoritative by our metadata signals. It struggled specifically on a subset of queries where an older but still-correct document scored lower on recency than a newer, superficially similar but actually wrong document, which is exactly the failure mode we'd expect given how hybrid scoring weights its signals.

The cross-encoder reranker corrected almost all of those specific failures, at a meaningfully higher latency cost per query. What surprised us slightly was how narrow the accuracy gap actually was outside of that specific failure mode — on the majority of queries where hybrid scoring already did fine, the cross-encoder didn't meaningfully improve on it, which reinforced our sense that a cross-encoder's real value is concentrated in a specific, identifiable class of hard cases rather than being a uniform improvement across every query.

LLM-as-judge reranking matched the cross-encoder's accuracy on most of the same hard cases, but at a cost and latency profile that made it hard to justify running on every query the way we'd run the other two approaches. Where it distinguished itself was on a small number of queries with genuinely unusual, context-dependent relevance criteria — cases where what counted as "the right document" depended on a nuance the cross-encoder's narrower training hadn't specifically seen enough examples of.

That comparison is what settled us on the hybrid-scoring-by-default, cross-encoder-on-a-flagged-subset approach we described briefly above, rather than picking a single approach to run uniformly across every query. Routing a query to the more expensive cross-encoder pass only when hybrid scoring's own confidence signal is ambiguous captures most of the accuracy benefit of always using a cross-encoder, at a small fraction of the aggregate latency and cost.

We'd recommend any team facing this same choice run a similarly direct, side-by-side comparison on their own actual query and document distribution before committing to one approach, because the specific tradeoffs depend heavily on how much vocabulary overlap exists across genuinely different documents in a given Knowledge Base — a smaller, more curated Knowledge Base may never need anything beyond hybrid scoring, while a large, fast-growing one is more likely to benefit from at least a targeted cross-encoder pass on ambiguous cases.

Common questions

Is there a reranking approach that works well specifically for very short documents, like single-sentence policy statements, where there's not much text to compare against? Short documents are actually where hybrid scoring tends to do relatively well, since there's less room for the kind of subtle semantic mismatch a cross-encoder is built to catch — the harder cases for all three approaches tend to involve longer documents covering multiple related but distinct topics.

How much does reranking quality depend on the quality of the underlying first-stage retrieval, versus being able to compensate for a weak first stage? Reranking can only improve on what the first-stage retrieval actually surfaces in the shortlist — if the correct document never makes it into the shortlist at all, no reranking approach can recover it, which is why first-stage retrieval quality still matters even once a strong reranker is in place.

Does the choice between these three approaches depend on the underlying generation model being used downstream, or is it independent? It's largely independent — the reranker's job is to hand the best-supported document to the generation step, and that job doesn't change based on which model receives the output, though a stronger generation model can occasionally paper over a mediocre reranking choice by reasoning more carefully with an imperfect source.

Is it possible to combine hybrid scoring and cross-encoder scoring into a single blended score, rather than choosing one or routing between them? Yes, and some teams do this — blending a cross-encoder's relevance score with hybrid signals like recency into a single combined ranking — though we've found routing (using hybrid scoring by default, cross-encoder only on ambiguous cases) simpler to reason about and tune than a blended score with multiple weighted components.

How often should the choice of reranking approach be revisited as a Knowledge Base grows? We revisit this at the same cadence as other retrieval-quality checks, roughly quarterly, since the vocabulary overlap and document-set characteristics that make one approach more or less suitable can shift meaningfully as new documentation categories get added over time.

Does reranking approach choice interact with which underlying model you use for generation? Only loosely — the reranker's job ends once it hands the best-ranked document to the generation step, and a better reranker helps regardless of which generation model receives its output, though a stronger generation model can sometimes partially compensate for a weaker reranker by reasoning more carefully about an imperfect source.

Is it worth reranking on every single query, or only ones flagged as ambiguous? We rerank every query by default, because the added latency is small relative to the cost of an occasional wrong-source answer on a query nobody thought to flag as ambiguous in advance — ambiguity is often only obvious in hindsight, once you already know an answer was wrong.

If you take away only one practical recommendation from this comparison, it's to test against your own data before committing to an approach, rather than assuming the tradeoffs described here transfer exactly to your specific Knowledge Base. The general shape of the tradeoffs — cross-encoders are the most accurate and slowest, LLM-as-judge is the most flexible and most expensive, hybrid scoring is the fastest and good enough for well-curated data — has held consistently across the comparisons we've run, but the specific point at which one approach clearly outperforms another depends enough on your document set's own characteristics that we'd treat this piece as a starting framework, not a substitute for running the comparison yourself.

One practical detail we'd add for anyone actually implementing the routing approach described above, between hybrid scoring by default and a cross-encoder on ambiguous cases: the signal used to decide "is this ambiguous enough to warrant the more expensive pass" matters as much as the reranking approaches themselves. We use the gap between the top two hybrid-scored candidates as our ambiguity signal — a large gap means the top candidate is confidently ahead and doesn't need a second opinion, while a narrow gap between the top two is exactly the situation where a cross-encoder's more careful judgment tends to change the outcome.

Getting that specific threshold right took its own round of tuning, separate from tuning either reranking approach individually, and it's worth calling out as its own decision rather than assuming it falls out naturally once you've picked your two reranking methods.