How we cut hallucinations with a custom reranker
Standard vector search gets you the documents that are semantically close to a question — not necessarily the ones that actually answer it. That gap is where most hallucinations start, and it's a subtler failure than most teams expect when they first plug in a vector database and call retrieval solved.
We spent a meaningful chunk of an engineering quarter on exactly this gap, and the fix ended up being a reranking layer rather than a bigger model or a better embedding.
The problem with retrieval alone
A question about refund windows and a document that merely mentions the word "refund" in passing can end up close together in embedding space, simply because they share vocabulary, not because one answers the other. Vector similarity is a proxy for relevance, and like most proxies, it's good enough most of the time and quietly wrong in a way that's hard to notice until you go looking for it specifically.
We found this by sampling a batch of answered questions and manually checking whether the source document retrieved actually supported the answer given, rather than just checking whether the final answer sounded right. A meaningful fraction of technically-correct-sounding answers were built on a document that was topically adjacent rather than actually authoritative — the model had filled the gap between a loosely relevant source and a specific question with its own confident extrapolation.
How the reranker works
We added a reranking pass between retrieval and generation: the top candidates from vector search get scored a second time by a smaller model trained specifically to judge relevance, not just similarity. The distinction matters because relevance and similarity are correlated but not identical — a document can be highly similar in language and still not actually contain the specific answer, and a reranker's whole job is catching that gap before generation ever sees the document.
In practice this means a question about refund windows doesn't get answered from a document that merely mentions the word "refund" in passing — it gets answered from the paragraph that actually states the policy, because that paragraph scores higher on the second pass even if it scored similarly, or even slightly lower, on the first vector-similarity pass.
The reranker itself doesn't need to be large. We use a compact model trained on exactly this narrow task — given a query and a candidate passage, score how well the passage actually answers the query — rather than a general-purpose model asked to reason about relevance from scratch each time. Purpose-built and small outperformed general-purpose and large on this specific task, which surprised us less once we understood how narrow the actual judgment being made really was.
What changed after we shipped it
The tradeoff is a few hundred milliseconds of added latency per query, from the extra scoring pass. Given how much it cut down on wrong-but-confident answers in our internal evals, it was an easy call — the added latency is imperceptible to a customer waiting for a chat response, while a wrong-but-confident answer is the kind of failure that erodes trust in a way that's hard to win back with a single correct answer afterward.
We measured the effect specifically on the class of questions that had been most prone to this failure — narrow policy questions where several documents use similar vocabulary but only one is actually authoritative. Reranking closed most of that specific gap, while leaving broader, less ambiguous questions essentially unaffected, which is roughly what we'd expect if the mechanism we identified was the right one.
The other effect, less expected going in, was on how confidently the model hedged. Once it was reliably being handed the actually-relevant document rather than an adjacent one, it needed to hedge less often, because the source material genuinely supported a direct answer more of the time. Less hedging read, to customers, as more competence — even though nothing about the underlying model's capability had changed, only the quality of what it was given to work with.
Common questions
Why not just fine-tune the retrieval embeddings instead of adding a whole second scoring pass? We tried this first, actually, and it helped somewhat but didn't close the gap the way reranking did. Embeddings are optimized for fast approximate similarity across a huge candidate set; asking them to also carry fine-grained relevance judgment asks them to do two jobs at once, and a dedicated second pass, freed from the speed constraints of the first, can afford to be more discriminating.
Does this approach still make sense for a small knowledge base, or is it overkill? It scales down reasonably well — even a small knowledge base can have several documents that share vocabulary without all being equally relevant to a specific question, and the reranker's job doesn't get meaningfully harder or easier based on corpus size, only based on how much vocabulary overlap exists between genuinely different topics.
A specific case that convinced the team
The clearest before-and-after example we have came from a batch of questions about a mid-tier subscription plan's specific feature limits — a topic where three separate documents in the Knowledge Base each mentioned the plan by name, but only one of them actually stated the current, correct limits, with the other two being either outdated or discussing a related but different plan tier entirely.
Before reranking, roughly a third of the sampled questions on this specific topic pulled from one of the two incorrect documents, producing confident, wrong answers about feature limits — not because the model was reasoning poorly, but because vector similarity genuinely couldn't distinguish which of the three similarly-worded documents was the current, authoritative one for this specific plan.
After adding the reranker, trained specifically to score how well a candidate passage answers a given query rather than just how similar it sounds, the same batch of questions pulled from the correct document in the vast majority of cases. The reranker had, in effect, learned to notice a distinction the embedding model couldn't: the correct document stated the limits directly and unambiguously in response to a question shaped like the ones being asked, while the other two documents merely mentioned the plan in passing as part of a different discussion.
We used this specific before-and-after comparison internally as the case that got the rest of the engineering team fully convinced the reranking layer was worth the added latency, more than any aggregate metric could. It's one thing to see an overall accuracy number improve by some percentage; it's another to see the exact same question, asked the exact same way, go from confidently wrong to reliably correct once a specific fix is in place.
The broader lesson we took from this case, beyond the specific fix, was about how to prioritize what to test for hallucination risk in the first place. Topics where multiple documents share vocabulary but differ in currency or specificity — old versus new pricing tiers, discontinued versus current policies, a general answer versus a specific exception — are exactly the topics most likely to produce this kind of confident, retrieval-driven wrongness, and are worth auditing specifically rather than waiting to discover the pattern by accident the way we did here.
We now run a standing audit, roughly quarterly, specifically looking for clusters of documents that share significant vocabulary while differing in some detail that actually matters to the answer — exactly the pattern that caused this case — rather than relying on customer complaints or the Content Gap Analyzer to surface it after the fact.
That audit has caught two or three similar latent risks since, in each case before they'd generated enough wrong answers to show up clearly in any aggregate metric, which is roughly the point: the failure mode described in this piece is quiet by nature, and the fix for that quietness is deliberately looking for it rather than waiting for it to become loud.
Common questions
Does adding a reranker ever make an answer worse rather than better, and how do you catch that if it happens? It can, if the reranker itself has a systematic bias, which is why we monitor its effect specifically on a held-out set of known-correct answers over time rather than assuming the initial evaluation result holds forever as the underlying document set changes and grows.
How do you decide the size of the initial candidate shortlist that gets reranked, versus the full document set? We use vector search to narrow a large document set down to a manageable shortlist first, then rerank only that shortlist — reranking against the entire document set directly would defeat the latency benefit of having a fast first-stage retrieval step at all.
Is there a risk that the reranker itself becomes stale, the same way underlying documentation can go stale? Yes, in a different way — the reranker's judgment of relevance can drift out of alignment with a Knowledge Base that's grown to cover new topics it wasn't originally trained against, which is why we retrain it periodically against a refreshed sample of real queries and documents rather than treating the original training run as permanent.
Can a reranker introduce its own kind of bias, favoring certain document styles or lengths over others regardless of actual relevance? This is a real risk with any learned reranking model, and it's part of why our evaluation process specifically checks for this pattern — comparing reranker scores across documents of different lengths and styles that are equally relevant, to confirm the reranker isn't simply preferring, say, longer documents by default.
What's the actual cost difference in practice between running hybrid scoring alone versus adding a full reranking pass? The added cost is dominated by the extra model inference per candidate, which scales with shortlist size — for a typical ten-to-twenty candidate shortlist, the added cost per query is small in absolute terms, though it adds up meaningfully at very high query volumes, which is part of why we default to hybrid scoring and reserve the full reranking pass for the cases described elsewhere in this piece where it earns its cost back.
How often does the reranker itself need retraining? Less often than the underlying retrieval embeddings, in our experience, because the task it's learning — judging relevance given a query and a passage — is more stable over time than the specific vocabulary of a growing knowledge base. We retrain on a slower cadence than we update the Knowledge Base itself.
What's the failure mode if the reranker itself gets something wrong? It demotes a genuinely relevant document below one that looked plausible but wasn't, which is a real risk, but one we've found easier to catch than the original failure — because a reranking mistake tends to produce an answer that's subtly off-topic rather than confidently wrong, which is a much easier pattern to catch in review than a confident hallucination built on a wrong-but-plausible source.
It's worth being clear that reranking doesn't eliminate hallucination risk entirely — it closes a specific, well-understood gap between similarity and relevance, which happens to be a common and consequential gap, but not the only source of a wrong answer. A model can still misread a correctly-retrieved document, or extrapolate slightly beyond what a document actually states, and neither of those failure modes is what a reranker is built to catch.
That's why we treat reranking as one layer in a broader defense against hallucination, alongside the confidence-tiered fallback system and the traceability work described elsewhere in this series, rather than as a complete solution on its own. Each of those pieces catches a different specific failure mode, and the honest accounting of what reranking actually fixes — retrieval choosing the wrong document among several plausible ones — is more useful than an overstated claim that it solves hallucination generally.
There's also a maintenance dimension worth naming plainly for anyone deciding whether to build this internally versus treating it as a solved problem to buy off the shelf. A reranker is a model with its own lifecycle — its own training data, its own evaluation set, its own drift over time as the underlying Knowledge Base grows into new topics the original training data didn't anticipate. Teams that treat it as a build-once component, the same way they might treat a static configuration file, tend to see its benefit erode gradually and invisibly over a year or two, in exactly the way the original hallucination problem was invisible until someone went looking for it specifically.
The team that owns this at our end reviews reranker performance on the same quarterly cadence as the other standing audits described throughout this series, checking a fresh sample of real queries against the reranker's current scoring behavior rather than trusting the original launch-time evaluation to remain representative forever. That review has caught drift twice so far — both times traced to new documentation categories that hadn't existed when the reranker was originally trained, and both times fixed with a targeted retraining pass rather than a full rebuild.