Reranking
Reranking reorders an initial set of retrieved results using a more precise (and more expensive) model, to surface the most relevant ones first.
Prerequisites
Why Add a Second Pass?
First-stage retrieval — vector or hybrid search — is built to be fast across a large collection, which usually means it makes some accuracy tradeoffs. A reranker takes the broad set of candidates that first stage returns and re-scores them using a slower, more precise model, keeping only the strongest matches.
Query
Starting SearchWhat the user is trying to find.
Broad Retrieval
Fast, Less PreciseBuilt for speed across a large collection.
Candidates
First-Pass ResultsA broader set than what actually gets used.
Reranker
Slower, More PreciseA second, more expensive scoring pass.
Best Context
Strongest MatchesOnly what survives the second pass.
Why It Can Help
First-stage retrieval typically scores a query against each document independently and quickly. A reranker can look at the query and a candidate document together, more carefully weighing how well they actually match — which tends to catch cases where something looked similar by embedding distance but isn't actually the most useful answer.
Warning
Reranking doesn't universally improve every workload. It adds latency and cost by running an extra model over the candidates, so it's worth measuring whether it actually improves your results before adding it as a fixed step.
Common Mistakes
Reranking a huge candidate set
Rerankers are more expensive per item than first-stage retrieval — they work best on a modest shortlist, not thousands of raw results.
Assuming reranking fixes a weak retriever
A reranker can only reorder what first-stage retrieval actually returned — it can’t surface a relevant document that was never retrieved at all.
Adding it without measuring the impact
The latency and cost are real; skipping evaluation means you can’t confirm the quality improvement is worth that tradeoff.
Interview Question
Why would you add a reranking step to a RAG system?
First-stage retrieval is built for speed across a large collection, which usually trades off some precision. A reranker takes that broader candidate set and re-scores it with a more precise, more expensive model, often looking at the query and each candidate together rather than independently — which can surface the genuinely most relevant results higher, even when they weren't the closest by initial similarity score. The tradeoff is added latency and cost, so it's worth validating that it actually improves results for your workload rather than adding it by default.
What an interviewer may ask next
- Why not just use the reranker for retrieval directly, skipping first-stage search?
- What are the latency and cost tradeoffs of reranking?
- How would you measure whether reranking is actually improving results?
Explain It in 30 Seconds
Reranking takes the broad set of candidates from first-stage retrieval and re-scores them with a more precise, more expensive model, keeping the strongest matches. It can catch cases where first-stage search missed the best answer due to its speed-focused tradeoffs, but it adds latency and cost — so it’s worth measuring whether it actually helps your specific workload rather than assuming it always will.