Retrieval
Retrieval is the step of finding the most relevant pieces of content for a given query, before any text is generated.
What Is Retrieval?
Retrieval is the search step: given a question, find the pieces of content most likely to help answer it. It happens before generation, and it is a separate concern from it — retrieval decides what the model gets to see; generation decides what the model says with it.
Question
Before GenerationWhat the model needs help answering.
Search
Separate ConcernDecides what the model gets to see, not what it says.
Candidate Documents
Possible MatchesContent likely to help answer the question.
Top Results
Handed to GenerationWhat actually reaches the model.
Key Idea
Retrieval and generation are different problems. A great retriever with a weak model still limits answer quality — and a great model fed irrelevant context will still produce a bad answer.
Ways to Search
- Lexical search
- Matches based on exact or near-exact keywords, similar to traditional full-text search — strong for exact terms, IDs, and rare vocabulary.
- Vector search
- Matches based on semantic similarity between embeddings, finding related meaning even without shared keywords.
- Hybrid search
- Combines lexical and vector search, aiming to get the strengths of both rather than relying on either alone.
"Top-k" refers to how many candidate results a retriever returns — for example, the 5 or 10 most relevant chunks. Choosing k is itself a tradeoff: too few risks missing the right context, too many risks diluting the prompt with less relevant material.
Query Formulation
The raw question a user types is not always the best search query. Systems often rewrite or expand the query first — correcting ambiguity, adding synonyms, or breaking a complex question into simpler sub-queries — before running the actual search.
Common Mistakes
Assuming vector search alone is always enough
Keyword and hybrid search often outperform pure vector search for exact terms, IDs, or rare vocabulary that embeddings can blur together.
Confusing a retrieval problem with a generation problem
If the model was never given the right context, no amount of prompt tweaking on the generation side will fix the answer.
Not inspecting what was actually retrieved
The fastest way to debug a bad answer is to look directly at the retrieved content, not just the final output.
Interview Question
What is retrieval, and how is it different from generation?
Retrieval is the search step in a RAG system — given a question, finding the most relevant content from a knowledge source, typically using lexical search, vector search, or a hybrid of both. Generation is the separate step where a language model produces an answer using that retrieved content as context. Keeping them distinct matters for debugging: if an answer is wrong because the right information was never retrieved, that's a retrieval problem, not something you can fix by adjusting how the model generates.
What an interviewer may ask next
- What is the difference between lexical, vector, and hybrid search?
- How would you choose the value of top-k?
- Why might a retriever return technically similar but unhelpful results?
Explain It in 30 Seconds
Retrieval is the step of finding the most relevant content for a question before any answer is generated — using lexical search, vector search, or a hybrid of both. It’s a distinct problem from generation: retrieval decides what the model gets to see, generation decides what it says with it. When an answer is wrong, checking what was actually retrieved is usually the fastest way to find out why.