Query Transformation
Query transformation rewrites or expands a user's query before retrieval to improve the chances of finding relevant results.
Prerequisites
Why Transform a Query?
Users don't always phrase questions in a way that retrieves well. A short, ambiguous, or conversationally-phrased query — "what about the second one?" referring to something several messages earlier — may not embed or keyword-match anything useful on its own. Query transformation rewrites the raw user query into one or more versions better suited for retrieval, before the retrieval step ever runs.
User Query
Raw PhrasingMay be short, ambiguous, or conversational.
Query Transformation
Rewritten for SearchRuns before retrieval ever happens.
Retrieval
Better MatchesWorks against a version suited for search.
Generation
Grounded AnswerBuilt from what retrieval actually found.
Common Techniques
- Rewriting — using conversation history to turn an ambiguous follow-up question into a clear, self-contained query.
- Expansion — adding related terms or synonyms to the query to broaden what a keyword search can match.
- Decomposition — splitting a complex, multi-part question into several simpler sub-queries that are each retrieved separately.
- Multi-query retrieval — generating several differently-phrased versions of the same query, retrieving for each, and merging the results, to reduce the risk of a single phrasing missing relevant content.
Tip
Query transformation is usually done by asking an LLM to rewrite the query — a small, fast model call before the main retrieval step.
A Real-World Example
In a multi-turn conversation, a user asks "How does chunking work?" and then follows up with "What about for code?" Retrieving on "What about for code?" alone would likely fail — it has almost no useful retrieval signal on its own. Rewriting it using the conversation history into "How does chunking work for source code?" gives the retrieval step something concrete to search on.
Common Mistakes
Retrieving on the raw conversational query in multi-turn chat
Follow-up questions often depend on earlier context that has to be folded in before retrieval can work well.
Over-expanding a query with too many unrelated terms
Excessive expansion can pull in unrelated content and dilute retrieval quality instead of improving it.
Adding query transformation to every request regardless of need
Transformation adds latency and an extra model call — simple, well-formed queries may not need it.
Assuming multi-query retrieval always improves results
More query variants mean more retrieval calls and more candidates to merge — the added complexity and cost should earn its keep in your evaluation.
Interview Question
What is query transformation, and why might a raw user query retrieve poorly?
Query transformation rewrites or expands a user's query before retrieval, because raw queries — especially short, ambiguous, or conversational follow-ups — often don't retrieve well on their own. Common techniques include rewriting a follow-up question using conversation history, expanding a query with related terms, decomposing a complex question into sub-queries, and multi-query retrieval, which generates several phrasings of the same query and merges the results. It's usually implemented as a small, fast LLM call before the main retrieval step, and it's most valuable in multi-turn conversations where a question depends on context from earlier messages.
What an interviewer may ask next
- Why would "what about the second one?" retrieve poorly without query transformation?
- What is multi-query retrieval, and why might it improve recall?
- When would query transformation not be worth the added latency?
Explain It in 30 Seconds
Query transformation rewrites or expands a user's query before retrieval runs, because raw queries — especially short or conversational follow-ups — often don't retrieve well as-is. Common techniques include rewriting using conversation history, expanding with related terms, decomposing complex questions, and multi-query retrieval, which merges results from several phrasings. It's typically done with a quick LLM call before the main retrieval step, and matters most in multi-turn conversations.