Vector Search
Vector search finds the most similar items to a query by comparing embeddings instead of matching exact keywords.
Prerequisites
What Is Vector Search?
Vector search — often called semantic search — finds the items most relevant to a query by comparing embeddings, not by matching literal words. The query is embedded the same way the stored content was, and the system finds the stored vectors closest to the query vector.
Query
Starting PointWhat the user is trying to find.
Embed Query
Same Vector SpaceEmbedded the same way the stored content was.
Compare Against Stored Vectors
Closeness CheckFinds which stored vectors sit nearest.
Nearest Neighbors
Approximate MatchIndexed for speed, trading a little accuracy.
Results
Top-K CandidatesSeveral close matches, not just one.
Key Idea
Vector search answers "what's closest in meaning" instead of "what contains these exact words."
How It Works
Comparing a query against every stored vector one by one — an exact nearest-neighbor search — becomes too slow as a collection grows into the millions. In practice, vector databases use approximate nearest neighbor (ANN) algorithms that index vectors for fast lookup, trading a small amount of accuracy for a large gain in speed.
- Exact search — compares the query against every stored vector; accurate but doesn't scale to large collections.
- Approximate nearest neighbor (ANN) — indexes vectors so most queries can skip comparing against everything, dramatically faster at large scale with a small accuracy tradeoff.
- Top-k retrieval — vector search typically returns the k closest items, not a single best match, since RAG and search applications usually want several candidates to work with.
Vector Search vs. Keyword Search
Matches exact terms
Strong for IDs, codes, rare terms
Misses paraphrases
Fast and simple
Matches meaning
Strong for paraphrases and concepts
Can miss exact rare terms
Requires embeddings and an index
Neither approach strictly dominates the other, which is exactly why hybrid search — combining both — often outperforms either alone.
Common Mistakes
Assuming vector search always outperforms keyword search
Exact identifiers, product codes, or rare technical terms are often matched better by keyword search.
Ignoring the accuracy tradeoff of approximate nearest neighbor search
ANN indexes trade a small amount of accuracy for large speed gains — this is usually the right tradeoff at scale, but it is a real tradeoff.
Retrieving only the single top result
A single closest match can be wrong or incomplete — retrieving several top-k candidates and letting a later step (like reranking) narrow them down is usually more robust.
Forgetting that the query and stored content must use the same embedding model
Embeddings from different models generally aren't comparable to each other.
Interview Question
What is vector search, and how is it different from keyword search?
Vector search finds the most relevant items to a query by comparing embeddings instead of matching exact words — the query gets embedded the same way as the stored content, and the system returns the closest vectors. At scale it uses approximate nearest neighbor algorithms rather than comparing against every stored vector, trading a small amount of accuracy for a large speed gain. It's different from keyword search in that it matches meaning rather than exact terms, which makes it strong at handling paraphrases but sometimes weaker on exact identifiers or rare terms — which is exactly why hybrid search, combining both, often works better than either alone.
What an interviewer may ask next
- Why do vector databases use approximate nearest neighbor search instead of exact search at scale?
- When would keyword search actually outperform vector search?
- Why does vector search typically return several top-k results instead of just one?
Explain It in 30 Seconds
Vector search finds the items most relevant to a query by comparing embeddings rather than matching exact words — it embeds the query and finds the closest stored vectors. At scale, it uses approximate nearest neighbor algorithms to stay fast, trading a small amount of accuracy for speed. It excels at matching meaning and paraphrases, but can miss exact rare terms that keyword search handles well — which is why the two are often combined in hybrid search.