Vector Databases in Production
Choosing a vector database in practice depends on scale, existing infrastructure, and whether filtering or hybrid search is required.
Prerequisites
Overview
The theory of a vector database — store embeddings, search by similarity — is the same everywhere. Choosing one in practice comes down to a handful of practical questions: managed or self-hosted, how much metadata filtering is needed, and whether it needs to plug into infrastructure a team already runs.
Where It Fits
Scale + Filtering Needs
Managed (Pinecone, Atlas)
Self-Hosted (Qdrant, Milvus)
Bolt-On (pgvector)
Key Points
- Managed vs. self-hosted
- A managed service (Pinecone) trades operational effort for less infrastructure control than self-hosting (Qdrant, Milvus, Weaviate).
- Filtering support
- Applications that need to filter by tenant, permission, or metadata before ranking need strong native filtering, not just raw similarity search.
- Bolt-on option
- pgvector avoids a second system entirely when PostgreSQL is already the primary database and scale is moderate.
Interview Question
What factors would actually drive your choice of vector database for a new RAG project?
Expected scale and query volume, how much metadata filtering is needed before ranking, whether the team wants to operate infrastructure or use a managed service, and whether an existing database (like PostgreSQL) can reasonably absorb the vector workload via something like pgvector rather than adding a new system.
Explain It in 30 Seconds
Choosing a vector database in practice depends less on which one is theoretically fastest and more on scale, filtering requirements, and whether a team wants a managed service, self-hosted infrastructure, or a bolt-on extension to a database they already run.
Real-World Stack
Technologies commonly used to implement this in production.