The Vector Database Landscape
Vector databases differ mainly along managed-vs-self-hosted and standalone-vs-bolt-on-to-an-existing-database lines.
Prerequisites
Overview
The underlying idea — storing embeddings and retrieving nearest neighbors — is shared across every vector database. The real-world choice between them mostly comes down to two axes: fully-managed versus self-hosted, and a purpose-built vector database versus a vector-search extension bolted onto a database you already run.
Where It Fits
Purpose-Built vs. Bolt-On
Pinecone/Qdrant vs. pgvector/OpenSearch.
Managed vs. Self-Hosted
Retrieval Backend Choice
Key Points
- Purpose-built options
- Pinecone (managed-only), Qdrant, Weaviate, and Milvus (open-source, self-hostable or managed) are built specifically for vector similarity search.
- Bolt-on options
- pgvector adds vector search to an existing PostgreSQL database; OpenSearch and Elasticsearch add it to an existing search cluster — useful when you’d rather not run a separate system.
- The real tradeoff
- A bolt-on keeps data co-located with what you already operate; a purpose-built database is optimized specifically for vector workloads at larger scale.
Interview Question
When would you choose pgvector over a purpose-built vector database like Pinecone?
When the application data and the embeddings naturally belong together and the scale doesn’t demand a specialized engine — pgvector avoids running and syncing a separate system. A purpose-built database like Pinecone makes more sense once vector search volume or scale outgrows what fits comfortably inside the existing operational database.
Explain It in 30 Seconds
Vector databases split along two axes — managed versus self-hosted, and purpose-built versus a vector-search extension bolted onto a database you already run — and the right choice depends on scale and how tightly the embeddings need to live alongside existing application data.
Real-World Stack
Technologies commonly used to implement this in production.