The finale of the Vector Database Showroom series.
The finale of the Vector Database Showroom series. Six parts ago, this showroom made you a promise: we won't tell you which database to take — we'll give you a method to figure out what YOU should take. Today the showro
The finale of the Vector Database Showroom series.
Six parts ago, this showroom made you a promise: we won't tell you which database to take — we'll give you a method to figure out what YOU should take. Today the showroom closes: one table, one tree, one checklist.
The roster, for the record: 🚙 pgvector, 🛴 Chroma, 🏎️ Qdrant, 🚐 Weaviate, 🚕 Pinecone, 🚄 Milvus — the station wagon, the scooter, the hatchback, the crossover, the taxi, and the train.
The cost-of-ownership table
| Vehicle | License / self-host | Sweet spot | Signature feature | Main pain | Ops burden | Managed by |
|---|---|---|---|---|---|---|
| 🚙 pgvector | PostgreSQL license, self-host anywhere | ~1M comfortable, ~10M tuned | ACID + joins — vectors live with your data | Vertical ceiling; RAM for the HNSW graph | Your usual Postgres routine | RDS, Aurora, Cloud SQL, AlloyDB, Azure, Supabase, Neon |
| 🛴 Chroma | Apache 2.0, embedded or server | Prototypes; single-digit millions | First search in 90 seconds, embeddings included | Single node by design; prototype-grade durability | None — until it's suddenly too many | Chroma Cloud |
| 🏎️ Qdrant | Apache 2.0, single binary → cluster | Millions → hundreds of millions (quantized) | Filterable HNSW — filters live inside the index | Manual snapshots; a tuning culture | Light-to-medium DevOps | Qdrant Cloud |
| 🚐 Weaviate | BSD-3-Clause, containers | Multi-tenant SaaS; scale via distributed mode | Factory vectorization + built-in BM25F hybrid + native tenancy | A service bay for all those options; the v3/v4 client split | Medium DevOps | Weaviate Cloud |
| 🚕 Pinecone | Proprietary — cloud only | Prototype → billions, if the wallet agrees | True zero ops + serverless | A closed box, and the meter (read units) | None — the dealer's problem | Pinecone only |
| 🚄 Milvus | Apache 2.0 — Lite / Standalone / Distributed | Hundreds of millions → billions | Capacity scales with object storage; GPU_CAGRA; DiskANN | The rails: etcd, Pulsar/Kafka, MinIO/S3 | Real ops muscle | Zilliz Cloud |
Two synthesis notes the table doesn't fit:
- Hybrid search out of the box: yes — Qdrant (dense + sparse, RRF/DBSF fusion), Weaviate (BM25F built in), Milvus (built-in BM25). No — pgvector, Chroma. Pinecone: you bring the sparse vectors yourself (the Qdrant model, not the built-in-BM25 model of Weaviate/Milvus) — check the docs for the current state before betting on it.
- Multi-tenancy, natively: Weaviate (tenant = shard), Pinecone (namespaces), Milvus (partition keys / separate databases). pgvector covers it with RLS; Qdrant expects payload filters; Chroma assembles nothing for you.
The recurring math of the series: 1M × 1536-dim × 4 bytes ≈ 6 GB of raw vectors — and an HNSW graph stores a copy of every vector, so roughly double it before picking instance sizes.
Honorable mentions, three lines: Elasticsearch/OpenSearch — if it's already your stack and the vector requirements are modest. Redis — when you need sub-millisecond and already run it. FAISS — an engine, not a car: blazingly fast, but no garage, no service, no keys — you build the car around it.
flowchart TD
START["You need vector search"] --> Q1{"Prototyping, learning,\nor a personal tool?"}
Q1 -- "Yes" --> CHROMA["🛴 Chroma"]
Q1 -- "No" --> Q2{"Data already in Postgres\nand under ~10M vectors?"}
Q2 -- "Yes" --> PGVECTOR["🚙 pgvector"]
Q2 -- "No" --> Q3{"Anyone on the team\nwants to own infrastructure?"}
Q3 -- "Yes — wants to own infra" --> Q4
Q3 -- "Must self-host (compliance)" --> Q4
Q4{"Hundreds of millions\nto billions of vectors?"}
Q4 -- "Yes" --> MILVUS["🚄 Milvus — or rent Zilliz Cloud"]
Q4 -- "No" --> Q5{"Many tenants, want\nvectorization + hybrid\nfrom the factory?"}
Q5 -- "Yes" --> WEAVIATE["🚐 Weaviate"]
Q5 -- "No" --> QDRANT["🏎️ Qdrant"]
Two honest notes on the tree:
- No Postgres yet? The wagon is still an option — managed Postgres is one click away (Part 1). The garage is not a prerequisite.
- Data can't leave your perimeter AND nobody wants ops? That's a hiring problem, not a database problem. No car in this showroom fixes it — and pretending otherwise is how two-week adventures begin.
The one-day test drive
Pick 2–3 finalists from the tree — not all six. Then:
Before you start
- One embedding model for all candidates — the honest-comparison rule since Part 3
- Real data: 10–50k real documents and 50–100 real queries from your product. Lorem ipsum measures lorem ipsum
- Deploy the gauge you'll actually ride: if production is a cluster, benchmark the cluster shape — the Milvus lesson
(2–3 hours)
- Create the schema and payload/metadata indexes before loading — the Qdrant and Milvus lessons
- Load the data. Then wait: settled segments, built indexes (benchmark traps #1 and #2). A benchmark on fresh data measures brute force, not your index
(2 hours) — the scorecard
Recall@10 vs exact ground truth. Brute-force the same vectors yourself — at 50k rows it takes seconds:
import numpy as np
# vectors: (N, 1536) of your uploaded embeddings, query: (1536,)
# ground truth must use the SAME metric you query with —
# for cosine, normalize first
scores = vectors @ query
truth_indices = set(np.argsort(scores)[-10:][::-1]) # exact top-10 INDICES into `vectors`
# Option 1: if your DB returns positions in the same order you uploaded,
# compare indices directly
recall = len(truth_indices & set(returned_indices)) / 10
# Option 2: if your DB returns IDs, map indices to IDs first —
# the upload order defines the mapping, keep it consistent
# ids = ["doc_0", "doc_1", ...] # index i -> id, fixed at upload time
# truth_ids = {ids[i] for i in truth_indices}
# recall = len(truth_ids & set(returned_ids)) / 10
- p95 latency at your realistic QPS, not the demo QPS
- RAM after indexing — expect roughly 2× the raw vectors (the series rule)
- Filtered search: does "similar AND year=2024" return the full LIMIT? (the Part 1 trap)
- Insert → immediately search: flaky? Now you know before production does (the taxi and the train)
- Kill the process mid-write (self-hosted) or test retry after errors (managed) — then check nothing was lost
- Export everything back out — measure the exit in hours, not promises
(1 hour) — the bill
- Managed (Pinecone): your real QPS × top_k → read units → dollars. Managed (Qdrant / Weaviate / Zilliz Cloud): RAM + CPU → monthly subscription. Either way, the meter never sleeps
- Self-hosted: the RAM you measured → the smallest instance that fits → × 12 months
- Fill the scorecard, then sleep on it:
| Candidate | recall@10 | p95, ms | RAM after index, GB | $/month at your QPS | Exit, hours | Gut feeling |
|---|---|---|---|---|---|---|
| ... |
Yes, "gut feeling" is a column. DX is a cost of ownership too — you'll pay it every working day.
Next morning: re-run the same queries cold. Numbers that only look good warm are a finding, not a footnote.
The method doesn't expire
Everything in this series was true at the time of writing; databases evolve — verify versions against the docs. The method doesn't expire: judge by cost of ownership, verify on your data, buy for the scale you have, re-check at the scale you reach.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.