Dev.to AI 🤖 Ai 👁 0 📖 5 min read

The finale of the Vector Database Showroom series.

The finale of the Vector Database Showroom series. Six parts ago, this showroom made you a promise: we won't tell you which database to take — we'll give you a method to figure out what YOU should take. Today the showro

The finale of the Vector Database Showroom series.

Six parts ago, this showroom made you a promise: we won't tell you which database to take — we'll give you a method to figure out what YOU should take. Today the showroom closes: one table, one tree, one checklist.

The roster, for the record: 🚙 pgvector, 🛴 Chroma, 🏎️ Qdrant, 🚐 Weaviate, 🚕 Pinecone, 🚄 Milvus — the station wagon, the scooter, the hatchback, the crossover, the taxi, and the train.

The cost-of-ownership table

Vehicle License / self-host Sweet spot Signature feature Main pain Ops burden Managed by
🚙 pgvector PostgreSQL license, self-host anywhere ~1M comfortable, ~10M tuned ACID + joins — vectors live with your data Vertical ceiling; RAM for the HNSW graph Your usual Postgres routine RDS, Aurora, Cloud SQL, AlloyDB, Azure, Supabase, Neon
🛴 Chroma Apache 2.0, embedded or server Prototypes; single-digit millions First search in 90 seconds, embeddings included Single node by design; prototype-grade durability None — until it's suddenly too many Chroma Cloud
🏎️ Qdrant Apache 2.0, single binary → cluster Millions → hundreds of millions (quantized) Filterable HNSW — filters live inside the index Manual snapshots; a tuning culture Light-to-medium DevOps Qdrant Cloud
🚐 Weaviate BSD-3-Clause, containers Multi-tenant SaaS; scale via distributed mode Factory vectorization + built-in BM25F hybrid + native tenancy A service bay for all those options; the v3/v4 client split Medium DevOps Weaviate Cloud
🚕 Pinecone Proprietary — cloud only Prototype → billions, if the wallet agrees True zero ops + serverless A closed box, and the meter (read units) None — the dealer's problem Pinecone only
🚄 Milvus Apache 2.0 — Lite / Standalone / Distributed Hundreds of millions → billions Capacity scales with object storage; GPU_CAGRA; DiskANN The rails: etcd, Pulsar/Kafka, MinIO/S3 Real ops muscle Zilliz Cloud

Two synthesis notes the table doesn't fit:

  • Hybrid search out of the box: yes — Qdrant (dense + sparse, RRF/DBSF fusion), Weaviate (BM25F built in), Milvus (built-in BM25). No — pgvector, Chroma. Pinecone: you bring the sparse vectors yourself (the Qdrant model, not the built-in-BM25 model of Weaviate/Milvus) — check the docs for the current state before betting on it.
  • Multi-tenancy, natively: Weaviate (tenant = shard), Pinecone (namespaces), Milvus (partition keys / separate databases). pgvector covers it with RLS; Qdrant expects payload filters; Chroma assembles nothing for you.

The recurring math of the series: 1M × 1536-dim × 4 bytes ≈ 6 GB of raw vectors — and an HNSW graph stores a copy of every vector, so roughly double it before picking instance sizes.

Honorable mentions, three lines: Elasticsearch/OpenSearch — if it's already your stack and the vector requirements are modest. Redis — when you need sub-millisecond and already run it. FAISS — an engine, not a car: blazingly fast, but no garage, no service, no keys — you build the car around it.

flowchart TD
    START["You need vector search"] --> Q1{"Prototyping, learning,\nor a personal tool?"}
    Q1 -- "Yes" --> CHROMA["🛴 Chroma"]
    Q1 -- "No" --> Q2{"Data already in Postgres\nand under ~10M vectors?"}
    Q2 -- "Yes" --> PGVECTOR["🚙 pgvector"]
    Q2 -- "No" --> Q3{"Anyone on the team\nwants to own infrastructure?"}
    Q3 -- "Yes — wants to own infra" --> Q4
    Q3 -- "Must self-host (compliance)" --> Q4
    Q4{"Hundreds of millions\nto billions of vectors?"}
    Q4 -- "Yes" --> MILVUS["🚄 Milvus — or rent Zilliz Cloud"]
    Q4 -- "No" --> Q5{"Many tenants, want\nvectorization + hybrid\nfrom the factory?"}
    Q5 -- "Yes" --> WEAVIATE["🚐 Weaviate"]
    Q5 -- "No" --> QDRANT["🏎️ Qdrant"]

Two honest notes on the tree:

  • No Postgres yet? The wagon is still an option — managed Postgres is one click away (Part 1). The garage is not a prerequisite.
  • Data can't leave your perimeter AND nobody wants ops? That's a hiring problem, not a database problem. No car in this showroom fixes it — and pretending otherwise is how two-week adventures begin.

The one-day test drive

Pick 2–3 finalists from the tree — not all six. Then:

Before you start

  • One embedding model for all candidates — the honest-comparison rule since Part 3
  • Real data: 10–50k real documents and 50–100 real queries from your product. Lorem ipsum measures lorem ipsum
  • Deploy the gauge you'll actually ride: if production is a cluster, benchmark the cluster shape — the Milvus lesson

(2–3 hours)

  • Create the schema and payload/metadata indexes before loading — the Qdrant and Milvus lessons
  • Load the data. Then wait: settled segments, built indexes (benchmark traps #1 and #2). A benchmark on fresh data measures brute force, not your index

(2 hours) — the scorecard

Recall@10 vs exact ground truth. Brute-force the same vectors yourself — at 50k rows it takes seconds:

import numpy as np

# vectors: (N, 1536) of your uploaded embeddings, query: (1536,)
# ground truth must use the SAME metric you query with —
# for cosine, normalize first
scores = vectors @ query
truth_indices = set(np.argsort(scores)[-10:][::-1])   # exact top-10 INDICES into `vectors`

# Option 1: if your DB returns positions in the same order you uploaded,
# compare indices directly
recall = len(truth_indices & set(returned_indices)) / 10

# Option 2: if your DB returns IDs, map indices to IDs first —
# the upload order defines the mapping, keep it consistent
# ids = ["doc_0", "doc_1", ...]          # index i -> id, fixed at upload time
# truth_ids = {ids[i] for i in truth_indices}
# recall = len(truth_ids & set(returned_ids)) / 10
  • p95 latency at your realistic QPS, not the demo QPS
  • RAM after indexing — expect roughly 2× the raw vectors (the series rule)
  • Filtered search: does "similar AND year=2024" return the full LIMIT? (the Part 1 trap)
  • Insert → immediately search: flaky? Now you know before production does (the taxi and the train)
  • Kill the process mid-write (self-hosted) or test retry after errors (managed) — then check nothing was lost
  • Export everything back out — measure the exit in hours, not promises

(1 hour) — the bill

  • Managed (Pinecone): your real QPS × top_k → read units → dollars. Managed (Qdrant / Weaviate / Zilliz Cloud): RAM + CPU → monthly subscription. Either way, the meter never sleeps
  • Self-hosted: the RAM you measured → the smallest instance that fits → × 12 months
  • Fill the scorecard, then sleep on it:
Candidate recall@10 p95, ms RAM after index, GB $/month at your QPS Exit, hours Gut feeling
...

Yes, "gut feeling" is a column. DX is a cost of ownership too — you'll pay it every working day.

Next morning: re-run the same queries cold. Numbers that only look good warm are a finding, not a footnote.

The method doesn't expire

Everything in this series was true at the time of writing; databases evolve — verify versions against the docs. The method doesn't expire: judge by cost of ownership, verify on your data, buy for the scale you have, re-check at the scale you reach.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.