Skip to primary content
Vector Database Deep Dive

Milvus: Distributed Vector Search at Billion Scale

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

Milvus is an open-source, distributed vector database created by Zilliz and governed under the LF AI and Data foundation. It stores high-dimensional embeddings and runs approximate nearest neighbor search across billions of vectors, with pluggable indexes including HNSW, IVF, DiskANN, and GPU graph indexes for scalar-filtered retrieval workloads.

LicenseApache 2.0
Index EngineKnowhere
Deployment ModesLite, Standalone, Distributed
GovernanceLF AI and Data
Problem & Purpose

What Milvus Solves in Production

Retrieval systems that start on a single-node vector store tend to break when embedding volume crosses into the hundreds of millions and query traffic grows concurrent. Recall degrades under naive sharding, index rebuilds block ingestion, and metadata filters push latency past service budgets. Milvus addresses this by separating compute from storage and letting query, data, and index roles scale independently on Kubernetes. Its pluggable index layer, Knowhere, lets teams pick the recall and cost tradeoff per collection rather than accepting one fixed algorithm. That flexibility is what makes billion-scale approximate nearest neighbor search operationally tractable.

Inside the Milvus Distributed Architecture

Anatomy Explainer

Core Component Component Parts:

1. Access Layer (Proxy) → View Definition
2. Coordinator Services → View Definition
3. Query Nodes → View Definition
4. Data and Index Nodes → View Definition
5. Storage Layer → View Definition
PART 1

Access Layer (Proxy)

Stateless entry point that validates requests, routes them, and merges results.

Technical Implementation:

The proxy handles client connections through the SDK or REST, performs load balancing across query nodes, and enforces consistency levels by managing timestamps before returning merged top-k results.

The five roles that separate compute, coordination, and storage.
Text alternative for screen readers & search engines
  • Part 1: Access Layer (Proxy) - Stateless entry point that validates requests, routes them, and merges results. [Tech: The proxy handles client connections through the SDK or REST, performs load balancing across query nodes, and enforces consistency levels by managing timestamps before returning merged top-k results.]
  • Part 2: Coordinator Services - Control plane that assigns tasks and maintains cluster topology. [Tech: Root, data, query, and index coordination responsibilities manage collection metadata, segment allocation, load balancing, and index build scheduling, with state persisted in etcd for failover recovery.]
  • Part 3: Query Nodes - Workers that load segments and execute vector plus scalar search. [Tech: Query nodes hold growing and sealed segments in memory, run ANN search through Knowhere, and evaluate boolean filter expressions during search rather than as a post-filter, scaling horizontally with traffic.]
  • Part 4: Data and Index Nodes - Workers that persist writes and build indexes asynchronously. [Tech: Data nodes consume the streaming log and flush sealed segments to object storage, while index nodes build IVF, HNSW, DiskANN, or GPU indexes on those segments without blocking ingestion or queries.]
  • Part 5: Storage Layer - External systems for metadata, streaming log, and persisted data. [Tech: etcd stores metadata and coordinates the cluster, a log broker such as Apache Pulsar or Kafka carries the write-ahead log, and object storage such as MinIO or S3 holds segment and index files.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Independent horizontal scaling: The compute and storage separation lets query, data, and index nodes scale on their own, so a single collection can span billions of vectors without a monolithic rebuild.
  • Pluggable index choice: Knowhere exposes IVF, HNSW, DiskANN, and GPU indexes per collection, letting teams tune the recall, memory, and cost tradeoff for each workload independently.
  • GPU acceleration: GPU_CAGRA and GPU_IVF_PQ built on NVIDIA RAFT deliver high queries-per-second and fast index builds for very large or high-throughput datasets.
  • Filtered vector search: Scalar boolean filters evaluated during search, not after, keep recall accurate even when metadata constraints are highly selective.
Specific Production Limits (Real Constraints)
  • Operational surface area: The distributed mode requires running etcd, a log broker, and object storage on Kubernetes, which is a heavier operational commitment than a single-binary store.
  • Memory-resident indexes: IVF and HNSW indexes are held in memory on query nodes, so large collections drive significant RAM cost unless DiskANN or quantization is used.
  • GPU cost and memory: GPU indexes need vectors resident in GPU memory and add hardware expense, so they only pay off above moderate throughput or dataset thresholds.
  • Eventual freshness by default: Lower consistency levels return recently written data with delay, so read-your-write scenarios require Strong or Session consistency and its added latency.
Production Implementation

How We Deploy Milvus in Production

Our team runs Milvus in distributed mode on Kubernetes with S3-compatible object storage and a dedicated log broker, sizing query nodes against measured concurrency rather than peak guesses. We pin the Milvus and pymilvus versions together, standardize on COSINE or inner-product metrics matched to the embedding model, and select HNSW or DiskANN per collection based on dataset size and latency budget. Consistency level, segment size, and index parameters are treated as tunable knobs validated against recall benchmarks before any promotion.

Milvus Retrieval Pipeline

Interactive Flow Diagram
Milvus Retrieval Pipeline From ingestion to filtered top-k retrieval in production. 1. Schema and Collection Define fields and metric 2. Ingest and Flush Stream writes to segments 3. Index Build Asynchronous indexing 4. Load and Search Filtered ANN query 5. Rerank and Serve Return to application
Stage 1: 1. Schema and Collection dim aligned to model output

Design the collection schema with primary key, vector dimension, and scalar metadata fields, choosing the metric type to match the embedding model.

From ingestion to filtered top-k retrieval in production.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Schema and Collection Design the collection schema with primary key, vector dimension, and scalar metadata fields, choosing the metric type to match the embedding model. dim aligned to model output
2 2. Ingest and Flush Insert embeddings through the SDK; data nodes consume the log and flush sealed segments to object storage for durability. batched upserts
3 3. Index Build Index nodes build HNSW, DiskANN, or GPU indexes on sealed segments without blocking ongoing ingestion or queries. M and efConstruction tuned
4 4. Load and Search Query nodes load segments into memory and run vector search with boolean scalar filters applied during retrieval. ef tuned to latency target
5 5. Rerank and Serve The proxy merges top-k results across shards under the chosen consistency level, feeding the RAG or search application. consistency per query
Production Configuration (Version Pinned):
# pymilvus==2.4.9, Milvus 2.4.x
from pymilvus import MilvusClient, DataType

client = MilvusClient(uri="http://milvus:19530", token="root:Milvus")

schema = client.create_schema(auto_id=False, enable_dynamic_field=True)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_field("doc_id", DataType.VARCHAR, max_length=128)

index_params = client.prepare_index_params()
index_params.add_index(
  field_name="embedding",
  index_type="HNSW",
  metric_type="COSINE",
  params={"M": 16, "efConstruction": 200},
)

client.create_collection(
  collection_name="clinical_chunks",
  schema=schema,
  index_params=index_params,
  consistency_level="Bounded",
)

res = client.search(
  collection_name="clinical_chunks",
  data=[query_vector],
  limit=10,
  search_params={"metric_type": "COSINE", "params": {"ef": 64}},
  filter='doc_id like "guideline_%"',
  output_fields=["doc_id"],
)
Delivering Commercial Impact

Services Engineered with Milvus

We build and operate Milvus-backed retrieval as part of these engagements.

Alternatives Evaluation

Milvus vs Alternative Vector Engines

How Milvus compares to Qdrant and Pinecone across the dimensions that shape a production choice.

Vector Database Comparison

Benchmark Matrix
Evaluation Metric Milvus Qdrant Pinecone
Billion-scale distributed search
Compute and storage separated, scales to billions Winner
Scales well, distributed sharding maturing
Managed scale, opaque internals
GPU indexing support
CAGRA and GPU IVF via NVIDIA RAFT Winner
CPU-focused, no native GPU index
Not exposed to users
Operational simplicity
Heavier stack: etcd, broker, object store
Single binary, lean deployment
Fully managed, no ops Winner
Metadata filtering ergonomics
Boolean expressions during search
Rich payload filters, strong ergonomics Winner
Metadata filters, less expressive
Illustrative relative suitability, defensible not benchmarked.
Text alternative for screen readers & search engines
  • Billion-scale distributed search: Milvus: Compute and storage separated, scales to billions vs Qdrant: Scales well, distributed sharding maturing vs Pinecone: Managed scale, opaque internals (Winning option: Milvus).
  • GPU indexing support: Milvus: CAGRA and GPU IVF via NVIDIA RAFT vs Qdrant: CPU-focused, no native GPU index vs Pinecone: Not exposed to users (Winning option: Milvus).
  • Operational simplicity: Milvus: Heavier stack: etcd, broker, object store vs Qdrant: Single binary, lean deployment vs Pinecone: Fully managed, no ops (Winning option: Pinecone).
  • Metadata filtering ergonomics: Milvus: Boolean expressions during search vs Qdrant: Rich payload filters, strong ergonomics vs Pinecone: Metadata filters, less expressive (Winning option: Qdrant).
Production Proof

Milvus in a Reference Architecture

Clinical RAG Retrieval

We used Milvus as the retrieval backbone for a clinical RAG system, indexing guideline and record embeddings with scalar filters that scope results to the right document type and patient context. The separation of index roles let ingestion continue while indexes rebuilt, and tuned HNSW parameters held query latency within the interactive budget clinicians needed.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is Milvus used for?↓

Milvus is a purpose-built vector database for storing and searching high-dimensional embeddings at scale. It powers retrieval-augmented generation, semantic search, recommendation, and image similarity systems. Applications query it for the nearest vectors to a given embedding, optionally filtered by scalar metadata fields.

Is Milvus open source and free?↓

Yes. Milvus is licensed under Apache 2.0 and is free to self-host with no usage restrictions. Zilliz, the company behind it, also offers Zilliz Cloud as a managed commercial service. The open-source core and the managed offering share the same engine and query semantics.

What indexes does Milvus support?↓

Milvus supports IVF_FLAT, IVF_PQ, IVF_SQ8, HNSW, SCANN, and DiskANN on CPU. It also provides GPU indexes including GPU_IVF_FLAT, GPU_IVF_PQ, GPU_BRUTE_FORCE, and GPU_CAGRA built on NVIDIA RAFT. Sparse and binary vector indexes are available for keyword and hash-based retrieval.

How does Milvus scale to billions of vectors?↓

The distributed deployment separates compute from storage and splits responsibilities across query, data, and index nodes coordinated by a set of coordinator services. Data is sharded into segments held in object storage, so query and index nodes scale horizontally and independently. This lets a single collection span billions of vectors across many pods.

What is the difference between Milvus Standalone and Distributed?↓

Standalone runs the full engine plus etcd, a message broker, and MinIO inside a single deployment, which suits development and moderate workloads. Distributed runs the same components as separately scalable services on Kubernetes for high throughput and large datasets. Milvus Lite is a lightweight embedded option for prototyping in Python.

Does Milvus support metadata filtering during search?↓

Yes. Milvus supports scalar filtering with a boolean expression language applied alongside vector search, so you can constrain results by numeric ranges, string matches, and JSON fields. Filtering is evaluated during the search rather than as a post-processing step. This keeps recall accurate when filters are selective.

What consistency levels does Milvus offer?↓

Milvus exposes four tunable consistency levels: Strong, Bounded, Session, and Eventually. Strong guarantees the freshest data at higher latency, while Bounded and Eventually trade freshness for speed. Session gives read-your-own-writes consistency within a client. You choose per collection or per query based on the freshness the workload needs.

What does Milvus use for storage and coordination?↓

Milvus depends on three external systems. etcd stores metadata and coordinates the cluster, a log broker such as Apache Pulsar or Kafka handles the streaming write-ahead log, and object storage such as MinIO or S3 holds persisted segments and index files. This separation lets the stateless compute nodes restart without data loss.

When should GPU indexing be used in Milvus?↓

GPU indexes such as GPU_CAGRA and GPU_IVF_PQ help when query throughput is very high or index build times on large datasets become a bottleneck. They deliver strong queries-per-second at scale but consume GPU memory and add operational cost. For moderate traffic, HNSW or DiskANN on CPU is often more economical.

Milvus vs Pinecone: which should I choose?↓

Milvus is open source and self-hosted, giving full control over indexes, hardware, and GPU acceleration, which suits teams with strong infrastructure ownership. Pinecone is fully managed and removes cluster operations at the cost of vendor lock-in and less index-level control. The decision usually turns on whether you want to run the platform yourself.