Milvus: Distributed Vector Search at Billion Scale
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Milvus is an open-source, distributed vector database created by Zilliz and governed under the LF AI and Data foundation. It stores high-dimensional embeddings and runs approximate nearest neighbor search across billions of vectors, with pluggable indexes including HNSW, IVF, DiskANN, and GPU graph indexes for scalar-filtered retrieval workloads.
What Milvus Solves in Production
Retrieval systems that start on a single-node vector store tend to break when embedding volume crosses into the hundreds of millions and query traffic grows concurrent. Recall degrades under naive sharding, index rebuilds block ingestion, and metadata filters push latency past service budgets. Milvus addresses this by separating compute from storage and letting query, data, and index roles scale independently on Kubernetes. Its pluggable index layer, Knowhere, lets teams pick the recall and cost tradeoff per collection rather than accepting one fixed algorithm. That flexibility is what makes billion-scale approximate nearest neighbor search operationally tractable.
Inside the Milvus Distributed Architecture
Anatomy ExplainerCore Component Component Parts:
Access Layer (Proxy)
Stateless entry point that validates requests, routes them, and merges results.
The proxy handles client connections through the SDK or REST, performs load balancing across query nodes, and enforces consistency levels by managing timestamps before returning merged top-k results.
Text alternative for screen readers & search engines
- Part 1: Access Layer (Proxy) - Stateless entry point that validates requests, routes them, and merges results. [Tech: The proxy handles client connections through the SDK or REST, performs load balancing across query nodes, and enforces consistency levels by managing timestamps before returning merged top-k results.]
- Part 2: Coordinator Services - Control plane that assigns tasks and maintains cluster topology. [Tech: Root, data, query, and index coordination responsibilities manage collection metadata, segment allocation, load balancing, and index build scheduling, with state persisted in etcd for failover recovery.]
- Part 3: Query Nodes - Workers that load segments and execute vector plus scalar search. [Tech: Query nodes hold growing and sealed segments in memory, run ANN search through Knowhere, and evaluate boolean filter expressions during search rather than as a post-filter, scaling horizontally with traffic.]
- Part 4: Data and Index Nodes - Workers that persist writes and build indexes asynchronously. [Tech: Data nodes consume the streaming log and flush sealed segments to object storage, while index nodes build IVF, HNSW, DiskANN, or GPU indexes on those segments without blocking ingestion or queries.]
- Part 5: Storage Layer - External systems for metadata, streaming log, and persisted data. [Tech: etcd stores metadata and coordinates the cluster, a log broker such as Apache Pulsar or Kafka carries the write-ahead log, and object storage such as MinIO or S3 holds segment and index files.]
Architectural Strengths & Specific Production Limits
- Independent horizontal scaling: The compute and storage separation lets query, data, and index nodes scale on their own, so a single collection can span billions of vectors without a monolithic rebuild.
- Pluggable index choice: Knowhere exposes IVF, HNSW, DiskANN, and GPU indexes per collection, letting teams tune the recall, memory, and cost tradeoff for each workload independently.
- GPU acceleration: GPU_CAGRA and GPU_IVF_PQ built on NVIDIA RAFT deliver high queries-per-second and fast index builds for very large or high-throughput datasets.
- Filtered vector search: Scalar boolean filters evaluated during search, not after, keep recall accurate even when metadata constraints are highly selective.
- Operational surface area: The distributed mode requires running etcd, a log broker, and object storage on Kubernetes, which is a heavier operational commitment than a single-binary store.
- Memory-resident indexes: IVF and HNSW indexes are held in memory on query nodes, so large collections drive significant RAM cost unless DiskANN or quantization is used.
- GPU cost and memory: GPU indexes need vectors resident in GPU memory and add hardware expense, so they only pay off above moderate throughput or dataset thresholds.
- Eventual freshness by default: Lower consistency levels return recently written data with delay, so read-your-write scenarios require Strong or Session consistency and its added latency.
How We Deploy Milvus in Production
Our team runs Milvus in distributed mode on Kubernetes with S3-compatible object storage and a dedicated log broker, sizing query nodes against measured concurrency rather than peak guesses. We pin the Milvus and pymilvus versions together, standardize on COSINE or inner-product metrics matched to the embedding model, and select HNSW or DiskANN per collection based on dataset size and latency budget. Consistency level, segment size, and index parameters are treated as tunable knobs validated against recall benchmarks before any promotion.
Milvus Retrieval Pipeline
Interactive Flow DiagramDesign the collection schema with primary key, vector dimension, and scalar metadata fields, choosing the metric type to match the embedding model.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Schema and Collection | Design the collection schema with primary key, vector dimension, and scalar metadata fields, choosing the metric type to match the embedding model. | dim aligned to model output |
| 2 | 2. Ingest and Flush | Insert embeddings through the SDK; data nodes consume the log and flush sealed segments to object storage for durability. | batched upserts |
| 3 | 3. Index Build | Index nodes build HNSW, DiskANN, or GPU indexes on sealed segments without blocking ongoing ingestion or queries. | M and efConstruction tuned |
| 4 | 4. Load and Search | Query nodes load segments into memory and run vector search with boolean scalar filters applied during retrieval. | ef tuned to latency target |
| 5 | 5. Rerank and Serve | The proxy merges top-k results across shards under the chosen consistency level, feeding the RAG or search application. | consistency per query |
# pymilvus==2.4.9, Milvus 2.4.x
from pymilvus import MilvusClient, DataType
client = MilvusClient(uri="http://milvus:19530", token="root:Milvus")
schema = client.create_schema(auto_id=False, enable_dynamic_field=True)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("embedding", DataType.FLOAT_VECTOR, dim=1024)
schema.add_field("doc_id", DataType.VARCHAR, max_length=128)
index_params = client.prepare_index_params()
index_params.add_index(
field_name="embedding",
index_type="HNSW",
metric_type="COSINE",
params={"M": 16, "efConstruction": 200},
)
client.create_collection(
collection_name="clinical_chunks",
schema=schema,
index_params=index_params,
consistency_level="Bounded",
)
res = client.search(
collection_name="clinical_chunks",
data=[query_vector],
limit=10,
search_params={"metric_type": "COSINE", "params": {"ef": 64}},
filter='doc_id like "guideline_%"',
output_fields=["doc_id"],
)Services Engineered with Milvus
We build and operate Milvus-backed retrieval as part of these engagements.
Milvus vs Alternative Vector Engines
How Milvus compares to Qdrant and Pinecone across the dimensions that shape a production choice.
Vector Database Comparison
Benchmark Matrix| Evaluation Metric | Milvus | Qdrant | Pinecone |
|---|---|---|---|
| Billion-scale distributed search | Compute and storage separated, scales to billions Winner | Scales well, distributed sharding maturing | Managed scale, opaque internals |
| GPU indexing support | CAGRA and GPU IVF via NVIDIA RAFT Winner | CPU-focused, no native GPU index | Not exposed to users |
| Operational simplicity | Heavier stack: etcd, broker, object store | Single binary, lean deployment | Fully managed, no ops Winner |
| Metadata filtering ergonomics | Boolean expressions during search | Rich payload filters, strong ergonomics Winner | Metadata filters, less expressive |
Text alternative for screen readers & search engines
- Billion-scale distributed search: Milvus: Compute and storage separated, scales to billions vs Qdrant: Scales well, distributed sharding maturing vs Pinecone: Managed scale, opaque internals (Winning option: Milvus).
- GPU indexing support: Milvus: CAGRA and GPU IVF via NVIDIA RAFT vs Qdrant: CPU-focused, no native GPU index vs Pinecone: Not exposed to users (Winning option: Milvus).
- Operational simplicity: Milvus: Heavier stack: etcd, broker, object store vs Qdrant: Single binary, lean deployment vs Pinecone: Fully managed, no ops (Winning option: Pinecone).
- Metadata filtering ergonomics: Milvus: Boolean expressions during search vs Qdrant: Rich payload filters, strong ergonomics vs Pinecone: Metadata filters, less expressive (Winning option: Qdrant).
Milvus in a Reference Architecture
We used Milvus as the retrieval backbone for a clinical RAG system, indexing guideline and record embeddings with scalar filters that scope results to the right document type and patient context. The separation of index roles let ingestion continue while indexes rebuilt, and tuned HNSW parameters held query latency within the interactive budget clinicians needed.
Read Reference Architecture →Frequently Asked Questions
What is Milvus used for?↓
Milvus is a purpose-built vector database for storing and searching high-dimensional embeddings at scale. It powers retrieval-augmented generation, semantic search, recommendation, and image similarity systems. Applications query it for the nearest vectors to a given embedding, optionally filtered by scalar metadata fields.
Is Milvus open source and free?↓
Yes. Milvus is licensed under Apache 2.0 and is free to self-host with no usage restrictions. Zilliz, the company behind it, also offers Zilliz Cloud as a managed commercial service. The open-source core and the managed offering share the same engine and query semantics.
What indexes does Milvus support?↓
Milvus supports IVF_FLAT, IVF_PQ, IVF_SQ8, HNSW, SCANN, and DiskANN on CPU. It also provides GPU indexes including GPU_IVF_FLAT, GPU_IVF_PQ, GPU_BRUTE_FORCE, and GPU_CAGRA built on NVIDIA RAFT. Sparse and binary vector indexes are available for keyword and hash-based retrieval.
How does Milvus scale to billions of vectors?↓
The distributed deployment separates compute from storage and splits responsibilities across query, data, and index nodes coordinated by a set of coordinator services. Data is sharded into segments held in object storage, so query and index nodes scale horizontally and independently. This lets a single collection span billions of vectors across many pods.
What is the difference between Milvus Standalone and Distributed?↓
Standalone runs the full engine plus etcd, a message broker, and MinIO inside a single deployment, which suits development and moderate workloads. Distributed runs the same components as separately scalable services on Kubernetes for high throughput and large datasets. Milvus Lite is a lightweight embedded option for prototyping in Python.
Does Milvus support metadata filtering during search?↓
Yes. Milvus supports scalar filtering with a boolean expression language applied alongside vector search, so you can constrain results by numeric ranges, string matches, and JSON fields. Filtering is evaluated during the search rather than as a post-processing step. This keeps recall accurate when filters are selective.
What consistency levels does Milvus offer?↓
Milvus exposes four tunable consistency levels: Strong, Bounded, Session, and Eventually. Strong guarantees the freshest data at higher latency, while Bounded and Eventually trade freshness for speed. Session gives read-your-own-writes consistency within a client. You choose per collection or per query based on the freshness the workload needs.
What does Milvus use for storage and coordination?↓
Milvus depends on three external systems. etcd stores metadata and coordinates the cluster, a log broker such as Apache Pulsar or Kafka handles the streaming write-ahead log, and object storage such as MinIO or S3 holds persisted segments and index files. This separation lets the stateless compute nodes restart without data loss.
When should GPU indexing be used in Milvus?↓
GPU indexes such as GPU_CAGRA and GPU_IVF_PQ help when query throughput is very high or index build times on large datasets become a bottleneck. They deliver strong queries-per-second at scale but consume GPU memory and add operational cost. For moderate traffic, HNSW or DiskANN on CPU is often more economical.
Milvus vs Pinecone: which should I choose?↓
Milvus is open source and self-hosted, giving full control over indexes, hardware, and GPU acceleration, which suits teams with strong infrastructure ownership. Pinecone is fully managed and removes cluster operations at the cost of vendor lock-in and less index-level control. The decision usually turns on whether you want to run the platform yourself.