Elasticsearch kNN: Lucene HNSW Vector Search With BM25 Hybrid Retrieval
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Elasticsearch kNN adds approximate nearest neighbor search to Elasticsearch using the dense_vector field type backed by Lucene HNSW graphs. It supports cosine, dot product, and L2 similarity, scalar and binary quantization, and combines with BM25 lexical scoring through reciprocal rank fusion for hybrid retrieval at index scale across sharded clusters.
What Elasticsearch kNN Solves in Production
Teams that already run Elasticsearch for logs, catalog search, or observability rarely want a second data store just for embeddings. Bolting a dedicated vector engine alongside an existing cluster duplicates ingest pipelines, permissions, and operational tooling. Elasticsearch kNN keeps dense vectors, BM25 text fields, filters, and aggregations in one index and one query. That matters most for retrieval augmented generation, where semantic recall alone misses exact identifiers, product codes, and rare terms that lexical matching catches. The reciprocal rank fusion retriever lets one request return a blended ranking without maintaining separate systems.
Inside an Elasticsearch kNN Index
Anatomy ExplainerCore Component Component Parts:
dense_vector field
The mapping type that stores per-document embeddings and enables ANN indexing.
Set index to true, choose dims up to 4096, and pick a similarity of cosine, dot_product, l2_norm, or max_inner_product. Vectors are stored per segment.
Text alternative for screen readers & search engines
- Part 1: dense_vector field - The mapping type that stores per-document embeddings and enables ANN indexing. [Tech: Set index to true, choose dims up to 4096, and pick a similarity of cosine, dot_product, l2_norm, or max_inner_product. Vectors are stored per segment.]
- Part 2: Lucene HNSW graph - The layered proximity graph that makes approximate nearest neighbor search fast. [Tech: Built per Lucene segment with m controlling connections per node (default 16) and ef_construction controlling build breadth (default 100). Segment merges rebuild graph links.]
- Part 3: Quantization codec - Compresses float vectors to shrink heap and page cache pressure. [Tech: int8_hnsw is the float default since 8.14, with int4_hnsw and bbq_hnsw available. Compressed vectors are searched first, then candidates are re-ranked against higher precision copies.]
- Part 4: knn query and num_candidates - The per-shard search that walks the graph and collects candidates. [Tech: Each shard returns num_candidates neighbors, the coordinating node reduces them to the top k. Optional filter clauses are applied as pre-filters during traversal to keep recall high.]
- Part 5: RRF retriever - Fuses lexical BM25 and vector result sets into one ranking. [Tech: The rrf retriever merges ranked lists by reciprocal rank with a tunable rank_constant and rank_window_size, avoiding score normalization across incompatible scoring functions.]
Architectural Strengths & Specific Production Limits
- One store for text and vectors: Embeddings, BM25 fields, structured filters, and aggregations live in the same index, so hybrid RAG needs no second database or separate sync pipeline.
- Native reciprocal rank fusion: The rrf retriever blends lexical and semantic rankings by rank position, which sidesteps the fragile score normalization required by linear fusion schemes.
- Filter-aware ANN: Filters are applied as pre-filters during HNSW traversal, so constrained queries keep recall high instead of losing candidates to post-filtering.
- Mature operations story: Sharding, replication, snapshots, role based access, and Kibana monitoring carry over directly, which shortens the path to production for existing Elastic users.
- Per-segment graph cost: HNSW graphs are built per Lucene segment, so heavy indexing and frequent merges add CPU overhead and force force-merge planning for stable latency.
- Heap and memory pressure: Vector graphs and quantized data compete for off-heap page cache, and undersized nodes see recall or latency degrade sharply as the working set exceeds RAM.
- License restrictions: Elastic License 2.0 and SSPL prohibit offering Elasticsearch as a managed service, which rules the project out for some vendors despite source availability.
- Tuning surface area: Getting good recall requires coordinated tuning of m, ef_construction, num_candidates, quantization, and shard count, which is more involved than a purpose-built vector engine default.
How We Deploy Elasticsearch kNN in Production
We reach for Elasticsearch kNN when a client already runs Elastic for search or observability and wants retrieval augmented generation without standing up a separate vector store. Our work starts with mapping design, choosing similarity to match the embedding model and selecting a quantization codec against a measured recall target. We size shards and heap for the vector working set, wire pre-filters for tenant and access control, and expose retrieval through the rrf retriever so lexical and semantic results arrive in one ranked response.
Elasticsearch kNN Retrieval Pipeline
Interactive Flow DiagramDocuments are embedded upstream and indexed with their text and dense_vector fields through the bulk API.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Embed and Ingest | Documents are embedded upstream and indexed with their text and dense_vector fields through the bulk API. | Batched bulk requests, pipeline preprocessing |
| 2 | 2. Build HNSW | Lucene builds an HNSW graph per segment using the configured m and ef_construction, applying quantization on write. | m 16, ef_construction 100 defaults |
| 3 | 3. Query Vectors | The query embedding is searched across shards, each returning num_candidates neighbors with any pre-filters applied. | num_candidates tuned for recall at k |
| 4 | 4. Fuse Rankings | The reciprocal rank fusion retriever merges the BM25 and vector result sets into a single ranked list. | rank_constant, rank_window_size |
| 5 | 5. Re-rank and Serve | Top candidates can be re-ranked with a cross encoder or ELSER before results return to the application. | p95 latency budget enforced |
# Elasticsearch 8.15 - dense_vector mapping and hybrid RRF search
PUT /products
{
"mappings": {
"properties": {
"title": { "type": "text" },
"embedding": {
"type": "dense_vector",
"dims": 1024,
"index": true,
"similarity": "cosine",
"index_options": {
"type": "int8_hnsw",
"m": 16,
"ef_construction": 100
}
}
}
}
}
POST /products/_search
{
"retriever": {
"rrf": {
"retrievers": [
{ "standard": { "query": { "match": { "title": "wireless headphones" } } } },
{ "knn": {
"field": "embedding",
"query_vector": [0.12, 0.07, 0.44],
"k": 50,
"num_candidates": 200
} }
],
"rank_window_size": 50,
"rank_constant": 20
}
}
}Services Engineered with Elasticsearch kNN
We design and operate Elasticsearch kNN retrieval as part of these engagements.
Elasticsearch kNN vs Alternative Vector Engines
How Elasticsearch kNN compares against purpose-built vector databases for hybrid retrieval workloads.
Elasticsearch kNN vs Qdrant vs Weaviate
Benchmark Matrix| Evaluation Metric | Elasticsearch kNN | Qdrant | Weaviate |
|---|---|---|---|
| Hybrid lexical and vector search | Native BM25 with RRF retriever Winner | Sparse and dense fusion | BM25 and vector hybrid module |
| Filtered vector query performance | Pre-filter during traversal | Payload index filtered HNSW Winner | Filtered ANN with inverted index |
| Schema, objects, and multi-tenancy | Index and mapping model | Collections and payloads | Class schema with native tenants Winner |
| Analytics and observability ecosystem | Kibana, aggregations, log stack Winner | Focused vector API | GraphQL and modules |
Text alternative for screen readers & search engines
- Hybrid lexical and vector search: Elasticsearch kNN: Native BM25 with RRF retriever vs Qdrant: Sparse and dense fusion vs Weaviate: BM25 and vector hybrid module (Winning option: Elasticsearch kNN).
- Filtered vector query performance: Elasticsearch kNN: Pre-filter during traversal vs Qdrant: Payload index filtered HNSW vs Weaviate: Filtered ANN with inverted index (Winning option: Qdrant).
- Schema, objects, and multi-tenancy: Elasticsearch kNN: Index and mapping model vs Qdrant: Collections and payloads vs Weaviate: Class schema with native tenants (Winning option: Weaviate).
- Analytics and observability ecosystem: Elasticsearch kNN: Kibana, aggregations, log stack vs Qdrant: Focused vector API vs Weaviate: GraphQL and modules (Winning option: Elasticsearch kNN).
Elasticsearch kNN in a Reference Architecture
For a healthcare clinical RAG platform we used Elasticsearch kNN to combine semantic recall over clinical notes with exact BM25 matching on codes and identifiers, fused through the RRF retriever. Pre-filters enforced tenant and record-level access during graph traversal, keeping restricted content out of candidate sets while preserving recall. Keeping vectors and structured filters in one store simplified the compliance and audit surface.
Read Reference Architecture →Frequently Asked Questions
What is kNN search in Elasticsearch?↓
kNN search finds the k approximate nearest neighbors to a query vector using the dense_vector field type. Elasticsearch builds a Lucene HNSW graph per segment and traverses it to return the closest vectors by the configured similarity. It is approximate, so recall is traded against latency through the num_candidates parameter.
Which HNSW parameters does Elasticsearch expose?↓
The index_options for an HNSW dense_vector accept m and ef_construction. The default m is 16 and the default ef_construction is 100. Raising m increases graph connectivity and recall at the cost of index size, while ef_construction controls build-time search breadth and indexing speed.
How does Elasticsearch do hybrid search?↓
Elasticsearch combines a lexical BM25 query and a vector knn query using the rrf retriever, which applies reciprocal rank fusion. Each result set is ranked independently and merged by rank position rather than by raw score, so lexical and semantic signals combine without score normalization. This avoids the calibration problems of linear score blending.
What is the maximum vector dimension in Elasticsearch?↓
A dense_vector field indexed for kNN supports up to 4096 dimensions in current 8.x and 9.x releases. Fields configured with index set to false, used only for exact brute-force scoring, can hold larger vectors. Most embedding models fit well within the indexed limit.
Does Elasticsearch quantize vectors?↓
Yes. Since 8.14 the default index type for float dense_vectors is int8_hnsw, which applies scalar quantization to roughly a quarter of the raw memory footprint. Later releases add int4_hnsw and bbq_hnsw, a Better Binary Quantization option that compresses further while re-ranking candidates against fuller precision to preserve recall.
Is Elasticsearch open source?↓
Elasticsearch is distributed under the Elastic License 2.0 and the Server Side Public License, not the Apache 2.0 license used before 7.11. Source is publicly available and self-hosting is free, but the licenses restrict offering Elasticsearch as a managed service. Elastic also restored an AGPL v3 option in 2024.
What similarity metrics does Elasticsearch support for vectors?↓
A dense_vector field can be configured with the similarity parameter set to cosine, dot_product, l2_norm, or max_inner_product. For dot_product Elasticsearch expects unit-length vectors, while cosine normalizes internally at some indexing cost. The choice must match how the embedding model was trained.
How does num_candidates affect kNN recall and latency?↓
num_candidates sets how many nearest neighbors each shard collects from its HNSW graph before returning the top k. Higher values improve recall by widening the graph traversal but increase CPU and latency per shard. The final results are the best k across all shard candidate lists.
Can Elasticsearch filter during kNN search?↓
Yes. The knn query and knn retriever accept a filter clause that is applied as a pre-filter during graph traversal, so only documents matching the filter are considered as neighbors. This avoids the recall loss of post-filtering, where excluded documents shrink an already fixed candidate set.