Skip to primary content
Vector Database Deep Dive

Elasticsearch kNN: Lucene HNSW Vector Search With BM25 Hybrid Retrieval

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

Elasticsearch kNN adds approximate nearest neighbor search to Elasticsearch using the dense_vector field type backed by Lucene HNSW graphs. It supports cosine, dot product, and L2 similarity, scalar and binary quantization, and combines with BM25 lexical scoring through reciprocal rank fusion for hybrid retrieval at index scale across sharded clusters.

LicenseElastic License 2.0, SSPL, AGPL v3
ANN AlgorithmLucene HNSW
Max Indexed Dims4096
Default Vector Indexint8_hnsw
Problem & Purpose

What Elasticsearch kNN Solves in Production

Teams that already run Elasticsearch for logs, catalog search, or observability rarely want a second data store just for embeddings. Bolting a dedicated vector engine alongside an existing cluster duplicates ingest pipelines, permissions, and operational tooling. Elasticsearch kNN keeps dense vectors, BM25 text fields, filters, and aggregations in one index and one query. That matters most for retrieval augmented generation, where semantic recall alone misses exact identifiers, product codes, and rare terms that lexical matching catches. The reciprocal rank fusion retriever lets one request return a blended ranking without maintaining separate systems.

Inside an Elasticsearch kNN Index

Anatomy Explainer

Core Component Component Parts:

1. dense_vector field → View Definition
2. Lucene HNSW graph → View Definition
3. Quantization codec → View Definition
4. knn query and num_candidates → View Definition
5. RRF retriever → View Definition
PART 1

dense_vector field

The mapping type that stores per-document embeddings and enables ANN indexing.

Technical Implementation:

Set index to true, choose dims up to 4096, and pick a similarity of cosine, dot_product, l2_norm, or max_inner_product. Vectors are stored per segment.

The dense_vector field, the Lucene HNSW graph, and the retrieval path that turns a query vector into ranked hits.
Text alternative for screen readers & search engines
  • Part 1: dense_vector field - The mapping type that stores per-document embeddings and enables ANN indexing. [Tech: Set index to true, choose dims up to 4096, and pick a similarity of cosine, dot_product, l2_norm, or max_inner_product. Vectors are stored per segment.]
  • Part 2: Lucene HNSW graph - The layered proximity graph that makes approximate nearest neighbor search fast. [Tech: Built per Lucene segment with m controlling connections per node (default 16) and ef_construction controlling build breadth (default 100). Segment merges rebuild graph links.]
  • Part 3: Quantization codec - Compresses float vectors to shrink heap and page cache pressure. [Tech: int8_hnsw is the float default since 8.14, with int4_hnsw and bbq_hnsw available. Compressed vectors are searched first, then candidates are re-ranked against higher precision copies.]
  • Part 4: knn query and num_candidates - The per-shard search that walks the graph and collects candidates. [Tech: Each shard returns num_candidates neighbors, the coordinating node reduces them to the top k. Optional filter clauses are applied as pre-filters during traversal to keep recall high.]
  • Part 5: RRF retriever - Fuses lexical BM25 and vector result sets into one ranking. [Tech: The rrf retriever merges ranked lists by reciprocal rank with a tunable rank_constant and rank_window_size, avoiding score normalization across incompatible scoring functions.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • One store for text and vectors: Embeddings, BM25 fields, structured filters, and aggregations live in the same index, so hybrid RAG needs no second database or separate sync pipeline.
  • Native reciprocal rank fusion: The rrf retriever blends lexical and semantic rankings by rank position, which sidesteps the fragile score normalization required by linear fusion schemes.
  • Filter-aware ANN: Filters are applied as pre-filters during HNSW traversal, so constrained queries keep recall high instead of losing candidates to post-filtering.
  • Mature operations story: Sharding, replication, snapshots, role based access, and Kibana monitoring carry over directly, which shortens the path to production for existing Elastic users.
Specific Production Limits (Real Constraints)
  • Per-segment graph cost: HNSW graphs are built per Lucene segment, so heavy indexing and frequent merges add CPU overhead and force force-merge planning for stable latency.
  • Heap and memory pressure: Vector graphs and quantized data compete for off-heap page cache, and undersized nodes see recall or latency degrade sharply as the working set exceeds RAM.
  • License restrictions: Elastic License 2.0 and SSPL prohibit offering Elasticsearch as a managed service, which rules the project out for some vendors despite source availability.
  • Tuning surface area: Getting good recall requires coordinated tuning of m, ef_construction, num_candidates, quantization, and shard count, which is more involved than a purpose-built vector engine default.
Production Implementation

How We Deploy Elasticsearch kNN in Production

We reach for Elasticsearch kNN when a client already runs Elastic for search or observability and wants retrieval augmented generation without standing up a separate vector store. Our work starts with mapping design, choosing similarity to match the embedding model and selecting a quantization codec against a measured recall target. We size shards and heap for the vector working set, wire pre-filters for tenant and access control, and expose retrieval through the rrf retriever so lexical and semantic results arrive in one ranked response.

Elasticsearch kNN Retrieval Pipeline

Interactive Flow Diagram
Elasticsearch kNN Retrieval Pipeline From embedding ingest to a fused hybrid ranking served to the application. 1. Embed and Ingest Bulk API 2. Build HNSW Per segment 3. Query Vectors knn retriever 4. Fuse Rankings rrf 5. Re-rank and Serve Optional
Stage 1: 1. Embed and Ingest Batched bulk requests, pipeline preprocessing

Documents are embedded upstream and indexed with their text and dense_vector fields through the bulk API.

From embedding ingest to a fused hybrid ranking served to the application.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Embed and Ingest Documents are embedded upstream and indexed with their text and dense_vector fields through the bulk API. Batched bulk requests, pipeline preprocessing
2 2. Build HNSW Lucene builds an HNSW graph per segment using the configured m and ef_construction, applying quantization on write. m 16, ef_construction 100 defaults
3 3. Query Vectors The query embedding is searched across shards, each returning num_candidates neighbors with any pre-filters applied. num_candidates tuned for recall at k
4 4. Fuse Rankings The reciprocal rank fusion retriever merges the BM25 and vector result sets into a single ranked list. rank_constant, rank_window_size
5 5. Re-rank and Serve Top candidates can be re-ranked with a cross encoder or ELSER before results return to the application. p95 latency budget enforced
Production Configuration (Version Pinned):
# Elasticsearch 8.15 - dense_vector mapping and hybrid RRF search
PUT /products
{
"mappings": {
  "properties": {
    "title": { "type": "text" },
    "embedding": {
      "type": "dense_vector",
      "dims": 1024,
      "index": true,
      "similarity": "cosine",
      "index_options": {
        "type": "int8_hnsw",
        "m": 16,
        "ef_construction": 100
      }
    }
  }
}
}

POST /products/_search
{
"retriever": {
  "rrf": {
    "retrievers": [
      { "standard": { "query": { "match": { "title": "wireless headphones" } } } },
      { "knn": {
          "field": "embedding",
          "query_vector": [0.12, 0.07, 0.44],
          "k": 50,
          "num_candidates": 200
      } }
    ],
    "rank_window_size": 50,
    "rank_constant": 20
  }
}
}
Delivering Commercial Impact

Services Engineered with Elasticsearch kNN

We design and operate Elasticsearch kNN retrieval as part of these engagements.

Alternatives Evaluation

Elasticsearch kNN vs Alternative Vector Engines

How Elasticsearch kNN compares against purpose-built vector databases for hybrid retrieval workloads.

Elasticsearch kNN vs Qdrant vs Weaviate

Benchmark Matrix
Evaluation Metric Elasticsearch kNN Qdrant Weaviate
Hybrid lexical and vector search
Native BM25 with RRF retriever Winner
Sparse and dense fusion
BM25 and vector hybrid module
Filtered vector query performance
Pre-filter during traversal
Payload index filtered HNSW Winner
Filtered ANN with inverted index
Schema, objects, and multi-tenancy
Index and mapping model
Collections and payloads
Class schema with native tenants Winner
Analytics and observability ecosystem
Kibana, aggregations, log stack Winner
Focused vector API
GraphQL and modules
Illustrative relative suitability by workload, not measured benchmarks.
Text alternative for screen readers & search engines
  • Hybrid lexical and vector search: Elasticsearch kNN: Native BM25 with RRF retriever vs Qdrant: Sparse and dense fusion vs Weaviate: BM25 and vector hybrid module (Winning option: Elasticsearch kNN).
  • Filtered vector query performance: Elasticsearch kNN: Pre-filter during traversal vs Qdrant: Payload index filtered HNSW vs Weaviate: Filtered ANN with inverted index (Winning option: Qdrant).
  • Schema, objects, and multi-tenancy: Elasticsearch kNN: Index and mapping model vs Qdrant: Collections and payloads vs Weaviate: Class schema with native tenants (Winning option: Weaviate).
  • Analytics and observability ecosystem: Elasticsearch kNN: Kibana, aggregations, log stack vs Qdrant: Focused vector API vs Weaviate: GraphQL and modules (Winning option: Elasticsearch kNN).
Production Proof

Elasticsearch kNN in a Reference Architecture

Clinical RAG Retrieval

For a healthcare clinical RAG platform we used Elasticsearch kNN to combine semantic recall over clinical notes with exact BM25 matching on codes and identifiers, fused through the RRF retriever. Pre-filters enforced tenant and record-level access during graph traversal, keeping restricted content out of candidate sets while preserving recall. Keeping vectors and structured filters in one store simplified the compliance and audit surface.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is kNN search in Elasticsearch?↓

kNN search finds the k approximate nearest neighbors to a query vector using the dense_vector field type. Elasticsearch builds a Lucene HNSW graph per segment and traverses it to return the closest vectors by the configured similarity. It is approximate, so recall is traded against latency through the num_candidates parameter.

Which HNSW parameters does Elasticsearch expose?↓

The index_options for an HNSW dense_vector accept m and ef_construction. The default m is 16 and the default ef_construction is 100. Raising m increases graph connectivity and recall at the cost of index size, while ef_construction controls build-time search breadth and indexing speed.

How does Elasticsearch do hybrid search?↓

Elasticsearch combines a lexical BM25 query and a vector knn query using the rrf retriever, which applies reciprocal rank fusion. Each result set is ranked independently and merged by rank position rather than by raw score, so lexical and semantic signals combine without score normalization. This avoids the calibration problems of linear score blending.

What is the maximum vector dimension in Elasticsearch?↓

A dense_vector field indexed for kNN supports up to 4096 dimensions in current 8.x and 9.x releases. Fields configured with index set to false, used only for exact brute-force scoring, can hold larger vectors. Most embedding models fit well within the indexed limit.

Does Elasticsearch quantize vectors?↓

Yes. Since 8.14 the default index type for float dense_vectors is int8_hnsw, which applies scalar quantization to roughly a quarter of the raw memory footprint. Later releases add int4_hnsw and bbq_hnsw, a Better Binary Quantization option that compresses further while re-ranking candidates against fuller precision to preserve recall.

Is Elasticsearch open source?↓

Elasticsearch is distributed under the Elastic License 2.0 and the Server Side Public License, not the Apache 2.0 license used before 7.11. Source is publicly available and self-hosting is free, but the licenses restrict offering Elasticsearch as a managed service. Elastic also restored an AGPL v3 option in 2024.

What similarity metrics does Elasticsearch support for vectors?↓

A dense_vector field can be configured with the similarity parameter set to cosine, dot_product, l2_norm, or max_inner_product. For dot_product Elasticsearch expects unit-length vectors, while cosine normalizes internally at some indexing cost. The choice must match how the embedding model was trained.

How does num_candidates affect kNN recall and latency?↓

num_candidates sets how many nearest neighbors each shard collects from its HNSW graph before returning the top k. Higher values improve recall by widening the graph traversal but increase CPU and latency per shard. The final results are the best k across all shard candidate lists.

Can Elasticsearch filter during kNN search?↓

Yes. The knn query and knn retriever accept a filter clause that is applied as a pre-filter during graph traversal, so only documents matching the filter are considered as neighbors. This avoids the recall loss of post-filtering, where excluded documents shrink an already fixed candidate set.