Redis Vector Search: In-Memory Vector Similarity for Low-Latency Retrieval
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Redis Vector Search is the vector similarity capability in the Redis query engine, part of Redis Stack and Redis 8. It stores embeddings on hash or JSON fields, indexes them with FLAT or HNSW, and serves approximate or exact KNN queries with L2, inner product, or cosine distance, all from in-memory data structures.
What Redis Vector Search Solves in Production
RAG and semantic search systems often need retrieval latency in the single-digit millisecond range while also serving caches, feature stores, and session state. Standing up a separate vector database alongside an existing Redis deployment adds another system to secure, monitor, and keep consistent. Redis Vector Search collapses that split by indexing embeddings on the same in-memory store teams already run. It answers hybrid KNN queries that combine similarity with tag and numeric filters, so retrieval and metadata scoping happen in one round trip. The result is fewer moving parts for latency-sensitive workloads where a network hop to a second store is the bottleneck.
Anatomy of a Redis Vector Index
Anatomy ExplainerCore Component Component Parts:
Vector Field
The embedding column declared on a hash or JSON schema.
Defined in FT.CREATE with TYPE FLOAT32 or FLOAT64, a fixed DIM, and a DISTANCE_METRIC of L2, IP, or COSINE. The metric and dimension are immutable after index creation.
Text alternative for screen readers & search engines
- Part 1: Vector Field - The embedding column declared on a hash or JSON schema. [Tech: Defined in FT.CREATE with TYPE FLOAT32 or FLOAT64, a fixed DIM, and a DISTANCE_METRIC of L2, IP, or COSINE. The metric and dimension are immutable after index creation.]
- Part 2: Index Algorithm - FLAT for exact search or HNSW for approximate graph search. [Tech: HNSW is tuned by M and EF_CONSTRUCTION at build time and EF_RUNTIME per query. FLAT scans all vectors for perfect recall on small sets.]
- Part 3: Keyspace Prefix - The key pattern the index tracks for automatic sync. [Tech: IndexDefinition binds a prefix such as doc: so any matching hash or JSON key is indexed, updated, or removed in real time without an explicit rebuild.]
- Part 4: Hybrid Filter - Attribute predicates applied around the KNN stage. [Tech: Tag, numeric, text, and geo fields are combined with the vector clause in one query string, scoping the neighbor search to a subset of the keyspace.]
- Part 5: Query Executor - The engine that runs KNN and range queries over the index. [Tech: FT.SEARCH and FT.AGGREGATE with query dialect 2 or higher accept the query vector as a bound byte parameter and return neighbors sorted by the computed distance field.]
Architectural Strengths & Specific Production Limits
- Low-latency retrieval: In-memory data structures and HNSW deliver single-digit millisecond KNN queries, which suits interactive RAG and recommendation paths.
- Workload consolidation: The same instance serves caching, session state, streams, and vectors, reducing the number of systems to operate and secure.
- Real-time upserts: Vectors live on ordinary keys, so writing or deleting a key updates the index immediately with no batch reindex step.
- Hybrid filtering: Tag, numeric, text, and geo predicates combine with KNN in a single query, scoping similarity search without a second lookup.
- Memory-bound capacity: The working set must fit in RAM across the cluster, so large vector corpora demand careful sizing or precision reduction.
- No native disk tiering: The open engine lacks larger-than-memory vector storage, unlike engines with disk-backed or tiered indexes.
- Immutable index config: Distance metric and dimension are fixed at FT.CREATE time, so changing them requires dropping and rebuilding the index.
- Licensing considerations: Redis uses RSALv2, SSPLv1, and AGPLv3 terms rather than a permissive license, which some deployments must review or route to the Valkey fork.
How We Deploy Redis Vector Search in Production
Our team uses Redis Vector Search where retrieval latency dominates and an existing Redis footprint is already in place. We pin the Redis Stack or Redis 8 version, define HNSW indexes with recall-tuned M and EF_CONSTRUCTION, and validate against a labeled ground-truth set before promotion. Ingestion writes embeddings and metadata onto prefixed hash keys so the index stays current, and we drive hybrid queries that scope KNN by tenant and document class. Memory sizing, cluster sharding, and eviction policy are locked down as part of the same capacity plan.
Redis Vector Retrieval Pipeline
Interactive Flow DiagramDocuments and queries are embedded to fixed-dimension FLOAT32 vectors by the chosen model.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Embed | Documents and queries are embedded to fixed-dimension FLOAT32 vectors by the chosen model. | DIM 384 to 3072 |
| 2 | 2. Ingest | Vectors and metadata are written to prefixed hash or JSON keys, which the index tracks automatically. | Real-time upsert |
| 3 | 3. Index | The vector field is indexed with tuned M and EF_CONSTRUCTION for the target recall. | M 16, EF 200 |
| 4 | 4. Query | FT.SEARCH combines tag and numeric filters with a KNN clause and a bound query vector. | Dialect 2+ |
| 5 | 5. Rank | Neighbors return sorted by the computed distance field for downstream reranking or context assembly. | p99 low ms |
# redis-py==5.0.8, Redis Stack 7.4 / Redis 8 query engine
import numpy as np
import redis
from redis.commands.search.field import VectorField, TagField
from redis.commands.search.index_definition import IndexDefinition, IndexType
from redis.commands.search.query import Query
r = redis.Redis(host='localhost', port=6379, decode_responses=False)
schema = (
TagField('category'),
VectorField(
'embedding', 'HNSW',
{'TYPE': 'FLOAT32', 'DIM': 1536, 'DISTANCE_METRIC': 'COSINE',
'M': 16, 'EF_CONSTRUCTION': 200},
),
)
r.ft('docs_idx').create_index(
schema,
definition=IndexDefinition(prefix=['doc:'], index_type=IndexType.HASH),
)
vec = np.random.rand(1536).astype(np.float32).tobytes()
q = (
Query('(@category:{clinical})=>[KNN 5 @embedding $vec AS score]')
.sort_by('score')
.return_fields('score', 'category')
.dialect(2)
)
res = r.ft('docs_idx').search(q, query_params={'vec': vec})
for d in res.docs:
print(d.id, d.score)Services Engineered with Redis Vector Search
We design and operate Redis-backed retrieval layers end to end.
Redis Vector Search vs Alternative Engines
How Redis compares to a dedicated vector engine and a relational option on the dimensions that matter for retrieval.
Vector Engine Comparison
Benchmark Matrix| Evaluation Metric | Redis Vector Search | Qdrant | pgvector |
|---|---|---|---|
| Query latency | In-memory, single-digit ms Winner | Optimized, low ms | Depends on Postgres load |
| Larger-than-memory scale | RAM-bound working set | Disk-tiered, quantized Winner | Disk-backed via Postgres |
| Operational consolidation | Shares cache and session store Winner | Separate dedicated service | Rides existing Postgres |
| Filtering and payload richness | Tag, numeric, text, geo | Rich payload filters Winner | Full SQL predicates |
Text alternative for screen readers & search engines
- Query latency: Redis Vector Search: In-memory, single-digit ms vs Qdrant: Optimized, low ms vs pgvector: Depends on Postgres load (Winning option: Redis Vector Search).
- Larger-than-memory scale: Redis Vector Search: RAM-bound working set vs Qdrant: Disk-tiered, quantized vs pgvector: Disk-backed via Postgres (Winning option: Qdrant).
- Operational consolidation: Redis Vector Search: Shares cache and session store vs Qdrant: Separate dedicated service vs pgvector: Rides existing Postgres (Winning option: Redis Vector Search).
- Filtering and payload richness: Redis Vector Search: Tag, numeric, text, geo vs Qdrant: Rich payload filters vs pgvector: Full SQL predicates (Winning option: Qdrant).
Redis Vector Search in a Reference Architecture
For a healthcare clinical RAG system, we needed retrieval fast enough for interactive clinician queries while scoping every search to the correct patient and document class. Redis Vector Search let us combine HNSW similarity with tag filters in a single round trip on infrastructure the team already ran. We tuned index parameters against a labeled ground-truth set to hold recall while keeping tail latency within the interactive budget.
Read Reference Architecture →Frequently Asked Questions
Is Redis a vector database?↓
Redis is a general-purpose in-memory data store that gained vector similarity search through the RediSearch query engine, shipped in Redis Stack and folded into core Redis 8. It indexes embedding fields and answers KNN and range queries, so it functions as a vector database while also serving caching, streams, and key-value workloads on the same instance.
Which index types does Redis Vector Search support?↓
Redis supports two vector index algorithms, FLAT and HNSW. FLAT performs exact brute-force search over every vector and suits small datasets or perfect-recall requirements. HNSW is a graph-based approximate index tuned by the M and EF_CONSTRUCTION build parameters and the EF_RUNTIME query parameter, trading a small recall loss for far lower latency at scale.
What distance metrics can Redis use for vector search?↓
Redis vector fields accept three distance metrics, L2 Euclidean, IP inner product, and COSINE. The metric is fixed at index creation time in the FT.CREATE vector field definition. Cosine is the common choice for normalized text embeddings, while inner product suits dot-product-trained models.
How do I run a KNN query in Redis?↓
You issue an FT.SEARCH with a vector query expression such as a KNN clause that references a query parameter holding the raw byte vector. The query returns the nearest neighbors ordered by the computed distance field. You bind the query vector as a FLOAT32 or FLOAT64 byte blob and set query dialect 2 or higher.
Can Redis combine vector search with metadata filters?↓
Yes. Redis supports hybrid queries that prefilter on tag, numeric, text, or geo fields before or alongside the KNN stage, expressed inside the same query string. This lets you scope a nearest-neighbor search to, for example, a tenant, a document category, or a date range without a separate lookup.
What are the memory implications of Redis Vector Search?↓
Because Redis holds data in RAM, every vector, its index structure, and metadata consume memory, and HNSW graphs add per-vector overhead on top of the raw float payload. Capacity planning is dominated by dimension count, vector precision, and dataset size, so large corpora need careful sizing, sharding across a cluster, or precision reduction.
How does Redis Vector Search compare to a dedicated vector database?↓
Redis excels at low-latency retrieval and consolidating cache, session, and vector workloads on one system, which reduces operational surface. Dedicated engines like Qdrant or Pinecone offer richer built-in quantization, disk-tiered storage, and larger-than-memory datasets. The tradeoff is Redis speed and simplicity versus purpose-built scale and storage efficiency.
What license does Redis use for vector search?↓
Redis relicensed in 2024 to a dual RSALv2 and SSPLv1 model, and Redis 8 in 2025 added AGPLv3 as a third option. The vector search functionality that lived in Redis Stack is available under these terms, and a fully open-source fork, Valkey, exists for teams needing a permissively licensed path.
Does Redis support larger-than-memory vector datasets?↓
Core Redis keeps data in RAM, so the working set must fit in memory across the cluster. There is no native disk-tiered vector storage in the open engine, unlike some competitors. Teams handle scale through cluster sharding, memory sizing, or offloading cold vectors to a secondary store.
How do I update or delete vectors in Redis?↓
Vectors live on ordinary hash or JSON keys, so you update an embedding by writing the field and delete it by removing the key. The vector index reflects these mutations automatically because it tracks the indexed keyspace prefix. This gives real-time upserts without a separate rebuild step.