Skip to primary content
Vector Database Deep Dive

Chroma: Lightweight Embedding Store for Local RAG Prototyping

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

Chroma is an open-source, Apache 2.0 licensed embedding database built for RAG prototyping. It stores document embeddings plus metadata in collections, indexes them with HNSW for approximate nearest neighbor search, and runs in-process or as a lightweight server. It bundles a default embedding function so teams can query semantically with minimal setup.

LicenseApache 2.0
IndexHNSW (hnswlib)
Default metricSquared L2
StorageSQLite plus local index
Problem & Purpose

What Chroma Solves in RAG Prototyping

Early RAG work stalls when teams spend days standing up infrastructure before they can test whether retrieval even helps their use case. Provisioning a managed vector service, wiring authentication, and choosing index parameters is heavy overhead when the real question is whether embeddings surface the right passages. Chroma removes that friction by running in-process with a persistent local store and a bundled default embedding function. A developer can index a corpus and run semantic queries in a few lines, iterate on chunking and metadata, and validate the approach before any production commitment. That fast feedback loop is where Chroma earns its place in the stack.

Inside a Chroma Collection

Anatomy Explainer

Core Component Component Parts:

1. Collection → View Definition
2. HNSW Index → View Definition
3. Embedding Function → View Definition
4. Metadata and Filter Layer → View Definition
5. Persistence Backend → View Definition
PART 1

Collection

The primary container that groups embeddings, documents, IDs, and metadata under a named, configurable unit.

Technical Implementation:

Each collection carries its own distance metric and embedding function. Records are keyed by unique string IDs, and add, update, upsert, and delete operate at the collection level.

The core components that turn documents into searchable embeddings.
Text alternative for screen readers & search engines
  • Part 1: Collection - The primary container that groups embeddings, documents, IDs, and metadata under a named, configurable unit. [Tech: Each collection carries its own distance metric and embedding function. Records are keyed by unique string IDs, and add, update, upsert, and delete operate at the collection level.]
  • Part 2: HNSW Index - The approximate nearest neighbor structure that makes similarity search fast over high dimensional vectors. [Tech: Built on hnswlib, the graph exposes tunable parameters such as construction_ef, search_ef, and M. Trade offs here set the balance between recall, memory footprint, and query latency.]
  • Part 3: Embedding Function - The pluggable component that converts raw text into vectors at insert and query time. [Tech: The default is all-MiniLM-L6-v2 run through ONNX Runtime, yielding 384 dimensions. It can be replaced with OpenAI, Cohere, or custom functions, or bypassed by supplying precomputed embeddings.]
  • Part 4: Metadata and Filter Layer - Structured attributes attached to each record that enable scoped, hybrid retrieval. [Tech: Metadata is stored per record and queried with where clauses on fields plus where_document text filters. Filters are applied alongside vector similarity to constrain the candidate set.]
  • Part 5: Persistence Backend - The durable storage that keeps collections and indexes across process restarts. [Tech: PersistentClient writes metadata and documents to an embedded SQLite database in a local directory, with the HNSW index persisted alongside. No external database server is needed for single node deployments.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Minimal setup: An in-process PersistentClient and a bundled default embedding function let teams index and query a corpus within minutes, with no external service to provision.
  • Clean Python-first API: The add, query, and where interface maps directly onto RAG workflows and integrates natively with LangChain and LlamaIndex.
  • Truly open source: Apache 2.0 licensing removes vendor lock in for self hosted deployments, and the core has been rewritten in Rust for better throughput.
  • Hybrid retrieval built in: Metadata and document filters combine with vector similarity out of the box, supporting tenant scoping and source constrained queries without extra tooling.
Specific Production Limits (Real Constraints)
  • Single node ceiling: Chroma has no native sharding or horizontal scaling, so very large corpora and high concurrency workloads outgrow its design and need a distributed engine.
  • Memory resident index: HNSW graphs are held in memory, so index size is bounded by available RAM and large collections can become costly to serve.
  • Limited operational tooling: Compared to managed engines, Chroma offers fewer built in features for replication, backups, role based access, and multi region durability.
  • Prototyping bias: The defaults optimize for quick starts, which means production reliability, monitoring, and tuning fall to the team rather than the platform.
Production Implementation

How We Deploy Chroma in Production

We treat Chroma as the retrieval layer for early stage RAG builds and internal tools, where the corpus is bounded and time to validation matters more than horizontal scale. Our pattern is to pin the client version, disable telemetry, supply explicit embedding functions rather than rely on defaults, and stand up a Chroma server so multiple application processes share one persistent store. We wrap ingestion in idempotent upserts keyed by stable document IDs, encode tenant and source into metadata for filtered retrieval, and instrument recall against a labeled query set. When a workload crosses Chroma’s single node limits, that same benchmark harness makes the migration to Qdrant or pgvector a measured decision rather than a guess.

Chroma RAG Ingestion and Retrieval Flow

Interactive Flow Diagram
Chroma RAG Ingestion and Retrieval Flow From raw documents to filtered semantic retrieval. 1. Chunk Split source documents 2. Embed Vectorize chunks 3. Upsert Write to collection 4. Query Filtered similarity search 5. Assemble Build LLM context
Stage 1: 1. Chunk Typical chunk 400 to 800 tokens

Documents are segmented into overlapping chunks sized to the embedding model context, preserving section and source references.

From raw documents to filtered semantic retrieval.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Chunk Documents are segmented into overlapping chunks sized to the embedding model context, preserving section and source references. Typical chunk 400 to 800 tokens
2 2. Embed Each chunk is passed through the configured embedding function to produce a dense vector, batched for throughput. 384 dims with MiniLM default
3 3. Upsert Vectors, documents, and metadata are upserted under stable IDs so re-ingestion updates rather than duplicates records. Idempotent by document ID
4 4. Query At request time the query is embedded and matched against the HNSW index with where filters scoping the candidate set. Single digit ms on small corpora
5 5. Assemble Top ranked passages are deduplicated and ordered, then passed as grounded context to the generation model. Top k typically 3 to 8
Production Configuration (Version Pinned):
# requirements: chromadb==1.0.12
import chromadb
from chromadb.config import Settings
from chromadb.utils import embedding_functions

# Connect to a running Chroma server (chroma run --path ./data)
client = chromadb.HttpClient(
  host="localhost",
  port=8000,
  settings=Settings(anonymized_telemetry=False),
)

# Pin the embedding function explicitly instead of relying on defaults
embed_fn = embedding_functions.SentenceTransformerEmbeddingFunction(
  model_name="all-MiniLM-L6-v2"
)

collection = client.get_or_create_collection(
  name="clinical_docs",
  embedding_function=embed_fn,
  metadata={"hnsw:space": "cosine"},
)

# Idempotent upsert keyed by stable IDs
collection.upsert(
  ids=["doc-1", "doc-2"],
  documents=["Patient intake summary...", "Discharge note..."],
  metadatas=[{"tenant": "clinic-a"}, {"tenant": "clinic-a"}],
)

# Filtered similarity query
results = collection.query(
  query_texts=["symptoms on admission"],
  n_results=5,
  where={"tenant": "clinic-a"},
)
print(results["ids"])
Alternatives Evaluation

Chroma vs Alternative Vector Engines

How Chroma stacks up against two common alternatives for RAG workloads.

Chroma, Pinecone, and pgvector Compared

Benchmark Matrix
Evaluation Metric Chroma Pinecone pgvector
Prototyping speed
In-process, minutes to first query Winner
Managed signup and API setup
Needs Postgres and extension setup
Horizontal scale
Single node ceiling
Fully managed, sharded at scale Winner
Scales with Postgres infrastructure
Operational simplicity
Self hosted, DIY ops
Serverless, no ops burden Winner
Leverages existing DBA workflows
Cost control (self hosted)
Free, Apache 2.0, no per query fee Winner
Usage based managed pricing
Free extension on your Postgres
Illustrative relative suitability across four practical dimensions.
Text alternative for screen readers & search engines
  • Prototyping speed: Chroma: In-process, minutes to first query vs Pinecone: Managed signup and API setup vs pgvector: Needs Postgres and extension setup (Winning option: Chroma).
  • Horizontal scale: Chroma: Single node ceiling vs Pinecone: Fully managed, sharded at scale vs pgvector: Scales with Postgres infrastructure (Winning option: Pinecone).
  • Operational simplicity: Chroma: Self hosted, DIY ops vs Pinecone: Serverless, no ops burden vs pgvector: Leverages existing DBA workflows (Winning option: Pinecone).
  • Cost control (self hosted): Chroma: Free, Apache 2.0, no per query fee vs Pinecone: Usage based managed pricing vs pgvector: Free extension on your Postgres (Winning option: Chroma).
Production Proof

Chroma in a Reference Architecture

Clinical RAG

On a healthcare clinical RAG engagement we used Chroma to prototype retrieval over de-identified clinical documents, iterating quickly on chunking strategy and metadata scoping per care setting. The fast local loop let us validate retrieval quality against clinician labeled queries before selecting the production store. Chroma’s filtered queries made tenant and source isolation straightforward during evaluation.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

Is Chroma free and open source?↓

Yes. Chroma is released under the Apache 2.0 license, so it is free to use commercially and to self host. The company also offers Chroma Cloud, a managed hosted service, but the core engine remains fully open source.

What indexing algorithm does Chroma use?↓

Chroma uses HNSW, Hierarchical Navigable Small World graphs, for approximate nearest neighbor search over embeddings. Index parameters such as the construction and search depth are configurable per collection. Metadata and documents are persisted alongside the vector index.

What distance metric does Chroma use by default?↓

The default distance function is squared L2, Euclidean distance. You can set the metric per collection to cosine or inner product through the collection configuration. Choosing the metric that matches how your embedding model was trained matters for retrieval quality.

How does Chroma store data on disk?↓

Using PersistentClient, Chroma writes to a local directory. Metadata, documents, and collection records are kept in a SQLite database, while the HNSW graph is persisted in the same path. There is no separate database server required for single node use.

Can Chroma run as a server?↓

Yes. You can start a Chroma server process and connect with HttpClient over REST, which lets multiple application processes share one store. This is still a single node design and is not a distributed, horizontally sharded cluster by default.

What embedding model does Chroma use by default?↓

By default Chroma uses the all-MiniLM-L6-v2 sentence-transformers model, executed through ONNX Runtime, producing 384 dimensional vectors. You can swap in OpenAI, Cohere, or any custom embedding function, or pass precomputed embeddings directly.

Is Chroma suitable for production at scale?↓

Chroma is excellent for prototyping, local development, and small to mid sized corpora. For very large collections, high concurrency, or multi region deployments, teams typically move to a purpose built engine such as Qdrant, Pinecone, or pgvector. Chroma's single node model is the main scaling constraint.

Does Chroma support metadata filtering?↓

Yes. Each stored record can carry a metadata dictionary, and queries accept where clauses that filter on metadata fields, plus where_document filters on document text. Filters combine with vector similarity so you can scope retrieval to a tenant, source, or date range.

How does Chroma compare to pgvector?↓

Chroma is a standalone embedding store with a Python first API and bundled embedding functions, ideal for quick RAG builds. pgvector adds vector search to existing PostgreSQL, which suits teams that want vectors alongside relational data under one transactional database. The right choice depends on your existing stack.