Chroma: Lightweight Embedding Store for Local RAG Prototyping
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Chroma is an open-source, Apache 2.0 licensed embedding database built for RAG prototyping. It stores document embeddings plus metadata in collections, indexes them with HNSW for approximate nearest neighbor search, and runs in-process or as a lightweight server. It bundles a default embedding function so teams can query semantically with minimal setup.
What Chroma Solves in RAG Prototyping
Early RAG work stalls when teams spend days standing up infrastructure before they can test whether retrieval even helps their use case. Provisioning a managed vector service, wiring authentication, and choosing index parameters is heavy overhead when the real question is whether embeddings surface the right passages. Chroma removes that friction by running in-process with a persistent local store and a bundled default embedding function. A developer can index a corpus and run semantic queries in a few lines, iterate on chunking and metadata, and validate the approach before any production commitment. That fast feedback loop is where Chroma earns its place in the stack.
Inside a Chroma Collection
Anatomy ExplainerCore Component Component Parts:
Collection
The primary container that groups embeddings, documents, IDs, and metadata under a named, configurable unit.
Each collection carries its own distance metric and embedding function. Records are keyed by unique string IDs, and add, update, upsert, and delete operate at the collection level.
Text alternative for screen readers & search engines
- Part 1: Collection - The primary container that groups embeddings, documents, IDs, and metadata under a named, configurable unit. [Tech: Each collection carries its own distance metric and embedding function. Records are keyed by unique string IDs, and add, update, upsert, and delete operate at the collection level.]
- Part 2: HNSW Index - The approximate nearest neighbor structure that makes similarity search fast over high dimensional vectors. [Tech: Built on hnswlib, the graph exposes tunable parameters such as construction_ef, search_ef, and M. Trade offs here set the balance between recall, memory footprint, and query latency.]
- Part 3: Embedding Function - The pluggable component that converts raw text into vectors at insert and query time. [Tech: The default is all-MiniLM-L6-v2 run through ONNX Runtime, yielding 384 dimensions. It can be replaced with OpenAI, Cohere, or custom functions, or bypassed by supplying precomputed embeddings.]
- Part 4: Metadata and Filter Layer - Structured attributes attached to each record that enable scoped, hybrid retrieval. [Tech: Metadata is stored per record and queried with where clauses on fields plus where_document text filters. Filters are applied alongside vector similarity to constrain the candidate set.]
- Part 5: Persistence Backend - The durable storage that keeps collections and indexes across process restarts. [Tech: PersistentClient writes metadata and documents to an embedded SQLite database in a local directory, with the HNSW index persisted alongside. No external database server is needed for single node deployments.]
Architectural Strengths & Specific Production Limits
- Minimal setup: An in-process PersistentClient and a bundled default embedding function let teams index and query a corpus within minutes, with no external service to provision.
- Clean Python-first API: The add, query, and where interface maps directly onto RAG workflows and integrates natively with LangChain and LlamaIndex.
- Truly open source: Apache 2.0 licensing removes vendor lock in for self hosted deployments, and the core has been rewritten in Rust for better throughput.
- Hybrid retrieval built in: Metadata and document filters combine with vector similarity out of the box, supporting tenant scoping and source constrained queries without extra tooling.
- Single node ceiling: Chroma has no native sharding or horizontal scaling, so very large corpora and high concurrency workloads outgrow its design and need a distributed engine.
- Memory resident index: HNSW graphs are held in memory, so index size is bounded by available RAM and large collections can become costly to serve.
- Limited operational tooling: Compared to managed engines, Chroma offers fewer built in features for replication, backups, role based access, and multi region durability.
- Prototyping bias: The defaults optimize for quick starts, which means production reliability, monitoring, and tuning fall to the team rather than the platform.
How We Deploy Chroma in Production
We treat Chroma as the retrieval layer for early stage RAG builds and internal tools, where the corpus is bounded and time to validation matters more than horizontal scale. Our pattern is to pin the client version, disable telemetry, supply explicit embedding functions rather than rely on defaults, and stand up a Chroma server so multiple application processes share one persistent store. We wrap ingestion in idempotent upserts keyed by stable document IDs, encode tenant and source into metadata for filtered retrieval, and instrument recall against a labeled query set. When a workload crosses Chroma’s single node limits, that same benchmark harness makes the migration to Qdrant or pgvector a measured decision rather than a guess.
Chroma RAG Ingestion and Retrieval Flow
Interactive Flow DiagramDocuments are segmented into overlapping chunks sized to the embedding model context, preserving section and source references.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Chunk | Documents are segmented into overlapping chunks sized to the embedding model context, preserving section and source references. | Typical chunk 400 to 800 tokens |
| 2 | 2. Embed | Each chunk is passed through the configured embedding function to produce a dense vector, batched for throughput. | 384 dims with MiniLM default |
| 3 | 3. Upsert | Vectors, documents, and metadata are upserted under stable IDs so re-ingestion updates rather than duplicates records. | Idempotent by document ID |
| 4 | 4. Query | At request time the query is embedded and matched against the HNSW index with where filters scoping the candidate set. | Single digit ms on small corpora |
| 5 | 5. Assemble | Top ranked passages are deduplicated and ordered, then passed as grounded context to the generation model. | Top k typically 3 to 8 |
# requirements: chromadb==1.0.12
import chromadb
from chromadb.config import Settings
from chromadb.utils import embedding_functions
# Connect to a running Chroma server (chroma run --path ./data)
client = chromadb.HttpClient(
host="localhost",
port=8000,
settings=Settings(anonymized_telemetry=False),
)
# Pin the embedding function explicitly instead of relying on defaults
embed_fn = embedding_functions.SentenceTransformerEmbeddingFunction(
model_name="all-MiniLM-L6-v2"
)
collection = client.get_or_create_collection(
name="clinical_docs",
embedding_function=embed_fn,
metadata={"hnsw:space": "cosine"},
)
# Idempotent upsert keyed by stable IDs
collection.upsert(
ids=["doc-1", "doc-2"],
documents=["Patient intake summary...", "Discharge note..."],
metadatas=[{"tenant": "clinic-a"}, {"tenant": "clinic-a"}],
)
# Filtered similarity query
results = collection.query(
query_texts=["symptoms on admission"],
n_results=5,
where={"tenant": "clinic-a"},
)
print(results["ids"])Services Engineered with Chroma
Teams engage us to design and harden the retrieval layer around Chroma and the data pipelines that feed it.
Chroma vs Alternative Vector Engines
How Chroma stacks up against two common alternatives for RAG workloads.
Chroma, Pinecone, and pgvector Compared
Benchmark Matrix| Evaluation Metric | Chroma | Pinecone | pgvector |
|---|---|---|---|
| Prototyping speed | In-process, minutes to first query Winner | Managed signup and API setup | Needs Postgres and extension setup |
| Horizontal scale | Single node ceiling | Fully managed, sharded at scale Winner | Scales with Postgres infrastructure |
| Operational simplicity | Self hosted, DIY ops | Serverless, no ops burden Winner | Leverages existing DBA workflows |
| Cost control (self hosted) | Free, Apache 2.0, no per query fee Winner | Usage based managed pricing | Free extension on your Postgres |
Text alternative for screen readers & search engines
- Prototyping speed: Chroma: In-process, minutes to first query vs Pinecone: Managed signup and API setup vs pgvector: Needs Postgres and extension setup (Winning option: Chroma).
- Horizontal scale: Chroma: Single node ceiling vs Pinecone: Fully managed, sharded at scale vs pgvector: Scales with Postgres infrastructure (Winning option: Pinecone).
- Operational simplicity: Chroma: Self hosted, DIY ops vs Pinecone: Serverless, no ops burden vs pgvector: Leverages existing DBA workflows (Winning option: Pinecone).
- Cost control (self hosted): Chroma: Free, Apache 2.0, no per query fee vs Pinecone: Usage based managed pricing vs pgvector: Free extension on your Postgres (Winning option: Chroma).
Chroma in a Reference Architecture
On a healthcare clinical RAG engagement we used Chroma to prototype retrieval over de-identified clinical documents, iterating quickly on chunking strategy and metadata scoping per care setting. The fast local loop let us validate retrieval quality against clinician labeled queries before selecting the production store. Chroma’s filtered queries made tenant and source isolation straightforward during evaluation.
Read Reference Architecture →Frequently Asked Questions
Is Chroma free and open source?↓
Yes. Chroma is released under the Apache 2.0 license, so it is free to use commercially and to self host. The company also offers Chroma Cloud, a managed hosted service, but the core engine remains fully open source.
What indexing algorithm does Chroma use?↓
Chroma uses HNSW, Hierarchical Navigable Small World graphs, for approximate nearest neighbor search over embeddings. Index parameters such as the construction and search depth are configurable per collection. Metadata and documents are persisted alongside the vector index.
What distance metric does Chroma use by default?↓
The default distance function is squared L2, Euclidean distance. You can set the metric per collection to cosine or inner product through the collection configuration. Choosing the metric that matches how your embedding model was trained matters for retrieval quality.
How does Chroma store data on disk?↓
Using PersistentClient, Chroma writes to a local directory. Metadata, documents, and collection records are kept in a SQLite database, while the HNSW graph is persisted in the same path. There is no separate database server required for single node use.
Can Chroma run as a server?↓
Yes. You can start a Chroma server process and connect with HttpClient over REST, which lets multiple application processes share one store. This is still a single node design and is not a distributed, horizontally sharded cluster by default.
What embedding model does Chroma use by default?↓
By default Chroma uses the all-MiniLM-L6-v2 sentence-transformers model, executed through ONNX Runtime, producing 384 dimensional vectors. You can swap in OpenAI, Cohere, or any custom embedding function, or pass precomputed embeddings directly.
Is Chroma suitable for production at scale?↓
Chroma is excellent for prototyping, local development, and small to mid sized corpora. For very large collections, high concurrency, or multi region deployments, teams typically move to a purpose built engine such as Qdrant, Pinecone, or pgvector. Chroma's single node model is the main scaling constraint.
Does Chroma support metadata filtering?↓
Yes. Each stored record can carry a metadata dictionary, and queries accept where clauses that filter on metadata fields, plus where_document filters on document text. Filters combine with vector similarity so you can scope retrieval to a tenant, source, or date range.
How does Chroma compare to pgvector?↓
Chroma is a standalone embedding store with a Python first API and bundled embedding functions, ideal for quick RAG builds. pgvector adds vector search to existing PostgreSQL, which suits teams that want vectors alongside relational data under one transactional database. The right choice depends on your existing stack.