Haystack: Building Composable RAG and LLM Pipelines in Production
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
Haystack is an open source Python framework from deepset for building retrieval augmented generation and LLM applications. It connects modular components such as retrievers, embedders, rankers, and generators into explicit pipeline graphs. Version 2.x runs on typed input output connections and supports document stores like OpenSearch, Qdrant, and pgvector.
What Haystack Solves in Production
RAG systems fail quietly when retrieval, prompting, and generation are tangled into one script that nobody can inspect or evaluate. Swapping a vector store or an LLM provider then means rewriting glue code and re-testing the whole flow. Prompt logic drifts, retrieval quality is never measured, and the same pipeline behaves differently in a notebook than in the service. Haystack addresses this by making each stage an explicit typed component and the whole flow an inspectable pipeline graph. That structure is what lets a team reason about, evaluate, and redeploy a RAG system without guesswork.
Anatomy of a Haystack Pipeline
Anatomy ExplainerCore Component Component Parts:
Component
The fundamental unit of work, a Python class decorated with @component that exposes a typed run method.
Inputs and outputs are declared with types so the pipeline can validate connections. Custom components subclass nothing special, they just implement run and declare output types.
Text alternative for screen readers & search engines
- Part 1: Component - The fundamental unit of work, a Python class decorated with @component that exposes a typed run method. [Tech: Inputs and outputs are declared with types so the pipeline can validate connections. Custom components subclass nothing special, they just implement run and declare output types.]
- Part 2: Pipeline - A directed graph that connects component outputs to inputs and executes them in dependency order. [Tech: Connections are made with pipeline.connect(sender.output, receiver.input). The graph serializes to and from YAML for reproducible deployment and inspection.]
- Part 3: Document Store - The storage abstraction over vector and keyword backends holding your indexed documents. [Tech: InMemoryDocumentStore ships in core, while OpenSearch, Qdrant, Weaviate, pgvector, and others come as separate integration packages implementing a common interface.]
- Part 4: Retriever - Fetches candidate documents for a query using lexical, embedding, or hybrid strategies. [Tech: BM25 retrievers score by keyword overlap, embedding retrievers use a text embedder and vector similarity, and a DocumentJoiner merges both for hybrid retrieval.]
- Part 5: PromptBuilder and Generator - Renders a Jinja2 prompt from retrieved documents and invokes the language model. [Tech: PromptBuilder templates iterate over documents, and generators such as OpenAIGenerator or HuggingFaceLocalGenerator return replies and metadata for downstream components.]
Architectural Strengths & Specific Production Limits
- Explicit pipeline graphs: Typed connections between components make data flow inspectable and catch mismatched wiring before runtime rather than deep inside a call stack.
- Backend portability: A common document store interface lets you move from InMemory to OpenSearch, Qdrant, or pgvector without changing retrieval logic.
- Serializable and deployable: Pipelines serialize to YAML and serve over HTTP via Hayhooks, so the same definition runs in development and production.
- Strong retrieval primitives: First class BM25, embedding, and hybrid retrievers with rerankers give you the building blocks RAG quality actually depends on.
- Python only: Haystack is a Python framework with no official runtime for other languages, so polyglot stacks must call it over a service boundary.
- Integration packaging overhead: Document store and model provider integrations ship as separate packages that version independently, so dependency pinning needs care.
- Smaller ecosystem than LangChain: Fewer third party wrappers and community tutorials exist compared to LangChain, so some niche integrations require writing custom components.
- Agent tooling is newer: Multi step tool calling and agent patterns are supported but less mature than the retrieval side, and evolve across minor releases.
How We Deploy Haystack in Production
Our team treats a Haystack pipeline as a versioned artifact, not a script. We pin haystack-ai and each integration package, define the pipeline once, and serialize it to YAML so staging and production run identical graphs. Retrieval quality is measured against a labeled evaluation set before any change ships, and pipelines are served either through Hayhooks or embedded in an ASGI service behind our own auth and observability. Document stores are chosen per workload, typically OpenSearch for hybrid search or pgvector when the data already lives in Postgres.
A RAG Pipeline End to End
Interactive Flow DiagramFile converters extract text and DocumentSplitter chunks it into passages with controlled overlap.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Ingest and preprocess | File converters extract text and DocumentSplitter chunks it into passages with controlled overlap. | Chunk size tuned to embedding model context |
| 2 | 2. Embed and index | A SentenceTransformers document embedder writes vectors into the document store via a DocumentWriter. | Batched embedding for throughput |
| 3 | 3. Retrieve | BM25 and embedding retrievers run in parallel and a DocumentJoiner merges the candidate sets. | top_k tuned for recall |
| 4 | 4. Build prompt | PromptBuilder renders retrieved passages and the user question into a grounded prompt. | Context trimmed to token budget |
| 5 | 5. Generate | The generator returns replies plus metadata, ready for citation extraction and response validation. | Groundedness checked before return |
# pip install haystack-ai==2.9.0
from haystack import Pipeline
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator
document_store = InMemoryDocumentStore()
template = """
Given the context, answer the question.
Context:
{% for doc in documents %}{{ doc.content }}
{% endfor %}
Question: {{ question }}
Answer:
"""
pipe = Pipeline()
pipe.add_component('text_embedder', SentenceTransformersTextEmbedder(model='sentence-transformers/all-MiniLM-L6-v2'))
pipe.add_component('retriever', InMemoryEmbeddingRetriever(document_store=document_store, top_k=5))
pipe.add_component('prompt_builder', PromptBuilder(template=template))
pipe.add_component('generator', OpenAIGenerator(model='gpt-4o-mini'))
pipe.connect('text_embedder.embedding', 'retriever.query_embedding')
pipe.connect('retriever.documents', 'prompt_builder.documents')
pipe.connect('prompt_builder.prompt', 'generator.prompt')
query = 'What is Haystack?'
result = pipe.run({'text_embedder': {'text': query}, 'prompt_builder': {'question': query}})
print(result['generator']['replies'][0])Services Engineered with Haystack
We design, build, and operate Haystack based RAG and agent systems end to end.
Haystack vs LangChain vs LlamaIndex
How Haystack stacks up against the two most common Python RAG and agent frameworks.
Framework Comparison Matrix
Benchmark Matrix| Evaluation Metric | Haystack | LangChain | LlamaIndex |
|---|---|---|---|
| Pipeline structure and inspectability | Explicit typed DAG, serializes to YAML Winner | Chains and LCEL, more implicit | Query engines and index abstractions |
| Integration and ecosystem breadth | Solid, curated integrations | Very broad third party surface Winner | Broad data connector library |
| Data ingestion and indexing for RAG | Strong retrievers and stores | Loaders and splitters | Purpose built indexing and connectors Winner |
| Deployment and serving | YAML serialization plus Hayhooks REST Winner | LangServe and custom services | Custom services or framework integration |
Text alternative for screen readers & search engines
- Pipeline structure and inspectability: Haystack: Explicit typed DAG, serializes to YAML vs LangChain: Chains and LCEL, more implicit vs LlamaIndex: Query engines and index abstractions (Winning option: Haystack).
- Integration and ecosystem breadth: Haystack: Solid, curated integrations vs LangChain: Very broad third party surface vs LlamaIndex: Broad data connector library (Winning option: LangChain).
- Data ingestion and indexing for RAG: Haystack: Strong retrievers and stores vs LangChain: Loaders and splitters vs LlamaIndex: Purpose built indexing and connectors (Winning option: LlamaIndex).
- Deployment and serving: Haystack: YAML serialization plus Hayhooks REST vs LangChain: LangServe and custom services vs LlamaIndex: Custom services or framework integration (Winning option: Haystack).
Haystack in a Reference Architecture
We built a document intelligence workflow for a fintech client on a Haystack retrieval pipeline, combining hybrid search over an OpenSearch document store with an LLM generator for grounded answers. The explicit pipeline structure let us evaluate retrieval quality independently and swap embedding models without disturbing the rest of the flow. Serialized pipeline definitions kept staging and production behavior consistent through iteration.
Read Reference Architecture →Frequently Asked Questions
What is Haystack used for?↓
Haystack is used to build production retrieval augmented generation systems, document search, question answering, and agentic LLM applications. It provides modular components that you wire into pipelines, so retrieval, prompting, and generation stay decoupled. It is maintained by deepset and released under the Apache 2.0 license.
Is Haystack free and open source?↓
Yes. The core framework is open source under the Apache 2.0 license and installable from PyPI as haystack-ai. deepset also offers a commercial platform, deepset Cloud, and a visual builder called deepset Studio for teams that want managed hosting and collaboration.
What is the difference between Haystack 1.x and 2.x?↓
Haystack 2.x is a ground up rewrite around a Component and Pipeline model with explicit typed connections between outputs and inputs. The legacy 1.x line shipped as farm-haystack and used Node based pipelines. New projects should use haystack-ai, since 2.x is the actively developed line.
How does Haystack compare to LangChain?↓
Haystack emphasizes explicit pipeline graphs with typed connections and strong retrieval primitives, which makes data flow easy to inspect and serialize to YAML. LangChain offers a broader integration surface and agent tooling. Teams often pick Haystack when RAG correctness and deployable pipeline structure matter more than breadth of third party wrappers.
Which document stores does Haystack support?↓
Haystack ships an InMemoryDocumentStore and provides integrations for OpenSearch, Elasticsearch, Qdrant, Weaviate, Pinecone, Chroma, and pgvector among others. Each store implements a common interface, so you can swap backends without rewriting the retrieval logic in your pipeline.
How do you deploy a Haystack pipeline as an API?↓
Haystack pipelines serialize to YAML and can be served over HTTP with Hayhooks, which wraps a pipeline as a REST endpoint. You can also embed the Pipeline object directly in a FastAPI or ASGI service. Both approaches keep the same pipeline definition across development and production.
Does Haystack support hybrid search?↓
Yes. You can combine a keyword BM25 retriever with an embedding retriever and merge results using a DocumentJoiner, then optionally rerank with a cross encoder ranker. Hybrid retrieval typically improves recall on queries where lexical overlap and semantic similarity each catch different relevant documents.
What Python version does Haystack require?↓
Haystack 2.x targets modern Python, generally 3.9 and newer at the time of writing, and is distributed on PyPI as haystack-ai. Integrations for individual document stores and model providers are published as separate packages, so you install only the backends your pipeline actually uses.