Skip to primary content
RAG Framework Deep Dive

Haystack: Building Composable RAG and LLM Pipelines in Production

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

Haystack is an open source Python framework from deepset for building retrieval augmented generation and LLM applications. It connects modular components such as retrievers, embedders, rankers, and generators into explicit pipeline graphs. Version 2.x runs on typed input output connections and supports document stores like OpenSearch, Qdrant, and pgvector.

LicenseApache 2.0
Maintainerdeepset
Packagehaystack-ai (2.x)
LanguagePython
Problem & Purpose

What Haystack Solves in Production

RAG systems fail quietly when retrieval, prompting, and generation are tangled into one script that nobody can inspect or evaluate. Swapping a vector store or an LLM provider then means rewriting glue code and re-testing the whole flow. Prompt logic drifts, retrieval quality is never measured, and the same pipeline behaves differently in a notebook than in the service. Haystack addresses this by making each stage an explicit typed component and the whole flow an inspectable pipeline graph. That structure is what lets a team reason about, evaluate, and redeploy a RAG system without guesswork.

Anatomy of a Haystack Pipeline

Anatomy Explainer

Core Component Component Parts:

1. Component → View Definition
2. Pipeline → View Definition
3. Document Store → View Definition
4. Retriever → View Definition
5. PromptBuilder and Generator → View Definition
PART 1

Component

The fundamental unit of work, a Python class decorated with @component that exposes a typed run method.

Technical Implementation:

Inputs and outputs are declared with types so the pipeline can validate connections. Custom components subclass nothing special, they just implement run and declare output types.

The core building blocks that compose into a retrieval and generation graph.
Text alternative for screen readers & search engines
  • Part 1: Component - The fundamental unit of work, a Python class decorated with @component that exposes a typed run method. [Tech: Inputs and outputs are declared with types so the pipeline can validate connections. Custom components subclass nothing special, they just implement run and declare output types.]
  • Part 2: Pipeline - A directed graph that connects component outputs to inputs and executes them in dependency order. [Tech: Connections are made with pipeline.connect(sender.output, receiver.input). The graph serializes to and from YAML for reproducible deployment and inspection.]
  • Part 3: Document Store - The storage abstraction over vector and keyword backends holding your indexed documents. [Tech: InMemoryDocumentStore ships in core, while OpenSearch, Qdrant, Weaviate, pgvector, and others come as separate integration packages implementing a common interface.]
  • Part 4: Retriever - Fetches candidate documents for a query using lexical, embedding, or hybrid strategies. [Tech: BM25 retrievers score by keyword overlap, embedding retrievers use a text embedder and vector similarity, and a DocumentJoiner merges both for hybrid retrieval.]
  • Part 5: PromptBuilder and Generator - Renders a Jinja2 prompt from retrieved documents and invokes the language model. [Tech: PromptBuilder templates iterate over documents, and generators such as OpenAIGenerator or HuggingFaceLocalGenerator return replies and metadata for downstream components.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Explicit pipeline graphs: Typed connections between components make data flow inspectable and catch mismatched wiring before runtime rather than deep inside a call stack.
  • Backend portability: A common document store interface lets you move from InMemory to OpenSearch, Qdrant, or pgvector without changing retrieval logic.
  • Serializable and deployable: Pipelines serialize to YAML and serve over HTTP via Hayhooks, so the same definition runs in development and production.
  • Strong retrieval primitives: First class BM25, embedding, and hybrid retrievers with rerankers give you the building blocks RAG quality actually depends on.
Specific Production Limits (Real Constraints)
  • Python only: Haystack is a Python framework with no official runtime for other languages, so polyglot stacks must call it over a service boundary.
  • Integration packaging overhead: Document store and model provider integrations ship as separate packages that version independently, so dependency pinning needs care.
  • Smaller ecosystem than LangChain: Fewer third party wrappers and community tutorials exist compared to LangChain, so some niche integrations require writing custom components.
  • Agent tooling is newer: Multi step tool calling and agent patterns are supported but less mature than the retrieval side, and evolve across minor releases.
Production Implementation

How We Deploy Haystack in Production

Our team treats a Haystack pipeline as a versioned artifact, not a script. We pin haystack-ai and each integration package, define the pipeline once, and serialize it to YAML so staging and production run identical graphs. Retrieval quality is measured against a labeled evaluation set before any change ships, and pipelines are served either through Hayhooks or embedded in an ASGI service behind our own auth and observability. Document stores are chosen per workload, typically OpenSearch for hybrid search or pgvector when the data already lives in Postgres.

A RAG Pipeline End to End

Interactive Flow Diagram
A RAG Pipeline End to End From raw documents to a grounded, cited answer. 1. Ingest and preprocess Converters and splitters 2. Embed and index Document embedder 3. Retrieve Hybrid retrieval 4. Build prompt Jinja2 template 5. Generate LLM generator
Stage 1: 1. Ingest and preprocess Chunk size tuned to embedding model context

File converters extract text and DocumentSplitter chunks it into passages with controlled overlap.

From raw documents to a grounded, cited answer.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Ingest and preprocess File converters extract text and DocumentSplitter chunks it into passages with controlled overlap. Chunk size tuned to embedding model context
2 2. Embed and index A SentenceTransformers document embedder writes vectors into the document store via a DocumentWriter. Batched embedding for throughput
3 3. Retrieve BM25 and embedding retrievers run in parallel and a DocumentJoiner merges the candidate sets. top_k tuned for recall
4 4. Build prompt PromptBuilder renders retrieved passages and the user question into a grounded prompt. Context trimmed to token budget
5 5. Generate The generator returns replies plus metadata, ready for citation extraction and response validation. Groundedness checked before return
Production RAG Pipeline (Version Pinned):
# pip install haystack-ai==2.9.0
from haystack import Pipeline
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.components.retrievers.in_memory import InMemoryEmbeddingRetriever
from haystack.components.embedders import SentenceTransformersTextEmbedder
from haystack.components.builders import PromptBuilder
from haystack.components.generators import OpenAIGenerator

document_store = InMemoryDocumentStore()

template = """
Given the context, answer the question.
Context:
{% for doc in documents %}{{ doc.content }}
{% endfor %}
Question: {{ question }}
Answer:
"""

pipe = Pipeline()
pipe.add_component('text_embedder', SentenceTransformersTextEmbedder(model='sentence-transformers/all-MiniLM-L6-v2'))
pipe.add_component('retriever', InMemoryEmbeddingRetriever(document_store=document_store, top_k=5))
pipe.add_component('prompt_builder', PromptBuilder(template=template))
pipe.add_component('generator', OpenAIGenerator(model='gpt-4o-mini'))

pipe.connect('text_embedder.embedding', 'retriever.query_embedding')
pipe.connect('retriever.documents', 'prompt_builder.documents')
pipe.connect('prompt_builder.prompt', 'generator.prompt')

query = 'What is Haystack?'
result = pipe.run({'text_embedder': {'text': query}, 'prompt_builder': {'question': query}})
print(result['generator']['replies'][0])
Delivering Commercial Impact

Services Engineered with Haystack

We design, build, and operate Haystack based RAG and agent systems end to end.

Alternatives Evaluation

Haystack vs LangChain vs LlamaIndex

How Haystack stacks up against the two most common Python RAG and agent frameworks.

Framework Comparison Matrix

Benchmark Matrix
Evaluation Metric Haystack LangChain LlamaIndex
Pipeline structure and inspectability
Explicit typed DAG, serializes to YAML Winner
Chains and LCEL, more implicit
Query engines and index abstractions
Integration and ecosystem breadth
Solid, curated integrations
Very broad third party surface Winner
Broad data connector library
Data ingestion and indexing for RAG
Strong retrievers and stores
Loaders and splitters
Purpose built indexing and connectors Winner
Deployment and serving
YAML serialization plus Hayhooks REST Winner
LangServe and custom services
Custom services or framework integration
Illustrative relative suitability by dimension, not measured benchmarks.
Text alternative for screen readers & search engines
  • Pipeline structure and inspectability: Haystack: Explicit typed DAG, serializes to YAML vs LangChain: Chains and LCEL, more implicit vs LlamaIndex: Query engines and index abstractions (Winning option: Haystack).
  • Integration and ecosystem breadth: Haystack: Solid, curated integrations vs LangChain: Very broad third party surface vs LlamaIndex: Broad data connector library (Winning option: LangChain).
  • Data ingestion and indexing for RAG: Haystack: Strong retrievers and stores vs LangChain: Loaders and splitters vs LlamaIndex: Purpose built indexing and connectors (Winning option: LlamaIndex).
  • Deployment and serving: Haystack: YAML serialization plus Hayhooks REST vs LangChain: LangServe and custom services vs LlamaIndex: Custom services or framework integration (Winning option: Haystack).
Production Proof

Haystack in a Reference Architecture

Fintech Document Automation

We built a document intelligence workflow for a fintech client on a Haystack retrieval pipeline, combining hybrid search over an OpenSearch document store with an LLM generator for grounded answers. The explicit pipeline structure let us evaluate retrieval quality independently and swap embedding models without disturbing the rest of the flow. Serialized pipeline definitions kept staging and production behavior consistent through iteration.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is Haystack used for?↓

Haystack is used to build production retrieval augmented generation systems, document search, question answering, and agentic LLM applications. It provides modular components that you wire into pipelines, so retrieval, prompting, and generation stay decoupled. It is maintained by deepset and released under the Apache 2.0 license.

Is Haystack free and open source?↓

Yes. The core framework is open source under the Apache 2.0 license and installable from PyPI as haystack-ai. deepset also offers a commercial platform, deepset Cloud, and a visual builder called deepset Studio for teams that want managed hosting and collaboration.

What is the difference between Haystack 1.x and 2.x?↓

Haystack 2.x is a ground up rewrite around a Component and Pipeline model with explicit typed connections between outputs and inputs. The legacy 1.x line shipped as farm-haystack and used Node based pipelines. New projects should use haystack-ai, since 2.x is the actively developed line.

How does Haystack compare to LangChain?↓

Haystack emphasizes explicit pipeline graphs with typed connections and strong retrieval primitives, which makes data flow easy to inspect and serialize to YAML. LangChain offers a broader integration surface and agent tooling. Teams often pick Haystack when RAG correctness and deployable pipeline structure matter more than breadth of third party wrappers.

Which document stores does Haystack support?↓

Haystack ships an InMemoryDocumentStore and provides integrations for OpenSearch, Elasticsearch, Qdrant, Weaviate, Pinecone, Chroma, and pgvector among others. Each store implements a common interface, so you can swap backends without rewriting the retrieval logic in your pipeline.

How do you deploy a Haystack pipeline as an API?↓

Haystack pipelines serialize to YAML and can be served over HTTP with Hayhooks, which wraps a pipeline as a REST endpoint. You can also embed the Pipeline object directly in a FastAPI or ASGI service. Both approaches keep the same pipeline definition across development and production.

Does Haystack support hybrid search?↓

Yes. You can combine a keyword BM25 retriever with an embedding retriever and merge results using a DocumentJoiner, then optionally rerank with a cross encoder ranker. Hybrid retrieval typically improves recall on queries where lexical overlap and semantic similarity each catch different relevant documents.

What Python version does Haystack require?↓

Haystack 2.x targets modern Python, generally 3.9 and newer at the time of writing, and is distributed on PyPI as haystack-ai. Integrations for individual document stores and model providers are published as separate packages, so you install only the backends your pipeline actually uses.