Skip to primary content
RAG Governance Deep Dive

TruLens for Enterprise AI: Architecture & Integration

Reviewed by Umar Abbas • Founder & Principal AI Architect

TruLens is an open-source evaluation and feedback governance library designed to assess LLM application quality and guard against hallucinations. By defining automated evaluation metrics called Feedback Functions (groundedness, context relevance, answer relevance), TruLens instruments RAG pipelines and agent systems to compute quantitative quality scores during offline testing and online monitoring.

Core MethodologyRAG Triad Metrics
Scoring EngineFeedback Functions
Storage EngineSQLite / PostgreSQL
LicenseMIT Open Source
Problem & Purpose

What TruLens Solves in RAG Quality & Hallucination Prevention

Deploying RAG pipelines without quantitative verification risks delivering hallucinated or ungrounded answers to users. TruLens solves this by establishing the RAG Triad framework—programmatically evaluating whether retrieved context is relevant, whether generated text is grounded strictly in context, and whether final responses answer user intent.

TruLens Evaluation & RAG Triad Architecture

Anatomy Explainer

TruLens Component Component Parts:

1. TruLlama / TruChain Instrumentor → View Definition
2. Feedback Function Engine → View Definition
3. RAG Triad Evaluator → View Definition
4. TruSession Database & Leaderboard → View Definition
5. TruLens Streamlit Dashboard → View Definition
PART 1

TruLlama / TruChain Instrumentor

Framework-native wrappers capturing intermediate retrieval chunks and final generation outputs.

Technical Implementation:

Instruments internal step execution graph without breaking application pipeline return types.

Architecture of TruLens showing TruLlama Wrapper, Feedback Function Evaluators, RAG Triad Calculator, and TruSession DB.
Text alternative for screen readers & search engines
  • Part 1: TruLlama / TruChain Instrumentor - Framework-native wrappers capturing intermediate retrieval chunks and final generation outputs. [Tech: Instruments internal step execution graph without breaking application pipeline return types.]
  • Part 2: Feedback Function Engine - Configurable evaluation modules executing LLM-as-a-judge or NLP similarity metrics. [Tech: Supports asynchronous worker execution to avoid slowing production response streams.]
  • Part 3: RAG Triad Evaluator - Calculates Context Relevance, Groundedness, and Answer Relevance normalized from 0.0 to 1.0. [Tech: Pinpoints exact failure modes (e.g., bad vector retrieval vs model hallucination).]
  • Part 4: TruSession Database & Leaderboard - Central registry storing evaluation runs, trace trees, and score leaderboards across iterations. [Tech: Enables side-by-side comparison of chunking strategy scores across experiment versions.]
  • Part 5: TruLens Streamlit Dashboard - Interactive UI exploring trace trees, hallucination highlights, and cost breakdown charts. [Tech: Launches locally via `tru.run_dashboard()` for rapid developer feedback.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Scientific RAG Triad Standard: Standardized, actionable metrics pinpointing exact failure modes.
  • Flexible Feedback Functions: Define custom evaluation logic using OpenAI, BERT, or local LLMs.
  • Framework Integration: First-class wrappers for LlamaIndex (TruLlama) and LangChain (TruChain).
  • Self-Hosted Data Control: Stores all evaluation data in local SQLite or PostgreSQL instances.
Specific Production Limits
  • Evaluation LLM API Costs: Running LLM-as-a-judge Feedback Functions on 100% of production traffic doubles model API calls.
  • Synchronous Latency Impact: Running Feedback Functions synchronously adds 500ms+ to query turnaround times.
  • Database Scaling: High-frequency production logging requires PostgreSQL tuning to handle rapid trace insertions.
Production Implementation

Production TruLens RAG Triad Evaluation Script

Python script configuring Feedback Functions (Groundedness, Context Relevance) and wrapping a LlamaIndex query engine.

TruLens RAG Triad Evaluation Flow

Interactive Flow Diagram
TruLens RAG Triad Evaluation Flow Pipeline: User Query -> TruLlama Wrapper -> Vector Retrieval -> Feedback Function Evaluator -> RAG Triad Score. 1. App Invocation tru_recorder.with_record() 2. Vector Search Context Retrieval 3. LLM Completion Model Answer 4. Feedback Engine Async Scorer Workers 5. Leaderboard Update TruSession Database
Stage 1: 1. App Invocation Span start

Triggers query execution and starts span recording.

Pipeline: User Query -> TruLlama Wrapper -> Vector Retrieval -> Feedback Function Evaluator -> RAG Triad Score.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. App Invocation Triggers query execution and starts span recording. Span start
2 2. Vector Search Retrieves top-k document chunks from vector index. Chunk extract
3 3. LLM Completion Generates final response from retrieved context. Token generation
4 4. Feedback Engine Executes Groundedness and Context Relevance Feedback Functions. Async eval
5 5. Leaderboard Update Records normalized RAG Triad scores (0.0 - 1.0) into database. Score log
Production TruLens RAG Triad Evaluation Script:
from trulens.core import TruSession, Feedback
from trulens.providers.openai import OpenAI as TruOpenAI
from trulens.apps.llama import TruLlama
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
import numpy as np

# Initialize TruLens session
session = TruSession()
session.reset_database()

# Set up OpenAI feedback provider for scoring
provider = TruOpenAI()

# Define RAG Triad Feedback Functions
f_groundedness = (
  Feedback(provider.groundedness_measure_with_cot_reasons)
  .on_input_to_context()
  .on_output()
)

f_context_relevance = (
  Feedback(provider.relevance_with_cot_reasons)
  .on_input()
  .on_input_to_context()
)

f_answer_relevance = (
  Feedback(provider.relevance_with_cot_reasons)
  .on_input()
  .on_output()
)

def build_evaluated_rag():
  documents = SimpleDirectoryReader("./data/legal_docs").load_data()
  index = VectorStoreIndex.from_documents(documents)
  query_engine = index.as_query_engine()

  # Wrap query engine with TruLlama recorder and feedback functions
  tru_query_engine = TruLlama(
      query_engine,
      app_id="legal_rag_v1",
      feedbacks=[f_groundedness, f_context_relevance, f_answer_relevance]
  )

  with tru_query_engine as recorder:
      response = query_engine.query("What is the termination clause penalty for breach?")
  
  print("Response:", str(response))
  print("Leaderboard Results:")
  print(session.get_leaderboard())

if __name__ == "__main__":
  build_evaluated_rag()
Performance & Benchmarks

TruLens Trade-Off & Benchmark Matrix

TruLens Trade-Off Matrix

Benchmark Matrix
Evaluation Metric TruLens Arize Phoenix LangSmith
RAG Triad Methodological Rigor
Formal Triad Spec (CoT Reasons) Winner
Ragas / Phoenix Eval
Custom Evaluator API
Feedback Function Customization
Flexible Decorator Selectors Winner
Eval Dataset Runners
Dataset Run Evaluator
3D Vector Embedding Space Visualizer
2D / Text Search
Native 3D UMAP Engine Winner
Embedding Search
Self-Hosted Data Control (Postgres)
Local SQLite / PostgreSQL Winner
In-Memory / Local
Commercial Cloud
Evaluating TruLens against Arize Phoenix and LangSmith across RAG Triad metrics, Feedback Functions, and local session governance.
Text alternative for screen readers & search engines
  • RAG Triad Methodological Rigor: TruLens: Formal Triad Spec (CoT Reasons) vs Arize Phoenix: Ragas / Phoenix Eval vs LangSmith: Custom Evaluator API (Winning option: TruLens).
  • Feedback Function Customization: TruLens: Flexible Decorator Selectors vs Arize Phoenix: Eval Dataset Runners vs LangSmith: Dataset Run Evaluator (Winning option: TruLens).
  • 3D Vector Embedding Space Visualizer: TruLens: 2D / Text Search vs Arize Phoenix: Native 3D UMAP Engine vs LangSmith: Embedding Search (Winning option: Arize Phoenix).
  • Self-Hosted Data Control (Postgres): TruLens: Local SQLite / PostgreSQL vs Arize Phoenix: In-Memory / Local vs LangSmith: Commercial Cloud (Winning option: TruLens).
Production Proof

TruLens Reference Architecture

Healthcare RAG Hallucination Prevention

Implemented TruLens RAG Triad scoring across a medical compliance QA agent. Reduced hallucination rates by 38% through automated Groundedness Feedback Function filtering across 2M queries.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is the RAG Triad evaluated by TruLens?↓

The RAG Triad consists of three core quality metrics: Context Relevance (retrieval quality), Groundedness (hallucination check), and Answer Relevance (query satisfaction).

How do Feedback Functions operate in TruLens?↓

Feedback Functions wrap model-based or programmatic scorers, evaluating prompt-response pairs asynchronously or synchronously during app execution.

Can TruLens run evaluation on local self-hosted open-weight LLMs?↓

Yes. TruLens Feedback Functions can use local HuggingFace or Ollama models to score groundedness without third-party API data egress.

How does TruLens store evaluation results and telemetry?↓

TruLens uses a local SQLite or PostgreSQL database backend (`TruSession`) to store detailed trace records, evaluation scores, and feedback logs.

Does TruLens support native LlamaIndex and LangChain instrumentation?↓

Yes. Wrappers like `TruLlama` and `TruChain` automatically instrument pipeline steps and compute RAG Triad scores.