Skip to primary content
Agent Framework Deep Dive

OpenAI Agents SDK: Handoffs, Guardrails, and Tracing in Production

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

OpenAI Agents SDK is a lightweight Python and TypeScript framework for building multi-agent applications. It provides a small set of primitives, agents, handoffs, guardrails, sessions, and built-in tracing, running a provider-agnostic agent loop. It evolved from the experimental Swarm project and works with OpenAI models plus many others through LiteLLM.

LicenseMIT
LanguagePython and TypeScript
First ReleaseMarch 2025
Installpip install openai-agents
Problem & Purpose

What the OpenAI Agents SDK Solves in Production

Teams building agents frequently end up hand-rolling the same plumbing, a loop that calls the model, parses tool calls, appends results, and decides when to stop. Add multiple specialist agents and you also need routing, state transfer, input validation, and observability into what the model actually did. Most heavyweight frameworks answer this with deep abstractions that are hard to debug when a run misbehaves. The OpenAI Agents SDK takes the opposite stance, offering a small number of composable primitives so the control flow stays legible. That legibility is what makes it maintainable once agents are handling real customer traffic.

Anatomy of an OpenAI Agents SDK Application

Anatomy Explainer

Core Primitive Component Parts:

1. Agent → View Definition
2. Runner → View Definition
3. Handoffs → View Definition
4. Guardrails → View Definition
5. Tracing → View Definition
PART 1

Agent

A configured LLM with instructions, tools, and optional handoffs and guardrails.

Technical Implementation:

Constructed with name, instructions, model, tools, handoffs, and output_type. The output_type field can be a Pydantic model to force structured, validated final output.

The five primitives that compose into a running multi-agent workflow.
Text alternative for screen readers & search engines
  • Part 1: Agent - A configured LLM with instructions, tools, and optional handoffs and guardrails. [Tech: Constructed with name, instructions, model, tools, handoffs, and output_type. The output_type field can be a Pydantic model to force structured, validated final output.]
  • Part 2: Runner - Executes the agent loop until a final output or a stop condition is reached. [Tech: Runner.run is async and Runner.run_sync is blocking. Each iteration sends messages, executes any tool calls, and repeats, bounded by max_turns to prevent runaway loops.]
  • Part 3: Handoffs - Delegation from one agent to another, surfaced to the model as tools. [Tech: Created with the handoff helper or by listing agents in the handoffs argument. On invocation, control and conversation history transfer to the target agent, enabling triage-and-specialist patterns.]
  • Part 4: Guardrails - Parallel validation functions that can halt a run via a tripwire. [Tech: Declared with input_guardrail and output_guardrail. Each returns a GuardrailFunctionOutput; setting tripwire_triggered raises an exception, letting you enforce policy before or after the agent responds.]
  • Part 5: Tracing - Built-in spans covering every step of a run for debugging and evaluation. [Tech: Enabled by default, grouped under a workflow name via RunConfig. Spans capture generations, tool calls, handoffs, and guardrails, and can be exported to third-party processors alongside the OpenAI Traces dashboard.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Minimal abstraction surface: The primitive count is small, so the control flow reads like ordinary Python and is straightforward to step through in a debugger.
  • Handoffs as first-class routing: Multi-agent triage falls out naturally because delegation is modeled as a tool call rather than bespoke routing logic you maintain yourself.
  • Observability out of the box: Tracing is on by default and exportable, giving span-level visibility into generations, tool calls, and guardrail decisions without extra wiring.
  • Provider flexibility: Optimized for OpenAI models but usable with many providers through LiteLLM, so you are not locked to a single vendor at the framework level.
Specific Production Limits (Real Constraints)
  • No built-in durable execution: There is no native workflow persistence or replay, so long-running or crash-recoverable agents need an external layer such as Temporal or your own checkpointing.
  • Graph control is implicit: Complex cyclic or branching state machines are less explicit than in graph-first frameworks, which can make intricate conditional flows harder to reason about.
  • OpenAI-centric defaults: Some features and ergonomics lean toward the Responses API and OpenAI models, and non-OpenAI setups can surface capability gaps around structured output or tooling.
  • Young ecosystem: The SDK is relatively new, so patterns, integrations, and community references are still maturing compared with longer-established frameworks.
Production Implementation

How We Deploy the OpenAI Agents SDK in Production

Our team treats the SDK as the orchestration core and keeps everything around it explicit and testable. We pin the SDK and model versions, wrap tools with typed Pydantic schemas, and enforce policy with input and output guardrails before any output reaches a user. Every run is grouped under a workflow name so traces are queryable, and we back stateful conversations with a session store. For anything long-running or transactional we add a durability layer outside the SDK rather than relying on in-process state.

OpenAI Agents SDK Production Pipeline

Interactive Flow Diagram
OpenAI Agents SDK Production Pipeline From request intake through traced, guarded execution. 1. Intake and Session Load context 2. Input Guardrails Validate request 3. Triage and Handoff Route work 4. Tool Execution Act 5. Output Guardrails and Trace Verify and record
Stage 1: 1. Intake and Session SQLite or custom session backend

Resolve or create a session so prior turns and user context are attached to the run.

From request intake through traced, guarded execution.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Intake and Session Resolve or create a session so prior turns and user context are attached to the run. SQLite or custom session backend
2 2. Input Guardrails Run PII and policy checks in parallel with the agent; a tripwire halts unsafe requests early. Fail-fast on tripwire
3 3. Triage and Handoff A triage agent inspects intent and hands off to the correct specialist agent as a tool call. Bounded by max_turns
4 4. Tool Execution Specialist agents call typed function tools and MCP servers to fetch data and take actions. Pydantic-validated I/O
5 5. Output Guardrails and Trace Output guardrails check the final answer, then spans are exported for evaluation and audit. Traces to dashboard plus processor
Production Configuration (Version Pinned):
# pip install "openai-agents==0.2.3"
from agents import (
  Agent, Runner, function_tool, handoff,
  input_guardrail, GuardrailFunctionOutput,
  RunConfig, SQLiteSession,
)
from pydantic import BaseModel

@function_tool
def lookup_invoice(invoice_id: str) -> str:
  return fetch_invoice(invoice_id)

class ComplianceCheck(BaseModel):
  is_allowed: bool
  reason: str

compliance_agent = Agent(
  name='Compliance',
  instructions='Flag requests that expose sensitive data.',
  output_type=ComplianceCheck,
  model='gpt-4.1-mini',
)

@input_guardrail
async def pii_guardrail(ctx, agent, user_input):
  result = await Runner.run(compliance_agent, user_input, context=ctx.context)
  check = result.final_output_as(ComplianceCheck)
  return GuardrailFunctionOutput(output_info=check, tripwire_triggered=not check.is_allowed)

billing_agent = Agent(
  name='Billing',
  instructions='Resolve billing questions using the invoice tool.',
  tools=[lookup_invoice],
  model='gpt-4.1',
)

triage_agent = Agent(
  name='Triage',
  instructions='Route the user to the correct specialist.',
  handoffs=[handoff(billing_agent)],
  input_guardrails=[pii_guardrail],
  model='gpt-4.1',
)

session = SQLiteSession('customer-42', 'conversations.db')
result = Runner.run_sync(
  triage_agent,
  'Why was I charged twice this month?',
  session=session,
  run_config=RunConfig(workflow_name='support'),
)
print(result.final_output)
Delivering Commercial Impact

Services Engineered with OpenAI Agents SDK

We design, build, and operate agent systems on the OpenAI Agents SDK for teams shipping to production.

Alternatives Evaluation

OpenAI Agents SDK vs Alternative Frameworks

How the SDK compares with two established agent orchestration frameworks on the tradeoffs that matter in production.

Agent Framework Comparison

Benchmark Matrix
Evaluation Metric OpenAI Agents SDK LangGraph CrewAI
Minimal abstraction and boilerplate
Small primitive set, readable loop Winner
Explicit graph wiring required
Role and task scaffolding
Explicit graph and cyclic control
Implicit via handoffs
First-class stateful graphs Winner
Sequential and hierarchical flows
Built-in tracing and observability
On by default, exportable Winner
Strong via LangSmith
Basic, add-on tooling
Role-based crew orchestration templates
Compose agents manually
Node-based, not role-first
Purpose-built crews and roles Winner
Illustrative relative suitability, not measured benchmarks.
Text alternative for screen readers & search engines
  • Minimal abstraction and boilerplate: OpenAI Agents SDK: Small primitive set, readable loop vs LangGraph: Explicit graph wiring required vs CrewAI: Role and task scaffolding (Winning option: OpenAI Agents SDK).
  • Explicit graph and cyclic control: OpenAI Agents SDK: Implicit via handoffs vs LangGraph: First-class stateful graphs vs CrewAI: Sequential and hierarchical flows (Winning option: LangGraph).
  • Built-in tracing and observability: OpenAI Agents SDK: On by default, exportable vs LangGraph: Strong via LangSmith vs CrewAI: Basic, add-on tooling (Winning option: OpenAI Agents SDK).
  • Role-based crew orchestration templates: OpenAI Agents SDK: Compose agents manually vs LangGraph: Node-based, not role-first vs CrewAI: Purpose-built crews and roles (Winning option: CrewAI).
Production Proof

OpenAI Agents SDK in a Reference Architecture

Fintech Document Automation

We used an OpenAI Agents SDK triage-and-specialist design to classify and extract fields from financial documents, with a triage agent handing off to focused extraction agents. Input guardrails screened for sensitive data before processing, and built-in tracing gave the team span-level visibility to debug misroutes. The small primitive set kept the pipeline auditable for compliance review.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is the OpenAI Agents SDK?↓

It is an open source framework from OpenAI for building agentic applications with a minimal set of primitives. The core building blocks are agents, tools, handoffs, guardrails, sessions, and tracing. It is a production-focused successor to the experimental Swarm project.

Is the OpenAI Agents SDK free and open source?↓

Yes. The SDK is released under the MIT license and is free to use. You still pay for underlying model inference through whichever provider you configure, whether that is OpenAI or another vendor via LiteLLM.

How do handoffs work in the OpenAI Agents SDK?↓

A handoff is exposed to the model as a tool call. When an agent decides to hand off, control and the conversation state transfer to the target agent, which continues the loop. This lets a triage agent route work to specialist agents without custom routing code.

What are guardrails in the OpenAI Agents SDK?↓

Guardrails are validation functions that run alongside an agent to check inputs or outputs. Input guardrails run in parallel with the agent, and if a guardrail sets its tripwire flag the run halts with an exception. They are commonly used for PII checks, topic filtering, and policy enforcement.

Does the OpenAI Agents SDK support non-OpenAI models?↓

Yes. Although it is optimized for OpenAI models and the Responses API, it is provider-agnostic. Through the LiteLLM integration you can point agents at Anthropic, Google, or other providers using standard model identifier strings.

How does tracing work in the OpenAI Agents SDK?↓

Tracing is built in and enabled by default. Each run records spans for agent steps, tool calls, handoffs, and guardrails, viewable in the OpenAI Traces dashboard. You can also route spans to external processors such as Logfire, AgentOps, or Braintrust.

What is the difference between the OpenAI Agents SDK and Swarm?↓

Swarm was an experimental, educational project that demonstrated agents and handoffs. The Agents SDK is its production-ready evolution, adding guardrails, sessions, integrated tracing, and typed tooling while keeping the same lightweight design philosophy.

How does the SDK handle memory and conversation state?↓

Sessions provide automatic conversation history across runs. Built-in options include an in-memory store and SQLiteSession backed by a database file, and you can implement a custom session backend for other stores. Without a session, each run starts without prior turns.