Skip to primary content
Agent Framework Deep Dive

smolagents: Building Code-Action Agents That Reason in Python

Reviewed by Umar Abbas • Founder & Principal AI Architect

Last reviewed: 14 August 2026

smolagents is a minimal open source agent library from Hugging Face where an agent chooses actions by writing Python code instead of emitting JSON tool calls. The CodeAgent generates code, runs it in a sandboxed executor, observes results, and iterates. It supports any LLM through pluggable model backends.

LanguagePython 3.10+
LicenseApache License 2.0
MaintainerHugging Face
Default agentCodeAgent
Problem & Purpose

What smolagents Solves in Production

Most agent stacks force the model to select tools by emitting structured JSON, which becomes verbose and step heavy the moment a task needs to combine several calls, filter results, or branch on an intermediate value. Each composition demands another model round trip, inflating latency and cost. smolagents attacks this by making the action a snippet of Python, so one generation can call multiple tools, loop, and post process in a single step. The cost is that generated code is untrusted and must be sandboxed. The framework answers that with sandboxed executors and a restricted local interpreter, keeping the mental model small while acknowledging the real security boundary.

Inside a smolagents CodeAgent

Anatomy Explainer

Core Component Component Parts:

1. CodeAgent → View Definition
2. Python Executor → View Definition
3. Model Backend → View Definition
4. Tool Abstraction → View Definition
5. AgentMemory → View Definition
PART 1

CodeAgent

The reasoning loop that prompts the model to emit Python, runs it, and feeds observations back until a final answer.

Technical Implementation:

Implements the ReAct style think, act, observe cycle with actions expressed as code blocks. Bounded by max_steps and carries the running AgentMemory.

The core objects that turn a prompt into executed, observed Python.
Text alternative for screen readers & search engines
  • Part 1: CodeAgent - The reasoning loop that prompts the model to emit Python, runs it, and feeds observations back until a final answer. [Tech: Implements the ReAct style think, act, observe cycle with actions expressed as code blocks. Bounded by max_steps and carries the running AgentMemory.]
  • Part 2: Python Executor - The component that actually runs generated code, either locally with restrictions or in a remote sandbox. [Tech: LocalPythonExecutor enforces an import allowlist and blocks unsafe builtins. executor_type can be set to e2b or docker for isolated remote execution.]
  • Part 3: Model Backend - A pluggable interface that abstracts the underlying LLM provider from the agent loop. [Tech: InferenceClientModel, LiteLLMModel, TransformersModel, and cloud backends share a common call signature, so providers swap without touching agent logic.]
  • Part 4: Tool Abstraction - Typed callables the model can invoke by name from inside generated code. [Tech: Defined via the tool decorator or the Tool class. Signature and docstring introspection builds the schema. Tools load from the Hub or other frameworks.]
  • Part 5: AgentMemory - The structured log of system prompt, task, and every step that grounds each new generation. [Tech: Stores action steps, code, tool outputs, and errors as replayable records, enabling inspection, logging, and OpenTelemetry style tracing.]
Production Evaluation

Architectural Strengths & Specific Production Limits

Core Strengths
  • Minimal core: The agent logic sits in a small, readable codebase, so teams can audit the full loop instead of trusting an opaque orchestration layer.
  • Code as action: Expressing actions in Python lets one step compose multiple tools, loops, and conditionals, often cutting the number of model round trips.
  • Model and tool agnostic: Pluggable backends reach Hugging Face Inference, LiteLLM providers, and local transformers, and tools can be pulled directly from the Hub.
  • First class sandboxing: Built in E2B and Docker executors treat generated code as untrusted by default, which is the correct posture for running model output.
Specific Production Limits (Real Constraints)
  • Execution risk: Because agents run generated Python, a misconfigured or absent sandbox is a direct remote code execution exposure on the host.
  • Weak explicit state control: There is no first class graph with named nodes and checkpoints, so deterministic, resumable workflows are harder to model than in LangGraph.
  • Smaller ecosystem: Prebuilt integrations, connectors, and community tooling are thinner than the larger LangChain and LlamaIndex ecosystems.
  • Model dependent reliability: Code action quality varies with the model, and weaker code models produce more failed snippets and retries, raising cost and latency.
Production Implementation

How We Deploy smolagents in Production

Our team reaches for smolagents when a task is genuinely tool composing rather than a single call, and when we want a loop we can read end to end. We standardize on a remote sandbox executor, pin the library and model versions, and constrain the import allowlist to the minimum a task needs. Every run emits structured traces through OpenTelemetry so we can inspect the exact code the model wrote at each step. We keep tools small and strongly typed, and we gate any action with side effects behind a human in the loop review.

smolagents Production Pipeline

Interactive Flow Diagram
smolagents Production Pipeline From typed tools to observed, sandboxed execution. 1. Define Tools Typed Python callables 2. Select Backend Model interface 3. Configure Sandbox Isolated executor 4. Run Loop Think, code, observe 5. Trace and Review Inspect and gate
Stage 1: 1. Define Tools Signature plus docstring schema

Wrap capabilities as decorated functions with clear type hints and docstrings so the model calls them accurately.

From typed tools to observed, sandboxed execution.
Text alternative for screen readers & search engines
Step Stage Name Function & Detail Metrics / SLA
1 1. Define Tools Wrap capabilities as decorated functions with clear type hints and docstrings so the model calls them accurately. Signature plus docstring schema
2 2. Select Backend Bind a model instance such as InferenceClientModel or LiteLLMModel, isolating provider choice from agent logic. One line provider swap
3 3. Configure Sandbox Set executor_type to e2b or docker and restrict additional_authorized_imports to the minimum required. Untrusted code boundary
4 4. Run Loop CodeAgent generates Python, executes it, and feeds outputs back, bounded by max_steps to prevent runaway loops. max_steps cap
5 5. Trace and Review Capture every step through OpenTelemetry and require human approval before side effecting actions reach real systems. Full step level traces
Production Configuration (Version Pinned):
# pip install smolagents==1.14.0
from smolagents import CodeAgent, InferenceClientModel, WebSearchTool, tool

# Provider stays isolated behind the model backend
model = InferenceClientModel(
  model_id='Qwen/Qwen2.5-Coder-32B-Instruct',
  provider='together',
)

@tool
def unit_price(total: float, quantity: int) -> float:
  '''Return the price per unit for a bulk order.'''
  return total / quantity

agent = CodeAgent(
  tools=[WebSearchTool(), unit_price],
  model=model,
  additional_authorized_imports=['statistics'],
  executor_type='e2b',   # run untrusted code in a remote sandbox
  max_steps=6,
  verbosity_level=1,
)

result = agent.run(
  'Find two supplier bulk prices online and pick the lower unit cost.'
)
print(result)
Delivering Commercial Impact

Services Engineered with smolagents

We design, harden, and ship code-action agents built on smolagents for enterprise workloads.

Alternatives Evaluation

smolagents vs Alternative Agent Frameworks

How smolagents compares with two widely used agent frameworks on the axes that matter in production.

Agent Framework Comparison

Benchmark Matrix
Evaluation Metric smolagents LangGraph AutoGen
Action expressiveness
Executable Python code actions Winner
JSON tool calls in graph nodes
JSON tool calls in conversations
Explicit state and control
Implicit loop, hierarchical only
Named nodes with checkpoints Winner
Conversation driven state
Multi agent conversation
Managed agents as tools
Graph orchestrated agents
Native group chat patterns Winner
Minimal footprint
Small readable core Winner
Moderate graph runtime
Broader conversation runtime
Illustrative relative suitability, not measured benchmarks.
Text alternative for screen readers & search engines
  • Action expressiveness: smolagents: Executable Python code actions vs LangGraph: JSON tool calls in graph nodes vs AutoGen: JSON tool calls in conversations (Winning option: smolagents).
  • Explicit state and control: smolagents: Implicit loop, hierarchical only vs LangGraph: Named nodes with checkpoints vs AutoGen: Conversation driven state (Winning option: LangGraph).
  • Multi agent conversation: smolagents: Managed agents as tools vs LangGraph: Graph orchestrated agents vs AutoGen: Native group chat patterns (Winning option: AutoGen).
  • Minimal footprint: smolagents: Small readable core vs LangGraph: Moderate graph runtime vs AutoGen: Broader conversation runtime (Winning option: smolagents).
Production Proof

smolagents in a Reference Architecture

Fintech Document Automation

For a fintech document automation engagement, we prototyped extraction and reconciliation logic as smolagents CodeAgent steps, letting a single generated snippet parse a document, call validation tools, and compute checks in one pass. Running inside an isolated sandbox with a tight import allowlist kept untrusted code contained while we iterated. The code action pattern reduced round trips on the multi step reconciliation tasks compared with a JSON tool calling baseline.

Read Reference Architecture →
Technical FAQ

Frequently Asked Questions

What is smolagents?↓

smolagents is a lightweight Python library from Hugging Face for building agents in about a thousand lines of core code. Its distinguishing idea is that agents express actions as executable Python code rather than as structured JSON tool calls. It ships with a CodeAgent for code actions and a ToolCallingAgent for traditional JSON tool calling.

How is smolagents different from LangChain or LangGraph?↓

LangChain and LangGraph focus on composable chains and explicit stateful graphs, with the agent emitting JSON to select tools. smolagents keeps the surface area small and makes code generation the primary action format, which lets a single step compose multiple tool calls, loops, and control flow. It trades fine grained graph control for brevity and expressiveness.

Why do smolagents write code instead of JSON tool calls?↓

Research such as the CodeAct work found that letting models write executable code can reduce the number of steps and improve success on multi tool tasks, because code naturally expresses composition, variables, and conditionals. A single generated snippet can chain several tools and process intermediate results without extra round trips. smolagents adopts this pattern as its default CodeAgent behavior.

Is smolagents safe to run in production?↓

Generated code is untrusted, so smolagents provides sandboxed execution paths through E2B and Docker in addition to a restricted local Python interpreter. The local executor limits imports to an allowlist and blocks dangerous builtins, but for real workloads a remote sandbox is strongly recommended. Treat every executor as a boundary and never run raw generated code on a privileged host.

Which models does smolagents support?↓

smolagents is model agnostic through pluggable backends. InferenceClientModel talks to Hugging Face Inference Providers, LiteLLMModel reaches OpenAI, Anthropic, and hundreds of others, TransformersModel runs local weights, and there are backends for Azure and Amazon Bedrock. You pass a model instance into the agent, so switching providers is a one line change.

What license does smolagents use?↓

smolagents is released under the Apache License 2.0, which permits commercial use, modification, and redistribution with attribution and a patent grant. That makes it straightforward to embed in commercial and enterprise systems. Always verify the license of the specific model and any tools you attach, since those carry their own terms.

Can smolagents run multi agent systems?↓

Yes. A managed agent can be wrapped and passed to a higher level agent as a callable tool, letting you compose a manager agent that delegates subtasks to specialized worker agents. This gives a simple hierarchical pattern without a separate orchestration graph. Complex branching topologies are still easier to express in graph first frameworks.

Does smolagents support tool calling without code execution?↓

Yes. If you prefer the classic pattern, ToolCallingAgent emits actions as JSON tool calls rather than Python, which suits environments where code execution is not acceptable. You keep the same tool definitions and model backends, only the action format changes. CodeAgent remains the recommended default for tasks that benefit from composition.

How do you define a tool in smolagents?↓

You decorate a typed Python function with the tool decorator or subclass the Tool class, and the framework introspects the signature and docstring to build the schema exposed to the model. Tools can also be loaded from the Hugging Face Hub or imported from other ecosystems. Clear type hints and docstrings directly improve how reliably the model calls each tool.

When should you not use smolagents?↓

It is a weaker fit when you need deterministic, auditable state machines with explicit checkpoints and rollback, where a graph first framework is clearer. It is also unsuitable if your security posture forbids executing model generated code and a sandbox is not available. For simple single prompt calls, a plain model client is lighter than any agent.