smolagents: Building Code-Action Agents That Reason in Python
Reviewed by Umar Abbas • Founder & Principal AI Architect
Last reviewed: 14 August 2026
smolagents is a minimal open source agent library from Hugging Face where an agent chooses actions by writing Python code instead of emitting JSON tool calls. The CodeAgent generates code, runs it in a sandboxed executor, observes results, and iterates. It supports any LLM through pluggable model backends.
What smolagents Solves in Production
Most agent stacks force the model to select tools by emitting structured JSON, which becomes verbose and step heavy the moment a task needs to combine several calls, filter results, or branch on an intermediate value. Each composition demands another model round trip, inflating latency and cost. smolagents attacks this by making the action a snippet of Python, so one generation can call multiple tools, loop, and post process in a single step. The cost is that generated code is untrusted and must be sandboxed. The framework answers that with sandboxed executors and a restricted local interpreter, keeping the mental model small while acknowledging the real security boundary.
Inside a smolagents CodeAgent
Anatomy ExplainerCore Component Component Parts:
CodeAgent
The reasoning loop that prompts the model to emit Python, runs it, and feeds observations back until a final answer.
Implements the ReAct style think, act, observe cycle with actions expressed as code blocks. Bounded by max_steps and carries the running AgentMemory.
Text alternative for screen readers & search engines
- Part 1: CodeAgent - The reasoning loop that prompts the model to emit Python, runs it, and feeds observations back until a final answer. [Tech: Implements the ReAct style think, act, observe cycle with actions expressed as code blocks. Bounded by max_steps and carries the running AgentMemory.]
- Part 2: Python Executor - The component that actually runs generated code, either locally with restrictions or in a remote sandbox. [Tech: LocalPythonExecutor enforces an import allowlist and blocks unsafe builtins. executor_type can be set to e2b or docker for isolated remote execution.]
- Part 3: Model Backend - A pluggable interface that abstracts the underlying LLM provider from the agent loop. [Tech: InferenceClientModel, LiteLLMModel, TransformersModel, and cloud backends share a common call signature, so providers swap without touching agent logic.]
- Part 4: Tool Abstraction - Typed callables the model can invoke by name from inside generated code. [Tech: Defined via the tool decorator or the Tool class. Signature and docstring introspection builds the schema. Tools load from the Hub or other frameworks.]
- Part 5: AgentMemory - The structured log of system prompt, task, and every step that grounds each new generation. [Tech: Stores action steps, code, tool outputs, and errors as replayable records, enabling inspection, logging, and OpenTelemetry style tracing.]
Architectural Strengths & Specific Production Limits
- Minimal core: The agent logic sits in a small, readable codebase, so teams can audit the full loop instead of trusting an opaque orchestration layer.
- Code as action: Expressing actions in Python lets one step compose multiple tools, loops, and conditionals, often cutting the number of model round trips.
- Model and tool agnostic: Pluggable backends reach Hugging Face Inference, LiteLLM providers, and local transformers, and tools can be pulled directly from the Hub.
- First class sandboxing: Built in E2B and Docker executors treat generated code as untrusted by default, which is the correct posture for running model output.
- Execution risk: Because agents run generated Python, a misconfigured or absent sandbox is a direct remote code execution exposure on the host.
- Weak explicit state control: There is no first class graph with named nodes and checkpoints, so deterministic, resumable workflows are harder to model than in LangGraph.
- Smaller ecosystem: Prebuilt integrations, connectors, and community tooling are thinner than the larger LangChain and LlamaIndex ecosystems.
- Model dependent reliability: Code action quality varies with the model, and weaker code models produce more failed snippets and retries, raising cost and latency.
How We Deploy smolagents in Production
Our team reaches for smolagents when a task is genuinely tool composing rather than a single call, and when we want a loop we can read end to end. We standardize on a remote sandbox executor, pin the library and model versions, and constrain the import allowlist to the minimum a task needs. Every run emits structured traces through OpenTelemetry so we can inspect the exact code the model wrote at each step. We keep tools small and strongly typed, and we gate any action with side effects behind a human in the loop review.
smolagents Production Pipeline
Interactive Flow DiagramWrap capabilities as decorated functions with clear type hints and docstrings so the model calls them accurately.
Text alternative for screen readers & search engines
| Step | Stage Name | Function & Detail | Metrics / SLA |
|---|---|---|---|
| 1 | 1. Define Tools | Wrap capabilities as decorated functions with clear type hints and docstrings so the model calls them accurately. | Signature plus docstring schema |
| 2 | 2. Select Backend | Bind a model instance such as InferenceClientModel or LiteLLMModel, isolating provider choice from agent logic. | One line provider swap |
| 3 | 3. Configure Sandbox | Set executor_type to e2b or docker and restrict additional_authorized_imports to the minimum required. | Untrusted code boundary |
| 4 | 4. Run Loop | CodeAgent generates Python, executes it, and feeds outputs back, bounded by max_steps to prevent runaway loops. | max_steps cap |
| 5 | 5. Trace and Review | Capture every step through OpenTelemetry and require human approval before side effecting actions reach real systems. | Full step level traces |
# pip install smolagents==1.14.0 from smolagents import CodeAgent, InferenceClientModel, WebSearchTool, tool # Provider stays isolated behind the model backend model = InferenceClientModel( model_id='Qwen/Qwen2.5-Coder-32B-Instruct', provider='together', ) @tool def unit_price(total: float, quantity: int) -> float: '''Return the price per unit for a bulk order.''' return total / quantity agent = CodeAgent( tools=[WebSearchTool(), unit_price], model=model, additional_authorized_imports=['statistics'], executor_type='e2b', # run untrusted code in a remote sandbox max_steps=6, verbosity_level=1, ) result = agent.run( 'Find two supplier bulk prices online and pick the lower unit cost.' ) print(result)
Services Engineered with smolagents
We design, harden, and ship code-action agents built on smolagents for enterprise workloads.
smolagents vs Alternative Agent Frameworks
How smolagents compares with two widely used agent frameworks on the axes that matter in production.
Agent Framework Comparison
Benchmark Matrix| Evaluation Metric | smolagents | LangGraph | AutoGen |
|---|---|---|---|
| Action expressiveness | Executable Python code actions Winner | JSON tool calls in graph nodes | JSON tool calls in conversations |
| Explicit state and control | Implicit loop, hierarchical only | Named nodes with checkpoints Winner | Conversation driven state |
| Multi agent conversation | Managed agents as tools | Graph orchestrated agents | Native group chat patterns Winner |
| Minimal footprint | Small readable core Winner | Moderate graph runtime | Broader conversation runtime |
Text alternative for screen readers & search engines
- Action expressiveness: smolagents: Executable Python code actions vs LangGraph: JSON tool calls in graph nodes vs AutoGen: JSON tool calls in conversations (Winning option: smolagents).
- Explicit state and control: smolagents: Implicit loop, hierarchical only vs LangGraph: Named nodes with checkpoints vs AutoGen: Conversation driven state (Winning option: LangGraph).
- Multi agent conversation: smolagents: Managed agents as tools vs LangGraph: Graph orchestrated agents vs AutoGen: Native group chat patterns (Winning option: AutoGen).
- Minimal footprint: smolagents: Small readable core vs LangGraph: Moderate graph runtime vs AutoGen: Broader conversation runtime (Winning option: smolagents).
smolagents in a Reference Architecture
For a fintech document automation engagement, we prototyped extraction and reconciliation logic as smolagents CodeAgent steps, letting a single generated snippet parse a document, call validation tools, and compute checks in one pass. Running inside an isolated sandbox with a tight import allowlist kept untrusted code contained while we iterated. The code action pattern reduced round trips on the multi step reconciliation tasks compared with a JSON tool calling baseline.
Read Reference Architecture →Frequently Asked Questions
What is smolagents?↓
smolagents is a lightweight Python library from Hugging Face for building agents in about a thousand lines of core code. Its distinguishing idea is that agents express actions as executable Python code rather than as structured JSON tool calls. It ships with a CodeAgent for code actions and a ToolCallingAgent for traditional JSON tool calling.
How is smolagents different from LangChain or LangGraph?↓
LangChain and LangGraph focus on composable chains and explicit stateful graphs, with the agent emitting JSON to select tools. smolagents keeps the surface area small and makes code generation the primary action format, which lets a single step compose multiple tool calls, loops, and control flow. It trades fine grained graph control for brevity and expressiveness.
Why do smolagents write code instead of JSON tool calls?↓
Research such as the CodeAct work found that letting models write executable code can reduce the number of steps and improve success on multi tool tasks, because code naturally expresses composition, variables, and conditionals. A single generated snippet can chain several tools and process intermediate results without extra round trips. smolagents adopts this pattern as its default CodeAgent behavior.
Is smolagents safe to run in production?↓
Generated code is untrusted, so smolagents provides sandboxed execution paths through E2B and Docker in addition to a restricted local Python interpreter. The local executor limits imports to an allowlist and blocks dangerous builtins, but for real workloads a remote sandbox is strongly recommended. Treat every executor as a boundary and never run raw generated code on a privileged host.
Which models does smolagents support?↓
smolagents is model agnostic through pluggable backends. InferenceClientModel talks to Hugging Face Inference Providers, LiteLLMModel reaches OpenAI, Anthropic, and hundreds of others, TransformersModel runs local weights, and there are backends for Azure and Amazon Bedrock. You pass a model instance into the agent, so switching providers is a one line change.
What license does smolagents use?↓
smolagents is released under the Apache License 2.0, which permits commercial use, modification, and redistribution with attribution and a patent grant. That makes it straightforward to embed in commercial and enterprise systems. Always verify the license of the specific model and any tools you attach, since those carry their own terms.
Can smolagents run multi agent systems?↓
Yes. A managed agent can be wrapped and passed to a higher level agent as a callable tool, letting you compose a manager agent that delegates subtasks to specialized worker agents. This gives a simple hierarchical pattern without a separate orchestration graph. Complex branching topologies are still easier to express in graph first frameworks.
Does smolagents support tool calling without code execution?↓
Yes. If you prefer the classic pattern, ToolCallingAgent emits actions as JSON tool calls rather than Python, which suits environments where code execution is not acceptable. You keep the same tool definitions and model backends, only the action format changes. CodeAgent remains the recommended default for tasks that benefit from composition.
How do you define a tool in smolagents?↓
You decorate a typed Python function with the tool decorator or subclass the Tool class, and the framework introspects the signature and docstring to build the schema exposed to the model. Tools can also be loaded from the Hugging Face Hub or imported from other ecosystems. Clear type hints and docstrings directly improve how reliably the model calls each tool.
When should you not use smolagents?↓
It is a weaker fit when you need deterministic, auditable state machines with explicit checkpoints and rollback, where a graph first framework is clearer. It is also unsuitable if your security posture forbids executing model generated code and a sandbox is not available. For simple single prompt calls, a plain model client is lighter than any agent.