← Back to research
•·20 min read·industry

Agent Frameworks

Compare 16 agent frameworks by execution model, state and recovery, approvals, language support, observability, licensing, and deployment responsibilities.

Key takeaways

  • Choose the execution contract first: a model-driven loop, explicit workflow, durable runtime, or embeddable harness.
  • Saved conversation history, persisted workflow state, and safe recovery of external side effects are different capabilities.
  • AutoGen remains in maintenance; Google ADK and Pydantic AI have moved beyond the prior report's version assumptions.
  • Open-source cores can coexist with paid hosting, tracing, gateways, and enterprise-licensed modules.

FAQ

What is an agent framework?

A library or developer toolkit that coordinates model calls, tools, state, and control flow inside an application. A hosted agent service can expose an SDK without giving the application the same execution ownership.

Which agent framework is best for production?

There is no universal winner in this review. Evaluate recovery, approvals, model compatibility, observability, deployment ownership, and actual task outcomes in the language and infrastructure the team operates.

Does checkpointing prevent duplicate actions?

Checkpointing preserves selected state; it does not by itself make an external side effect safe to repeat. Test the exact recovery boundary and implement idempotency where the business operation requires it.

Are the frameworks free to operate?

Open-source code does not include inference, storage, compute, or every hosted feature. Mastra also explicitly separates its Apache-2.0 core from enterprise directories with different production-use terms.

Executive Summary

This comparison covers 16 agent frameworks, reviewed September 15–16, 2026. It retains the prior thirteen and adds smolagents, Pi, and Letta. These expose generated-code loops, an embeddable session harness, and a stateful harness with local or remote SDK execution. The scope now includes both libraries and open-source harness toolkits; a minimal abstraction or a server interface is not itself grounds for exclusion.[1][2][3][4][5]

The useful decision is what the team wants to control. A model-driven loop lets the model choose tools. An explicit graph or event workflow controls the progression of work. A durable runtime adds recovery machinery. A harness packages more of the surrounding context, sessions, and tools. These responsibilities overlap, but they are not interchangeable.

Several earlier assumptions need updating: AutoGen is maintained without new feature development; Google ADK's Python line is now 2.x; Pydantic AI 2.x is released; and Mastra documents an Apache-2.0 core with separately licensed enterprise directories. Those facts matter more to an adoption decision than an undated star ranking.[6][7][8][9]


Scope and Inclusion Criteria

The main matrix compares code-first toolkits with an open-source core, an agent loop or orchestration primitives, and documented multi-model paths. Products remain relevant either for new implementation or, in AutoGen's case, for operating and migrating an installed application. A published library is not assumed to be production-proven merely because it is downloadable.

This is a selected landscape, not an exhaustive directory. Voice-agent APIs have additional realtime transport and audio requirements. Commercial deployment services and provider-specific hosted agent APIs are discussed as adjacent choices rather than counted as open-source orchestration toolkits. LangChain, LangGraph, and the associated higher-level Deep Agents offering are treated as one ecosystem entry; their different responsibilities are identified below.[10][11]

Tembo receives a separate platform-and-SDK comparison because operating coding-agent sessions is a relevant alternative to building their orchestration. Disclosure: Ry Walker is Tembo's founder and CEO. Its inclusion in the discussion does not imply that Tembo is an in-process agent-loop library.[12][13]


Adjacent: Hosted Multi-Agent Research

SpaceXAI / Grok API offers a hosted research loop through the beta grok-4.20-multi-agent model. Its documented configurations use four or sixteen agents; it supports selected managed tools and remote MCP, but not client-side custom function tools. The readable result exposes the leader's tool calls and answer. Worker state can be preserved as encrypted content, not a complete readable trace of every worker's execution.[14]

That is a service to call from an application, not an open-source orchestration toolkit the application operates. It remains outside this report's 16-member matrix. The distinction changes the evaluation: an API consumer chooses the request and permitted service features, while a framework user also designs control flow, custom tool execution, recovery, and deployment.

For a proposed research workflow, compare the hosted result against a single-agent baseline using citation quality, task completion, latency, and total usage. More workers are not independent evidence of better answers. If the application must stop before a consequential custom action or inspect every worker's tool trace, the documented API limits matter before any quality comparison. This focused September 16 addition is a source review, not a new benchmark.

Comparison Matrix

Languages identify documented implementations, not a promise of feature parity. Follow the linked profile for background and the cited primary source for the current mechanism.

FrameworkLanguage / execution modelImportant boundary
AgnoPython agents, teams, and workflows; AgentOS runtime/UISDK, runtime, storage, and control UI are separate pieces of the operating stack.[15]
AutoGenPython/.NET event-driven core; AgentChat abstractionsMaintenance mode; Microsoft directs new projects toward Agent Framework.[6]
Strands AgentsPython and TypeScript model-driven harnessIn-process loop with provider/tool adapters; Bedrock defaults do not require every deployment to use a hosted control plane.[16]
Cargo AIRust runtime with JSON-defined agents and native build workflowDeclarative configuration and actions; scheduling is delegated to external operating tools.[17]
CrewAIPython role-based Crews plus event-driven FlowsA Crew's agent collaboration differs from a Flow's explicit state and control logic.[18]
Google ADKPython plus sibling language SDKs; agents, graph workflows, tasksADK 2 changes APIs and state representations; language implementations need separate compatibility checks.[19]
LangChain / LangGraphPython/TypeScript agent abstractions and stateful graph orchestrationLangGraph can be used without LangChain; persistent storage must be configured.[10][11][20]
LlamaIndexData/retrieval toolkit plus Python event-driven WorkflowsFramework, workflow package, and hosted document services are distinct products.[21][22]
MastraTypeScript agents and graph-based workflowsSuspension snapshots require configured storage; enterprise directories have different license terms.[23][24][9]
Microsoft Agent FrameworkPython/.NET agents and graph workflows; separate Go previewGo does not yet implement every .NET feature or orchestration pattern.[25][26]
OpenAI Agents SDKPython/TypeScript agents, tools, handoffs, and approvalsApplication owns execution and state; provider adapters do not guarantee every OpenAI-specific capability elsewhere.[27][28]
Pydantic AIPython typed agents, capability composition, graph and runtime integrationsDurable backends and the separate Harness library add responsibilities beyond a basic Agent.run.[29][30]
Vercel AI SDKTypeScript provider interface, ToolLoopAgent, and UI streamingSDK code, Vercel AI Gateway, and a tool's execution environment are separate choices.[31]
smolagentsPython code actions or structured tool callsLocalPythonExecutor is explicitly not an isolation boundary.[1][32]
PiTypeScript stateful core and embeddable session harnessChoose low-level messages/tools/hooks or the fuller coding-agent session/resource layer.[2][3]
LettaStateful open-source harness with JavaScript/TypeScript Agent SDK and App ServerLocal subprocess, remote execution host, and cloud state are distinct choices; local tools do not imply local-only memory.[4][5][33]

Framework Profiles and Design Tradeoffs

Agno

Agno exposes agents, teams, and workflows through its SDK, then adds an AgentOS runtime and control UI. The current repository describes database-backed sessions, memory, knowledge, traces, scheduling, and access controls. That makes it relevant when a team wants a coherent application runtime rather than assembling every operating surface itself.[15]

Evaluate the actual deployment composition: where the runtime runs, which database persists state, how user identity reaches tool calls, and which controls protect each tenant. The presence of a runtime does not remove those design decisions.

AutoGen

AutoGen's repository states that it is in maintenance mode, community-managed, and no longer adding new features. Its Core and AgentChat packages remain relevant to existing systems; the maintainers recommend Microsoft Agent Framework for new users. AutoGen Studio is explicitly described as a prototyping tool rather than a production application.[6]

The practical choice is migration cost versus the value of new capabilities. Do not describe an operating AutoGen installation as automatically broken, or start a greenfield project on it without considering the stated maintenance direction.

Strands Agents

Strands packages a model-driven agent loop in Python and TypeScript. The current harness-sdk repository documents tools, MCP, structured output, memory/session integrations, hooks, tracing, and controls for turn limits, budgets, cancellation, and stop reasons. The default model path uses Amazon Bedrock, while other providers are supported.[16]

The relevant tradeoff is how much workflow structure the application needs around that loop. A model-directed next action is convenient for open-ended work; business invariants still need explicit application controls. A TypeScript implementation should be checked against its own API and release rather than presumed identical to Python.

Cargo AI

Cargo AI uses JSON definitions for inputs, runtime variables, schemas, and actions, with run and hatch workflows and native-platform targets. Its documented providers include OpenAI, Anthropic, Gemini, xAI, Mistral, and Ollama. It can operate locally without an account, while scheduling uses external tools such as the operating system's scheduler.[17]

It is useful to evaluate when reviewable definitions and a compiled artifact are central requirements. This review does not establish the same breadth of independent production experience as for larger ecosystems; the maintainer's availability claim is not a substitute for deployment testing.

CrewAI

CrewAI separates Crews—agents with roles, tools, and tasks—from Flows that coordinate state, branching, and events in Python. A Flow can invoke an individual agent or a Crew, so applications can combine controlled business logic with more autonomous collaboration. The commercial AMP suite is a separate operating layer around the framework.[18]

Its @persist decorator can save Flow state through the default SQLite backend or a custom persistence implementation. Current documentation distinguishes resuming under the same state ID from forking a saved snapshot into a new ID. That distinction matters for audit history and should not be hidden behind a single “durable” checkmark.[34]

Google ADK

Google ADK's Python repository now documents the 2.x architecture, including graph workflows, dynamic nodes, task delegation, tool confirmation, and model-independent paths. The latest release checked was 2.9.1 on September 15. Sibling Java/Kotlin, Go, and TypeScript SDKs are linked, but that is not proof of identical coverage.[19][7]

The 2.x transition changes agent APIs, event models, and session schemas. The repository notes that 2.x sessions are readable by sufficiently recent 1.x versions, not arbitrary older installations. Existing adopters should test stored sessions and deployment integrations alongside source-code migration.[19]

LangChain and LangGraph

LangChain supplies higher-level agent construction and integrations. LangGraph provides lower-level stateful orchestration and can be used independently. The current LangChain repository positions Deep Agents as a more opinionated layer with planning, subagents, and filesystem tools; it should not be confused with the entire LangGraph API.[10][11]

The key operational choice is persistence. In-memory checkpoints disappear with the process; a persistent checkpointer such as PostgreSQL changes that boundary. Thread-scoped checkpoint state and cross-thread stores serve different purposes, and retained state needs its own lifecycle and cleanup policy.[20]

LlamaIndex

LlamaIndex remains a framework for data-connected applications, with connectors, retrieval, and agent integrations. The current repository also states that the company's primary focus has shifted toward document products such as LlamaParse and LiteParse. That is a positioning change, not a declaration that the open-source framework has disappeared.[21]

Its Workflows package uses typed events to connect asynchronous steps: a step receives an event and returns the event type that activates another step. Branches, loops, and concurrent work follow that event structure. Workflows can be installed separately or accessed through the core framework; a retrieval-heavy team can assess this control flow without assuming every hosted document product is required.[22]

Mastra

Mastra is a TypeScript framework with agents, graph-based workflows, memory, integrations, and development tooling. Workflow composition includes sequences, branches, and parallel steps. Its suspension API saves a workflow snapshot to configured storage, allowing a run to wait for human input and resume at a specified step.[23][24]

That makes storage configuration part of the approval workflow, not optional implementation trivia. Licensing also needs a module-level check: Apache-2.0 covers the ordinary core, while ee/ directories use a separate enterprise license.[9]

Microsoft Agent Framework

Microsoft Agent Framework offers Python and .NET agents, middleware, sessions, graph workflows, checkpointing, streaming, and OpenTelemetry integration. The repository links a separate durable-execution extension rather than equating every in-memory agent call with an externally managed durable workflow.[25]

The Go implementation is a public preview. Its README explicitly lists missing capabilities, including handoff orchestration and Foundry-hosted deployment; it also says .NET has broader product integrations. Language availability and feature parity therefore need separate rows in an internal evaluation.[26]

OpenAI Agents SDK

The OpenAI Agents SDK runs the agent loop inside the application's process. Python and TypeScript APIs supply agents, tools, handoffs, state, guardrails, and approvals, while the application owns hosting, tool implementation, and persistence. Non-OpenAI model adapters exist, but advanced features can depend on the Responses API path.[27][28]

The managed Agents API is a different choice: OpenAI operates the session orchestration and harness, with execution environments supplied through hosted or self-hosted paths. Its current overview lists US-only data residency and no zero-data-retention support. Those service constraints should not be incorrectly attributed to the open-source SDK as if the two were the same product.[35]

Pydantic AI

Pydantic AI centers on typed outputs, dependency injection, tools, and composable capabilities. Its current architecture also includes a separate Harness library for memory, subagents, context management, and packaged coding/research agents. The 2.43.0 release checked on September 16 means the prior “v2 beta” framing is stale.[29][8]

Durability is supplied through explicit runtime integrations. Current documentation lists co-maintained Temporal, DBOS, Prefect, Restate, and AWS Lambda durable-function paths, plus external integrations. Choose that operating dependency deliberately; type validation and an external workflow engine solve different problems.[30]

Vercel AI SDK

Vercel AI SDK combines a provider interface, agent tool loops, typed streaming, and UI integration. Its current README shows ToolLoopAgent, framework-specific UI packages, and both gateway model strings and direct provider packages. It is a useful starting point when an application must connect a tool-using backend to an interactive web UI.[31]

The SDK is not the same product as AI Gateway, and a shell tool still needs an execution environment. Decide separately how requests are routed, how tool code runs, and which infrastructure hosts the application.

smolagents

smolagents makes generated Python a first-class action format while retaining a conventional structured-tool agent. Its adapters cover multiple model and tool ecosystems, and its executor choices can move code outside the application process.[1]

Its local interpreter is explicitly not a security boundary. A team choosing code actions should evaluate the external sandbox and permitted tools as part of the framework decision, rather than assuming an import allowlist provides host isolation.[32]

Pi

Pi's lower-level agent core exposes stateful messages, event streams, tool execution, context transformation, and before/after tool hooks. Its fuller SDK adds sessions, compaction, model switching, resource discovery, and embedding in custom applications. The core and coding-session layers are separate choices within the same ecosystem.[2][3]

The documented tool-execution contract is particularly concrete: parallel execution is the default, preflight hooks can block a call, and completion events can arrive in a different order from persisted tool-result messages. An integration that displays progress or attaches side effects should honor those ordering semantics.[2]

Letta

Letta's current developer surface extends beyond the stateful server described in the earlier comparison. Its open-source harness manages persistent agent identity, memory, tools, skills, and subagents. The JavaScript/TypeScript Agent SDK creates agents and sessions, sends messages, and streams responses through a cloud backend, a local Letta Code subprocess, or a remote App Server. The documented local path requires Node.js 22.19 or later.[4][5]

App Server separates the execution host from the state backend: tool execution can occur on a laptop while conversation history and memory use cloud storage. A local state backend is a different configuration. Python clients currently use the App Server WebSocket protocol directly rather than the new JavaScript/TypeScript SDK. This distinction matters when embedding a personal assistant, choosing deployment responsibilities, or assessing where private context is retained.[33][5]


State, Recovery, and Human Approval

Three requirements are often collapsed into one feature:

RequirementWhat to verify
Conversation continuityWhich messages and context survive between turns, and who can read them?
Workflow recoveryWhich step and state snapshot survive a process restart or deployment?
Safe external effectsWhat prevents a repeated tool invocation from creating a duplicate business action?

A saved transcript alone does not answer the latter two. LangGraph documents checkpointers and stores; Mastra persists suspension snapshots; CrewAI persists selected Flow state; Pydantic AI integrates with separate durable engines. They provide different building blocks, so reproduce the application's failure boundary instead of comparing a Boolean “memory” column.[20][24][34][30]

For example, consider a workflow that drafts a repository change, waits for approval, then creates a pull request. A useful acceptance test stops the worker after approval is saved but before the PR response is recorded. On restart, the application must know whether to resume, retry, reconcile an already-created PR, or ask for intervention. Idempotency belongs at the external-operation boundary, not solely in the model prompt.

OpenAI's approval mechanism illustrates the separation. A tool requiring approval produces an interruption and resumable state; the application approves or rejects it, then resumes that state. Input guardrails cover the first agent, output guardrails cover the final output, and function-tool guardrails cover the tools they are attached to. Agent-level checks therefore do not automatically validate every intermediate action.[36]

First-hand recovery experience

In LangGraph issue 7361, sardismart reported on March 31, 2026 that resuming with a particular checkpoint configuration replayed work unexpectedly. On April 29 the same reporter said upgrading to 1.1.9 appeared to resolve it; a June 8 participant described the same improvement in their FastAPI/PostgreSQL setup. This is a versioned account with a reported fix, not a current general failure claim.[37]

The lesson is to test the installed package versions, nested workflow shape, and resume arguments. No independent cross-framework load or recovery benchmark was conducted for this report.


Observability, Data Flow, and Cost

Tracing can reveal model calls, tool results, and state transitions, but the trace destination becomes part of the application's data flow. OpenAI's standard server-side SDK tracing is enabled by default and records model/tool activity, handoffs, and guardrails. Its SDK supports controls for reducing or disabling tracing. Hosted MCP and SDK-managed local/private MCP also have different connectivity ownership.[38]

Pydantic AI exposes OpenTelemetry instrumentation that can use Logfire or another backend. CrewAI distinguishes its default telemetry from optional share_crew detail, and its README documents an opt-out. Agno likewise documents telemetry behavior and an opt-out. Review these settings together with application logging and provider retention, not just the framework's license.[29][18][15]

Open-source code is only one part of the operating bill. Budget model calls, retries, storage, workers, sandbox compute, traces, and any commercial control plane. Inference platforms and agent sandboxes are separate procurement decisions. A framework can be inexpensive to install and expensive to operate if its workflow repeatedly calls costly models or retains excessive state.

Mastra provides a specific licensing example: its repository excludes ee/ directories from the ordinary Apache-2.0 grant. The enterprise license, effective August 24, 2026, requires a written agreement for production use of that enterprise code. This qualification is more useful than saying every feature in every listed repository is freely usable in production.[9][39]


Tembo: Operating Agent Workflows Through a Platform SDK

Tembo addresses a related build-versus-operate decision. Its current platform runs coding-agent sessions with repository context, tools, integrations, and reviewable outputs; it also offers self-hosted deployment. Its TypeScript SDK exposes typed API access for initiating agent work, with retries, error types, and logging controls. This is application integration with a platform rather than importing an agent loop that executes entirely inside the caller's process.[12][13]

The current Agents documentation describes a practical workflow: choose a template or write instructions, select a harness and model, attach schedule/event/webhook triggers, configure integrations, and optionally select a prepared project. Macro invocations can start a configured agent on demand. These mechanisms are relevant when the desired outcome is a recurring engineering workflow rather than a new general-purpose agent runtime.[40]

Team requirementFramework responsibilityTembo platform question
Custom application reasoningDefine the loop, state, tools, and recovery contractDoes an existing coding harness and its available configuration fit?
Scheduled engineering workOperate triggers, workers, credentials, and execution environmentsDo documented schedules, event hooks, projects, and integrations cover the workflow?
Human reviewBuild the approval state and application interfaceDoes the session/output review process match the team's controls?
Deployment ownershipHost the application and selected supporting servicesCompare managed and self-hosted platform arrangements and their operational terms

This is an editorial responsibility comparison based on the documented platform. No custom framework-to-Tembo deployment was executed for this review. Ry Walker's founder/CEO relationship is disclosed above.

Agent Studio: a separate control plane for team agents

Tembo Agent Studio is a distinct MIT-licensed project, not an old name for the commercial platform's Agents configuration. It operates Pydantic AgentSpec and Cargo AI definitions through a self-hosted web application, Rust API, and Postgres. Teams can hand-author specs without a Tembo subscription; the commercial coding-agent service is optional for natural-language authoring and improvement requests. This is a way to operate supported agents without building the entire shared application around a framework.[41]

The relevant comparison is ownership of the surrounding workflow: Git definitions, shared users, schedules, run history, and promotion. A promoted stable version is a database snapshot, while the draft follows the repository's default branch. Automated runs default to stable after the first promotion, but schedules can opt into draft. Review those settings when moving an experiment into recurring use.[42] Agent Studio belongs in the team-agent platform comparison; it is adjacent to this report's sixteen framework members because its primary offering is the operating application around supported runtimes.


How to Choose

Start with a workload that exposes the difficult parts, then narrow by language and ownership:

  • Typed Python application contracts: evaluate Pydantic AI alongside the typed output and tool facilities of other candidates; determine whether a separate durable engine is needed.
  • Explicit stateful orchestration: evaluate LangGraph, Mastra workflows, Microsoft Agent Framework, Google ADK, CrewAI Flows, or LlamaIndex Workflows against the same interruption and restart scenario.
  • Model-directed tool loops: compare Strands, OpenAI Agents SDK, Agno, smolagents, and Pi's core on tool control, state, and provider behavior.
  • Interactive TypeScript applications: compare Vercel AI SDK's UI/streaming path, Mastra's workflow/runtime choices, and Pi's embedding interfaces.
  • Persistent assistant identity and memory: evaluate Letta's harness and SDK while explicitly selecting execution placement and local or cloud state.
  • Existing AutoGen deployments: evaluate migration effort and maintenance needs before changing frameworks solely to follow a category trend.
  • Configured engineering automation: compare building the orchestration with operating the workflow through Tembo's platform and SDK.
  • Shared agents from versioned specifications: evaluate Agent Studio when its supported runtimes fit and the missing work is ownership, scheduling, and promotion rather than a new agent loop.

These are starting points, not mutually exclusive rankings. A credible evaluation should demonstrate one successful task, one rejected tool call, one provider failure, one worker restart, and one delayed approval. Record the final output, external side effects, trace data, and complete cost. That evidence explains the choice more clearly than repository popularity or a feature checklist.


Research by Ry Walker Research • methodology

Sources