← Back to research
•·19 min read·industry

AI Agent Sandboxes

Compare 23 AI agent sandbox tools by isolation, persistence, network controls, deployment, workflow, and current pricing. Verify VM, egress, and credentials separately.

Key takeaways

  • A VM boundary, an egress policy, and scoped credentials solve different problems; verify all three.
  • Filesystem persistence and memory recovery differ by product, runtime class, and lifecycle event.
  • Compare complete agent workflows such as Tembo with execution APIs according to how much system your team intends to build.
  • Idle compute savings do not eliminate storage, subscription, inference, or operational costs.

FAQ

Which sandbox is best for an AI agent?

Start with the workload: execution API, complete team workflow, GPU environment, or self-operated runtime. There is no verified universal performance or security winner across these products.

Does persistence mean a running process resumes?

Not necessarily. Some products preserve only files, while others retain memory in specific pause modes. Cold sleep, expiration, deletion, and documented fallback behavior can change what returns.

Does a microVM prevent data exfiltration?

It isolates guest execution from the host, but allowed network destinations and credentials still grant authority. Check effective egress rules, application authentication, and provider-side permissions.

Where does Tembo fit in the comparison?

Tembo is a coding-agent workflow platform with dedicated VM sessions, integrations, and review workflows, plus commercial self-hosting. It is relevant when a team wants the surrounding delivery system included.

Executive Summary

Choosing an AI agent sandbox now means choosing several things at once: an execution boundary, a persistence model, a network and credential policy, and how much agent workflow to build yourself. A product can preserve files without preserving running processes, and a separate VM does not prevent misuse of credentials intentionally granted to its guest.

This report compares 23 tools and platforms, checked September 15, 2026. It retains the previous nineteen entries, adds Deno Sandbox, DeepInfra, Nebius Sandboxes, and Tembo, and separates managed execution, self-operated infrastructure, and provider abstractions. Those layers solve related problems but are not interchangeable.

The most consequential changes are:

  • Runtime choice is becoming explicit. Daytona documents container, VM, and GPU classes with different boundaries. Modal's new VM Sandboxes are beta and CPU-only; its established gVisor path remains relevant for GPU work.[1][2][3]
  • Persistence needs a precise definition. E2B pause/resume can retain memory; Vercel's default persistence saves files; Cloudflare's durable sandbox identity does not preserve a stopped container's filesystem. These are materially different recovery contracts.[4][5][6]
  • Company and security status changed. Baseten announced its acquisition of Blaxel on September 10 and says the product continues. The researcher who reported AgentCore DNS exfiltration now says that path was fixed on April 15; it should not be presented as a current unpatched defect.[7][8]
  • A complete platform is a separate buying decision. Tembo combines isolated coding sessions with repository, ticket, review, and team workflows. That matters when the goal is deploying working agents rather than building a service around a sandbox API.[9][10]

Disclosure: Ry Walker is CEO and co-founder of Tembo. Its capabilities and limitations below use the same first-party evidence standards as the other products. This report includes no hands-on performance comparison or independent security audit.


Scope and Selection Criteria

The comparison covers environments for executing agent-generated code, systems for operating those environments, and abstractions used to select an execution provider. It preserves the existing broader infrastructure scope, including experimental projects, while labeling their readiness and role. Inclusion is not a production recommendation.

Deno is a newly profiled execution service. DeepInfra documents an execution SDK/API, and Nebius documents a beta cloud execution API through Contree; both meet the same programmable-execution criterion.[11][12] Tembo joins as a workflow platform with an explicit VM execution layer. Local command wrappers belong in Local Agent Sandboxes; low-level components such as libkrun are contextual building blocks, not extra members. Docker Sandboxes belongs in that local comparison because its primary documented workflow runs coding-agent CLIs on the developer's computer.[13][10][14]

Discovery also reviewed GitHub Copilot CLI's cloud sandbox preview. GitHub documents an interactive-only cloud session: it cannot currently combine cloud mode with the CLI's programmatic options. It is useful context for people moving a Copilot session off their laptop, while this matrix focuses on execution APIs, programmable infrastructure, and broader automated agent workflows. It remains outside the member count; its published scope, not vendor affiliation, drives that choice.[15]

The matrix favors documented boundaries, state, deployment, and workflow over vendor startup rankings. Cold creation, cached restoration, filesystem cloning, and memory restoration measure different operations. There is no supported basis here for declaring one provider universally fastest or safest.


Adjacent: Model-Managed Python Execution

SpaceXAI / Grok API exposes a code-execution tool for Python calculations and analysis with common packages such as NumPy and pandas. Its documented contract is temporary execution without state persistence between requests and without external network or filesystem access. Passing earlier conversation context is not evidence that the previous Python process or files survive.[16]

This can be sufficient when the task is calculating from supplied data inside a model response. It does not establish a reusable repository environment, an arbitrary system-package installation path, or a long-running service endpoint. Nor does the generic word “sandbox” establish a specific hypervisor or an independently tested isolation guarantee. These are limits of the published contract, not results of an escape test.

For a coding workflow that needs private packages, prepared repositories, and subsequent review, evaluate the actual execution products in the matrix. Tembo's dedicated Linux sessions and project caches, for example, address a broader agent workflow than a model-managed Python calculation.[10] The founder disclosure above applies. This September 16 contextual addition leaves the 23-member matrix and the other providers' September 15 review baseline unchanged.

Comparison Matrix

Managed execution and workflow platforms

PlatformExecution and deploymentState and distinguishing tradeoff
E2BManaged Firecracker microVMs; enterprise deployment in the customer's cloudPause/resume can save disk and memory; timeout behavior must be configured. Open runtime also has an evaluation-oriented local package.[17][4][18]
DaytonaContainer, Linux/Windows VM, and GPU classesFiles survive stop/start for container and VM classes; memory pause/fork belongs to VM classes. GPU stop deletes the ephemeral sandbox.[1][19]
ModalEstablished gVisor execution plus beta VM SandboxesCPU/GPU platform, volumes and snapshots; beta VM runtime has a real Linux kernel but currently no GPU support.[3][2]
SpritesManaged Firecracker machinesPersistent writable filesystem and disk checkpoints; cold sleep discards process memory. Services can restart on wake.[20][21]
Vercel SandboxManaged Firecracker VMs, custom images and privileged guest workloadsFilesystem autosave on stop and restoration on resume; snapshots and beta Drives provide other storage paths.[5]
Cloudflare Sandbox SDKWorkers coordinate Containers, each with its own VMDurable Objects retain identity, not the stopped container's files. Use bucket mounts or backups; 1.0 remains a preview.[22][6][23]
AWS AgentCore Code InterpreterManaged sessions for Python, JavaScript, TypeScript, shell, and filesSession state lasts until expiry; customer-owned EFS/S3 Files mounts can persist and share data through a VPC.[24][25]
Google Agent SandboxGemini Enterprise Agent Platform sandbox environmentsCode execution, computer use, and custom Linux containers; VPC Service Controls and CMEK are documented. Do not infer a specific hypervisor from the sandbox label.[26][27]
BlaxelManaged microVMs; now part of BasetenStandby restores memory and filesystem; snapshot storage is billed. Agent Drive is a private preview, not a generally available shared filesystem.[28][7]
NorthflankManaged or customer-cloud platform with workload-specific Kata, Firecracker, or gVisor optionsPersistent volumes, CPU/GPU workloads, and deployment tooling; runtime selection changes the boundary.[29][30]
RunloopManaged microVM devboxes, enterprise VPC deploymentDevelopment environments, blueprints, disk snapshots, and evaluation workflows; suspend/resume is a Pro-plan feature.[31][32][33]
CodeSandbox SDK / Together SandboxCodeSandbox-derived microVM infrastructure under TogetherMemory/filesystem hibernation, cloning, Dev Containers, and previews; current docs still distinguish Together custom plans from CodeSandbox self-service.[34][35]
Deno SandboxManaged Linux Firecracker microVMs through DeployHost-scoped secret substitution, writable volumes, prepared filesystem roots, and Deploy integration; paid plan required.[36][37][38]
DeepInfraManaged Linux microVMs through Kata/QEMU/KVMOnly /workspace survives a clean stop/start; failures lose it. Five active sandboxes per account; stopped environments expire after seven days.[11]
Nebius Sandboxes / ContreeBeta cloud API with documented VM-level isolation and OCI image importBranch commands from saved filesystem states. Beta permits 50 simultaneous operations; untagged, unreferenced checkpoint images may be deleted after 180 days.[12][39]
TemboCoding-agent workflow platform with dedicated Linux VM sessions; commercial self-hostingRepositories, tools, integrations, and review workflow are included. Completed sessions are ephemeral; reusable project caches are an explicit persistence exception.[9][10][40]

Infrastructure you operate and provider abstractions

ToolRole and boundaryOperational qualification
OpenShellApache-2.0 policy runtime with container and MicroVM compute driversAlpha; backend and policy choices matter. HTTP inspection defaults to audit and Landlock can degrade unless made a hard requirement.[41][42][43]
OpenSandboxApache-2.0 execution APIs and SDKs with Docker/Kubernetes backendsSupports additional secure runtimes, but an API-compatible deployment does not automatically select Firecracker or Kata. Operators supply and secure the infrastructure.[44]
MicrosandboxApache-2.0 local microVM runtime and embeddable SDKsv0.6.18 documents Linux, macOS, and Windows support and no required long-running server; explicitly beta. Host virtualization prerequisites still apply.[45]
AIO SandboxApache-2.0 Docker image bundling browser, terminal, files, editor, Jupyter, and MCPA useful tool environment, not a distinct VM boundary. The documented quickstart disables seccomp; API authentication is optional unless configured.[46]
ComputeSDKMIT unified TypeScript API and multi-provider benchmark harnessThe selected provider supplies isolation, persistence, credentials, and the bill. API portability does not make those contracts equivalent.[47]
QuiltLinux container runtime with namespaces, cgroups, OCI images, daemon/CLI, and volumesPublic repository remains available. This review verified the runtime, not current managed-service terms; a quiet repository is not proof of shutdown.[48]
ZerobootApache-2.0 KVM/Firecracker snapshot-forking prototypeREADME explicitly says not production hardened; no guest networking and one vCPU. Useful systems research, not a default production recommendation.[49]

Docker packaging and Kubernetes scheduling describe how workloads are distributed; they do not by themselves establish the host-isolation boundary. OpenSandbox's configurable backends and AIO's container quickstart make that distinction concrete.[44][46]


Persistence: What Actually Comes Back?

Files, memory, and external effects are separate

A filesystem snapshot can restore installed tools and source files without restoring an in-flight process. Sprites explicitly distinguishes warm sleep from cold sleep: only the warm case retains memory, and checkpoints recover writable disk state. Runloop's snapshot guide likewise says its snapshots are disk-only.[20][21][32]

E2B supports memory-inclusive pause/resume, but its current persistence guide describes a rollout caveat: automatic pause retries memory preservation for up to two minutes before falling back to filesystem-only preservation. The default timeout kills a sandbox unless pause is selected. A recovery design must handle that documented fallback, reconnect clients, and distinguish a paused sandbox from a deleted one.[4]

Daytona's class distinction is equally consequential. Container stop/start keeps files but does not resume process memory; VM classes support memory pause and hot-snapshot forks. GPU-class environments are deleted when stopped, so durable results need volumes or external storage.[19]

Vercel now saves the filesystem by default and documents stopped/resumed environments. Cloudflare starts a fresh container after idle shutdown; keepAlive prevents idle sleep but not other restarts. A stable sandbox ID is an address, not evidence that its previous files remain.[5][6]

Nebius Contree branches from filesystem checkpoints, not a documented snapshot of running process memory. Its SDK defaults ordinary runs to disposable execution and sessions to persistence; use disposable=False deliberately when a result must seed another branch. Preserved environment variables become part of the resulting image, so checkpoints deserve the same credential review as other persisted artifacts.[39][50]

State outside the VM needs its own policy

Deno's writable volumes have a single attached sandbox at a time, while snapshots can seed multiple roots. AWS's EFS and S3 Files access points deliberately support shared access across sessions. These are different collaboration models: isolated working copies versus shared mutable data.[51][25]

Tembo documents session teardown alongside persistent cached repositories and dependencies for prepared projects. Its public pricing lists pausable/resumable sessions, but that does not establish the same memory-fork contract as a sandbox SDK. Evaluate the application workflow and its cache lifecycle explicitly.[10][52]

Restoring a local checkpoint is not a rollback of an already sent email, pushed commit, database update, or payment. Treat external actions as a separate approval, idempotency, and audit problem. That is an architectural recommendation, not an additional feature attributed to any provider.


Network and Credential Controls

VM isolation does not set an egress policy

Runloop devboxes have unrestricted outbound networking by default; hostname policies are optional. Deno likewise permits outbound access when allowNet is omitted. Sprites offers DNS-based egress policies, but outbound networking is otherwise open. The presence of a microVM says nothing about whether the guest can reach an attacker-controlled server.[53][37][54]

Daytona's defaults depend on account tier and class, with explicit domain/CIDR policies available. Avoid copying a “restricted by default” statement from an introductory tier into a production threat model without checking the effective policy.[1]

OpenShell makes another distinction visible: observing HTTP requests is not the same as enforcing HTTP rules. The released security model documents audit defaults and optional hard failure when Landlock is unavailable. Enabling an enforcement feature and testing its failure behavior are separate from installing the runtime.[43]

DeepInfra documents public-internet egress, blocked inbound connections, and blocked SMTP ports; port exposure remains roadmap work. Its built-in network restrictions do not amount to a documented per-domain allowlist.[11]

Placeholder credentials still grant authority

Deno's secrets API substitutes real credentials for selected destinations; ordinary injected environment variables remain readable by guest code. Cloudflare can keep credentials in a Worker-side outbound handler, while an environment variable placed inside the container is exposed to its processes. Choose the intended mechanism, and scope the external account's permissions as well as its allowed hosts.[37][22]

Cloudflare also documents application responsibilities: sandbox IDs are not authentication secrets, preview URLs can act as bearer access, and separate customers should receive separate sandbox instances. A correct VM boundary does not fix an application that attaches two users to the same environment.[22]

The AgentCore incident illustrates the layers. BeyondTrust originally demonstrated DNS exfiltration despite restricted network mode; its April 22 update says the DNS issue was fixed April 15. That historical result does not establish a hypervisor escape or justify treating the old DNS path as still open. The research also makes execution-role authority part of the evaluation: permitted cloud APIs can carry data even when an arbitrary internet destination is blocked.[8]


Pricing and Licensing

Public USD terms checked September 15, 2026. These are billing structures and selected rates, not an equal-workload cost ranking. CPU utilization, allocated memory, persistence, region, egress, subscription allowances, and agent inference all affect the result.

PlatformPublished billing modelCost detail that changes the comparison
E2BHobby without a subscription; Pro $150/month, both plus usageHobby has a one-time $100 credit and one-hour sessions; Pro extends standard sessions to 24 hours. Enterprise terms differ.[55]
Daytona$0.0504/vCPU-hour, $0.0162/GiB-hour, storage separatelyWindows surcharge and GPU class/rates are separate; a stopped environment's retained storage is still a resource.[56]
ModalStarter $0 subscription with $30 monthly compute; paid tiers plus usagePhysical-core, memory, GPU, and volume meters differ. Its CPU unit is a physical core representing two vCPUs, so compare units carefully.[57]
SpritesMonthly plans with included allowances and usage metersCPU, memory, hot storage, cold storage, and model usage are distinct. Sleeping does not mean retained storage is free.[58]
Vercel SandboxUsage across CPU, provisioned memory, creation, egress, and storageSelected default-region rates: $0.128/active CPU-hour and $0.0212/GB-hour memory. Memory has a one-minute minimum; region and plan matter.[59]
Cloudflare Sandbox SDKUnderlying Containers usage plus other Cloudflare servicesInclude Workers, Durable Objects, optional logs, network, and persistence services; the open SDK is not free hosted compute.[60]
AWS AgentCore Code InterpreterActive CPU and peak-consumed memory, per secondListed rates include $0.0895/vCPU-hour and $0.00945/GB-hour; network and customer-owned persistent storage add costs.[61][25]
Google Agent SandboxAllocated Agent Compute and Agent MemoryCurrent pricing lists $0.085/vCPU-hour and $0.009/GiB-hour for runtime and sandbox environments. The earlier free-preview assumption is obsolete.[62]
BlaxelUsage-based compute, billed by allocated RAM while active$0.0000115/GB-RAM-second; standby memory/filesystem snapshots $0.20/GB-month. “Zero standby compute” does not mean zero storage cost.[28]
NorthflankResource-based managed or BYOC deploymentAccount for services/jobs, volumes, builds, egress, GPUs, and the underlying cloud allocation. Obtain a topology-specific estimate.[63]
RunloopBasic has no subscription; Pro $250/month plus usagePublished CPU $0.108/hour and memory $0.0252/GB-hour; Pro adds suspend/resume and custom benchmarks. Storage continues while suspended.[31]
Together / CodeSandboxTogether custom terms or CodeSandbox self-service VM creditsDocs list $0.01486/credit and minute-rounded runtime; a 2-core/4-GB Nano is $0.1486/hour under that self-service table.[34]
Deno SandboxPro $20/month or Builder $200/month plus overagesShared Deploy allowances; $0.10/CPU-hour, $0.025/GiB-hour, and $0.20/GiB-month volume overages. Free excludes Sandbox.[38]
DeepInfraPer-second while creating, starting, running, or stoppingNo minimum; stopped is unbilled. Obtain current sizes/rates from the account catalog; boot and stop time are billed.[11]
Nebius SandboxesBeta; no sandbox-specific public tariff verifiedThe reviewed overview documents limits but no billable resource rates. Obtain beta billing and retention terms before extrapolating inference prices to execution.[12]
TemboFree trial allowance; Pro $60/month or Max $200/month with equivalent monthly usageVM compute and Tembo Gateway inference draw from the same allowance. VM rate is $0.0403/vCPU-hour + $0.0130/GiB-hour; BYOK changes inference billing, not VM billing.[64]

For the self-operated entries, the software license does not pay for hosts, images, orchestration, networking, updates, or on-call work. OpenShell, OpenSandbox, Microsandbox, AIO, and Zeroboot publish Apache-2.0 code; ComputeSDK is MIT. Quilt's repository includes an MIT license file. Those licenses do not turn experimental components into supported managed services.[41][44][45][46][49][47][65]

E2B also publishes its Apache-2.0 runtime, but its single-machine Embed package is explicitly an evaluation package rather than a production deployment pattern. Tembo self-hosting is a licensed commercial platform with customer responsibilities for infrastructure, secrets, backups, and operations. These are different ways to obtain control of deployment.[18][40]


Choosing for a Real Workflow

Building a product around execution APIs

Shortlist E2B, Deno, Vercel, Blaxel, or Cloudflare against the application's actual requirements: guest toolchain, exposed services, network policy, latency distribution, state retention, and supported regions. Deno adds a direct Deploy path; Cloudflare fits a Worker-coordinated application; Vercel combines custom VM environments with its sandbox APIs. These are conditional fits, not a universal ranking.[36][23][5]

ComputeSDK can reduce adapter work and provides an open benchmark harness. Its common API should be treated as a portability layer; repeat tests for provider-specific timeouts, mounts, credentials, and restore behavior after switching backends.[47]

Deploying coding agents for a team

Tembo is relevant when the desired outcome includes agent selection, prepared repositories, ticket-triggered work, shared sessions, and review. Buying those workflows changes the build-versus-buy calculation compared with purchasing only a VM API. Self-hosting can move the platform into infrastructure the team controls, with operational ownership remaining explicit.[9][40]

Runloop is relevant when the team is building and evaluating its own coding agents: blueprints, reusable devboxes, and benchmark workflows address that work. Together Sandbox emphasizes interactive development environments, cloning, browser sessions, and previews. Neither a snapshot primitive nor a benchmark platform should be assumed to include every repository-to-review workflow of a complete agent application.[33][34]

GPU, customer-cloud, and runtime-control requirements

For GPU execution, evaluate the actual Daytona class, Modal's established execution path, or Northflank's GPU deployment. Modal's beta VM Sandboxes currently exclude GPUs, so “Modal supports GPUs” is insufficient to validate a VM-specific design.[1][2][57][30]

For existing cloud governance, AgentCore and Google Agent Sandbox deserve evaluation alongside customer-cloud offerings. Check identity, persistent mounts, permitted network paths, image restrictions, and costs. Google's custom-container guide excludes images requiring root or restricted system resources; that matters for toolchains that assume unrestricted Linux administration.[25][26][27]

OpenSandbox and OpenShell suit teams prepared to operate runtime and policy infrastructure. Microsandbox is useful for embedding local VMs; AIO packages agent-facing tools but needs an independently appropriate deployment boundary. Quilt and Zeroboot remain visible for their designs, with the verification and prototype limits above. Do not transfer a managed provider's operational guarantees to a self-hosted README example.[44][42][45][46][48][49]


Developer Evidence and Evaluation Method

Independent reports can identify failure modes without establishing a reliability ranking. In a February 9 Deno follow-up, nihakue confirmed a missing CLI option had been fixed, described remaining volume/snapshot workflow friction, and reported a useful SSH-and-Claude setup. It is dated experience with a particular workflow, not proof that an early defect persists today.[66]

The AgentCore research is stronger evidence for the specific tested network/credential behavior, including the later fix, than for a broad vendor security grade. Vendor-supplied latency numbers and customer logos answer different questions. This review did not find a common independent benchmark that establishes the best end-to-end choice across all 23 entries.[8]

A useful pilot should record the runtime/version, region, image, dependency cache, concurrency, and workload. Measure cold starts separately from resumes; interrupt the environment during a write; verify which files and processes return; test blocked destinations and authorized API misuse; and record the bill after idle storage and inference. These are proposed evaluation steps, not tests performed for this report.


Outlook

Persistence, credential proxies, and deployment options are becoming more widely documented, but their contracts remain different. Blaxel's acquisition adds an inference-platform owner to the landscape; it does not establish a delivered combined product or a reason to assume all other providers must consolidate.[7]

The practical choice is still workload-specific: buy an execution API when building the surrounding agent system is intentional, buy a workflow platform when team delivery is the objective, and operate a runtime when control justifies the engineering responsibility. Preserve that distinction as serverless execution and agent platforms continue to overlap.


Research by Ry Walker Research • methodology

Sources