← Back to research
•·26 min read·industry

Foundation Lab Coding Agents

Compare 16 selected foundation lab coding agents across local and cloud workflows, pricing, licensing, and security, including ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder.

Key takeaways

  • The 16 evaluated products add ZCode, IBM Bob, CodeBuddy, TRAE/Trae Agent, and Qoder; this remains a selected comparison, not a complete market census.
  • Claude Pro includes Claude Code, while current Gemini CLI access documentation conflicts with Google's older transition announcement.
  • First-party clients can serve other labs' models, and a terminal interface can control remote execution.
  • Compare inference, compute, account terms, and approval boundaries separately; a seat price or sandbox label does not settle the operating cost or risk.

FAQ

What is a foundation lab coding agent?

This report uses the term for a publicly usable coding agent with a downloadable client, developed within a company or corporate group that also develops foundation models. The client need not use only that company's models.

Which foundation lab coding agent is best?

There is no independently measured winner in this report. Compare representative repository tasks, execution location, account eligibility, billed usage, and the work a maintainer must do to accept the result.

Does Claude Pro include Claude Code?

Yes: Anthropic's pricing and installation documentation checked September 16, 2026 include Claude Code with Pro. The previous claim that it had left the plan was incorrect.

Has Gemini CLI shut down for individuals?

Current authentication and quota docs describe individual access, despite Google's May announcement of a June 18 sunset. The conflicting documentation was not resolved through account testing, so verify your login and quota instead of assuming either universal shutdown or a confirmed rollback.

Executive Summary

Foundation lab coding agents are no longer a choice between one terminal and one proprietary model. The useful comparison is between the agent that plans and edits, the environment that executes its commands, and the account that pays for inference. A lab-built client can offer third-party models; a terminal client can delegate work to a cloud machine. Qwen Code, Grok Build, Kiro, and GitHub Copilot all demonstrate why company ownership alone does not tell you which model handles a task.[1][2][3][4]

This report compares 16 selected products, including platform-led Kiro and GitHub Copilot, the remotely executed Jules service, and the September 23, 2026 expansion of ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder. It is an evaluated set, not a complete market census. It covers local iteration, cloud delegation, source availability, access, and operational limits. It does not rank model intelligence or claim an independently measured productivity winner.

Checked September 23, 2026. September 16 corrections still stand: Claude Pro currently includes Claude Code; and current Gemini CLI documentation advertises individual access despite Google's older announcement that it would end. The Google conflict is explained below rather than treated as a verified service rollback. The new members are individual profiles, not a claim that every regional SKU or research CLI is a drop-in replacement for Claude Code.[5][6][7][8]


Market Definition and Coverage

The inclusion criteria are a publicly usable coding agent, a downloadable developer client, and first-party development by a company that also develops foundation models. The company relationship can be direct or through a product organization in the same corporate group. A proprietary model is not required for every request, and that relationship is not evidence of superior quality.

This includes two platform-led entrants previously excluded incorrectly. Kiro says it is built and operated within AWS; Amazon develops Nova models, but Kiro's published model choices include other providers. Microsoft explicitly describes training MAI-Code within the Copilot harness, while Copilot offers a multi-provider model catalog. Neither Amazon nor Microsoft can accurately be excluded as a company that does not develop models.[9][10][3][11][4]

Jules qualifies even though its CLI controls hosted sessions: the downloadable-client criterion does not require local execution. Google has three entries because Gemini CLI, Antigravity CLI, and Jules have distinct documented interfaces and operating models; their counts do not represent three different labs.[12][13][14]

This expansion closes the September 16 coverage follow-up for five named products. Individual profiles now exist for ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder. The set is still not a census: other labs, regional SKUs, and acquired IDEs remain out of scope unless they meet the criteria and are added later.[15][16][17][18][19][20]

Other products, including Cursor, remain useful alternatives in a broader buying decision. Corporate origin and current ownership are separate questions; this selected set is not a census of acquired products either. See AI Coding Assistants and Cloud Coding Agent Platforms for broader coverage.


Market Map

“Local” below describes code execution, not an assurance that prompts or files stay on the machine. Cloud model calls, telemetry, connected tools, and account policies require separate review.

ProductDeveloperMain coding workflowImportant distinction
Claude CodeAnthropicTerminal, desktop, IDE, and delegated cloud sessionsRemote Control drives a local session; web tasks run remotely[21]
CodexOpenAILocal CLI/editor work, review, and cloud handoffAPI-key usage and ChatGPT sign-in have different entitlements[22][23]
Gemini CLIGoogleOpen-source terminal agent with file, shell, search, and MCP toolsCurrent individual-access docs conflict with the May transition announcement[14][6]
Antigravity CLIGoogleTerminal interface sharing the Antigravity agent coreAccount authentication or an explicitly configured Gemini API-key route[13][24]
JulesGoogleDelegate a repository task to a short-lived cloud VMJules Tools manages remote sessions; it does not install a local agent runtime[25][12]
Grok BuildxAI / SpaceXAILocal terminal agent, scripting, and ACP clientsSeparate from the Grok web/mobile application builder with the same name[2][26]
Qwen CodeAlibaba / QwenCLI, desktop, editor, and self-run web interfaceMulti-provider authentication; experimental web serving is not a managed cloud service[27][1]
Kimi CodeMoonshot AITerminal agent and editor integrationOfficial account benefits and Open Platform API access are separate paths[28][29]
Mistral VibeMistral AIVibe Code CLI, VS Code, or web sessionsVibe is the wider product; Code is its coding mode[30][31]
KiroAWS / AmazonSpecs, local IDE/CLI, and managed cloud sessionsAWS origin does not mean Nova-only execution[9][32][3]
GitHub CopilotGitHub / MicrosoftIDE and CLI assistance plus Actions-based cloud tasksMicrosoft's own MAI models coexist with other labs' models[33][34][11]
ZCodeZ.ai / ZhipuDesktop ADE; GitHub also documents a CLI/TUIFirst-party ZCode Agent plus GLM Coding Plan; September 2026 snapshot-upload incident[15]
IBM BobIBMDesktop IDE and Bob ShellLocal execution, SaaS inference; remote hosted agents were undated in V2[16]
CodeBuddyTencent CloudPlugin, IDE, and CLI with local web UIDistinct from WorkBuddy; Hy4 preview was time-boxed[17]
TRAEByteDanceCommercial TraeCode IDE vs MIT Trae Agent CLIDo not copy the research CLI license onto TraeCode[18]
QoderBright Zenith / Alibaba Cloud CNDesktop, IDE, CLI, cloud agentsInternational Qoder and Qoder CN are regional siblings, not one SKU[19][20]

Product Assessments

Claude Code: one engine, different execution locations

Anthropic's strongest practical distinction is continuity across terminal, desktop, editor, and cloud surfaces. A developer can steer a local session from another device or delegate a task that continues in Anthropic-managed infrastructure. Those are different deployment decisions. Team and Enterprise customers also have a documented self-hosted environment route for cloud sessions; that does not make every consumer session self-hosted.[21]

Claude Code remains available with Pro, Max, Team, Enterprise, or Console access. The CLI also supports documented provider routes through Bedrock, Google's Agent Platform, and Microsoft Foundry. The previous claim that the $20 Pro plan had lost Claude Code is not supported by the current pricing and installation documentation.[35][5]

Choose it when: the Claude workflow and its local/cloud handoff fit the team. Check first: whether the chosen surface, provider, and plan expose the specific capability required.

Codex: local tooling and delegated work

OpenAI's CLI can inspect and edit a checkout, run commands, review changes against a commit or branch, invoke skills and MCP servers, and run noninteractively through codex exec. It also exposes cloud-task handoff. Native installers cover macOS, Linux, and Windows; “Codex” should not be reduced to its earliest cloud-only presentation.[22]

Account choice matters. ChatGPT plans bundle Codex access, while API-key usage is billed separately and does not include cloud features such as GitHub review or Slack integration. Current model guidance points ChatGPT-sign-in users toward GPT-5.6 Sol and dates GPT-5.5 retirement from ChatGPT, Work, and Codex to October 14, 2026; the API is excluded from that retirement. Avoid hard-coding the older model as the permanent default.[23][36]

Choose it when: terminal automation and moving between local work and delegated tasks are important. Check first: account entitlements, configured permissions, and model settings in saved automations.

Gemini CLI: open source, with an unresolved access-documentation conflict

Gemini CLI remains a published Apache-2.0 terminal agent with file operations, shell execution, Google Search grounding, and MCP integration. Current authentication documentation describes personal Google sign-in, Gemini API keys, and Vertex AI. Its quota page lists 1,000 daily model requests for the individual Google-account tier, with higher Pro and Ultra allowances.[14][6][7]

The conflict is material: Google's May 19 announcement still says free individual, Pro, and Ultra Gemini CLI serving would end June 18. The authentication page updated August 17 and the current quota page describe those access paths again. This review did not test every account or establish a dated reversal announcement. Treat current documentation as the published setup path, verify your actual account, and do not repeat an unconditional “Gemini CLI is shut down” claim.[8][6][7]

Choose it when: an inspectable Google-oriented terminal client matters. Check first: login eligibility and live quota before committing a workflow to an advertised free allowance.

Antigravity CLI: Google's shared agent platform in the terminal

Antigravity CLI shares its agent core, settings, and permissions with Antigravity 2.0. It supports terminal work, SSH use, and headless workflows; a one-time migration path imports Gemini CLI configuration. Background coordination does not by itself mean commands execute on a cloud VM.[13]

Current installation docs cover macOS, Linux, and Windows. They also document direct Gemini API-key use, but exporting GEMINI_API_KEY alone is insufficient: the CLI's configuration must select the Gemini model provider. This is a meaningful alternative to browser sign-in for CI. Permission rules distinguish file access, shell commands, URLs, MCP tools, and unsandboxed execution.[24][37]

Choose it when: a common Antigravity configuration across graphical and terminal work is valuable. Check first: the authentication route and permission defaults, especially workspace writes that are automatically allowed.

Jules: a remote task service with a terminal control surface

Jules prepares a short-lived Ubuntu VM, clones a repository, installs dependencies, and works on the delegated task. A validated setup script can produce an environment snapshot for later tasks. Jules Tools creates and monitors remote sessions and pulls changes back into a local workflow; its installation does not move that execution onto the developer's laptop.[25][12]

Plan review is not an unconditional approval barrier. The planning guide says a plan can auto-approve after a timer when the user leaves. Teams requiring an explicit human decision should verify that configuration rather than infer it from a “plan first” description. Higher task allowances are tied to eligible individual Google accounts, an important purchasing constraint for company adoption.[38][39]

Choose it when: bounded GitHub tasks can run away from the workstation. Check first: reproducible environment setup, review behavior, and account eligibility.

Grok Build: distinguish the CLI from the application builder

The current CLI supports interactive terminal use, headless scripts, ACP clients, and custom model endpoints. Its documentation identifies Grok 4.6 as the model powering Build, but the harness is not limited to that endpoint. Browser login and API-key authentication are separate setup options.[2]

The May CLI announcement offered an early beta to SuperGrok and X Premium Plus subscribers. The August 19 “for everyone” expansion explicitly describes generating and publishing applications through Grok on web, iOS, and Android. It is not evidence that every free account receives the terminal CLI's subscription entitlement.[40][26]

The security distinction is especially concrete: the sandbox is off by default, and Plan Mode blocks ordinary edit tools but does not prevent shell writes or impose the same edit gate on subagents. Neither a plan nor a worktree should be sold as an isolation guarantee.[41][42]

Choose it when: Grok-oriented terminal work or its configurable model interface fits. Check first: CLI access and the sandbox profile, independently of the chat builder's availability.

Qwen Code: first-party does not mean single-provider

Qwen Code offers a local CLI, desktop app, editor integrations, headless operation, and an experimental qwen serve web/daemon interface. Its extensibility includes skills, subagents, and MCP. The self-run web interface should not be counted as a vendor-operated asynchronous task service.[27]

The current authentication menu supports Alibaba ModelStudio, built-in third-party providers, and custom compatible endpoints. ModelStudio separates an individual Coding Plan from a usage-based Token Plan for teams and standard API access. The repository explicitly dates the former free Qwen OAuth tier's discontinuation to April 15, 2026; this particular free-tier removal remains supported, unlike the previous article's blanket pricing narrative.[1]

Choose it when: modifying an open harness or switching provider protocols is important. Check first: the exact endpoint, regional plan, and sandbox configuration. Its container mode mounts both the workspace and ~/.qwen, so the authentication directory remains part of the exposure analysis.[43]

Kimi Code: a configurable client with account-specific benefits

The current MoonshotAI/kimi-code repository documents a single-binary terminal client, hooks, plugins, MCP configuration, parallel subagents, and ACP integration. It runs on macOS, Linux, and Windows, with Git Bash required for the Windows shell. The client can use compatible providers as well as Moonshot's own models.[28]

Kimi membership also supplies a coding service usable from official clients and supported third-party tools. Its account benefits have weekly and rolling five-hour controls, shared device usage, and a monthly membership credit pool. The benefits page explicitly limits that subscription benefit to personal development and directs enterprise needs to the Open Platform. That restriction is distinct from the MIT license on the client.[29][44][45]

Choose it when: the Kimi account workflow or configurable client is useful. Check first: the relevant commercial access path; do not assume an individual membership is an enterprise inference agreement.

Mistral Vibe: Code is now one mode of a broader product

Mistral's August rebrand places Work, Chat, and Code under Vibe. Vibe Code runs locally through its CLI or VS Code extension and remotely through web sessions. This preserves the useful terminal coding client while adding a deployment decision that the older terminal-only profile missed.[31][30]

The open-source CLI supports configurable models and providers, skills, hooks, subagents, and MCP. Its current default accept-edits agent automatically permits file edits; the ask, plan, and auto-approve profiles differ. The repository targets Unix environments and says Windows works but is not its official support target. These specifics matter more than a generic “human approves every action” label.[46]

Choose it when: the Mistral ecosystem and inspectable CLI fit the organization. Check first: local versus web execution, OS support, and the shared subscription usage settings.

Kiro: specs and multiple models inside an AWS-built workflow

Kiro's spec workflow turns a request into requirements, design, and implementation tasks; hooks support recurring development actions. Its current product family includes local IDE and CLI use alongside cloud sessions. Those cloud sessions persist after a disconnect, and the web workflow can coordinate GitHub and GitLab repositories.[9][32]

Kiro is included because of who builds it, not because Nova is its exclusive or preferred coding model. Published choices span OpenAI, Anthropic, and open-weight providers, alongside Auto routing. Its commercial IDE/CLI license is separate from the licenses of bundled open-source components.[3][47]

Choose it when: explicit specifications and an AWS-managed adoption path are useful. Check first: enterprise versus social-login data settings and cloud availability. The FAQ says IAM Identity Center organizations must enable cloud sessions and notes limits on regional availability and cloud governance controls.[32]

GitHub Copilot: a platform-led member with distinct local and cloud agents

GitHub Copilot's own CLI supports repository work, custom instructions, MCP, hooks, skills, and specialized agents. Its cloud agent instead runs tasks in an ephemeral GitHub Actions environment. Microsoft's MAI-Code work gives this product a real model-development relationship, but it does not make Copilot's other supported models Microsoft-trained.[33][34][11]

The cloud workflow has concrete limits: one repository and one branch per task, at most one PR, and a maximum 59-minute session. It can research, plan, and iterate before a PR is created; treating every session as an immediately opened PR misses the current workflow.[34]

Choose it when: GitHub-native issue, branch, and review workflows are the center of development. Check first: the target surface's model policy and billing; the CLI, editor, and cloud service should not be assigned one undifferentiated execution boundary.

ZCode: first-party GLM ADE, local trust boundary

ZCode is Z.ai's self-developed agent inside a desktop ADE, with GLM-5.3 as the default family and optional third-party endpoints. Install docs ship macOS, Windows, and Linux desktop builds; the public repository also documents a CLI/TUI. Execution is local or SSH/WSL/Docker, with permission modes rather than a VM sandbox. The GLM Coding Plan (from $18/month in current overview docs) is a separate subscription that can also feed Claude Code or OpenCode; that is not the same as running ZCode Agent.[15]

Choose it when: a GLM-native desktop harness is the requirement. Check first: client version against the September 2026 workspace-snapshot incident, and whether you are buying the ADE or only a GLM plan for another CLI.

IBM Bob: enterprise IDE/Shell, not a shipped remote agent

IBM Bob is generally available as a desktop IDE and Bob Shell, with IBM SaaS inference and Bobcoin metering. Launch materials describe Granite plus Claude and Mistral routing. V2 documented Ask/Plan/Agent modes and approvals. Remote hosted-agent execution was an undated "further out" item on June 24, 2026. Individual list prices on bob.ibm.com are $20 / $60 / $200 per month.[16]

Choose it when: IBM i, Z, or governed Java modernization is in scope. Check first: local execution plus Bobcoin burn, not a cloud-task SLA.

CodeBuddy: Tencent coding client, not WorkBuddy

CodeBuddy is Tencent Cloud's coding agent with plugin, IDE, and CLI surfaces, including a localhost web UI. Keep it distinct from WorkBuddy. Hunyuan Hy4's August 28 two-week preview is not list pricing. International USD cards were not verified live in this expansion.[17]

Choose it when: Tencent/Hunyuan first-party coding surfaces matter. Check first: region, SKU, and whether any model preview is still running.

TRAE: commercial IDE versus MIT research CLI

TRAE on trae.ai is TraeCode/TraeWork. Trae Agent on GitHub is a MIT Python CLI with BYO keys and optional Docker. The repo links to trae.ai; that is not proof the IDE is that codebase. Do not assign MIT or Docker to TraeCode.[18]

Choose it when: you either want ByteDance's closed IDE or an inspectable research CLI, and you know which. Check first: live commercial terms; they were JS-gated in earlier audits.

Qoder: Bright Zenith internationally, Alibaba Cloud in China

Qoder on qoder.com is Bright Zenith in Singapore. Qoder CN is Alibaba Cloud's renamed Tongyi Lingma. Similar desktop/IDE/CLI names; separate legal entities, invoices, and data regions.[19][20]

Choose it when: you are buying a specific region's catalog. Check first: operator and fapiao/USD path. For an Alibaba open-source CLI, start with Qwen Code.


Pricing: Compare the Meter, Not Just the Seat

Published entry points below were checked September 16, 2026, with ZCode, Bob, CodeBuddy, TRAE, and Qoder prices dated September 23, 2026. USD prices are monthly unless stated otherwise; regional availability, taxes, and account terms can differ. A request, task, credit, and token are different units.

ProductPublished access or priceWhat determines sustained usage
Claude CodePro $20/month, or $200 billed annually; Max from $100/monthPlan limits; Console/API usage is a different billing route[5][35]
CodexFree and Go access; Plus $20/month; Pro from $100/monthTask/model/context usage; separately metered API keys lack bundled cloud features[23]
Gemini CLICurrent docs list free individual access and Pro/Ultra, plus API/Vertex routesDaily model requests or API usage; verify the documentation conflict described above[7][6]
Antigravity CLIPlatform lists $0 individual access, Pro/Ultra upgrades, and organization routesRate limits/AI credits or configured API consumption; models vary by route[48][24]
JulesFree: 15 tasks/rolling day, three concurrent; Pro: 100/15; Ultra: 300/60Task and concurrency allowances, rather than an unlimited background worker[39]
Grok BuildCLI launch: SuperGrok/X Premium Plus; current CLI also accepts an API keySubscription entitlement or API use; free web/mobile builder access does not settle CLI eligibility[40][2][26]
Qwen CodeFree client; no continuing free Qwen OAuth tierIndividual Coding Plan, team Token Plan, standard API, or selected provider costs[1]
Kimi CodePublished membership tiers: RMB ¥49/99/199/699 per monthShared monthly credits, weekly/five-hour controls, optional Extra Usage; enterprise path differs[49][44]
Mistral VibeFree limited coding; Pro $14.99/month; Team $24.99/user/monthIncluded usage and optional per-token PAYG across Vibe, Studio, and API[50][51]
KiroFree 50 credits; Pro $20/1,000; Pro+ $40/2,000; Pro Max $100/5,000; Power $200/10,000Model/task-weighted credits; paid add-ons or enabled enterprise overages at $0.04/credit[52]
GitHub CopilotFree; Pro $10; Pro+ $39; Max $100; Business $19/seat; Enterprise $39/seatToken-based AI credits and GitHub Actions minutes for cloud tasks, rather than unlimited agent work[4][53][34]
ZCodeApp free; GLM Coding Plan from $18/month; 5-day trial with conflicting daily token figuresCredits with 5-hour and weekly caps; China BigModel CNY plans are a different checkout
IBM BobTrial 50 Bobcoins / 30 days; Pro $20; Pro+ $60; Ultra $200Bobcoins, not tokens; IBM product page says "per instance" for the same dollars
CodeBuddyRegional catalogs; Hy4 two-week preview expiredDo not use WorkBuddy promotions or expired Hy4 access as list price[17]
TRAECommercial catalog JS-rendered; not reprinted hereTrae Agent CLI is BYO keys, not a TraeCode seat
QoderInternational credit plans documented around $20 / $60 / $200; CN is CNYSeparate Bright Zenith vs Alibaba Cloud billing[19][20]

The maintainable budget is a workload budget: subscription or provider inference, execution infrastructure, retries, and human review. GitHub explicitly converts model-token spending into AI credits at $0.01 per credit; Kiro describes credits as task-dependent units affected by model choice. Comparing their credit counts directly would be misleading.[53][3]

One first-hand illustration of the purchasing problem is Reddit user Old-Glove9438, who described confusion between Mistral's displayed API, Vibe, and PAYG allowances in a thread marked one month old when reviewed on September 16, 2026. Subsequent replies included a claimed interface update and further disagreement about charges. That is an individual usability report, not proof of an overbilling defect. Use the current billing documentation and the actual account dashboard, not a forum screenshot, to establish entitlements.[54][51]

Source Availability and Security Boundaries

The Codex, Gemini CLI, Qwen Code, and Mistral Vibe repositories carry Apache-2.0 licenses; Kimi Code carries MIT. Those licenses apply to the software, not free hosted inference, every model's weights, or the associated cloud service. Claude Code's repository instead points to Anthropic's commercial terms; Kiro licenses its IDE and CLI as AWS Content. ZCode published an Apache-2.0 tree after the snapshot incident; that grant does not automatically cover every desktop binary. Trae Agent is MIT; TraeCode is commercial. IBM Bob and CodeBuddy are commercial SaaS/clients.[55][56][57][58][45][59][47][15][18]

A useful security review separates four controls:

ControlWhat it establishesWhat to verify
Tool approvalWhether a requested action may startAutomatic modes, remembered grants, hooks, and child-agent behavior
Filesystem/process isolationWhat an approved command can accessWritable mounts, sensitive files, sandbox availability, and escape approvals
Network policyWhich services commands or tools may contactShell-child egress versus in-process API and MCP traffic
Change reviewWhether generated work is accepted into the productTests, diff review, repository checks, and merge permissions

These are analysis categories, not a certification score. Current examples show why a single sandbox checkbox is insufficient:

  • Codex: default local permission mode applies sandboxing to spawned commands; approvals are a separate control. Its implementation differs across macOS, Linux/WSL2, and native Windows.[60]
  • Claude Code: the Bash sandbox covers commands and children on macOS, Linux, and WSL2, with configurable unsandboxed fallback. The docs warn that an unavailable sandbox can fall back to unsandboxed execution unless failIfUnavailable is enforced.[61]
  • Qwen Code: macOS Seatbelt and Docker/Podman have different dependencies and access profiles; the default Seatbelt profile allows outbound networking.[43]
  • Grok Build: sandboxing is opt-in; child-network restrictions are Linux-only, and model/API web traffic is a separate path. Its planning gate also leaves shell and subagent exceptions.[41][42]
  • Kimi Code: “Always Ask,” “Ask When Needed,” and “Never Ask” are materially different. Never Ask also automatically handles sensitive-file and plan-exit approvals; an unattended mode is not an isolation mechanism.[62]

For detailed runtime choices, see Local Agent Sandboxes. A separate Git worktree avoids mixing ordinary edits between tasks, but still needs an appropriate execution boundary.


How to Choose and Evaluate

The following shortlist is editorial judgment based on the documented workflows, not a benchmark ranking.

Starting requirementEvaluate firstWhy
Work in an existing terminal/editor workflowClaude Code, Codex, Gemini CLI, Grok BuildNative repository interaction; compare actual account and permission behavior
Modify or inspect the agent clientCodex, Gemini CLI, Qwen Code, Kimi Code, Mistral VibePublished permissive client licenses
Switch among provider APIs in one local harnessQwen Code, Grok Build, Kimi CodeDocumented configurable provider paths
Delegate a bounded GitHub taskJules or Copilot cloud agent; also evaluate Claude/Codex cloud routesRemote execution with reviewable changes; setup and surface limits differ
Turn requirements into structured engineering workKiroSpecs are a central product workflow
Stay within an existing Mistral account and tooling setupVibe CodeLocal/editor and remote coding surfaces with shared billing controls
GLM-native desktop ADEZCodeFirst-party ZCode Agent; verify client version after the snapshot incident
IBM i / Z / governed JavaIBM BobLocal IDE/Shell; not a dated remote-agent runtime
Tencent/Hunyuan coding clientCodeBuddyDistinct from WorkBuddy; verify region
ByteDance research CLI vs commercial IDETrae Agent vs TraeCodeDifferent licenses and runtimes
Regional Qoder purchaseQoder or Qoder CNBright Zenith vs Alibaba Cloud; not Qwen Code
Run different labs' agents through a team operating layerTembo, alongside the relevant lab clientsCommon execution, triggers, and review workflow; see the scope distinction below

A practical evaluation should use representative repository work, not a toy prompt. For example, choose a real dependency upgrade with a failing regression test. Prepare a clean branch, reproducible dependency installation, the expected test command, and a concrete acceptance condition. Give each candidate the same starting commit and available credentials.

Record the result at each stage: could it prepare the environment, identify the failing behavior, implement a bounded fix, run the relevant tests, and produce a diff a maintainer accepts? Keep failed runs and manual interventions. Measure wall time, billed usage, and review time separately; a faster generation with expensive rework is not automatically the better tool.

For a remote candidate, also test a disconnect and return, a missing secret, a denied network request, and a task that outlasts its allowed runtime. For a local candidate, confirm which processes and files its configured sandbox actually covers. This is a proposed evaluation procedure; no comparative execution benchmark was run for this report.

Tembo: an Adjacent Operating Platform

Disclosure: Ry Walker is Tembo's co-founder and CEO. Tembo is relevant when the buying question moves from “which lab client?” to “how does the team run and supervise these clients?” It does not meet this report's model-developer criterion and is therefore adjacent rather than an additional lab member.

Tembo's documented workflow runs agents in cloud environments, accepts work from sources such as Slack, Linear, GitHub, schedules, and webhooks, and returns output for review. The September 15 sandbox image inventory explicitly includes Claude Code, Codex, Gemini CLI, and Grok, while warning that installation does not guarantee every agent is enabled in every session. ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder were not in that inventory.[63][64]

That creates a concrete choice: run a client's native local/cloud workflow, or use a common operating layer for agent selection, execution, and team visibility. Tembo documents workspace defaults, per-session overrides, provider credentials, shared MCP configuration, and administrative model availability. It is useful to evaluate for recurring maintenance across teams; it does not remove the need to validate each agent's configuration or provider terms.[65]

Costs also remain distinct. Tembo's dollar allowance covers gateway inference and VM compute; its pricing page says BYOK and ChatGPT/Codex OAuth inference are not charged through Tembo, while their VM compute still consumes allowance. That statement does not grant portability to every lab subscription. Kiro, for example, explicitly restricts routing subscription requests outside its native interfaces through third-party automation harnesses.[66][52]

Assessment

First-party development is a useful category boundary, not a quality ranking. The decision should follow the repository workflow, execution location, account terms, and controls that a team can actually operate. Keep the client choice separable from inference and deployment where the product supports it, and recheck the account-specific details before standardizing a team on any advertised allowance.


Research by Ry Walker Research • methodology

Sources