Key takeaways
- The 16 evaluated products add ZCode, IBM Bob, CodeBuddy, TRAE/Trae Agent, and Qoder; this remains a selected comparison, not a complete market census.
- Claude Pro includes Claude Code, while current Gemini CLI access documentation conflicts with Google's older transition announcement.
- First-party clients can serve other labs' models, and a terminal interface can control remote execution.
- Compare inference, compute, account terms, and approval boundaries separately; a seat price or sandbox label does not settle the operating cost or risk.
FAQ
What is a foundation lab coding agent?
This report uses the term for a publicly usable coding agent with a downloadable client, developed within a company or corporate group that also develops foundation models. The client need not use only that company's models.
Which foundation lab coding agent is best?
There is no independently measured winner in this report. Compare representative repository tasks, execution location, account eligibility, billed usage, and the work a maintainer must do to accept the result.
Does Claude Pro include Claude Code?
Yes: Anthropic's pricing and installation documentation checked September 16, 2026 include Claude Code with Pro. The previous claim that it had left the plan was incorrect.
Has Gemini CLI shut down for individuals?
Current authentication and quota docs describe individual access, despite Google's May announcement of a June 18 sunset. The conflicting documentation was not resolved through account testing, so verify your login and quota instead of assuming either universal shutdown or a confirmed rollback.
Executive Summary
Foundation lab coding agents are no longer a choice between one terminal and one proprietary model. The useful comparison is between the agent that plans and edits, the environment that executes its commands, and the account that pays for inference. A lab-built client can offer third-party models; a terminal client can delegate work to a cloud machine. Qwen Code, Grok Build, Kiro, and GitHub Copilot all demonstrate why company ownership alone does not tell you which model handles a task.[1][2][3][4]
This report compares 16 selected products, including platform-led Kiro and GitHub Copilot, the remotely executed Jules service, and the September 23, 2026 expansion of ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder. It is an evaluated set, not a complete market census. It covers local iteration, cloud delegation, source availability, access, and operational limits. It does not rank model intelligence or claim an independently measured productivity winner.
Checked September 23, 2026. September 16 corrections still stand: Claude Pro currently includes Claude Code; and current Gemini CLI documentation advertises individual access despite Google's older announcement that it would end. The Google conflict is explained below rather than treated as a verified service rollback. The new members are individual profiles, not a claim that every regional SKU or research CLI is a drop-in replacement for Claude Code.[5][6][7][8]
Market Definition and Coverage
The inclusion criteria are a publicly usable coding agent, a downloadable developer client, and first-party development by a company that also develops foundation models. The company relationship can be direct or through a product organization in the same corporate group. A proprietary model is not required for every request, and that relationship is not evidence of superior quality.
This includes two platform-led entrants previously excluded incorrectly. Kiro says it is built and operated within AWS; Amazon develops Nova models, but Kiro's published model choices include other providers. Microsoft explicitly describes training MAI-Code within the Copilot harness, while Copilot offers a multi-provider model catalog. Neither Amazon nor Microsoft can accurately be excluded as a company that does not develop models.[9][10][3][11][4]
Jules qualifies even though its CLI controls hosted sessions: the downloadable-client criterion does not require local execution. Google has three entries because Gemini CLI, Antigravity CLI, and Jules have distinct documented interfaces and operating models; their counts do not represent three different labs.[12][13][14]
This expansion closes the September 16 coverage follow-up for five named products. Individual profiles now exist for ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder. The set is still not a census: other labs, regional SKUs, and acquired IDEs remain out of scope unless they meet the criteria and are added later.[15][16][17][18][19][20]
Other products, including Cursor, remain useful alternatives in a broader buying decision. Corporate origin and current ownership are separate questions; this selected set is not a census of acquired products either. See AI Coding Assistants and Cloud Coding Agent Platforms for broader coverage.
Market Map
“Local” below describes code execution, not an assurance that prompts or files stay on the machine. Cloud model calls, telemetry, connected tools, and account policies require separate review.
| Product | Developer | Main coding workflow | Important distinction |
|---|---|---|---|
| Claude Code | Anthropic | Terminal, desktop, IDE, and delegated cloud sessions | Remote Control drives a local session; web tasks run remotely[21] |
| Codex | OpenAI | Local CLI/editor work, review, and cloud handoff | API-key usage and ChatGPT sign-in have different entitlements[22][23] |
| Gemini CLI | Open-source terminal agent with file, shell, search, and MCP tools | Current individual-access docs conflict with the May transition announcement[14][6] | |
| Antigravity CLI | Terminal interface sharing the Antigravity agent core | Account authentication or an explicitly configured Gemini API-key route[13][24] | |
| Jules | Delegate a repository task to a short-lived cloud VM | Jules Tools manages remote sessions; it does not install a local agent runtime[25][12] | |
| Grok Build | xAI / SpaceXAI | Local terminal agent, scripting, and ACP clients | Separate from the Grok web/mobile application builder with the same name[2][26] |
| Qwen Code | Alibaba / Qwen | CLI, desktop, editor, and self-run web interface | Multi-provider authentication; experimental web serving is not a managed cloud service[27][1] |
| Kimi Code | Moonshot AI | Terminal agent and editor integration | Official account benefits and Open Platform API access are separate paths[28][29] |
| Mistral Vibe | Mistral AI | Vibe Code CLI, VS Code, or web sessions | Vibe is the wider product; Code is its coding mode[30][31] |
| Kiro | AWS / Amazon | Specs, local IDE/CLI, and managed cloud sessions | AWS origin does not mean Nova-only execution[9][32][3] |
| GitHub Copilot | GitHub / Microsoft | IDE and CLI assistance plus Actions-based cloud tasks | Microsoft's own MAI models coexist with other labs' models[33][34][11] |
| ZCode | Z.ai / Zhipu | Desktop ADE; GitHub also documents a CLI/TUI | First-party ZCode Agent plus GLM Coding Plan; September 2026 snapshot-upload incident[15] |
| IBM Bob | IBM | Desktop IDE and Bob Shell | Local execution, SaaS inference; remote hosted agents were undated in V2[16] |
| CodeBuddy | Tencent Cloud | Plugin, IDE, and CLI with local web UI | Distinct from WorkBuddy; Hy4 preview was time-boxed[17] |
| TRAE | ByteDance | Commercial TraeCode IDE vs MIT Trae Agent CLI | Do not copy the research CLI license onto TraeCode[18] |
| Qoder | Bright Zenith / Alibaba Cloud CN | Desktop, IDE, CLI, cloud agents | International Qoder and Qoder CN are regional siblings, not one SKU[19][20] |
Product Assessments
Claude Code: one engine, different execution locations
Anthropic's strongest practical distinction is continuity across terminal, desktop, editor, and cloud surfaces. A developer can steer a local session from another device or delegate a task that continues in Anthropic-managed infrastructure. Those are different deployment decisions. Team and Enterprise customers also have a documented self-hosted environment route for cloud sessions; that does not make every consumer session self-hosted.[21]
Claude Code remains available with Pro, Max, Team, Enterprise, or Console access. The CLI also supports documented provider routes through Bedrock, Google's Agent Platform, and Microsoft Foundry. The previous claim that the $20 Pro plan had lost Claude Code is not supported by the current pricing and installation documentation.[35][5]
Choose it when: the Claude workflow and its local/cloud handoff fit the team. Check first: whether the chosen surface, provider, and plan expose the specific capability required.
Codex: local tooling and delegated work
OpenAI's CLI can inspect and edit a checkout, run commands, review changes against a commit or branch, invoke skills and MCP servers, and run noninteractively through codex exec. It also exposes cloud-task handoff. Native installers cover macOS, Linux, and Windows; “Codex” should not be reduced to its earliest cloud-only presentation.[22]
Account choice matters. ChatGPT plans bundle Codex access, while API-key usage is billed separately and does not include cloud features such as GitHub review or Slack integration. Current model guidance points ChatGPT-sign-in users toward GPT-5.6 Sol and dates GPT-5.5 retirement from ChatGPT, Work, and Codex to October 14, 2026; the API is excluded from that retirement. Avoid hard-coding the older model as the permanent default.[23][36]
Choose it when: terminal automation and moving between local work and delegated tasks are important. Check first: account entitlements, configured permissions, and model settings in saved automations.
Gemini CLI: open source, with an unresolved access-documentation conflict
Gemini CLI remains a published Apache-2.0 terminal agent with file operations, shell execution, Google Search grounding, and MCP integration. Current authentication documentation describes personal Google sign-in, Gemini API keys, and Vertex AI. Its quota page lists 1,000 daily model requests for the individual Google-account tier, with higher Pro and Ultra allowances.[14][6][7]
The conflict is material: Google's May 19 announcement still says free individual, Pro, and Ultra Gemini CLI serving would end June 18. The authentication page updated August 17 and the current quota page describe those access paths again. This review did not test every account or establish a dated reversal announcement. Treat current documentation as the published setup path, verify your actual account, and do not repeat an unconditional “Gemini CLI is shut down” claim.[8][6][7]
Choose it when: an inspectable Google-oriented terminal client matters. Check first: login eligibility and live quota before committing a workflow to an advertised free allowance.
Antigravity CLI: Google's shared agent platform in the terminal
Antigravity CLI shares its agent core, settings, and permissions with Antigravity 2.0. It supports terminal work, SSH use, and headless workflows; a one-time migration path imports Gemini CLI configuration. Background coordination does not by itself mean commands execute on a cloud VM.[13]
Current installation docs cover macOS, Linux, and Windows. They also document direct Gemini API-key use, but exporting GEMINI_API_KEY alone is insufficient: the CLI's configuration must select the Gemini model provider. This is a meaningful alternative to browser sign-in for CI. Permission rules distinguish file access, shell commands, URLs, MCP tools, and unsandboxed execution.[24][37]
Choose it when: a common Antigravity configuration across graphical and terminal work is valuable. Check first: the authentication route and permission defaults, especially workspace writes that are automatically allowed.
Jules: a remote task service with a terminal control surface
Jules prepares a short-lived Ubuntu VM, clones a repository, installs dependencies, and works on the delegated task. A validated setup script can produce an environment snapshot for later tasks. Jules Tools creates and monitors remote sessions and pulls changes back into a local workflow; its installation does not move that execution onto the developer's laptop.[25][12]
Plan review is not an unconditional approval barrier. The planning guide says a plan can auto-approve after a timer when the user leaves. Teams requiring an explicit human decision should verify that configuration rather than infer it from a “plan first” description. Higher task allowances are tied to eligible individual Google accounts, an important purchasing constraint for company adoption.[38][39]
Choose it when: bounded GitHub tasks can run away from the workstation. Check first: reproducible environment setup, review behavior, and account eligibility.
Grok Build: distinguish the CLI from the application builder
The current CLI supports interactive terminal use, headless scripts, ACP clients, and custom model endpoints. Its documentation identifies Grok 4.6 as the model powering Build, but the harness is not limited to that endpoint. Browser login and API-key authentication are separate setup options.[2]
The May CLI announcement offered an early beta to SuperGrok and X Premium Plus subscribers. The August 19 “for everyone” expansion explicitly describes generating and publishing applications through Grok on web, iOS, and Android. It is not evidence that every free account receives the terminal CLI's subscription entitlement.[40][26]
The security distinction is especially concrete: the sandbox is off by default, and Plan Mode blocks ordinary edit tools but does not prevent shell writes or impose the same edit gate on subagents. Neither a plan nor a worktree should be sold as an isolation guarantee.[41][42]
Choose it when: Grok-oriented terminal work or its configurable model interface fits. Check first: CLI access and the sandbox profile, independently of the chat builder's availability.
Qwen Code: first-party does not mean single-provider
Qwen Code offers a local CLI, desktop app, editor integrations, headless operation, and an experimental qwen serve web/daemon interface. Its extensibility includes skills, subagents, and MCP. The self-run web interface should not be counted as a vendor-operated asynchronous task service.[27]
The current authentication menu supports Alibaba ModelStudio, built-in third-party providers, and custom compatible endpoints. ModelStudio separates an individual Coding Plan from a usage-based Token Plan for teams and standard API access. The repository explicitly dates the former free Qwen OAuth tier's discontinuation to April 15, 2026; this particular free-tier removal remains supported, unlike the previous article's blanket pricing narrative.[1]
Choose it when: modifying an open harness or switching provider protocols is important. Check first: the exact endpoint, regional plan, and sandbox configuration. Its container mode mounts both the workspace and ~/.qwen, so the authentication directory remains part of the exposure analysis.[43]
Kimi Code: a configurable client with account-specific benefits
The current MoonshotAI/kimi-code repository documents a single-binary terminal client, hooks, plugins, MCP configuration, parallel subagents, and ACP integration. It runs on macOS, Linux, and Windows, with Git Bash required for the Windows shell. The client can use compatible providers as well as Moonshot's own models.[28]
Kimi membership also supplies a coding service usable from official clients and supported third-party tools. Its account benefits have weekly and rolling five-hour controls, shared device usage, and a monthly membership credit pool. The benefits page explicitly limits that subscription benefit to personal development and directs enterprise needs to the Open Platform. That restriction is distinct from the MIT license on the client.[29][44][45]
Choose it when: the Kimi account workflow or configurable client is useful. Check first: the relevant commercial access path; do not assume an individual membership is an enterprise inference agreement.
Mistral Vibe: Code is now one mode of a broader product
Mistral's August rebrand places Work, Chat, and Code under Vibe. Vibe Code runs locally through its CLI or VS Code extension and remotely through web sessions. This preserves the useful terminal coding client while adding a deployment decision that the older terminal-only profile missed.[31][30]
The open-source CLI supports configurable models and providers, skills, hooks, subagents, and MCP. Its current default accept-edits agent automatically permits file edits; the ask, plan, and auto-approve profiles differ. The repository targets Unix environments and says Windows works but is not its official support target. These specifics matter more than a generic “human approves every action” label.[46]
Choose it when: the Mistral ecosystem and inspectable CLI fit the organization. Check first: local versus web execution, OS support, and the shared subscription usage settings.
Kiro: specs and multiple models inside an AWS-built workflow
Kiro's spec workflow turns a request into requirements, design, and implementation tasks; hooks support recurring development actions. Its current product family includes local IDE and CLI use alongside cloud sessions. Those cloud sessions persist after a disconnect, and the web workflow can coordinate GitHub and GitLab repositories.[9][32]
Kiro is included because of who builds it, not because Nova is its exclusive or preferred coding model. Published choices span OpenAI, Anthropic, and open-weight providers, alongside Auto routing. Its commercial IDE/CLI license is separate from the licenses of bundled open-source components.[3][47]
Choose it when: explicit specifications and an AWS-managed adoption path are useful. Check first: enterprise versus social-login data settings and cloud availability. The FAQ says IAM Identity Center organizations must enable cloud sessions and notes limits on regional availability and cloud governance controls.[32]
GitHub Copilot: a platform-led member with distinct local and cloud agents
GitHub Copilot's own CLI supports repository work, custom instructions, MCP, hooks, skills, and specialized agents. Its cloud agent instead runs tasks in an ephemeral GitHub Actions environment. Microsoft's MAI-Code work gives this product a real model-development relationship, but it does not make Copilot's other supported models Microsoft-trained.[33][34][11]
The cloud workflow has concrete limits: one repository and one branch per task, at most one PR, and a maximum 59-minute session. It can research, plan, and iterate before a PR is created; treating every session as an immediately opened PR misses the current workflow.[34]
Choose it when: GitHub-native issue, branch, and review workflows are the center of development. Check first: the target surface's model policy and billing; the CLI, editor, and cloud service should not be assigned one undifferentiated execution boundary.
ZCode: first-party GLM ADE, local trust boundary
ZCode is Z.ai's self-developed agent inside a desktop ADE, with GLM-5.3 as the default family and optional third-party endpoints. Install docs ship macOS, Windows, and Linux desktop builds; the public repository also documents a CLI/TUI. Execution is local or SSH/WSL/Docker, with permission modes rather than a VM sandbox. The GLM Coding Plan (from $18/month in current overview docs) is a separate subscription that can also feed Claude Code or OpenCode; that is not the same as running ZCode Agent.[15]
Choose it when: a GLM-native desktop harness is the requirement. Check first: client version against the September 2026 workspace-snapshot incident, and whether you are buying the ADE or only a GLM plan for another CLI.
IBM Bob: enterprise IDE/Shell, not a shipped remote agent
IBM Bob is generally available as a desktop IDE and Bob Shell, with IBM SaaS inference and Bobcoin metering. Launch materials describe Granite plus Claude and Mistral routing. V2 documented Ask/Plan/Agent modes and approvals. Remote hosted-agent execution was an undated "further out" item on June 24, 2026. Individual list prices on bob.ibm.com are $20 / $60 / $200 per month.[16]
Choose it when: IBM i, Z, or governed Java modernization is in scope. Check first: local execution plus Bobcoin burn, not a cloud-task SLA.
CodeBuddy: Tencent coding client, not WorkBuddy
CodeBuddy is Tencent Cloud's coding agent with plugin, IDE, and CLI surfaces, including a localhost web UI. Keep it distinct from WorkBuddy. Hunyuan Hy4's August 28 two-week preview is not list pricing. International USD cards were not verified live in this expansion.[17]
Choose it when: Tencent/Hunyuan first-party coding surfaces matter. Check first: region, SKU, and whether any model preview is still running.
TRAE: commercial IDE versus MIT research CLI
TRAE on trae.ai is TraeCode/TraeWork. Trae Agent on GitHub is a MIT Python CLI with BYO keys and optional Docker. The repo links to trae.ai; that is not proof the IDE is that codebase. Do not assign MIT or Docker to TraeCode.[18]
Choose it when: you either want ByteDance's closed IDE or an inspectable research CLI, and you know which. Check first: live commercial terms; they were JS-gated in earlier audits.
Qoder: Bright Zenith internationally, Alibaba Cloud in China
Qoder on qoder.com is Bright Zenith in Singapore. Qoder CN is Alibaba Cloud's renamed Tongyi Lingma. Similar desktop/IDE/CLI names; separate legal entities, invoices, and data regions.[19][20]
Choose it when: you are buying a specific region's catalog. Check first: operator and fapiao/USD path. For an Alibaba open-source CLI, start with Qwen Code.
Pricing: Compare the Meter, Not Just the Seat
Published entry points below were checked September 16, 2026, with ZCode, Bob, CodeBuddy, TRAE, and Qoder prices dated September 23, 2026. USD prices are monthly unless stated otherwise; regional availability, taxes, and account terms can differ. A request, task, credit, and token are different units.
| Product | Published access or price | What determines sustained usage |
|---|---|---|
| Claude Code | Pro $20/month, or $200 billed annually; Max from $100/month | Plan limits; Console/API usage is a different billing route[5][35] |
| Codex | Free and Go access; Plus $20/month; Pro from $100/month | Task/model/context usage; separately metered API keys lack bundled cloud features[23] |
| Gemini CLI | Current docs list free individual access and Pro/Ultra, plus API/Vertex routes | Daily model requests or API usage; verify the documentation conflict described above[7][6] |
| Antigravity CLI | Platform lists $0 individual access, Pro/Ultra upgrades, and organization routes | Rate limits/AI credits or configured API consumption; models vary by route[48][24] |
| Jules | Free: 15 tasks/rolling day, three concurrent; Pro: 100/15; Ultra: 300/60 | Task and concurrency allowances, rather than an unlimited background worker[39] |
| Grok Build | CLI launch: SuperGrok/X Premium Plus; current CLI also accepts an API key | Subscription entitlement or API use; free web/mobile builder access does not settle CLI eligibility[40][2][26] |
| Qwen Code | Free client; no continuing free Qwen OAuth tier | Individual Coding Plan, team Token Plan, standard API, or selected provider costs[1] |
| Kimi Code | Published membership tiers: RMB ¥49/99/199/699 per month | Shared monthly credits, weekly/five-hour controls, optional Extra Usage; enterprise path differs[49][44] |
| Mistral Vibe | Free limited coding; Pro $14.99/month; Team $24.99/user/month | Included usage and optional per-token PAYG across Vibe, Studio, and API[50][51] |
| Kiro | Free 50 credits; Pro $20/1,000; Pro+ $40/2,000; Pro Max $100/5,000; Power $200/10,000 | Model/task-weighted credits; paid add-ons or enabled enterprise overages at $0.04/credit[52] |
| GitHub Copilot | Free; Pro $10; Pro+ $39; Max $100; Business $19/seat; Enterprise $39/seat | Token-based AI credits and GitHub Actions minutes for cloud tasks, rather than unlimited agent work[4][53][34] |
| ZCode | App free; GLM Coding Plan from $18/month; 5-day trial with conflicting daily token figures | Credits with 5-hour and weekly caps; China BigModel CNY plans are a different checkout |
| IBM Bob | Trial 50 Bobcoins / 30 days; Pro $20; Pro+ $60; Ultra $200 | Bobcoins, not tokens; IBM product page says "per instance" for the same dollars |
| CodeBuddy | Regional catalogs; Hy4 two-week preview expired | Do not use WorkBuddy promotions or expired Hy4 access as list price[17] |
| TRAE | Commercial catalog JS-rendered; not reprinted here | Trae Agent CLI is BYO keys, not a TraeCode seat |
| Qoder | International credit plans documented around $20 / $60 / $200; CN is CNY | Separate Bright Zenith vs Alibaba Cloud billing[19][20] |
The maintainable budget is a workload budget: subscription or provider inference, execution infrastructure, retries, and human review. GitHub explicitly converts model-token spending into AI credits at $0.01 per credit; Kiro describes credits as task-dependent units affected by model choice. Comparing their credit counts directly would be misleading.[53][3]
One first-hand illustration of the purchasing problem is Reddit user Old-Glove9438, who described confusion between Mistral's displayed API, Vibe, and PAYG allowances in a thread marked one month old when reviewed on September 16, 2026. Subsequent replies included a claimed interface update and further disagreement about charges. That is an individual usability report, not proof of an overbilling defect. Use the current billing documentation and the actual account dashboard, not a forum screenshot, to establish entitlements.[54][51]
Source Availability and Security Boundaries
The Codex, Gemini CLI, Qwen Code, and Mistral Vibe repositories carry Apache-2.0 licenses; Kimi Code carries MIT. Those licenses apply to the software, not free hosted inference, every model's weights, or the associated cloud service. Claude Code's repository instead points to Anthropic's commercial terms; Kiro licenses its IDE and CLI as AWS Content. ZCode published an Apache-2.0 tree after the snapshot incident; that grant does not automatically cover every desktop binary. Trae Agent is MIT; TraeCode is commercial. IBM Bob and CodeBuddy are commercial SaaS/clients.[55][56][57][58][45][59][47][15][18]
A useful security review separates four controls:
| Control | What it establishes | What to verify |
|---|---|---|
| Tool approval | Whether a requested action may start | Automatic modes, remembered grants, hooks, and child-agent behavior |
| Filesystem/process isolation | What an approved command can access | Writable mounts, sensitive files, sandbox availability, and escape approvals |
| Network policy | Which services commands or tools may contact | Shell-child egress versus in-process API and MCP traffic |
| Change review | Whether generated work is accepted into the product | Tests, diff review, repository checks, and merge permissions |
These are analysis categories, not a certification score. Current examples show why a single sandbox checkbox is insufficient:
- Codex: default local permission mode applies sandboxing to spawned commands; approvals are a separate control. Its implementation differs across macOS, Linux/WSL2, and native Windows.[60]
- Claude Code: the Bash sandbox covers commands and children on macOS, Linux, and WSL2, with configurable unsandboxed fallback. The docs warn that an unavailable sandbox can fall back to unsandboxed execution unless
failIfUnavailableis enforced.[61] - Qwen Code: macOS Seatbelt and Docker/Podman have different dependencies and access profiles; the default Seatbelt profile allows outbound networking.[43]
- Grok Build: sandboxing is opt-in; child-network restrictions are Linux-only, and model/API web traffic is a separate path. Its planning gate also leaves shell and subagent exceptions.[41][42]
- Kimi Code: “Always Ask,” “Ask When Needed,” and “Never Ask” are materially different. Never Ask also automatically handles sensitive-file and plan-exit approvals; an unattended mode is not an isolation mechanism.[62]
For detailed runtime choices, see Local Agent Sandboxes. A separate Git worktree avoids mixing ordinary edits between tasks, but still needs an appropriate execution boundary.
How to Choose and Evaluate
The following shortlist is editorial judgment based on the documented workflows, not a benchmark ranking.
| Starting requirement | Evaluate first | Why |
|---|---|---|
| Work in an existing terminal/editor workflow | Claude Code, Codex, Gemini CLI, Grok Build | Native repository interaction; compare actual account and permission behavior |
| Modify or inspect the agent client | Codex, Gemini CLI, Qwen Code, Kimi Code, Mistral Vibe | Published permissive client licenses |
| Switch among provider APIs in one local harness | Qwen Code, Grok Build, Kimi Code | Documented configurable provider paths |
| Delegate a bounded GitHub task | Jules or Copilot cloud agent; also evaluate Claude/Codex cloud routes | Remote execution with reviewable changes; setup and surface limits differ |
| Turn requirements into structured engineering work | Kiro | Specs are a central product workflow |
| Stay within an existing Mistral account and tooling setup | Vibe Code | Local/editor and remote coding surfaces with shared billing controls |
| GLM-native desktop ADE | ZCode | First-party ZCode Agent; verify client version after the snapshot incident |
| IBM i / Z / governed Java | IBM Bob | Local IDE/Shell; not a dated remote-agent runtime |
| Tencent/Hunyuan coding client | CodeBuddy | Distinct from WorkBuddy; verify region |
| ByteDance research CLI vs commercial IDE | Trae Agent vs TraeCode | Different licenses and runtimes |
| Regional Qoder purchase | Qoder or Qoder CN | Bright Zenith vs Alibaba Cloud; not Qwen Code |
| Run different labs' agents through a team operating layer | Tembo, alongside the relevant lab clients | Common execution, triggers, and review workflow; see the scope distinction below |
A practical evaluation should use representative repository work, not a toy prompt. For example, choose a real dependency upgrade with a failing regression test. Prepare a clean branch, reproducible dependency installation, the expected test command, and a concrete acceptance condition. Give each candidate the same starting commit and available credentials.
Record the result at each stage: could it prepare the environment, identify the failing behavior, implement a bounded fix, run the relevant tests, and produce a diff a maintainer accepts? Keep failed runs and manual interventions. Measure wall time, billed usage, and review time separately; a faster generation with expensive rework is not automatically the better tool.
For a remote candidate, also test a disconnect and return, a missing secret, a denied network request, and a task that outlasts its allowed runtime. For a local candidate, confirm which processes and files its configured sandbox actually covers. This is a proposed evaluation procedure; no comparative execution benchmark was run for this report.
Tembo: an Adjacent Operating Platform
Disclosure: Ry Walker is Tembo's co-founder and CEO. Tembo is relevant when the buying question moves from “which lab client?” to “how does the team run and supervise these clients?” It does not meet this report's model-developer criterion and is therefore adjacent rather than an additional lab member.
Tembo's documented workflow runs agents in cloud environments, accepts work from sources such as Slack, Linear, GitHub, schedules, and webhooks, and returns output for review. The September 15 sandbox image inventory explicitly includes Claude Code, Codex, Gemini CLI, and Grok, while warning that installation does not guarantee every agent is enabled in every session. ZCode, IBM Bob, CodeBuddy, TRAE, and Qoder were not in that inventory.[63][64]
That creates a concrete choice: run a client's native local/cloud workflow, or use a common operating layer for agent selection, execution, and team visibility. Tembo documents workspace defaults, per-session overrides, provider credentials, shared MCP configuration, and administrative model availability. It is useful to evaluate for recurring maintenance across teams; it does not remove the need to validate each agent's configuration or provider terms.[65]
Costs also remain distinct. Tembo's dollar allowance covers gateway inference and VM compute; its pricing page says BYOK and ChatGPT/Codex OAuth inference are not charged through Tembo, while their VM compute still consumes allowance. That statement does not grant portability to every lab subscription. Kiro, for example, explicitly restricts routing subscription requests outside its native interfaces through third-party automation harnesses.[66][52]
Assessment
First-party development is a useful category boundary, not a quality ranking. The decision should follow the repository workflow, execution location, account terms, and controls that a team can actually operate. Keep the client choice separable from inference and deployment where the product supports it, and recheck the account-specific details before standardizing a team on any advertised allowance.
Research by Ry Walker Research • methodology
Sources
- [1] Qwen Code — Authentication and provider plans
- [2] xAI — Grok Build overview
- [3] Kiro — Models, routing, and relative usage costs
- [4] GitHub — Copilot plans
- [5] Anthropic — Claude plans and pricing
- [6] Gemini CLI — Authentication
- [7] Gemini CLI — Quotas and pricing
- [8] Google Developers Blog: Transitioning Gemini CLI to Antigravity CLI
- [9] AWS — About Kiro
- [10] AWS — Amazon Nova models and services
- [11] Microsoft AI — Training MAI models for GitHub Copilot and Excel
- [12] Google — Jules Tools CLI reference
- [13] Google — Antigravity CLI overview
- [14] Gemini CLI GitHub Repository
- [15] Z.ai — ZCode Agent documentation
- [16] IBM — Bob global availability, April 28, 2026
- [17] Tencent — Productivity agent suite and CodeBuddy, June 5, 2026
- [18] ByteDance — Trae Agent research-oriented coding agent
- [19] Qoder — About Bright Zenith
- [20] Alibaba Cloud — Qoder CN product and billing documentation
- [21] Anthropic — Claude Code platforms and deployment
- [22] OpenAI — Codex CLI
- [23] OpenAI — Pricing and account entitlements
- [24] Google — Antigravity CLI installation and authentication
- [25] Google — Jules environment setup
- [26] xAI — Grok Build for everyone, web and mobile
- [27] Qwen Code GitHub Repository
- [28] Kimi Code GitHub Repository
- [29] Kimi Code — Membership service guide
- [30] Mistral — Vibe Code overview
- [31] Mistral — Le Chat is now Vibe, August 12, 2026
- [32] Kiro — Frequently asked questions
- [33] GitHub — About Copilot CLI
- [34] GitHub — About Copilot cloud agent
- [35] Anthropic — Set up Claude Code
- [36] OpenAI — Models and retirement guidance
- [37] Google — Antigravity CLI permissions
- [38] Google — Jules plan review
- [39] Google — Jules usage limits
- [40] xAI — Grok Build CLI
- [41] xAI — Grok Build sandbox
- [42] xAI — Grok Build Plan Mode
- [43] Qwen Code — Sandboxing
- [44] Kimi Code — Benefits, limits, and permitted use
- [45] Kimi Code MIT license
- [46] Mistral Vibe GitHub Repository
- [47] Kiro — IDE and CLI license
- [48] Google — Antigravity pricing
- [49] Kimi — Membership tiers and pricing
- [50] Mistral — Pricing
- [51] Mistral — Subscriptions, shared usage, and PAYG
- [52] Kiro — Plans, credits, and subscription restrictions
- [53] GitHub — Copilot individual billing and AI credits
- [54] Old-Glove9438 on Reddit — Mistral billing-interface confusion, reviewed September 16, 2026
- [55] Codex Apache-2.0 license
- [56] Gemini CLI Apache-2.0 license
- [57] Qwen Code Apache-2.0 license
- [58] Mistral Vibe Apache-2.0 license
- [59] Claude Code repository license
- [60] OpenAI — Sandboxing
- [61] Anthropic — Claude Code sandboxing
- [62] Kimi Code — Interaction and approval modes
- [63] Tembo — Coding agents and execution workflows
- [64] Tembo — Built-in sandbox skills and tools, September 15 image inventory
- [65] Tembo — Model settings and provider credentials
- [66] Tembo — Pricing, allowances, and BYOK