← Back to research
•·22 min read·industry

In-House Coding Agents

Comparison of 15 in-house coding-agent cases: architecture, verification, public availability, measured outcomes, and the practical build-versus-buy decision.

Key takeaways

  • Internal agents differ in purpose and availability; a PR share, review-coverage figure and model-quality improvement are not a common leaderboard.
  • Durable task records, prepared execution, relevant context and meaningful verification explain more than the choice of chat interface.
  • Spotify now offers Honk commercially, while Zup adds a documented local-execution case to the comparison.
  • Evaluate Tembo and other available platforms against the responsibilities your team would otherwise operate itself.

FAQ

Why build an in-house coding agent?

A concrete integration, execution or verification requirement may justify owning part of the system. That decision includes ongoing maintenance, recovery, security and evaluation, not only the initial prototype.

Which of these systems can another team use?

Roast is a public framework and StrongDM publishes selected components/specifications; Honk is offered through Fleetshift in Spotify Portal. Most other cases describe private deployments rather than downloadable products.

Do higher agent PR shares prove productivity gains?

No. Definitions, task mix, user selection and review policies differ; accepted outcomes, quality, rework and cost are needed to interpret the figures.

Is there a minimum company size for building?

The reviewed evidence supports no universal headcount or code-size threshold. Compare a specific requirement and full operating cost with a fair pilot of available alternatives.

Executive Summary

An in-house coding agent is an engineering system around a model: it assembles context, executes tools, records work, applies company rules, and returns something people can evaluate. Building that system can be valuable when a team needs integration or control that available products do not provide. It also creates an ongoing platform responsibility.

This report compares 15 documented implementations and enabling frameworks. Their purposes differ: Stripe automates bounded coding tasks, Meta runs ranking experiments, Cloudflare reviews changes, and Shopify publishes workflow primitives. Treating their PR counts, accuracy improvements, and review coverage as one leaderboard would conceal those differences.

Three developments materially change the comparison. Spotify now offers Honk through Fleetshift in Portal; Uber publishes a broader software-factory and cost-measurement account; and Zup has described its local-execution CodeGen architecture in a first-person research report.[1][2][3] Publicly described internal systems, reusable source code, and purchasable products remain separate availability categories.

The central build-versus-buy question is which responsibilities your team should own. Tembo is discussed as a commercial alternative in the dedicated section below, with my relationship disclosed there. It is adjacent to this matrix, rather than counted as an internal company implementation.

Scope and Evidence

A member needs a concrete software-engineering workflow and an accessible first-person technical account or implementation. The scope includes code generation, testing, review, maintenance, and ML engineering; it does not require every system to generate an entire application. Open frameworks born from company work are included with their narrower role stated explicitly.

The matrix contains the 15 members listed in this report's membership metadata. Supporting infrastructure, commercial alternatives, and the limited-evidence note are not additional members. Company-reported results are dated observations, not independent performance tests. This refresh checked sources on September 15, 2026 and did not access private deployments.

Comparison Matrix

ImplementationPrimary work and executionVerification or controlPublic availability
Stripe MinionsBounded changes in prepared development boxes; Goose forkDeterministic workflow steps, CI, human reviewInternal service; engineering accounts[4][5]
Ramp InspectCollaborative background coding in prepared environmentsFull-stack reproduction, tests and reviewInternal; independent reimplementations are separate[6]
Coinbase ForgeSlack/GitHub/Linear requests to proposed changesReturn to conversation for reviewInternal; runtime details incomplete[7]
Spotify HonkFleet migrations and background changesSeparate verification facilities and CIInternal origin; now offered through Fleetshift/Portal[8][1]
Harvey SpectreShared durable runs with ephemeral sandbox workersScoped tools, persisted artifacts and review surfacesInternal engineering platform[9]
Browserbase bbSlack coding, investigation and business workflowsIntegration proxy, scoped sessions and skillsInternal system described publicly[10]
Cloudflare internal agentsShared context/access plus CI review orchestrationSpecialized reviewers, risk tiers and human overrideInternal integration; underlying components public[11][12]
Uber coding agentsPortfolio across coding, review, CI and maintenanceWorkload evaluations, quality and cost measuresInternal systems and engineering accounts[2]
Abnormal AI agentsTicket investigation, then requested implementationHuman assesses recommendation before second triggerInternal; narrow published workflow[13]
Meta REAMulti-week ads-ranking experimentsApproved plan, compute budget and production oversightInternal ML engineering system[14]
OpenAI harness caseAgent-generated product in a purpose-built repositoryAgent review, UI/observability access and encoded rulesInternal product experiment; techniques published[15]
StrongDM FactorySpecification-driven software generationBehavioral scenarios and dependency replicasFactory case; public components/specifications[16][17]
Bitrise agentCustom agent and evaluation infrastructureProgrammatic checks plus task-specific judgesInternal implementation supporting product features[18]
Shopify RoastRuby orchestration of models, commands and agentsWorkflow-defined checks and control flowPublic MIT framework; not a complete hosted platform[19][20]
Zup CodeGenLocal CLI execution with central orchestrationApproval/planning modes and tool policiesInternal experience report; public artifact not established[3]

Company Implementations

Stripe Minions: automate the path around the agent

Stripe's February account describes a Goose-derived agent running in the same prepared EC2 development boxes used by engineers. A warm pool brings up environments quickly. Slack is one entry point, with CLI and web access also documented. The published result reached more than 1,300 merged PRs per week by February 19: humans reviewed them, but did not write the code counted in that measure.[4][5] That population is narrower than all Stripe changes involving AI.

The transferable mechanism is a blueprint combining agent work with deterministic operations. Let the agent investigate and implement; let ordinary code handle repeatable steps such as linting and pushing. Stripe limits full CI iteration before returning unresolved work to a human. Its QA development boxes do not receive production services, real user data, or arbitrary network egress.[5] Those are Stripe's published boundaries, not properties of every background agent.

Context preparation matters too. Rules are scoped to directories or file patterns. Toolshed exposes nearly 500 internal/SaaS tools, while each agent gets a curated subset. Hydrating linked context before model execution can save the agent from spending its first turns rediscovering the request.[5] A large tool catalog is useful only if the right capabilities are discoverable and appropriately scoped.

Ramp Inspect: reproduce the environment, preserve the work

Ramp's architecture account describes prepared full-stack development environments and a shared control plane for sessions. Early sandbox startup can overlap with someone typing a request; repository synchronization still has to finish before editing. Collaboration, browser inspection and session history are part of the interface, rather than separate terminal conventions.[6]

The worked cases show why that investment matters. A February security campaign generated reproduction tests intended to fail before a fix and pass afterward. Ramp also describes a March Sheets-maintenance loop that turns monitoring signals into investigation and proposed repairs.[21][22] The useful comparison is the complete chain from evidence to check to review, not just whether an agent can produce a patch.

Ramp's August builder interview reports 75% of merged PRs originating from Inspect, while earlier accounts used different dates and percentages. Its retrospective January figure differs from the original January post, so this report does not construct a precise growth curve from them.[23][6] The profile retains the richer security, integration-generation, context and collaboration examples.

Coinbase Forge: carry the original report into the review

Linear's current case study describes Forge accepting requests from Slack, GitHub and Linear, then returning proposed fixes for review. Requirements and delivery context live in Linear; the conversation remains the place people inspect the result.[7]

A March demonstration captured a trade-form bug report, summarized it, created an issue and invoked an internal bot. The published episode writeup explicitly says it did not show a completed PR. It also identifies the adoption-analysis CSV as invented demonstration data.[24] Neither example supports an automated-success rate or actual employee-level productivity calculation.

Coinbase's separate Mux report covers parallel agents with individual worktrees, branches and terminals. Its April snapshot reports 5,068 merged PRs, and explicitly warns that its users likely skew toward engineers who were already productive and enthusiastic about AI.[25] Mux's PR ratio is not a causal Forge multiplier. A good implementation preserves provenance from report through tests, review and accepted change, even when several agents work concurrently.

Spotify Honk: fleet transformations need explicit verifiers

Spotify's documented Honk workflow combines a coding agent with fleet-management infrastructure. Its feedback-loop account separates lightweight checks, independent verification and the normal CI system.[8] A verifier should establish a required outcome, such as whether the old API remains, rather than merely asking the editing agent whether it succeeded.

The April dataset migration is concrete: identify downstream consumers, provide old-to-new dataset mappings, update supported pipelines, and validate the resulting changes. Spotify reported 240 automated PRs across a target population of roughly 1,800 downstream pipelines and estimated ten engineering weeks saved. Some pipeline types required different handling.[26] The estimate remains the team's report; coverage of the target set and correctness of individual transformations are different questions.

Honk is no longer an internal-only purchasing comparison. Spotify markets it through Fleetshift in Portal.[1] August's Xirp announcement describes a separate environment for multi-harness sessions and Portal context sharing.[27] Neither announcement proves feature parity with every internal Honk workflow. Evaluate the actual commercial configuration.

Harvey Spectre: durable runs outlive their workers

Harvey's April technical account makes the run record durable, while workers are disposable. Ownership, history, attachments, artifacts and provider-session references survive the execution process. A follow-up restores context in a fresh worker rather than reviving the old container.[9]

Slack, web, CLI and cron feed the same run model. The harness handles provider adapters, progress, timeouts, cost accounting and PR post-processing; workers receive scoped repository and tool access rather than direct control-plane database access.[9]

This provides a useful recovery design to compare with persistent environments. Persisting a conversation is not the same as persisting every process or filesystem mutation. Decide which state must survive and how it will be reconstructed. Spectre's documentation describes an internal engineering platform; it is not a public license to Harvey's legal-product infrastructure, and no comparable PR-success rate is established.

Browserbase bb: separate capabilities from the conversation loop

Browserbase describes bb as an OpenCode loop extended by on-demand skills and typed service wrappers. A Slack thread retains a workspace; tasks range from code changes to incident investigation and customer-context queries. Its create-PR skill connects an issue to a branch and proposed change.[10]

The interesting boundary is the integration proxy. Most real credentials remain in serverless services, which check the session's allowed methods. Selected integrations use network credential brokering. Background webhooks receive explicit scopes; interactive sessions are broader. Read-only data roles and domain restrictions remain relevant even when credentials are hidden.[10]

Browserbase's stated feature-request coverage concerns scanning support/meeting inputs, not autonomously implementing every requested feature. No general code-success rate follows from that number. For builders, the practical lesson is to inspect authorization at the service boundary as well as instructions inside the agent.

Cloudflare: coordinate reviews and make overrides visible

Cloudflare combines third-party coding tools with internal identity, model routing, a service catalog, MCP access and repository instructions. Its April adoption figures describe that wider stack, not the share of code written by one agent.[11]

The CI reviewer runs specialized agents through OpenCode and consolidates their findings. Risk tiers change the review effort; generated noise is filtered; previous findings inform subsequent runs. Serious findings can block a merge, with an explicit human override recorded in telemetry.[12]

From March 10 to April 9, Cloudflare reported 131,246 review runs across 48,095 MRs, averaging $1.19 per review. The 0.6% override rate is not a measured false-positive rate. Coverage applies to repositories using the standard CI integration.[12][11] The authors also identify limits around architecture, downstream consumers and subtle concurrency bugs. A reviewer seeing a diff cannot establish every property of the deployed system.

Uber: a portfolio with workload-specific economics

Uber's August 27 engineering account reports that more than 70% of PRs were attributed to local or cloud agents. Its managed work spans review, CI recovery, end-to-end changes, on-call investigation and maintenance. Model choices are evaluated against actual workloads, quality, cost and reliability, while shared context and tool-access mechanisms reduce repeated discovery.[2] That broad attribution measure should not be compared directly with fully unattended PRs from another company.

The earlier uReview account describes separate generation, grading and filtering stages. Uber measured useful feedback and whether findings were addressed, alongside a curated issue dataset. Its August 2025 estimate of 1,500 saved hours weekly relied on a modeled human-review duration, rather than observed cash savings.[28]

The builders later explained why they retired developer-years-saved as an ROI metric: maintenance of assumptions, employee interpretation and weak connection to business outcomes. Their newer direction emphasizes feature delivery with supporting flow and quality measures.[29] That is a valuable correction to the temptation to rank internal agents by estimated labor replacement.

Abnormal AI: make investigation a separate decision

Abnormal's June account describes two explicit transitions. One ticket tag asks an agent to investigate the cause and recommend a fix. After a person reviews that recommendation, another tag requests implementation and a PR.[13]

This is a concrete way to stage autonomy without requiring every task to enter one long write-capable loop. The investigator can establish that the report needs more information, or that a configuration change is preferable to code. The current source does not establish the deployment provider or a company-wide PR share. Earlier headline percentages should not substitute for those missing details.

Meta REA: long experiments require a different runtime

REA is Meta's ranking-engineering system for ads models, built on its Confucius infrastructure. It combines historical experiment knowledge with research-derived hypotheses, proposes a plan for approval, and operates within a compute budget. Hibernate/resume behavior supports long-running experiments; validation, combination and exploitation phases structure the search.[14]

Meta's March report describes twice the accuracy improvement over baseline approaches in initial work across six models, not a doubling of absolute model accuracy. It also reports higher engineering output and a changed staffing footprint. These are company observations within a specialized ML workflow, not a general software-engineer multiplier.[14]

The transferable question is how an agent waits for external evidence and knows when to stop. A ranking experiment may need hours or days to return a useful signal; repeatedly prompting a model while it waits adds neither evidence nor progress. Human experiment and production oversight remains part of the described system.

OpenAI harness engineering: make the application inspectable

OpenAI's February report describes an internal product experiment with about a million lines across code, infrastructure and documentation, all agent-written. Humans specified work and shaped the environment; they could review PRs, though review increasingly happened between agents. This was one purpose-built repository, not a claim about all OpenAI engineering.[15]

The concrete mechanisms are useful: a short instruction file points to structured repository knowledge; the application starts separately per worktree; agents can inspect the UI, logs and metrics; architectural rules become automated checks. Recurring cleanup tasks address drift.[15]

The authors explicitly caution that end-to-end behavior depends on that investment and that long-term architectural coherence remains an open question. “No manually written code” therefore says who produced the artifacts. It does not mean no human prioritization, judgment or accountability.

StrongDM Factory: move validation outside the implementation

StrongDM's small AI team's factory account adopts a deliberately strong constraint: people specify behavior without manually writing or traditionally reviewing the generated implementation. Scenarios and replicas of third-party dependencies support behavioral evaluation.[16][30] This is the team's reported practice, not proof that every StrongDM project follows it.

The Digital Twin Universe idea is to reproduce externally observable dependency behavior so tests can run repeatedly without live-service limits. Keeping evaluation scenarios separate from the coding workspace can make them a more useful check than tests the implementation agent can freely rewrite.[30] The replica must itself be validated: passing a mistaken model of an API does not prove compatibility with the real API.

Public components include CXDB, a context store, and an open Attractor specification for graph-structured coding workflows.[17][31] A specification and community implementations should not be described as the complete internal factory. Nor should the founder's provocative daily token-spend target become a demonstrated ROI threshold.

Bitrise: evaluate the feature, then own the necessary loop

Bitrise's November 2025 account describes a custom Go agent and an evaluation system that provisions Docker environments, applies task setup and patches, runs agents in parallel, and combines programmatic checks with task-specific judges. Results feed a database and dashboard.[18]

That supports different checks for different outputs. A build repair can be tested by executing the build; a proposed review comment needs a different assessment. Programmatic checkpoints inside the workflow were a major reason to own the implementation, alongside provider-independent logging and orchestration.[18]

The historical comparison used older agent/model versions. In particular, its archived Go OpenCode project is not today's separate TypeScript OpenCode implementation. The post is evidence for Bitrise's design decision at that time, not a current ranking of Claude Code, Codex or Gemini. A reusable evaluation harness retains value as those products change.

Shopify Roast: encode a workflow without building a whole platform

Roast is a public Ruby DSL that combines model calls, local agents, commands, iteration and collection processing. The current README makes Pi the default agent provider, with Claude Code also supported. Providers and credentials remain separate from the workflow definition.[19] Its code is MIT-licensed.[20]

A small illustrative workflow uses a deterministic input step, an agent investigation and a model summary:

execute do
  cmd(:files) { "git diff --name-only HEAD~1..HEAD" }

  agent(:review) do
    "Inspect these changed files and explain concrete risks:\n" + cmd!(:files)
  end

  chat(:summary) do
    "Summarize the findings for the reviewer:\n" + agent!(:review).response
  end
end

This adapts the documented cog interfaces; it was not executed for this report.[19] It illustrates orchestration, not a safety policy: the chosen agent's filesystem access, permissions and execution environment still need configuration. Add meaningful checks to the workflow before treating generated findings as accepted work.

Zup CodeGen: local execution, central orchestration

Zup's April experience report describes a CLI executor, a FastAPI backend and a Maestro agent loop. State and event history support reconnection. The authors emphasize targeted edits, understandable tool errors and consistent permissions across tools with overlapping capabilities.[3]

The distinction between instructions and enforcement matters. The paper explicitly places read-before-edit in the prompt; it is not a hard tool precondition. Planning and approval modes offer human checkpoints, while local execution means the developer's machine remains part of the boundary.[3]

The report supplies engineering observations, not a controlled productivity benchmark. Its profile also distinguishes this internal system from similarly named StackSpot products. That prevents a published architecture from becoming an unsupported claim about a downloadable product.

Reusable Architecture Patterns

These cases support a set of design questions rather than a universal stack. Slack is useful when it carries the original context, but CLI, CI, experiments and schedules also initiate work. Some systems keep warm environments; others deliberately restore into fresh workers. Local execution remains a valid category member.

ResponsibilityConcrete decision to make
Request and contextPreserve the original evidence and identify the correct repository, version and owner
ExecutionChoose local or hosted workers and document the actual filesystem, network and credential boundary
Durable stateSeparate the task record from process lifetime; define what survives a restart
ToolsMake allowed operations explicit and enforce consistent policies across equivalent capabilities
VerificationChoose tests, judges, previews or experiments appropriate to the requested outcome
Review and deliveryIdentify who can approve, merge, deploy, override and audit each transition
EconomicsMeasure accepted outcomes alongside quality, latency, failed attempts and operating effort

A worked migration evaluation

Suppose twenty services must replace an old client API. An evaluation inspired by these cases can make the work inspectable without prescribing one vendor:

  1. Define the target set. Record which repositories use the API and which versions are eligible. Explicitly exclude unsupported variants.
  2. Prepare one representative environment. Confirm that the existing application and checks run before asking an agent to change anything. Otherwise a failed test cannot distinguish a regression from a broken baseline.
  3. Specify the transformation. Provide an old/new example, compatibility requirements and stopping conditions. Keep the original request visible in the resulting PR.
  4. Use separate checks. Confirm both that the new behavior works and that the deprecated call is gone where required. A successful build alone may establish neither.
  5. Pilot and classify failures. Separate context mistakes, tool failures, incorrect edits and inadequate tests. Fix the recurring cause before expanding the fleet.
  6. Expand with review capacity. Track accepted migrations, rollback/rework, queue time and total effort. Twenty generated PRs are not twenty completed migrations.

This is an evaluation design, not a measured result. It preserves the most useful fleet-workflow lesson: the agent performs one part of a system whose inputs and outcomes must remain understandable.

Reusable implementations are starting points

Background Agents independently implements a Ramp-inspired pattern. Its current README documents multiple harness/provider paths and an explicit trusted-organization, single-tenant security model; shared repository access is not checked per user.[32] It is not Ramp's source release.

LangChain's Open SWE independently packages related ideas using Deep Agents and LangGraph. Its current repository includes multiple sandbox backends and workflow components; the MIT code and the license-key requirements of its documented production Agent Server deployment are separate considerations.[33] Reuse can reduce implementation effort, but it does not validate a deployment's permissions, reliability or economics.

Build, Buy, or Combine

There is no supported universal threshold of a thousand engineers or ten million lines of code. Size affects the investment, but the decision also depends on unusual infrastructure, task repetition, existing platform capability, failure cost and the alternatives available now.

Build the necessary parts when a concrete constraint survives a fair product evaluation: inaccessible internal context, a proprietary execution environment, a review system vendors cannot integrate with, or a verification loop central to the work. Budget for migrations, incident response, observability and ongoing model/tool changes, not just the prototype.

Buy or extend a platform when its execution and workflow model fit the task, and the team's scarce effort is better spent on repository readiness, context and acceptance checks. Buying orchestration does not outsource engineering judgment; it can change where that judgment needs to be implemented.

Combine approaches when a managed runtime covers execution but company-specific checks or tool integrations remain essential. Inspect the extension boundary: configuration is useful only if it can enforce the behavior the organization requires.

Tembo as a commercial alternative

Disclosure: Ry Walker is Tembo's co-founder and CEO.

Tembo offers background coding agents with selectable harnesses, prepared project environments, scheduled work and event/webhook triggers.[34] Its Agent Actions documentation describes workplace-tool context, sessions spanning repositories and git providers, and PR or merge-request artifacts.[35]

That makes it relevant to the request-to-reviewed-change portion of this category. A team can evaluate those provided capabilities against building its own worker lifecycle, trigger handling and coordination surface. Tembo is not a documented equivalent of Meta's ranking experiments or StrongDM's dependency replicas, and the sources do not establish integrations with the private agents above.

A useful pilot gives the available options the same task, repository state, checks and allowed credentials. Compare preparation and recovery effort as well as the successful run. Record reviewer time, abandoned attempts, rework, model/compute charges and recurring platform maintenance. The decision should turn on the actual workflow rather than a comparison between a mature internal deployment and a vendor's easiest demo.

Interpreting the Numbers

Keep the denominator visible. A PR share can describe fully agent-authored changes, any attributed assistance, opened drafts or merged work. Review coverage measures how broadly a check runs. Model-quality improvements describe a different output again.

ObservationWhat it cannot establish alone
More PRs per active userA causal productivity gain; early adopters and task mix may differ
More agent-authored codeCorrectness, customer value, or less review effort
Lower cost per model requestLower cost per accepted change if retries or task complexity rise
Few human overridesA measured false-positive or false-negative rate
Estimated time savedEquivalent headcount reduction or realized financial savings

Coinbase's Mux authors explicitly flag selection bias, and Uber's measurement presentation explains why activity metrics and saved-time estimates failed to answer some business questions.[25][29] These are reasons to improve measurement, not reasons to dismiss every reported gain.

A practical scorecard combines accepted outcomes with lead time, escaped defects, rollback/rework, reviewer effort and fully loaded operation. Keep a stable task cohort when comparing changes to the agent. When the work itself changes, explain that change before attributing the difference to a model or framework.

Limited Evidence and Outlook

Google Agent Smith appeared in the previous edition. This review did not locate an accessible first-person technical account sufficient to verify its architecture or attributed code-share claim, so it is retained here as a limited-evidence link rather than a matrix member. That is a sourcing boundary, not a claim that Google stopped internal agent work. Speculation about unnamed Amazon or Netflix systems is likewise not counted.

The practical outlook is conditional. More reusable infrastructure may reduce the cost of building, while more capable products may reduce the reasons to build. The enduring work is making company context accessible, defining real execution boundaries and checking outcomes. Which layer a team should own must be revisited as both its needs and the available products change.


Research by Ry Walker Research • methodology

Sources