← Back to research
•·10 min read·company

Ramp Inspect

Ramp's internal coding-agent platform combines prepared sandboxes, company context, and human review. New accounts describe 75% of merged PRs and specialized maintenance workflows.

Key takeaways

  • Ramp reported 75% of merged PRs originating in Inspect sessions in an August 2026 interview; that is adoption evidence, not a measured productivity uplift.
  • Prepared development environments, executable tests, browser feedback, and company context are central to Inspect's design.
  • Security repairs, monitoring-driven fixes, and integration generation now provide concrete examples of specialized workflows built around Inspect.
  • Inspect remains internal; the open-source Background Agents project is an independent implementation, while Tembo is a commercial alternative.

FAQ

What is Ramp Inspect?

Inspect is Ramp's internal background coding-agent platform. It runs coding sessions in prepared remote development environments and connects them to Ramp's tools and organizational context.

What results has Ramp reported?

An August 25, 2026 interview with Ramp's builders reported 75% of merged PRs coming from Inspect sessions and roughly 90% for Inspect's own repository. These company-reported shares do not establish how much of each change the agent completed or the net productivity gain.

Is Ramp Inspect open source or available to buy?

The reviewed sources describe Inspect as internal tooling, not a product subscription or public source release. Cole Murray's Background Agents is a separate open-source project inspired by its architecture.

Does Inspect eliminate human review?

No. Ramp's security and monitoring case studies explicitly retain engineer review before merging; the integration factory also hands evidence to engineers for review and curation.

Executive Summary

Inspect is Ramp's internal background coding-agent platform. Its central design choice is to give an agent a prepared development computer and relevant company context, then let it iterate using tests and other feedback. Ramp's builders reported in an August 25, 2026 interview that 75% of merged PRs originated in Inspect sessions. That indicates substantial use inside Ramp, not an independently measured 75% gain in engineering productivity.[1]

The platform now supports specialized workflows beyond an engineer typing a coding request: Ramp has published examples involving security fixes, monitoring-driven maintenance, and an integration factory. These examples retain review and verification boundaries; “background” does not mean every generated change is automatically accepted.[2][3][4]

See in-house coding agents for the category and Background Agents for a separate open-source implementation inspired by Inspect.

Product and Architecture

A prepared computer for every session

Modal's February case study describes a sandbox containing application services such as Postgres, Redis, Temporal, and RabbitMQ, with OpenCode as the coding runtime. A hosted VS Code server and terminal allow manual work, while Chromium and a streamed desktop support frontend inspection. It is a development environment with an agent inside it, rather than just a place to generate a patch.[5]

LayerDocumented responsibility
Session stateCloudflare Durable Objects, per-session SQLite, and Agents SDK streaming[6]
ExecutionModal sandboxes containing the repository and development services[5]
Company contextCode, documentation, product specifications, and connected internal tools[7]
Human interfaceWeb, Chrome extension, and Slack; sessions can be shared between colleagues[7]
AcceptanceReviewable changes and supporting evidence; review remains part of the published workflows[2][4]

Startup and continuation

Ramp prepares repository images on a 30-minute cadence, installing dependencies and running initial builds before a user requests work. New sandboxes start from filesystem snapshots, then synchronize recent repository changes. Modal also describes queues and locks for coordinating inputs and sessions, plus child sessions for parallel work.[5]

The original architecture describes warming a sandbox while a user types, permitting early reads while synchronization completes, and blocking writes until the repository is current. Follow-up prompts can queue behind ongoing execution; users can also stop a run. These details explain the engineering work behind a responsive interface.[6]

Prepared snapshots move cost and latency out of the interactive path; they do not eliminate them. An implementation still needs image refreshes, stale-image handling, dependency updates, concurrency limits, and a way to recover work after a failed session. Evaluate those operations alongside the visible startup time.

Context, collaboration, and attribution

Linear's customer account describes Inspect using product specifications, feedback, and roadmap history through Linear's API. It gives the example of a designer starting a dashboard change from a Figma file and handing the session to an engineer. The shared session preserves the work and discussion across that handoff.[7]

Ramp's January design also covers voice input, mobile web use, React-aware browser selection, MCP, and custom tools. It distinguishes code-push credentials from GitHub PR creation using the initiating user's token, so attribution supports the normal review process.[6]

“Full context” should not be read as unrestricted production access. The August interview specifically describes a sanitized, read-only production database replica for debugging. Tool access is an engineered permission decision, not a property conferred by a sandbox.[1]

Worked Examples

Security findings that must reproduce

Ramp's February security experiment combined focused vulnerability detectors, adversarial validators, integration tests, and an early Inspect-based repair agent. The test needed to fail before the fix and pass afterward; findings that already passed could be discarded as false positives. Ramp reported nearly 100 issues patched in the experiment, with a human reviewing and landing the resulting changes.[2]

The report also describes a failed approach: reproducing complex conditions against a live deployment proved difficult, while code-based integration tests worked better. This is a useful distinction for anyone adopting the pattern. Reproduction is part of the task design, not something to assume the model will improvise reliably.[2]

For your own evaluation, seed a known regression and a superficially similar non-bug. Check that the system produces a meaningful failing test for the first and abstains on the second. Counting how many suspected findings it creates would reward the wrong behavior.

Monitoring that leads to a proposed fix

Ramp Labs connected Inspect to Ramp Sheets maintenance. An initial broad nightly QA pass repeatedly explored similar paths. The next design generated monitors from merged changes; a Datadog alert started an agent with specific failure context. The agent reproduced the issue, proposed a fix, and notified the team. Ramp reported 40 real bugs caught during the first week.[3]

Noise became the next problem. The team added triage that could adjust or remove noisy monitors and recorded PR links to prevent duplicate work. It retained engineer review before merge and warned that generated monitoring should not replace the instrumentation the team already trusted.[3]

The transferable pattern is an observable event, a bounded investigation, a reproducible failure, and a proposed repair. A team copying it should measure false alerts and duplicate sessions, not only fixes. Otherwise an apparently productive agent can consume more on-call attention than it saves.

An integration factory with review evidence

The August integration-factory post describes a custom Inspect agent that researches a provider API, writes a connector, tests with supplied test credentials, repairs failures, and opens a PR. Ramp reported 75 integrations shipped through the broader system. The output includes request/response evidence and recordings so reviewers can inspect what was exercised.[4]

The architecture confines a new provider to its own module, uses synthetic values and test credentials, and keeps secrets out of the model context. The customer-facing side generates deterministic integration scripts, so a model is used to construct the workflow rather than interpret every future execution.[4]

That is a stronger worked example than a general claim that agents can build anything. A connector has a describable interface and observable responses. A change to shared business logic may need a different validation plan. The evidence required for approval should match the failure modes of the artifact being generated.

Adoption and Evidence Quality

ObservationSource periodWhat it establishes
Around 30% of frontend/backend merged PRsJanuary 2026 architecture postEarly company-reported adoption[6]
More than half of merged PRs; over 80% of Inspect written using InspectFebruary 19, 2026 Modal case studyVendor-published account of Ramp's usage[5]
75% of merged PRs; roughly 90% for Inspect's repositoryAugust 25, 2026 builder interviewLater reported PR-origin shares[1]
One million cumulative sessions crossed in JulyAugust 2026 interviewVolume of sessions, not one million completed changes[1]

The August interview's historical adoption timeline differs from the January post. This table preserves each account's date and definition rather than synthesizing a precise growth curve. PR origin, percentage of code, and percentage of developer effort are different measurements.

The same interview says engineers often use Inspect to start larger changes and then continue locally. The automation is therefore part of a mixed human-and-agent workflow, even where PR attribution points to Inspect.[1]

This research did not access Ramp's private metrics or run the private service. Company and infrastructure-vendor accounts support the described design and adoption claims; they are not independent reliability or security audits.

Availability, Open Source, and Build Versus Buy

Inspect remains an internal platform in the sources reviewed on September 15, 2026. There is no public Inspect subscription price or release license to compare. A team implementing a similar system pays for engineering ownership, compute, model usage, identity and secrets management, integrations, and review capacity.

Cole Murray's Background Agents repository is a separate, MIT-licensed implementation inspired by Inspect. It uses Cloudflare coordination and remote sandboxes with OpenCode. Its availability makes the pattern inspectable and adaptable; it does not mean Ramp open-sourced Inspect or independently validate Ramp's production outcomes.[8]

Tembo as a commercial alternative

Disclosure: Ry Walker is Tembo's co-founder and CEO.

Tembo belongs in this build-versus-buy decision. It provides background coding agents with selectable harnesses, prepared project environments, scheduled runs, event triggers, and webhooks. Its agent documentation includes Slack and issue-related triggers, which overlap with the work-routing problem Inspect solves internally.[9]

Tembo also documents cross-repository sessions, connected workplace tools through MCP, and PR or merge-request output across affected repositories.[10] The choice is how much company-specific workflow the team needs to own. Building an Inspect-like system allows deep control of the development image, interfaces, and internal context. A commercial platform can supply common execution and orchestration capabilities, leaving the team to validate its own environment, permissions, and acceptance criteria.

A fair pilot would use the same issue, repositories, test suite, and allowed credentials. Compare correct merged output, review time, environment setup effort, failure recovery, and recurring cost. The sources do not establish that either approach is universally cheaper or that Tembo integrates with Ramp Inspect.

Practitioner Discussion

In the January Hacker News discussion, memset identified themselves as a Ramp employee and praised the system's ability to complete features. ColinEberhardt emphasized that apparently one-shot results could contain an iterative feedback loop. yoav questioned whether remote infrastructure justified its cost compared with local agents. These are useful perspectives, with the employee relationship made explicit, rather than a representative user survey.[11]

Other participants described difficulty reproducing the setup when their own development environment was poorly documented. That observation points to a practical prerequisite: a remotely reproducible environment benefits both human developers and agents, and building one may be more work than installing an agent runtime.[11]

Strengths, Cautions, and Fit

Inspect is a useful reference for teams whose development work spans services, product context, observability, and several kinds of contributor. Its public case studies show how a common execution platform can support distinct workflows instead of forcing every task through an identical prompt.

The main costs are ownership and review. Someone must maintain snapshots, credentials, tool interfaces, session recovery, and metrics. More parallel sessions can increase spending and review load. A sandbox separates execution, but a credential with broad privileges still carries those privileges into the session. “Unlimited concurrency” and “instant startup” are vendor framings, not an operating guarantee established here.

A small codebase with infrequent background work may not justify that platform investment. Conversely, a complex organization should not assume purchasing a tool eliminates the need to define access rules or reliable checks. Start with one workflow whose success and failure are observable, then decide whether expanding the platform is justified by the measured result.

Inspect's most useful lesson is that agents benefit from the same operational foundations as engineers: a working environment, relevant context, fast feedback, and a clear acceptance process. Its reported adoption makes the system worth studying; the worked examples explain what to evaluate before copying it.


Research by Ry Walker Research • methodology