← Back to research
•·9 min read·company

OpenAI Harness Engineering

OpenAI's Harness team shipped ~1M lines of code with 0 manually written — ~1,500 PRs merged by 3-7 engineers at 3.5 PRs/engineer/day; OpenAI estimates ~1/10th the time to write by hand.

Key takeaways

  • ~1M lines of code shipped with 0 manually written over 5 months
  • 3.5 PRs per engineer per day; team grew 3→7; ~1,500 PRs; OpenAI estimates ~1/10th the time to write by hand
  • Agent-to-agent review handles most review effort; human review remains optional
  • The practice spun out Symphony, announced as an open-source spec on April 27, 2026 (~27.4K GitHub stars as of September 23, 2026)

FAQ

What is OpenAI Harness Engineering?

OpenAI's internal engineering team that built a product with ~1M lines of code and zero manually written code. Codex agents run autonomously for 6+ hours per task, and agent-to-agent review handles most review effort, with optional human review.

How does Harness manage agent context?

Harness uses AGENTS.md as a table of contents pointing to a structured docs/ directory — not an encyclopedia file. This gives agents navigable, structured knowledge bases instead of monolithic context dumps.

What is the team's throughput?

Starting with 3 engineers growing to 7, the team merged ~1,500 PRs at 3.5 PRs per engineer per day. OpenAI estimates the product was built in about 1/10th the time it would have taken to write the code by hand. PR throughput is not a replacement-team headcount equivalent.

Citation review, September 23, 2026: dropped 3-10x / 21-70 engineer-equivalent economics (throughput is not replacement-team headcount); sourced Symphony's April 27, 2026 announcement and 500% first-three-weeks claim; added OpenAI's ~1/10th-time estimate; dropped unsourced Codex org-wide PR-lift copy.

Executive Summary

OpenAI's Harness Engineering team represents the most extreme publicly documented case of agent-driven development. Over 5 months starting August 2025, a team of 3 engineers (growing to 7) shipped approximately 1 million lines of code with zero manually written code and ~1,500 merged PRs at 3.5 PRs per engineer per day. OpenAI estimates the team built this in about 1/10th the time it would have taken to write the code by hand. Codex agents run autonomously for 6+ hours per task, and almost all review effort is handled agent-to-agent; humans may still review PRs.[1]

The practice has since produced a public artifact: OpenAI open-sourced Symphony, the issue-tracker-driven agent orchestrator that grew out of this work, and announced it as an open-source spec on April 27, 2026 (~27.4K GitHub stars as of September 23, 2026).[2][3] This profile covers the internal practice; see the Symphony profile for the open-source project.

AttributeValue
CompanyOpenAI
TypeInternal methodology
Agent RuntimeCodex
Public DocumentationFebruary 2026 (blog post)
Open-Source SpinoffSymphony (announced April 27, 2026)
HeadquartersSan Francisco, CA

Product Overview

Harness is not a product — it's a methodology and team at OpenAI that treats "no manually-written code" as a core philosophy. Engineers act as orchestrators, delegating all implementation to Codex agents. The key innovation is not just using agents for coding, but building an entire engineering practice around the assumption that humans never write code directly.

Key Capabilities

CapabilityDescription
Zero manual codeAll code written by Codex agents, no exceptions
6+ hour autonomyAgents run for extended periods without human intervention
Agent-to-agent reviewMost review effort handled by agents; human review optional
Structured knowledge baseAGENTS.md as table of contents, docs/ directory as encyclopedia
UI verificationChrome DevTools Protocol wired into agent runtime
Observability accessLogQL and PromQL exposed directly to agents

Technical Architecture

Context Management

The Harness team's key insight on context management: AGENTS.md should be a table of contents, not an encyclopedia.[1] Rather than stuffing all project knowledge into a single file, they maintain a structured docs/ directory with AGENTS.md serving as a navigable map.

AGENTS.md (table of contents)
    ↓
docs/
  ├── architecture.md
  ├── conventions.md
  ├── api-reference.md
  └── ...

This pattern gives agents structured, discoverable knowledge without overwhelming their context windows.

Agent-to-Agent Code Review

Harness reduces required human review by having agents review each other's code. OpenAI explicitly says humans may review PRs, but are not required to; agent reviews and responses to human or agent feedback remain part of the workflow.[1]

UI Verification

The Chrome DevTools Protocol is wired directly into the agent runtime, allowing Codex agents to:

  • Render and inspect UI components
  • Verify visual correctness
  • Test interactive behavior
  • Debug rendering issues

Observability

Agents have access to an ephemeral observability stack local to each isolated worktree:[1]

  • LogQL — Query logs in real time
  • PromQL — Query metrics and alerting data

This lets agents inspect the isolated application's behavior; the report does not describe unrestricted access to production telemetry.


Results

Key Metrics

All figures are as disclosed in OpenAI's February 2026 blog post; OpenAI has not published updated numbers for this team since.

MetricValue
Lines of code~1,000,000
Manually written code0
Time period5 months (Aug 2025 start)
PRs merged~1,500
Starting team size3 engineers
Final team size7 engineers
PRs per engineer per day3.5
OpenAI time-to-write estimate~1/10th by hand
Agent autonomy per task6+ hours

Throughput Analysis

OpenAI discloses 3.5 PRs per engineer per day, a team that grew from 3 to 7, and ~1,500 merged PRs over five months, plus an estimate that the work took about 1/10th the time to write by hand.[1] Those are throughput and elapsed-time claims. They are not a measured 3-10x capacity multiplier or a 21-70 engineer replacement-team equivalent: PR volume is not headcount economics.

Since Publication

Martin Fowler's site published an analysis (updated April 2026) that adds detail on how the team keeps a million agent-written lines coherent: a layered architecture enforced by custom linters and structural tests, plus recurring "garbage collection" passes that scan for drift and have agents suggest fixes. It quotes the team's conclusion: "Our most difficult challenges now center on designing environments, feedback loops, and control systems."[4]

On April 27, 2026, OpenAI announced Symphony — the orchestration layer that grew out of this practice — as an open-source spec. The 500% increase in landed PRs is a Symphony figure for some internal teams in the first three weeks of use, not a Harness Engineering result.[2] As of September 23, 2026 the repo has ~27.4K GitHub stars.[3]


Key Insights

1. AGENTS.md as Table of Contents

The most transferable insight: structure your knowledge base as a navigable directory, not a monolithic file. AGENTS.md points to relevant docs, and agents can drill into what they need.

2. No Manual Code as Philosophy

This isn't "use agents when convenient" — it's "never write code manually, period." This constraint forces the team to invest in agent infrastructure, context management, and workflow design.

3. Agent-to-Agent Review Works

By shifting most review effort to agents, the team reduces its dependence on mandatory human review. The workflow still accepts human feedback and permits human inspection.

4. Extended Autonomy is Viable

6+ hour autonomous agent sessions demonstrate that modern agents can handle complex, multi-step tasks without human intervention. This is significantly longer than most reported agent session lengths.


Strengths

  • Unprecedented scale — ~1M LOC with zero manual code is the most extreme case documented
  • Proven throughput — 3.5 PRs/engineer/day sustained over months
  • Full autonomy — 6+ hour sessions and agent-to-agent review, with optional human review
  • Transferable insights — AGENTS.md pattern, docs/ structure, observability access are universally applicable
  • Officially documented — Published by OpenAI with analysis by Martin Fowler
  • Dogfooding — OpenAI eating their own cooking with Codex validates the product

Cautions

  • OpenAI advantage — Team has privileged access to Codex capabilities and can directly influence product direction
  • New product context — Building greenfield is easier for agents than modifying legacy code
  • Small team — 3-7 engineers may not represent patterns that scale to larger organizations
  • Codex-specific — Architecture and workflow designed around Codex's specific capabilities
  • Survivorship bias — We see the successful project, not the failed attempts or rejected approaches

Competitive Positioning

vs. Other In-House Agents

SystemComparison
Stripe MinionsMinions require human review; Harness uses agent-to-agent review
StrongDM FactoryStrongDM eliminates review; Harness makes human review optional and relies mainly on agent review
Ramp InspectInspect augments human engineers; Harness replaces manual coding entirely

Unique Position

Harness represents the most aggressive position on the "agent autonomy spectrum":

  • Conservative: Agents write code, humans review (Stripe, Ramp)
  • Moderate: Agents write code, behavioral validation replaces review (StrongDM)
  • Radical: Agents write code, agents review code, humans orchestrate (Harness)

What Developers Say

The Hacker News thread on the harness engineering post (296 points, 206 comments) was substantive and split — practitioners validated the harness concept while pushing back on the headline metrics:

"I've found keeping file sizes small has been important for agentic coding not just to maintain human readability, but also for optimizing agent performance, precisely because it limits the amount of incidental context they load" — stult[5]

"Agents help a ton with the discovery, but the act of building a product needs a deeper level of thought and validation to make it actually better than what came before." — krackers

"Yeah I cannot see how 'we shipped 1 million lines of code in three weeks' is... something to be proud of haha" — Aperocky

The recurring skeptical theme: lines of code and PR counts are input metrics, not evidence of product quality — and the post describes a greenfield beta, not a battle-tested production system.


Bottom Line

OpenAI's Harness Engineering is a proof-of-concept for fully agent-driven software development. The "no manual code" philosophy, agent-to-agent review, and 6+ hour autonomy sessions represent the frontier of what's possible today.

Key metrics: ~1M LOC, 0 manual, ~1,500 PRs, 3.5 PRs/engineer/day, team 3→7, OpenAI estimate ~1/10th the time to write by hand.

Architecture pattern: AGENTS.md as TOC → structured docs/ → Codex for all implementation → agent-to-agent review → Chrome DevTools for UI verification → LogQL/PromQL for observability.

Recommended study for: Engineering leaders interested in the upper bound of agent-driven development. The AGENTS.md-as-TOC pattern is immediately applicable regardless of scale.

Not recommended for: Teams expecting to replicate this without OpenAI-level access to frontier models and infrastructure.

Outlook: The practice is already escaping OpenAI's walls — Symphony, the orchestrator that grew from this work, was announced as an open-source spec on April 27, 2026 (~27.4K stars as of September 23, 2026).[2][3] If Harness-style development becomes viable outside OpenAI, the economics of software engineering change fundamentally. The constraint is model capability — as frontier models improve, this approach becomes more accessible.

Tembo: Tembo is a commercial coding-agent orchestrator — isolated cloud sessions, selectable harnesses, and human PR review — complementary to Harness's internal methodology and Symphony's open-source spec, not a replica of OpenAI's agent-to-agent review practice.[6] Disclosure: Ry Walker is Tembo's co-founder and CEO.


Research by Ry Walker Research • methodology

Disclosure: Author is CEO of Tembo, which offers agent orchestration as an alternative to building in-house.