The middle of the software factory is already automatable. Spin up the environment, write the change, run the tests, run an adversarial review, move the card. What is not automated, and should not be waved through on hope, is the production cut. That decision is still a human queue in most companies, which means it becomes the bottleneck the moment the rest of the line speeds up.
The right shape for that gate is a classifier, not another reviewer. Most changes are boring in a way that is easy to name. No schema change. No public API change. No security-sensitive change. If a card clears those checks, and the automated review already beat it up, it should ship. If it trips any of them, a person looks. The point is not to remove judgment. The point is to spend judgment where the blast radius actually lives.
This only works if the earlier stations are real software, not a prompt hoping for the best. No Human Should See the PR Until It Has Been Beat Up is the precondition. A classifier sitting on top of an unreviewed diff is just a faster way to ship mistakes. The factory has stations: draft, test, review, security, ship. Agents staff the stations. The classifier is the policy at the last one.
Two gaps remain once you see the line this way. The funnel in, meaning what is worth building, held at the level of a business outcome rather than a pile of tickets. And the production cut, meaning which finished cards are allowed to leave without a person. Everything between those two should be background execution with a reviewable trail. That is the pattern that holds in an enterprise, because a leader can inspect the policy and the output without sitting in the path of every card.
Build the classifier with criteria your security and platform owners will sign. Then measure how much of the line still waits on a human who is not adding judgment. That wait is the next thing to remove.
Key takeaways
- The production decision should be a classifier with explicit risk criteria, not a standing human queue.
- Auto-ship belongs only to changes with no schema, public API, or security blast radius, and only after automated review.
- Once the middle of the line is automated, the remaining gaps are what enters and what is allowed to leave.
FAQ
Which changes should skip a person at the production gate?
Changes with no schema change, no public API change, and no security-sensitive change, after automated tests and adversarial review have already run. Anything that trips those criteria gets a person.
Does a ship classifier remove human judgment?
No. It concentrates judgment on the policy and on the changes with real blast radius. Routine cards should not wait in a queue for a review that adds nothing.
Related Essays
No Human Should See the PR Until It Has Been Beat Up
As agents multiply code volume, reading diffs stops scaling. The answer is adversarial verification, where agents try to break the change before a human ever reviews it.
Merging Code Nobody Understands Is the Real Risk
Autonomous merge sounds like speed until nobody in the org knows what the product is anymore. Human review is how a company keeps a mental model of its own software.
Agents Start Broken and That Is the Point
New agents should deploy in supervised mode by default because no agent is production-ready on day zero.