AI Agents Write 741% More Code, but Review Limits What Ships
AI agents have sharply increased code output without a comparable rise in software shipped: developers using autonomous agents wrote 741% more code, but shipped only 30% more software. Laurie Voss, Arize AI’s head of developer relations and npm co-founder, argues that human review is now the bottleneck—and that asking people to review more code is not a scalable fix. The answer, she says, is to build automated review systems that test more than whether code passes, while keeping human judgment and production monitoring where automated checks fall short.

Code output has outrun the capacity to trust it
Laurie Voss frames the new constraint in software development as a widening gap between code generation and code review. Developers using autonomous agents wrote 741% more code, according to a study that tracked more than 100,000 GitHub developers and matched them with telemetry on when they began using AI. But software shipped rose only 30%.
Voss says the study’s authors identified review as the bottleneck. Producing plausible code has become much cheaper; deciding whether that code is safe and sound enough to trust has not. The mismatch is especially consequential in sensitive systems, where a change can have a large blast radius.
The cost of producing plausible code has collapsed. The cost of knowing whether to trust it has not.
The gap is visible in the scale of work agents can now take on. Voss cited a reported migration of a 50-million-line Ruby codebase in a day, work that had been estimated to take a team more than two months. Bun also reported moving more than a million lines from Zig to Rust in six days. At smaller scales, developers are already generating whole applications and merging code they have not read, hoping it works.
That practice is not a stable answer. But neither is simply asking people to review more. A Cisco study conducted over 10 months examined 2,500 reviews and 3.2 million lines of code. Voss summarized its finding this way: reviewers become much less effective after reading more than 400 lines in one sitting, and effectiveness drops sharply when they review more than 450 lines in an hour. At that pace, she calculated, a 10,000-line pull request could take three or four working days to review properly. One agent can produce that much code in a pull request; a developer may run several agents at once.
The constraint is not just that code review takes time. It is that attention degrades at the volumes agents make possible. “You can’t just review harder,” Voss said.
Passing tests answers a narrower question than “would you merge this?”
Some developers are trying to remove human inspection from the process altogether. Voss pointed to arguments from Peter Steinberger and Andrej Karpathy that the key is designing the loops that prompt and check agents, rather than keeping a person in the loop at every step. OpenAI has described an internal product built from an empty repository with no manually written code. Five months in, the project had about a million lines of code and roughly 1,500 merged pull requests, built by three engineers. Agents also wrote the scaffolding used to review other agents. OpenAI’s stated policy was that humans could review pull requests, but were not required to.
Voss sees the experiment as evidence that agent-to-agent review is possible, not proof that it is solved. OpenAI did not disclose what the product did or release the system as a reusable implementation. The important caveat is what the review loop can see, and whether it can judge more than whether the code passes tests.
For years, tests have served as the main proxy for code quality in benchmarks. A METR study tested that proxy against human maintainers. Four active maintainers from projects used by SWE-bench reviewed pull requests that had already passed the benchmark’s grader. The maintainers judged only about half of those pull requests good enough to merge.
The rejected work was not necessarily functionally incorrect: it had passed the tests. The concerns included code quality and changes that quietly broke other code outside the test suite. METR noted that the agents did not get a chance to revise their work after receiving feedback. Voss acknowledged that limitation, but argued that allowing the iteration would put a human back into the loop the experiment was supposed to assess.
Cognition’s FrontierCode benchmark asks a more demanding question: would a maintainer merge the change? More than 20 maintainers built 150 tasks from their own repositories, each involving more than 40 hours of expert work. Its rubric covers behavioral correctness, regression, safety, scope, discipline, test quality and maintainability.
The two reported Claude 3.5 figures differ: Voss said the model scored 88% on SWE-bench Pro, while the slide shown during the talk gave its score as 80.3%. On the hardest slice of FrontierCode, Voss said, it scored 29.3%. GPT-4o scored below 6% on FrontierCode. The comparison suggests a large gap between passing tests and producing code maintainers would accept.
A mergeability benchmark would matter beyond evaluation. Voss argues that once maintainability, scope, regression and safety can be measured reliably, those measures become training targets. Compilers and test suites helped coding models improve because they provide cheap ways to verify outputs. A benchmark for mergeability could become the next such signal. In that sense, the people who define what a good review means would also shape the default behavior of future models.
There is precedent for that feedback loop. Voss pointed to OpenAI’s CriticGPT, trained to catch bugs in model-written code. In the setup she described, human reviewers working with the model outperformed either humans or the model alone. The resulting review process helped produce a training signal for improving models.
Automated review is already in production—and its main job is avoiding false alarms
Automated code review is not a distant possibility. Voss said GitHub Copilot’s reviewer had performed 60 million reviews and accounted for more than one in five code reviews on GitHub. Cursor has also published details of its review system, and those details point to a central operational problem: false positives.
Cursor’s first reviewer ran eight passes over a diff and shuffled the order of reviewers to examine the code again. The order changed the results. The goal was not simply to find more issues, but to filter findings that were wrong. If an automated reviewer repeatedly flags acceptable code, developers stop paying attention to it.
Voss cited separate research from Peking University that tested multiple review passes, keeping findings on which the passes agreed. The researchers found that this approach improved review quality by as much as 44%. For Voss, the repeated rediscovery of multi-pass review reflects how damaging false positives are to a reviewer’s usefulness.
Cursor’s later system reasons over the diff, calls tools and chooses where to investigate. One adjustment stood out to Voss: the model had to be told to trust the code less. Its default tendency was to see code that looked fine and approve it. Effective review, she argued, requires the opposite stance—assume there may be a problem and investigate.
Review is also beginning to merge with repair. Cursor’s reviewer can spawn a fix agent based on its findings, which returns a patch for a person to approve. Cursor has also said it wants the reviewer to run the code and demonstrate that a reported bug is real. That blurs the boundary between identifying a problem and rewriting the code, even if a human still approves the resulting diff.
Other vendors are building around similar feedback. Voss said CodeRabbit had reviewed more than 13 million pull requests. Greptile builds a repository graph so the reviewer can account for changes in distant parts of the codebase. Graphite uses developers’ acceptance and rejection of suggestions to construct its evaluation set. Across these systems, she said, a shared measure of success is whether a human accepts the answer. Cursor calls this its resolution rate, which it had raised from 52% to more than 70%.
That acceptance data gives review systems a way to learn what developers consider useful. It also helps explain why false positives matter so much: a reviewer that users routinely reject has less chance of influencing what gets fixed.
Removing the reviewer does not remove the need for human-designed checks
Two prominent experiments show what it means to take people out of the immediate review loop—and what remains in it. In one, Anthropic’s Nicholas Carlini used 16 agents to build a C compiler from scratch in Rust. Across roughly 2,000 sessions, the compiler was able to compile the Linux kernel without a human approving code as it was written.
But the system was not without human oversight. People wrote the system that reviewed the code, the test harness and the feedback mechanisms that checked whether it did what it was supposed to do. Carlini’s warning, as Voss relayed it, was that watching tests pass can make it easy to assume the job is finished when it is not.
The Bun migration makes the distinction concrete. Agents ported about a million lines from Zig to Rust in six days, and no human read the entire diff. The existing test suite passed at 99.8%, providing a real check on behavior at the project’s public interface. Voss said the ported code had 13,044 unsafe blocks, compared with about 73 or 74 in a human-written Rust codebase of similar size.
| Codebase | Unsafe blocks |
|---|---|
| Bun's agent-built port | 13,044 |
| Comparable human-written Rust project | About 73–74 |
An unsafe block is a place where the author asserts, rather than proves, that memory is being handled correctly. Tests can check the behavior they were designed to check; they cannot certify thousands of such assertions if the test suite was not built to examine them. In Voss’s account, the migration shows both the power and the limit of automated testing: humans can be removed from line-by-line review while leaving unresolved questions beneath the tested interface.
OpenAI’s system similarly did not eliminate review so much as relocate it. Codex reviews its changes, brings in more agents to review those reviews, and repeats the process until the reviewers are satisfied. The system can boot and run a copy of Codex to inspect its own interface and test whether a bug appears fixed. OpenAI also exposed its logging stack to the agent.
Even there, some cleanup initially remained manual. Voss said the team spent Fridays removing “AI slop” by hand, then trained agents to identify and clean it up when that work stopped scaling. Her reading is that review became an engineered system, not a task that disappeared.
The experience of Dex Horthy points to the cost of skipping inspection in a working system. After spending about six months advocating that teams ship without reading agent-written code, Horthy publicly reversed his position: “Please, please read the code.” He said the experiment had not ended well and that the team had to rip out and replace large parts of the system. For Voss, this was not a benchmark result; it was a report from someone who tried the approach on real software and changed course.
Tests also lack organizational context. Voss, citing an argument from investor Sarah Guo, noted that a passing test does not establish that a change was the right one. A module may exist because external users depend on it, or because an obscure cron job relies on it. That context may not be represented in the test suite.
The human role moves up the stack, but production remains the final check
Voss sees human judgment surviving wherever correctness is not cheap to verify, the potential blast radius is large, or someone must take responsibility for the result. In security-sensitive environments, she said, proposals to eliminate human review tend to meet immediate resistance. The work changes level: instead of inspecting every line, people design and tune the systems that inspect code and define what counts as good.
That shift raises a second-order problem: who reviews the reviewers? Voss cited a warning in Anthropic’s automated security reviewer documentation that the action is not hardened against prompt-injection attacks and should only be used on trusted pull requests. She said the reviewer can be talked out of its findings by the code it is reviewing.
A study she discussed found that vulnerable code disguised with an innocent commit message fooled an autonomous review agent in 88% of attempts. The same attacks passed a human reviewer only 35% of the time. Removing the human, Voss argued, does not merely remove one checkpoint; it can remove the reviewer that is harder to fool.
There is also uncertainty about how to grade automated reviewers. Voss said benchmark setups have been found leaking answers, and researchers disagree about how to evaluate review quality. When pre-merge review is handled by machines, production becomes the final reviewer that cannot be automated away.
After deployment, the question is not only whether the tests passed, but what the system actually did in the real world. Voss described observing the system’s trajectory—its step-by-step behavior—as essential when shipping automatically reviewed code. Production monitoring and evaluations, in her view, become a necessary backstop when humans no longer inspect every change before it ships.
Her practical recommendation is to “stop reviewing PRs”—not because human judgment is dispensable, but because line-by-line review is the wrong place to spend all of it. Teams should build a review harness that encodes their standards, company context and domain knowledge, then use agents against that system. Code review is changing from a person reading every diff into a set of benchmarks, rules, tools and evaluations that humans design and supervise.
The productivity gain, Voss concluded, will not come simply from generating more code. It will come from being able to explain, with evidence, why the code that shipped can be trusted.
