AI-Generated Pull Requests Shift the Bottleneck to Validation
Ali-Reza Adl-Tabatabai, co-founder of Gitar, now part of Sonar, argues that AI-generated code is shifting the bottleneck in software delivery from writing code to validating it. More and larger pull requests put pressure on CI and code review, leaving teams to slow delivery or risk rubber-stamping defects, he says. His proposed response is agentic validation: software that reviews pull requests, diagnoses and fixes CI failures, and can approve and merge changes under rules set by the team.

AI-generated code moves the bottleneck to validation
? ali-tabatabai argues that code generation makes a centralized part of software delivery—CI and code review—more important and harder to manage. That phase enforces security, compliance, and code-quality gates before changes reach production. It is also on the critical path: delays there slow delivery.
As engineering teams grow, builds take longer, reviews span teams and time zones, and the systems behind validation become more expensive to integrate and operate. Failures add an “outer loop” measured in hours or days. Developers lose time and flow, while platform teams must maintain fragmented tools, bespoke configuration, and scripts—often with constrained headcount.
AI-generated code intensifies the load: more pull requests, larger changes, and more opportunities for defects. Adl-Tabatabai describes the resulting choice as either slowing down for careful review or rubber-stamping changes and risking production incidents. Both, he says, undermine developer productivity and morale.
Either you slow down and ask developers to review every PR very carefully, or you rubber stamp PRs and you risk incidents in production.
Validation automation spans review, repair, and merge
The proposed response is an agent that handles the validation workflow, from reviewing a pull request to fixing issues and, under defined conditions, approving and merging it. The goal is to produce green pull requests ready to merge, while keeping review comments focused on issues that matter.
The review can include custom repository rules and checks. In the demonstration, the bot flags a missing issue link and a pull-request description that does not follow the repository’s required format. Teams can also configure it to block a merge while code-review findings remain unresolved. These controls are separable: teams can use the review without blocking, or make unresolved findings a condition for merging.
The agent also analyzes CI failures and summarizes likely causes. One displayed example traces a build failure to a TypeScript type mismatch: a function returns a string where a number is expected. Adl-Tabatabai says the system is particularly useful for identifying flaky tests, and teams can configure automatic retries for them.
For findings in code review or CI, users can request a specific fix or configure the agent to keep working through issues until the pull request is green. The demonstrated code-review fix, for example, adds checks that numeric inputs are finite and non-negative, then adjusts the return value to match the expected type. The bot presents a proposed change that can be applied from the pull request. Adl-Tabatabai describes this as a loop: the agent can address review findings and CI failures, rerun validation, and continue until the issues are resolved and the PR is green.
The final step is conditional automation: rules determine which green pull requests the agent may approve and merge. A displayed example shows an agent-approved PR merged after meeting specified criteria, including passing the pipeline and required approvals. Review, blocking, fixing, and merging are distinct controls; teams can decide which steps to delegate.
Trust determines how far teams delegate
Adl-Tabatabai describes users building trust by first assessing whether the agent’s reviews are accurate and raise important issues. Teams can then enable blocking, use automatic fixes, and eventually set rules for approval and merging. In his account, greater trust allows users to unlock more automation and get more value from the system.
The sequence matters because approving and merging gives the agent more authority than commenting or proposing a fix. Teams can keep that authority bounded by setting conditions for which PRs qualify, such as requiring a green pipeline and required approvals. In this way, the workflow controls let teams delegate progressively while retaining rules around when a change can reach the main branch.
Processing every PR can reveal patterns across the team
Because the system processes each pull request, it can also surface operational insights. For platform teams, categorizing CI failures can help distinguish recurring test-flakiness problems from infrastructure issues or other sources of failure. A displayed dashboard groups failures such as linting, unit tests, type checking, builds, and integration tests.
For engineering leaders, categorizing pull requests offers a view of what work is being delivered: feature development, fixes, chores, refactors, tests, and other types. The demonstrated dashboard covered 6,620 reviewed pull requests; features accounted for 40%, fixes 19%, and chores 16%. Adl-Tabatabai framed these classifications as a way to understand the mix of work across an engineering organization.
The system combines workflow orchestration with analysis
Gitar’s architecture has a control plane for orchestrating pull-request workflows; a custom agent runtime, or harness, for coordinating multiple agents, context, memory, tool calls, and integrations; and an LLM proxy for routing requests across models and handling failover.
Adl-Tabatabai says the team built its own harness to tailor agent execution to validation work, including its context management and tool use. He says that design lets the team optimize token cost, precision, and coverage, and deliver outcomes at a fixed price per pull request.
Following Gitar’s acquisition by Sonar, the team has been integrating it with SonarQube. Adl-Tabatabai sees potential in combining agents with program-analysis techniques, including taint, control-flow, and data-flow analysis, as well as software composition analysis. He expects combining those methods with agents to improve precision and coverage at a favorable price point. He also attributes better quality and security, faster delivery, higher developer productivity and sentiment, and greater leverage for platform teams to the broader automation approach.