Maintainer merge decisions run about 24 percentage points below automated benchmark scores on SWE-bench Verified, according to a study by the AI research nonprofit METR.

The more striking number is the baseline. Even original, human-written reference patches, the ones the benchmark itself considers correct, only get merged by real maintainers 68% of the time.

Some of this gap is about AI. A lot of it is just what code review actually is.

This guide covers agentic AI coding specifically, building on the ticket-to-PR pattern already outlined in Tecla's Agentic AI in Software Development guide: what a benchmark score actually tells you, and why review stays the gate no team drops.

What Is Agentic AI Coding?

Agentic AI coding is software that plans and carries out a coding task on its own: reading a ticket, writing the change, running tests, and opening a pull request, rather than suggesting a line for a developer to accept or reject.

That's a meaningfully different job than autocomplete. A suggestion tool proposes the next few lines and waits. An agent takes a task from assignment to an open PR, making dozens of small decisions along the way that a developer would otherwise make one at a time.

None of that changes who's accountable for what ships. It changes where a developer's attention goes: from writing the first draft to deciding whether the draft is actually right.

From Benchmark to Merged: Why the Gap Exists

SWE-bench Verified tests whether a patch makes a specific, pre-written test suite pass. That's a real, useful signal. It's also a narrower question than the one a maintainer actually asks.

METR's study broke down why passing patches still got rejected.

The reasons, in rough order of severity: code quality that didn't match the repository's conventions, and changes that broke other code the test suite didn't happen to cover.

The most serious category was core functionality that didn't actually solve the underlying problem, despite passing the specific test written for it.

That last category is the sharpest reminder that a test passing and a problem being solved are not automatically the same thing.

The study also found later model generations improving mainly on code quality rather than raw correctness. What's improving fastest isn't whether the fix works. It's whether the patch looks like it belongs in the codebase it's touching.

Perceive ticket, issue
›
Retrieve codebase, conventions
›
Reason write, test change
↓
Human gate
›
Act open PR, merge
›
Verify
Verify feeds back into Perceive, refining the next attempt if rejected

The Write-Test-Ship Workflow, Step by Step

The workflow below covers the coding loop itself. The sections after it dig into what a benchmark score does and doesn't tell you, and why even a perfect test pass rate wouldn't remove the need for review.

The workflow

What it does: takes an assigned ticket, writes the change against the actual codebase, runs it against the existing test suite, and opens a pull request only once the change passes and follows the repository's own conventions.
1
Ticket or issue assigned with a description of the desired outcome
2
Codebase and existing patterns retrieved for context
3
Code change drafted against the actual repository, not an isolated snippet
4
Existing test suite run and the change revised if it fails
5
Human gate: developer reviews the pull request for quality and correctness before merge
6
Merge outcome and rejection reason, if any, logged for the next attempt
The stack: an agentic coding tool integrated with the repository, issue tracker, and CI pipeline, running against the project's existing test suite and style conventions.
Why it works: a well-scoped ticket with a clear success condition is exactly the shape of task current agents handle well, and running against real tests catches the failures a person would also catch.
Production concern: a passing test suite doesn't guarantee the change is actually correct, only that it satisfies whatever the test suite happened to check, which is precisely the gap that separates a benchmark score from a merge.

What SWE-bench Actually Measures

SWE-bench Verified draws from real GitHub issues in real open-source repositories, which is exactly why it's treated as one of the more meaningful coding benchmarks available.

What it measures precisely is whether a generated patch passes a specific, pre-selected set of tests. It doesn't measure whether a maintainer would want that patch in their codebase, and METR's research shows those two things diverge by a wide, measurable margin.

Why Even Human Patches Get Rejected

The most useful number in METR's study might be the human baseline: original patches that were already merged into the repository, resubmitted blind, only got approved again 68% of the time.

Some of that is genuine reviewer subjectivity, not a defect in the patch itself. Roughly 85% of those same human patches were rated as making at least 80% of the progress toward a mergeable state.

The last mile of review has always had some judgment call built into it, for AI and humans alike.

The Time Horizon Illusion

METR also measured something called time horizon: the length of task, by typical human completion time, that a model can complete with 50% reliability.

By the automated grader, one leading model's time horizon was around 50 minutes. By maintainer review, the same model's time horizon was closer to 8 minutes, roughly a sevenfold difference.

Whatever number a benchmark reports about task length, treat it as a ceiling, not an estimate of what a maintainer would actually accept.

The Developer's Role in Review

The developer's job doesn't shrink as agents write more of the first draft.

It shifts toward the exact judgment calls METR's rejection categories describe: does this actually solve the problem, does it fit how this codebase works, did it break something the tests didn't check.

Implementation: Guardrails Specific to Coding

Every guardrail below exists because a passing test and a mergeable change are two different bars, and only one of them is something an agent can fully verify on its own.

LayerWhat it doesCoding-specific example
System promptSets the non-negotiables up front"Match existing repository conventions, not just a passing test"
Input filtersBlock or sanitize out-of-scope requestsTreat ticket text and issue comments as context, not instructions to execute literally
Tool-call gatekeepersCap what actions an agent can takeDrafting and testing allowed; merging to the main branch always needs a human
Output checksScan before the action executesBlock any PR that touches code outside the ticket's stated scope without flagging it
Human-in-the-loopRequires approval for high-impact actionsA developer reviews and approves every pull request before merge

Rolling This Out: What to Expect

Pick a narrow, well-scoped ticket type first, ideally one with a clear, verifiable success condition, rather than open-ended feature work.

Track actual merge rate and post-merge defect rate, not the agent's own test-pass rate. Those numbers diverge, and the divergence is exactly what a benchmark score won't show you.

Expect rejection reasons to shift over time from correctness toward convention. Newer models tend to solve the stated problem more reliably; the harder remaining gap is usually fitting how a specific codebase actually works.

The Team Behind Production Agentic AI

A benchmark score tells you almost nothing about whether an agent will fit a specific codebase's conventions and history. Learning that takes deliberate measurement against your own repositories, not a leaderboard.

Tecla's Agentic AI services design, build, and operate this workflow directly, the same write-test-ship systems above, running in your stack with the evals and guardrails production requires.

Or bring the expertise in-house: AI engineers who've worked on live coding systems, past the demo stage.

Tecla runs a network of senior engineers across the US and Latin America, built over more than a decade, with a top 3% acceptance rate and first candidates in 3 to 5 business days.

FAQ

What is agentic AI coding?

It's software that plans and carries out a coding task on its own: reading a ticket, writing the change, running tests, and opening a pull request, rather than suggesting code for a developer to accept line by line.

Do coding agent benchmark scores predict real-world results?

Not directly. A study by the AI research nonprofit METR found maintainer merge decisions run about 24 percentage points below automated benchmark scores on SWE-bench Verified, and even human-written reference patches only get merged 68% of the time.

Why do agent-written pull requests get rejected if they pass their tests?

The most common reasons are code quality that doesn't match repository conventions, changes that break other code the automated tests didn't cover, and core functionality that doesn't actually solve the underlying issue despite passing the test written for it.

Does agentic AI coding replace developers?

No. It absorbs the first draft of well-scoped changes. A developer still reviews every pull request, and that review is where most of the gap between a passing test and a mergeable change actually gets caught.

What are the risks of agentic AI coding?

The main risk is treating a passing automated test as equivalent to a production-ready change. Research shows that gap is real and substantial, which is why code review before merge remains the primary control regardless of how capable the underlying model is.

How should a team start with agentic AI coding tools?

Start with a narrow, well-scoped ticket type where success is easy to verify, and track actual merge rate and post-merge defects, not just whether the agent's own tests passed, before expanding its scope.
Gino Ferrand
By 
Gino Ferrand
Gino Ferrand
Gino is an expert in global recruitment having spent the last 10 years leading Tecla and helping world-class tech companies in the U.S. hire top talent in Latin America.
Categories
AI Production Insights
Insights
Reviews
Recruiting
Case Studies
LATAM Reports
Management
Mobile Hero Image
Combine AI speed with LatAm engineering talent.
Software Developer
We map what you have and scope the AI transformation your business needs.
Get free agentic AI audit
Go to Top

Hire the best AI-driven tech talent with Tecla

Premium, vetted, time-zone aligned.

Checkmark
Checkmark
Checkmark
By submitting, you are agreeing to our Privacy Policy and Terms of Service
Thank you!
Someone from our team will be in touch within 24 business hours.
Something went wrong while submitting, please try again
x
X

Tell us where you're stuck

Checkmark
Checkmark
-
No commitment. We'll follow up within 1 business day.
By submitting, you are agreeing to our Privacy Policy and Terms of Service
Thank you!
Someone from our team will be in touch within 1 business day.
Something went wrong while submitting, please try again
X