Your pilot isn't stuck on reliability. It's stuck on a number no one will name.

Eighty-five percent of enterprises are piloting AI agents and only five percent have shipped one. That trap has a name now, AI pilot purgatory: the state where a working proof of concept keeps running through review cycles without ever reaching production. Everyone blames reliability, and reliability is real. But the deeper reason so many pilots never end is that almost no one decided, before they started, what failure rate would be good enough to ship. A pilot with no defined finish line runs forever, always almost ready. The blocker is not that the agent cannot be trusted. It is that no one will put their name on how much failure is acceptable.

There is a statistic making the rounds that should stop every technology leader cold. By Cisco's data, cited by Amazon's AGI autonomy director at a conference last month, 85 percent of enterprises are piloting AI agents and only 5 percent have moved them into production. Eighty points of enterprises, stuck between “we’re trying it” and “we shipped it.”

The standard explanation, and the one Amazon offered, is reliability. The agents work in the demo and fall apart in the wild, so they never earn the trust required to go live. That is true, and it is worth taking seriously. But I want to name the thing underneath it, because after a decade of watching companies try to put new systems into production, I have seen this exact gap before, and it is rarely a pure technology problem. It is a decision problem wearing a technology problem's clothes.

Here is the decision almost nobody makes: what failure rate is good enough to ship.

A pilot with no finish line runs forever

Think about what a pilot actually is. It is a test. And a test needs a passing grade defined in advance, or it is not a test, it is an open-ended science project. Most enterprise agent pilots are launched without anyone answering the one question that would let the pilot end: how well does this have to work before we turn it on for real?

Without that number, the pilot cannot conclude, because there is always a failure you can point to as the reason to wait. The agent gets the serial-number extraction right 96 times out of 100, and someone asks about the other four. It handles the common cases and stumbles on an edge case, and the edge case becomes the reason to keep piloting. Every system, including the humans it would replace, fails sometimes. But when you never decided what failure rate was acceptable, every observed failure reads as proof it is not ready yet. The bar is invisible, so nothing ever clears it.

This is why pilots do not fail so much as they fail to end. They do not blow up. They just quietly continue, quarter after quarter, always improving, never shipping, until the budget runs out or the sponsor loses interest. Most of the AI pilot to production gap is made of exactly this: not agents that failed, but agents that were never given a bar to pass. Gartner expects more than 40 percent of agentic AI projects to be cancelled by 2027, citing escalating costs and unclear business value. Most of those will not die because the technology failed. They will die because nobody ever set the threshold that would have let them succeed.

Every industry that ships anything real solved this first

The reason this feels hard is that AI is being held to a standard we apply to almost nothing else: perfection. And the moment you look at any other field that puts consequential systems into the world, you see they abandoned that standard long ago, on purpose, as the price of shipping at all.

Aviation is the clearest case. It is the safest complex system humans have built, and it did not get there by demanding zero risk. It got there by defining, explicitly and numerically, an acceptable level of safety before certifying anything. The governing principle is stated plainly in the industry’s own guidance: absolute safety is unachievable, so the field sets a target level of safety instead, a specific tolerable probability of failure, and certifies systems against it. U.S. aviation fixes the probability of a catastrophic failure condition at no more than one in a billion per flight hour. Nuclear regulators set explicit core-damage frequencies. These are not aspirations. They are numbers, decided in advance, that let an engineer say “this is good enough to operate” and mean it.

Notice what that number does. It converts an unwinnable argument about whether something is safe into an answerable question about whether it clears the bar. It gives the pilot a finish line. Software has its own version, the “five nines” of uptime, which is nothing more than a decision to accept about five minutes of downtime a year rather than chase an impossible zero. Every industry that ships real systems made peace with a defined, non-zero failure rate. AI agents are the first technology in a while that a lot of companies are trying to deploy without doing that, and then wondering why deployment never happens.

The number is hard because it means owning the failures

If setting a threshold is so clarifying, why does almost no one do it? Not because it is technically difficult. Because it is uncomfortable in a way that has nothing to do with engineering.

Defining an acceptable failure rate means writing down, in advance, that your system will fail a known percentage of the time, and that you decided to ship it anyway. It means owning those failures before they happen, with your name on the decision. That is a very different act from saying “we need it to be more reliable,” which sounds responsible and commits you to nothing. “Make it more reliable” is the safe thing to say in the meeting. “I am willing to ship at a 2 percent error rate and I will answer for the 2 percent” is the thing almost no one wants to say, because the moment you say it, the failures become yours.

So “it’s not reliable enough yet” becomes a place to hide. It is technically always true, since nothing is perfectly reliable, and it never requires anyone to accept responsibility for a live system that will sometimes be wrong. The pilot becomes a way to postpone the decision indefinitely while looking diligent. The reliability conversation, real as it is, quietly serves as cover for a decision nobody wants to sign.

What I'd actually do about it

If I had an agent stuck in pilot, I would stop asking “is it reliable enough yet” and force the prior question: what would good enough look like, as a number, and who owns it.

Concretely. Before the pilot, not after, write down the acceptable failure rate for the specific workflow, in the same terms you would judge a human doing the job, because the honest benchmark is almost never perfection, it is the error rate you already tolerate from people and have stopped noticing. Measure the pilot against that number and nothing else, so a stumble on a known edge case is scored against the bar rather than treated as an automatic veto. Name an owner who has the authority to say “it cleared the bar, we are shipping,” and give them the air cover to be wrong sometimes, because a team punished for the first live failure will keep everything in pilot forever. And design for the failures you have accepted rather than pretending they will not come: the checkpoints, the escalation paths, the rollback, so that the tolerated failure is contained instead of catastrophic. That last part is the real reliability work, and notice it only becomes possible after you have decided what failure you are engineering around.

None of this makes the agent smarter or even more reliable. It makes the organization capable of deciding, which is the thing that was actually missing. It is telling that the practitioners who study how teams escape pilot purgatory keep landing on the same first move: a written exit criterion, an accuracy threshold defined up front against a real sample, before the build. The teams that get out are the ones that decided what passing meant before they started. The ones that stay stuck are still arguing about it.

The honest objection

The strongest pushback is that some workflows genuinely cannot tolerate the failure rates today's agents produce, and that for those, staying in pilot is not cowardice, it is correct. If an agent is touching funds movement, or medical decisions, or anything where a single failure is catastrophic and irreversible, then “we need it to be more reliable” is not a dodge, it is the right answer, and the acceptable failure rate really might be close enough to zero that current agents cannot clear it. That is real, and I do not want to wave it away.

But notice that even in those cases, the discipline holds: you still have to decide the number. “This workflow requires a failure rate we cannot yet hit, so we are not shipping” is a defined threshold and a real decision. It ends the pilot honestly, with a clear reason and a clear bar to clear later, instead of leaving it to drift. The failure mode I am describing is not companies that consciously decide the stakes are too high. It is the far larger group piloting low-stakes, high-volume workflows, the internal QA checks and data extraction and routine triage, where the tolerable failure rate is plainly not zero, and who still cannot ship because they never let themselves say so.

So I will concede what I cannot prove. Maybe reliability improves fast enough that the threshold question gets easier and more pilots clear the bar on capability alone. But I would not bet the next two years on the models saving you from the decision. The companies that break out of pilot purgatory will not be the ones that waited for a perfect agent. They will be the ones with the nerve to write down what good enough means, put a name next to it, and ship the imperfect thing on purpose.

 

Gino Ferrand is the founder and CEO of Tecla, which builds and operates AI systems for U.S. companies and staffs the senior engineering teams behind them, across the U.S. and Latin America. He writes Founder's View, a weekly operator's take on the AI news that actually changes how technology companies build. This piece began as an issue of his newsletter, Redeployed. Talent is everywhere; opportunity is not.

‍

Gino Ferrand
By 
Gino Ferrand
Gino Ferrand
Gino is an expert in global recruitment having spent the last 10 years leading Tecla and helping world-class tech companies in the U.S. hire top talent in Latin America.
Categories
AI Production Insights
Insights
Reviews
Recruiting
Case Studies
LATAM Reports
Management
Mobile Hero Image
Combine AI speed with LatAm engineering talent.
Software Developer
See how much you'll save with AI-enhanced nearshore teams
Calculate my Savings
Go to Top

Hire the best AI-driven tech talent with Tecla

Premium, vetted, time-zone aligned.

Checkmark
Checkmark
Checkmark
By submitting, you are agreeing to our Privacy Policy and Terms of Service
Thank you!
Someone from our team will be in touch within 24 business hours.
Something went wrong while submitting, please try again
x
X

Tell us where you're stuck

Checkmark
Checkmark
-
No commitment. We'll follow up within 1 business day.
By submitting, you are agreeing to our Privacy Policy and Terms of Service
Thank you!
Someone from our team will be in touch within 1 business day.
Something went wrong while submitting, please try again
X