Enterprise AI agents: 86% shipped. Only 34% trust them.
Enterprise AI agents don't fail on model quality. 86% got past pilot, only 34% are trusted, and the 88% failure stat everyone quotes has no source at all.

There is a number running the enterprise AI conversation right now. 88% of AI agent pilots never reach production. It is in board decks, vendor pitches, and roughly every LinkedIn post about agentic AI written since May.
I went looking for the study. It does not exist.
That matters more than the number does. Because the reason enterprise AI agents stall in production is the same reason that statistic spread unchecked: almost nobody in this market verifies anything before they build on it.
Key takeaways:
- The widely quoted "88% of agent pilots never reach production" traces to no primary source. Every citation credits Anaconda and Forrester. Neither company publishes it.
- Forrester Consulting's actual survey of 409 IT decision-makers found the near opposite: 86% of organisations have moved beyond the pilot stage. Only 34% trust what their agents do.
- MIT's Project NANDA found 95% of GenAI initiatives return nothing, and attributed it explicitly to a learning gap, not model quality or regulation.
- Gartner expects over 40% of agentic AI projects to be cancelled by end of 2027 on cost, unclear value, and weak risk controls.
- The bottleneck on enterprise AI agents is not intelligence. It is people who can define success criteria, wire data access, and build evaluation before anyone writes an agent.
The number everyone quotes, and nobody can source
Start with the claim. 88% of AI agent pilots die before production.
Follow the citations. Every article credits "a 2026 joint study by Forrester Research and Anaconda." Some add that it was "replicated independently by a16z and an MIT Sloan CIO panel." None of them link it.
So I checked the companies directly. Anaconda does not publish that figure. Forrester's own 2026 state-of-the-market post on agentic AI, written by VP Brian Hopkins in June, says something different and much vaguer: three quarters of enterprise leaders say they are adopting agentic AI, and "only a small minority have it running in meaningful production beyond 'agentish' chatbots."
A small minority. No percentage. Certainly no 88.
It gets worse when you read the aggregators against each other. One publishes a full methodology for the 88%: a survey of 3,700 practitioners, with stalled pilots splitting into still-in-evaluation (34%), cancelled (29%), and paused (25%). Those sum to exactly 88. Another site says the figure actually came from a 650-person survey run by a marketing agency in February 2026. A third says it "originated in Anaconda and Forrester research" and moves on.
Three incompatible origin stories for one number. That is not research. That is a rumour with a decimal point.
What Forrester actually published
Here is the part that should end the 88% for good.
In July 2026, Boomi commissioned Forrester Consulting to survey 409 director-level-and-above IT and technology decision-makers across North America, Europe, and APAC. The published finding is that 86% of organisations have moved beyond the AI agent pilot stage.
Eighty-six percent past pilot. Against a viral claim of 88% never getting there. The most-cited number in enterprise AI is a near-perfect inversion of the closest real number attributed to the same firm.
I am not going to assert that someone flipped one into the other. I cannot prove that. What I can say is that the industry adopted a statistic that contradicts the published research it claims to come from, and nobody checked for four months.
Read that study properly, though, and the real finding is sharper than the fake one. Of those 86% running agents beyond pilot, only 34% trust the actions their agents are taking.
The gap is not deployment. The gap is trust.
Worth naming the bias: Boomi sells the integration layer the study concludes you need. Vendor-commissioned research points where the vendor wants it to point. The sample size, methodology, and date are published, which is 86% more than the number it replaces offers.
Why the credible numbers still disagree
Line up the studies on enterprise AI agents that do show their working, and they still do not agree with each other.
- Forrester Consulting, n=409, July 2026: 86% beyond pilot.
- LangChain's State of Agent Engineering, n=1,340, surveyed 18 November to 2 December 2025: 57.3% have agents in production, up from 51% the year before.
- MIT Project NANDA, January to June 2025: only 5% of custom enterprise AI tools reach production.
- Gartner, June 2025: over 40% of agentic AI projects will be cancelled by the end of 2027.
Five percent to eighty-six percent. That is not measurement error. That is four organisations using four definitions of the word "production" and never agreeing one.
MIT counted custom-built enterprise tools with measurable KPIs, six months post-pilot. Forrester counted anything past the pilot stage. LangChain surveyed a sample that was 63% technology companies, which is the population most likely to succeed at this.
Gartner adds the reason the definitions stay loose. It calls the practice "agent washing," the rebranding of assistants, RPA, and chatbots as agents without substantial agentic capability. Gartner estimates only about 130 of the thousands of agentic AI vendors are real.
When the category itself is undefined, every number in it is soft. Including the ones I just cited.
What actually stops enterprise AI agents
Across every study with a published methodology, one finding holds. The blocker is not model capability.

MIT was explicit that the divide "does not seem to be driven by model quality or regulation, but seems to be determined by approach." Their diagnosis was a learning gap: deployed systems that do not retain feedback, adapt to context, or improve over time.
Gartner's cancellation forecast lists escalating costs, unclear business value, and inadequate risk controls. Not accuracy.
The Forrester Consulting data puts a price on it. Organisations in the bottom quartile for operational readiness, what the study calls "agentic chaos," are shipping anyway. 77% of them move into production without governance, integration, or API management in place, and carry an average of $2.1 million in added cost from compliance fines, lost customers, downtime, and rework.
The top quartile behaves differently. 55% report high confidence in their agents' decisions, against 22% in the chaos group. They also report the returns: 59% saw productivity gains, 51% increased innovation, 46% built reusable capabilities.
And the LangChain data shows where the discipline actually breaks. Only 52.4% of teams run offline evaluations on test sets. Only 37.3% run online evaluations. Roughly one in four does both.
Most teams shipping autonomous software into production have no systematic way to know when it degrades.
That is the whole story. Not intelligence. Instrumentation.
The skill nobody is hiring for
The job posting says AI engineer and lists model APIs, vector databases, and framework names. Which is the same category error I keep seeing everywhere else: AI tools are not a strategy. Everyone has the same tools. Nobody has the same operating discipline.
The work that decides whether the agent survives contact with a business is different, and it happens before anyone writes a prompt:
- Define what success means in numbers, with a baseline from before the agent existed.
- Map which systems and data the agent needs, then get real access, which is usually a procurement and security problem, not a technical one.
- Build the evaluation harness first. Offline on a test set, online in production.
- Name one person accountable for the thing after launch.
- Decide in advance what result kills the project.

None of that is model work. All of it is engineering discipline applied to a probabilistic system. It is the same discipline I wrote about in building an AI content pipeline that outlasts the model: lock the contracts between stages, and the system survives the model changing underneath it.
It is also, precisely, the discipline that would have caught the 88%.
Notice the symmetry. A field that cannot verify its own headline statistic is the same field that ships agents without baselines. Both failures come from the same missing habit: check the claim before you build on it.
This is why I keep arguing that the constraint on enterprise AI is a skills constraint, not a model constraint. I made the macro version of that case in reskill or die. The micro version is this: the market is short of people who can turn an agent into a governed, measured, owned system.
At Experis Academy, that is what the AI Engineering Academy is built to produce. The Recruit, Train, Deploy track puts experienced software developers through a ten-week lab-based immersion covering AI systems foundations, LLM and knowledge engineering, RAG optimisation, integration patterns, and deployment platforms, assessed before and after, ending in a capstone project. The user-training track runs the same logic at organisational level across three stages: literacy, productivity, then agentic AI, closing on a code-along where teams prototype agents against their own real processes.
Ten-week immersion, not a certificate. Experis Academy has run programmes since 2019 and reports over 90% retention two years after deployment, at a 96% placement rate into IT roles. Placement and retention are the numbers that matter, because they are the ones measured after the training ends.
The point is not the curriculum. The point is that "productive from day one" is a claim you can only make if you have defined what productive means and measured it. Same standard the agents should be held to.
What I check before anyone writes an agent
I run a sales organisation and I build these systems. That combination is why the gap is obvious to me from both sides. The people funding agents cannot measure them. The people building agents were never asked to.
Five questions, before a line of code:
- What number moves, and what is it today? No baseline, no project. Gartner's May 2026 survey of 227 chief sales officers found 31% cited difficulty proving ROI of AI tools as a top challenge to hitting 2026 objectives. That is a measurement failure, and it starts on day one.
- What does the agent need to read and write, and do we have that access? Answer it before the build, not during.
- How do we know when it degrades? If the answer is "someone will notice," there is no answer.
- Who owns it in six months? Name a person.
- What result would make us kill this? Decide while it is still cheap to walk away.
Anything that clears those five is worth building. Anything that does not is a demo with a budget attached.
The 88% was never the problem. Quoting it without checking was. Enterprise AI agents will keep stalling as long as the people funding them accept numbers they have not verified, and the people building them ship systems they cannot measure. Both halves of that are trainable. Neither is a model problem.
More on what I work on is on my about page.