Manifesto

Why 95% of enterprise AI pilots fail, and what they all have in common

MIT NANDA found 95% of enterprise AI pilots deliver no measurable impact. The failure is structural, not technical. Here is the missing layer every pilot skips.

Ataberk Taçar 12 min read
Cover art for an essay on why 95% of enterprise AI pilots fail

Key takeaways

  • MIT NANDA found 95% of enterprise AI pilots deliver no measurable P&L impact, and Gartner projects more than 40% of agentic AI projects will be cancelled by 2027.
  • The model is almost never the problem. Pilots break because the company cannot tell the agent how the work actually gets done.
  • Roughly 80% of processes are undocumented and 42% of essential expertise lives only in employees' heads, so the substrate agents need does not exist anywhere they can reach.
  • Demos pass because they are scoped to the documented 20% of cases; production is the other 80%.
  • A large multi-brand insurance group cut a renewal workflow from six minutes to one after mining the underwriting expertise first.

Every enterprise AI leader in America is now living the same story.

The board asked for an AI strategy. The strategy shipped. Pilots were funded. Vendors were selected. Slides were made. And a year later, when the CFO asks what it earned, the honest answer is: not much.

You are not imagining this pattern. MIT's NANDA initiative studied enterprise AI programs and found that 95% of them delivered no measurable P&L impact. Gartner projects that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing unclear business value and rising costs. RAND puts the failure rate of AI projects at roughly 80%. McKinsey's State of AI report keeps landing on the same gap: adoption is up, measurable earnings impact is flat.

Something is systemically broken. It is not the model. It is not the vendor. It is not the pilot team.

It is the layer they all keep skipping.

The failure everyone blames on the wrong thing

When a pilot fails, the post-mortem always lands on the same three suspects: bad data, weak change management, or the wrong use case.

Those are real problems. They are also symptoms.

I spent two years inside real enterprises, under deadline, shipping agentic systems that had to work on real cases before anyone got paid. My team and I built 32 agentic AI products across 8 industries and deployed more than 300 agents into production. That work forced us to sit at the failure point of dozens of pilots, the moment where the demo passes and the production case does not.

The pattern is not what the industry keeps repeating.

The model was almost never the problem. Frontier models are more than good enough for the tasks enterprises want to automate. The integration was slow and unglamorous, but it was solvable. Change management was uncomfortable, but people accept tools that make their day easier.

What broke the pilot, every time, was that the company could not tell us how the work actually got done.

They had documents. They had process maps. They had dashboards and BI tools and knowledge bases. They had a very organized picture of the official process, the one the SOP describes.

The real process lived in one or two people.

The senior underwriter who could tell you why one policy renewal was approved and a nearly identical one was not. The planner who knew which supplier to trust when the ERP said everything was fine. The claims adjuster who could look at a submission for ten seconds and know which field was wrong. You could watch what they did. You could not see why.

Without the why, the agent was useless on every real case.

That is what 95% of enterprise AI pilots have in common. Not a technology gap. A knowledge gap.

The missing leg

Every working enterprise agent stands on two legs.

The first is knowledge, the operating expertise of your best people, connected to your existing systems, and kept current as the business changes. The second is action, multi-agent execution that runs the work, with humans in the loop, under governance.

The entire enterprise AI market is building the action leg. Agent builders. Orchestration frameworks. Copilots. Autonomous agent platforms. All of them assume the knowledge is already there, waiting to be picked up. They give you the executor. They do not give you the brain the executor is supposed to think with.

The knowledge leg is missing.

This is the substrate problem I keep coming back to. Every AI approach reads from a substrate. A copilot reads from your open documents. Enterprise search reads from your indexed content. RAG reads from your vector store. Process mining reads from your event logs. Each of these substrates is real, and each of them is missing the same thing.

None of them contain the tacit operating expertise that makes the work function.

Research suggests that 42% of essential expertise lives only in employees' heads (Panopto). Roughly 80% of business processes are undocumented (Tallyfy). Knowledge workers spend about 19% of their time searching for information they should already have (McKinsey). Add those together and you get a picture of an enterprise where the most valuable operational knowledge is not stored anywhere your AI system can reach.

You cannot fix that with a bigger context window. You cannot fix it with a better model. You cannot fix it with cleaner data pipelines. The content is not somewhere your infrastructure can find, because it was never written down in the first place.

That is why the pilot fails. The action leg is executing on air.

What the demo hid

I have watched dozens of enterprise pilot demos in the last two years. They almost always look great.

The agent handles the FAQ. It drafts a passable email. It summarizes a policy. It answers a customer question using the documented policy manual. Everyone in the room nods. Someone says "this is going to be transformational."

Then production starts.

The pilot hits its first real case. Not the tutorial case that lives in the training slides, the actual case that walks in the door on Tuesday morning. A policy renewal with three exceptions. A claim with a missing field. A supplier issue where the ERP says everything is fine but the planner knows otherwise. The kind of case that made your top operator valuable in the first place.

The agent gives a confident, wrong answer. Or it retrieves the official policy and applies it precisely, missing every exception the experienced operator would have caught in three seconds. Or it hallucinates something plausible that a compliance officer will not accept.

The pilot team is shocked. The vendor blames the data. The vendor is not entirely wrong, but the data is not the root cause. The root cause is that the agent was never given the real process, only the documented one. And documentation does not describe the real process. It describes the idealized workflow the company put on paper.

This is why pilots pass demos and fail production. The demo is scoped to the documented 20% of cases. Production is the other 80%.

Why agentic AI is failing at exactly the same rate

Everyone in enterprise AI knows the RAG-pilot failure story. What is less discussed is that the same pattern is now repeating with agentic AI, with even higher stakes.

Gartner projects that more than 40% of agentic AI projects will be cancelled by 2027. If anything, that number is generous. The projects the industry is running today assume a level of operational context that the enterprise does not have documented.

An agent that only answers questions can fail quietly. An agent that takes actions inside SAP, or drafts a policy renewal, or files a claim, cannot. The failure gets a name, a case number, and a chain of accountability. This is why agentic pilots stall and get pulled: the leadership team is not willing to hand a critical process to a system that clearly does not know the exceptions.

They are right to be cautious. But cautiousness is not a strategy. What kills the agentic project is not the risk. It is that nobody built the layer that would have made the risk manageable.

An agent needs three things to run a real process. It needs to know what the process actually is, including the exceptions. It needs to know which systems hold which fields and whether those fields are fresh. It needs to know when to stop and ask a human. All three of those live in your experienced operators' heads.

Skip the knowledge leg and you do not just get a lower-accuracy agent. You get an agent that has no idea what it does not know. That is the shape of an agentic pilot that gets cancelled.

The substrate problem compounds with every retirement

There is a second dynamic that makes this worse over time, and it is one CAIOs rarely price into the strategy.

The people who hold the real process are leaving.

Median tenure in the U.S. private sector is 3.5 years (BLS). In manufacturing, Deloitte and the Manufacturing Institute project 2.1 million jobs will go unfilled by 2030. In insurance, the wave of senior underwriter and adjuster retirements has already started. The senior operators who quietly hold the real process together are a finite resource, and the resource is depleting.

Every departure removes a copy of operating knowledge that was never written down. The new hire learns the documented 20% and improvises the rest. Quality drifts. Exception handling degrades. And the enterprise loses, quietly, the exact substrate its future AI agents were supposed to run on.

This is why the timing question matters. In three years, the people who could have been mined for operating expertise will be gone. The enterprises that captured that expertise now will have a working knowledge leg. The ones that did not will be running agentic systems on top of a hollowed-out organization.

The failure rate of enterprise AI pilots is not a snapshot. It is a leading indicator of a bigger structural gap.

Why bigger models and bigger context windows will not save the pilot

There is a comforting story going around in the enterprise AI community: as models get better and context windows get longer, the pilot problem will solve itself.

It will not.

A million-token context window does not know your operating exceptions. It has nowhere to read them from, because they were never written. A more accurate model does not know that the ERP field for supplier reliability is stale and should be ignored. A better retrieval system does not know that the underwriter always checks a specific third-party report before approving a case above a certain threshold.

The problem is not that the model cannot reason. The problem is that the substrate it reads from does not contain the reasoning.

RAG cannot solve this because the content does not exist. Enterprise search cannot solve it because the content does not exist. Copilots cannot solve it because the content does not exist. Every layer the market is investing in is a layer that presumes the knowledge is already accessible somewhere, and it is not.

You cannot retrieve what was never written. You have to extract it from the people who hold it.

That is a different discipline. It is not a better model. It is not a better pipeline. It is a different substrate.

What actually works, build the knowledge leg first

Here is the pattern of the pilots that do not fail.

They do not start with the agent. They start with the expert.

Sit with a senior operator. Ask them not "what do you do", that produces the SOP again, but "why this case and not that one." Pull out the decision logic. Pull out the exceptions. Pull out the trusted sources. Pull out the escalation moments and the reasons behind them. Separate personal habit from operating logic. Validate it against a second expert who does the same job. Package the result in a form that a machine can read, a structured, machine-readable expert profile, not a transcript, not a wiki page.

This is what expertise mining is. It is the extraction discipline that gives an agent the substrate it needs to handle real cases, not just easy ones.

Once that captured expertise is in place, you can wire it into the systems the process actually touches, the ERP, the CRM, the policy engine, the claims platform. You know which field holds what, whether it is fresh, and what an agent is allowed to do with it. You know when the agent should draft, when it should read, and when it should stop and ask a human.

Then, and only then, you turn on the action leg. Governed agent operations, real specialists, each with a narrow job, that ask for approval before anything that matters. Human in the loop by default. A continuous learning loop so the captured expertise stays current as the business changes.

That is a pilot that survives production. Not because the model is better, but because the substrate is real.

We have watched this pattern hold across sectors. A large multi-brand insurance group we work with had an insurance renewal workflow that took six minutes per case. After mining the underwriting expertise and wiring the captured process into their systems, the same workflow runs in one minute end-to-end. The model is the same. The staff is the same. What changed is that the agent finally had access to the real process instead of the documented shell of it. That is what building the knowledge leg first looks like from the outside.

What the missing layer costs you

The CAIO who is reading this already knows the pilot fatigue number. It shows up in every board deck as "learnings we will apply to phase two." The uncomfortable truth is that phase two is going to fail the same way if you do not fix the substrate.

Every wasted pilot burns two things you cannot get back. It burns budget the board expected to see returned as measurable impact. And it burns political capital inside the company, because every failed pilot makes the next one harder to fund and harder to get people to adopt.

The real cost is not the pilot. It is the strategic delay. The enterprises that build the knowledge leg now will spend the next three years turning captured expertise into compounding operating leverage. The ones that skip it will spend the next three years re-piloting the action leg on top of the same missing substrate, and getting the same result at a bigger scale.

The pattern is not new. Every major enterprise wave, ERP, cloud, data warehousing, had a first phase where companies mistook infrastructure for outcome. AI is now in that phase. The teams that pull ahead are the teams that recognize the missing layer for what it is.

What a CAIO should do this quarter

Stop starting new agent pilots on top of missing knowledge.

The first move is a diagnostic, not a purchase. Pick one process that matters, one where an experienced operator is the reason it works, and where their departure would create real operational risk. Sit with that person. Extract the real operating logic, not the documented one. Separate what the SOP says from what actually happens.

You will discover, very quickly, that most of your critical processes have a substrate problem. The documented version is thin. The real version lives in a handful of people. That is not a failure of your documentation team. It is the natural state of every enterprise, and it is why 95% of pilots fail.

Once you have that diagnostic in hand, the sequencing changes. You build the knowledge leg first, one expert, one process at a time. You wire the captured expertise into the systems the process touches. You put a governed action layer on top, with humans in the loop. And you let the compounding start.

That is the pilot that survives. That is the AI investment that shows up in the P&L instead of the learnings slide.

The missing layer is not exotic. It is not a new model. It is not a new framework. It is the operating expertise your best people already have, captured, structured, and finally made durable so your AI can actually use it.

Mine the expertise. Then operate differently.

Share this post

Continue reading

Cover art comparing enterprise search assistants with expertise-driven agents
Manifesto

Why Glean, Notion AI, and Copilot miss the real work

Ask a CAIO what the AI stack looks like and you will hear a familiar list: Glean, Notion AI, Copilot. Ask what the business impact has been and the answer gets quieter. People search a little faster, notes come out cleaner, and none of that is the real work.

Cover art for operating knowledge and its four components
Manifesto

What is operating knowledge, and why your documents do not have it

Ask a COO for a copy of how the work is done and you will get documents. Then sit next to the person who actually runs that work and watch which field they check first, which one they ignore, and which exception they escalate. None of that is in the documents, and that gap has a name.

Cover art for expertise mining vs RAG comparison
Manifesto

Expertise mining vs RAG, different problem, different tool

Every few weeks, someone in an enterprise AI meeting asks the same question: we already have a RAG stack, why would we also need expertise mining? It is a fair question. The answer is that the two are not competitors, they are different tools for different jobs, sitting on completely different substrates.