Point of view
Why 95% of GenAI pilots fail, and what the surviving 5% did differently
The 95% figure is contested and imperfect, and it still describes your pilot. The failure is not model quality. It is that nobody instrumented the workflow the tool was supposed to change, so no result could ever have been measured.

MIT's NANDA study found 95% of enterprise generative AI pilots produced no measurable P&L return, against US$30–40 billion invested. The cause it identifies is a learning and workflow gap, not model quality: tools that never entered the workflow they were bought to change.
- MIT NANDA, July 2025: 95% of enterprise GenAI pilots returned no measurable P&L impact.
- The base is 52 executive interviews, 153 survey responses and 300+ public deployments: small, and the report says so.
- Roughly 5% of custom enterprise AI tools reach production; over half of budget went to sales and marketing.
- The 95% figure is contested on its narrow success definition; the criticism attacks precision, not direction.
- The surviving 5% bought externally, chose back-office friction, and measured workflow change instead of licence adoption.
The evidence
95% of enterprise generative AI pilots produced no measurable P&L return, against an estimated US$30–40 billion invested.
The finding rests on 52 structured executive interviews, 153 survey responses and 300+ public deployments.
Roughly 5% of custom enterprise AI tools reach production, and more than half of GenAI budget went to sales and marketing despite better back-office returns.
The report's own base is described as directionally accurate from individual interviews rather than official company reporting: a limitation critics have pressed hard.
Over 90% of surveyed organisations reported employees using personal LLM tools for work outside any sanctioned deployment.
What did the MIT study actually measure?
MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, analysed more than 300 public AI deployments, ran 52 structured executive interviews and surveyed 153 leaders. It found that 95% of enterprise generative AI pilots produced no measurable profit-and-loss return, against an estimated US$30–40 billion of enterprise spend.
Read that sentence precisely, because most coverage does not. The study did not find that the models failed. It did not find that 95% of pilots produced nothing anyone valued. It found that 95% produced nothing that appeared in the P&L within roughly six months of deployment. That is a narrower claim than the headline, and a more uncomfortable one. A pilot that made people feel faster while moving no financial number is exactly the pilot that gets renewed in year two and quietly cancelled in year three.
We lead with this number in every first meeting. Not because it is a scare statistic (it is not, and we will take it apart below), but because it is the only honest starting position. If you have run a pilot that went nowhere, you are not an outlier. You are the base rate.
95% of enterprise generative AI pilots delivered no measurable P&L return, against an estimated US$30–40 billion invested. >MIT NANDA, The GenAI Divide: State of AI in Business 2025 (2025)
The study defines success narrowly: deployment past the pilot stage with measurable KPIs, and ROI observed about six months after. That definition is the reason the number is high, and it is also the reason the number is useful to a transformation owner. It measures the thing your CFO measures.
Why is it a learning gap and not a model gap?
MIT's own diagnosis is not about model capability. It is that the deployed systems could not retain feedback, could not adapt to the context of the work, and could not survive contact with a process that already existed. The tool sat beside the workflow instead of inside it.
This matches what a licence-utilisation audit looks like from the inside. People try the tool. They get an answer that is 80% right. Correcting the last 20% costs more than doing the task the old way, because the correction is not captured anywhere and the same 20% is wrong again tomorrow. Nothing in the deployment closes that loop. So usage decays, and the dashboard records adoption while the work is unchanged. That distinction (access versus change) is what we call workflow absorption, and it is the metric almost nobody instruments.
The model was never the constraint. In 2026 the frontier models are more than good enough for most back-office work. The constraint is that a workflow is a social object: it has owners, exceptions, handoffs, unwritten rules and a person who knows why step four exists. None of that is in the licence.
Where did the money go, and where were the returns?
More than half of enterprise GenAI budget went to sales and marketing functions, while the better returns MIT observed sat in back-office automation (MIT NANDA figure breakdown, AIGL, 2025). Roughly 5% of custom enterprise AI tools reached production at all.
The reason is not stupidity. It is visibility. Sales and marketing pilots are demonstrable. You can show a board a generated campaign. Back-office pilots are boring, and boring is where the friction is. Invoice exceptions, vendor reconciliation, RFQ drafting, EDI failures: high-frequency, rule-dense, already partly instrumented, and nobody is watching. That is precisely why they are measurable, and measurability is the whole game.
This is why our first paid engagement is a diagnostic and not a build. A two-week Adoption Diagnostic exists to find the boring workflow with a countable baseline, not the impressive one. An illustrative example of that shape is our engagement scenario for vendor invoice reconciliation: a method page, not a client result.
What did the surviving 5% do differently?
Three traits, in the order they matter.
They bought rather than built. MIT distinguishes organisations that purchased tools and partnered externally from those that built internally, and the purchase-and-partner group outperformed. Secondary coverage has attached a specific multiplier to that gap; we could not confirm the multiplier in the primary deck, so we will not quote one. The direction is what the study supports, and the direction is enough.
They picked high-friction back-office work. Not the demo-friendly workflow. The one with a queue.
They measured change, not access. Seats provisioned, logins, prompt counts and completion certificates all rose in the failing 95% too. What separated the 5% was that somebody could state what the workflow cost before and what it cost after. If you cannot produce that pair of numbers, you do not have a measurement problem. You have no measurement.
None of the three is a technology decision. All three are decided before a line of code is written, which is the actual finding hiding inside the headline.
Has the 95% figure been disputed?
Yes. We are going to state the criticism at full strength, because we quote this number constantly and you should know what it is worth.
The base is small. Fifty-two interviews is a thin foundation for a headline about the enterprise economy, and the report itself characterises its interview data as directionally accurate from individual interviews rather than official company reporting (Marketing AI Institute, 26 August 2025). The success definition is narrow (deployment past pilot with measurable KPIs and ROI observed at roughly six months), which excludes efficiency gains, deflected hiring, error reduction and anything with a longer payback. Critics have also pointed out that larger surveys asking differently shaped questions report far more positive outcomes. The report was not peer reviewed, and for a period was gated behind a form.
All of that is fair. Here is why we still open with it.
The criticism attacks the precision of the number, not its direction. If the true rate of pilots with no measurable P&L impact is 60% rather than 95%, nothing about the method changes: you still cannot tell which side of the line you are on without a baseline, and you still have no baseline. Every critique above is, in fact, an argument that measurement in this field is poor, which is our argument.
So treat 95% as a prior, not a verdict. A vendor who quotes it without the caveat is doing the same selective-evidence thing that produced the failing pilots.
Over 90% of surveyed organisations had employees using personal LLM tools for work, outside any sanctioned deployment. >MIT NANDA figure breakdown, AIGL (2025)
This is the number that should worry a GCC most. The sanctioned pilot returned nothing measurable while an unsanctioned, unlogged, ungoverned shadow deployment absorbed real work. Your data went somewhere. It is not in the dashboard.
What does this mean for a GCC transformation owner?
You are probably being asked to show returns on a spend that was approved on a slide. The 95% figure is not an argument for stopping. It is an argument for changing the order of operations.
Name one workflow. Measure it for a week before anything is built: cycle time per instance, rework rate, escalation rate. Decide in advance what result would cause you to kill the project. Then build the smallest thing that touches that workflow end to end, with an evaluation harness attached from day one so you can tell a regression from noise. If the number does not move, you have learned something cheap and reportable, which is a materially better outcome than a renewed licence and a shrug.
We will also tell you which workflows to leave alone. That is not modesty; it is the first deliverable. A vendor who cannot name what should not be automated is selling you the thing that fails at the base rate.
One disclosure, since this page is arguing about honesty: Chokmah is a new practice and has no completed client engagements to point at. Everything above is method and published evidence. Hold us to the evidence.
Find out which side of the 95% your pilot is on
Two weeks, 8–12 interviews, two shadowed workflows and a licence-utilisation audit. The output names three workflows to automate and the ones to leave alone.
Book an Adoption Diagnostic · See the free AI Reality Check
Frequently asked questions
What is the MIT NANDA report?
The GenAI Divide: State of AI in Business 2025 is a report from MIT's Project NANDA, published July 2025. It analysed 300+ public AI deployments, 52 structured executive interviews and 153 survey responses, and found that 95% of enterprise generative AI pilots produced no measurable P&L return.
Is the 95% figure about all AI or just generative AI?
Generative AI only. The study covers enterprise GenAI pilots: chat assistants, copilots and custom LLM tools deployed between roughly 2023 and 2025. It does not measure classical machine learning, forecasting, computer vision or automation that predates the current generation of models.
Has the 95% figure been disputed?
Yes. Critics argue the success definition is narrow (measurable KPIs and ROI at about six months), the interview base of 52 is small, and the report was not peer reviewed. The criticism targets the precision of the number rather than its direction, and the underlying point (that almost nobody baselined the workflow) is untouched by it.
What did the successful 5% do differently?
They bought and partnered externally rather than building alone, they targeted high-friction back-office workflows instead of demo-friendly customer-facing ones, and they measured workflow change rather than licence adoption. All three are decisions made before any code is written.
Does this mean we should not invest in AI?
No. It means the sequence matters more than the spend. Name one workflow, measure it before you build, define kill criteria in advance, then build the smallest end-to-end thing with an evaluation harness attached. The failure rate is a measurement failure at least as much as an execution one.
Key terms
- Workflow absorptionWorkflow absorption measures whether AI has actually changed how work runs (steps redesigned, cycle time reduced, errors cut) as opposed to adoption, which only counts access such as seats and logins.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
- Agentic AIAgentic AI is software that uses a language model to plan and carry out multi-step tasks by calling tools, observing the results, and choosing its next action in a loop.
More points of view
- Measure the workflow, not the loginsLicence logins, completions and prompt counts measure that a tool was opened, not that the work changed. The 95% failure rate is a measurement failure as much as an execution one: organisations counted adoption and never instrumented absorption.
- The workflows you should not automateNaming what to leave alone is the first deliverable, not a disclaimer. A vendor who cannot tell you which workflows to keep human is selling you the thing that fails 95% of the time, and the refusal list is the most useful page we can hand a transformation owner.
- AI made senior developers 19% slower, and they believed it sped them upThe one rigorous randomised trial of AI coding assistants found experienced developers got slower, not faster, while feeling faster. You cannot manage a productivity claim you have not measured, because the people doing the work are not reliable witnesses to their own speed.
Frequently asked questions
The GenAI Divide: State of AI in Business 2025 is a report from MIT's Project NANDA, published July 2025. It analysed 300+ public AI deployments, 52 structured executive interviews and 153 survey responses, and found that 95% of enterprise generative AI pilots produced no measurable P&L return.
Generative AI only. The study covers enterprise GenAI pilots: chat assistants, copilots and custom LLM tools deployed between roughly 2023 and 2025. It does not measure classical machine learning, forecasting, computer vision or automation that predates the current generation of models.
Yes. Critics argue the success definition is narrow (measurable KPIs and ROI at about six months), the interview base of 52 is small, and the report was not peer reviewed. The criticism targets the precision of the number rather than its direction, and the underlying point (that almost nobody baselined the workflow) is untouched by it.
They bought and partnered externally rather than building alone, they targeted high-friction back-office workflows instead of demo-friendly customer-facing ones, and they measured workflow change rather than licence adoption. All three are decisions made before any code is written.
No. It means the sequence matters more than the spend. Name one workflow, measure it before you build, define kill criteria in advance, then build the smallest end-to-end thing with an evaluation harness attached. The failure rate is a measurement failure at least as much as an execution one.
Bring the evidence to your team
We walk in with the failure rates, then the method. Book a free AI Reality Check.