Agentic engineering
The pilot-to-production checklist for enterprise AI
Most AI pilots never reach production because they never had an owner, a baseline, or an evaluation harness. Eleven checks, and the exit criteria for killing a pilot cleanly.


AI pilots stall at production because the pilot never had an owner, a baseline, or an evaluation harness. MIT's 2025 study found only about 5% of custom enterprise AI tools reach production. A pilot is production-ready when its outputs are scored against fixed cases and a named person owns the metric.
- MIT NANDA, 2025: only about 5% of custom enterprise AI tools reach production.
- Gartner expects over 40% of agentic AI projects to be cancelled by end-2027.
- The three named cancellation causes are escalating costs, unclear value, and inadequate risk controls.
- Production-readiness is a checklist, not a demo: owner, baseline, harness, failure log, rollback.
- A pilot needs exit criteria written before it starts, so killing it is a decision, not a defeat.
AI pilots stall at production because the pilot never had an owner, a baseline, or an evaluation harness. MIT's 2025 study found only about 5% of custom enterprise AI tools reach production. A pilot is production-ready when its outputs are scored against fixed cases, its failure modes are documented, and a named person owns the metric.
Key takeaways
- MIT NANDA, 2025: only about 5% of custom enterprise AI tools reach production.
- Gartner expects over 40% of agentic AI projects to be cancelled by end-2027.
- Production-readiness is a checklist (owner, baseline, harness, failure log, rollback) not a demo.
- Exit criteria written before the pilot starts are what let you kill it cleanly.
Why do AI pilots stall before production?
Because the demo answered the wrong question. A pilot demo shows that the system can produce a good output on a chosen example. Production asks something else entirely: can the organisation depend on this output, at volume, on the days the input is ugly, without a human checking every one? Almost every stalled pilot stalled because nobody built the evidence needed to answer that second question.
The scale of the gap is the headline of MIT's NANDA study: only about 5% of custom enterprise AI tools reach production (MIT NANDA, July 2025). The study locates the cause in an organisational learning gap rather than in model capability: the tools worked in isolation and the organisation could not absorb them. We take that finding apart in why 95% of GenAI pilots fail.
Gartner frames the same wall from the forward-looking side: it expects over 40% of agentic AI projects to be cancelled by the end of 2027, naming escalating costs, unclear business value and inadequate risk controls (Gartner, 25 June 2025). Every item on the checklist below exists to defuse one of those three causes before it detonates at the production gate.
The eleven-item readiness checklist
A pilot is production-ready when every one of these is true. Not most. Every one.
- A named owner who runs the workflow. Not IT, not a central AI team: the person accountable for the process metric.
- A pre-baseline. Cycle time, rework rate and escalation rate for the workflow, measured before the pilot touched it.
- A versioned golden-case set. Real inputs with known-good outputs, drawn from actual work, under version control.
- An evaluation harness that scores every run against those cases and reports task success, tool-call correctness, cost per run and failure mode.
- A documented failure-mode inventory. The ways it goes wrong, ranked by cost, each with a detection method.
- A human-in-the-loop design placing a checkpoint wherever an error is both expensive and hard to reverse: see human in the loop.
- A rollback plan that returns the workflow to its manual state the same day, tested at least once.
- A cost ceiling per successful task, monitored, with an alert before it is breached.
- A data and access boundary documented: what the system may read, what it may not, and who approved it.
- An audit trail capturing inputs, outputs and the human decisions around them, retained per policy.
- Written exit criteria, in both directions: the thresholds that justify go-live and the thresholds that stop the project.
Items 3, 4 and 5 are the ones pilots most often skip, because they are the least fun and the most load-bearing.
What an evaluation harness must cover
A harness that reports one number (accuracy) is not a harness. It is a thermometer. A production harness scores four axes, because those are the four ways an agent disappoints you in production:
| Axis | Question it answers | Why a benchmark cannot answer it |
|---|---|---|
| Task success | Did the run achieve the workflow's goal? | Model benchmarks test generic tasks, not yours |
| Tool-call correctness | Did it call the right tool with the right arguments? | Agents fail on integration, not on language |
| Cost per successful task | What does one good outcome cost? | Cost per token hides the retries behind a success |
| Failure mode | When it failed, how? | You cannot install a control for a failure you have not named |
The harness runs before and after every change: model version, prompt, tool, data. Without it you cannot tell a regression from ordinary variance, and that distinction is the entire difference between a system you can improve and one you can only hope about. The mechanics are in how to evaluate an AI agent in production.
Who owns the workflow after go-live?
The same person who owned the pilot. Ownership does not transfer to IT at go-live; if it does, the workflow slowly rots because nobody with authority over the process is watching the metric.
This is why the owner has to be the workflow owner from day one. Only they can accept the residual risk of depending on an AI output. Only they can change the steps around it: the approval that used to sit downstream, the format the next team expects. A central team can build the thing. It cannot own the consequence of running it, and production is the business of owning consequences.
The governance items procurement will ask for
Before go-live, procurement and risk will ask a predictable set of questions, and the pilot that anticipated them ships months earlier than the one that starts assembling answers at the gate: the data boundary and who approved it, the audit-trail retention, the human-in-the-loop checkpoints, the rollback plan, and the cost controls. These are checklist items 7 through 10, which is not a coincidence. They are on the checklist because they are the questions that stop pilots at the last step. A standing AI governance framework turns each of these from a scramble into a lookup.
When to kill a pilot instead
Kill it when it crosses the stop-side exit criteria you wrote at the start: task success below the floor after fair effort, rework rate that erases the time saved, or a cost per successful task the workflow cannot bear. Kill it also when the honest answer to "who depends on this in production" is nobody.
Killing a pilot on pre-agreed criteria is not a failure of the programme. It is the programme working. A practice that cannot kill a pilot will run the pilot-forever model instead, and pilot-forever is how unclear business value (Gartner's second named cause) becomes a permanent line item.
What this means for a GCC transformation owner
The production gate is where your credibility is priced. Walk up to it with a demo and you will be sent back to gather the evidence you should have gathered in week one. Walk up with the eleven items complete and the conversation is short.
So run the pilot as if the checklist were the deliverable, because it is. The working system is necessary and not sufficient; the owner, the baseline, the harness and the exit criteria are what convert a working system into one the organisation is allowed to depend on. That conversion is the whole job, and it is what a workflow sprint is built to leave behind.
Sources
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- AIGL, State of AI in Business 2025: figure breakdown, 2025. https://www.aigl.blog/state-of-ai-in-business-2025/
- Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Related reading: why 95% of GenAI pilots fail · what an evaluation harness is · scope a workflow sprint
Read next
- Why 95% of GenAI pilots fail, and what the surviving 5% did differentlyThe 95% figure is contested and imperfect, and it still describes your pilot. The failure is not model quality. It is that nobody instrumented the workflow the tool was supposed to change, so no result could ever have been measured.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
- AI governance frameworkAn AI governance framework is the documented set of policies, roles, controls and records that determine who may deploy an AI system, on what data, with what testing, and who is accountable when it fails.
More from the blog
- How to evaluate an AI agent in productionModel benchmarks do not transfer to your workflow. Score every run against versioned cases on four axes, build a golden set from real tickets, and catch regressions between model versions.
- How to choose the first workflow to automate (and three you should not)The first workflow decides whether the whole programme survives. A four-axis scoring grid, the case for back-office over customer-facing, and the three types to leave alone.
- Writing an AI business case that survives CFO reviewAI business cases fail because they cannot be falsified. Name one workflow, state a measured pre-baseline, and write the kill criteria in advance. A one-page template.
Frequently asked questions
A proof of concept tests whether a thing is technically possible; a pilot tests whether it is operationally worth doing on a real workflow with real users. A POC can succeed and still tell you nothing about production. Confusing the two is a common way to declare victory early: the demo works, so the project is assumed done.
Long enough to accumulate a statistically meaningful number of real instances on the target workflow, which for a high-frequency back-office process is usually four to eight weeks. Time-box it and write the exit criteria at the start. An open-ended pilot with no end date is the pilot-forever model, and it is a way of avoiding the production decision indefinitely.
MIT's NANDA study found only about 5% of custom enterprise AI tools reach production. The gap is rarely model quality. It is the absence of an owner, a pre-baseline, and an evaluation harness: the operational scaffolding that lets an organisation trust an output enough to depend on it. Those are checklist items, not research problems.
The person who owns the workflow the pilot changes, not IT and not a central AI team. Only the workflow owner can accept the risk of depending on an AI output, redesign the surrounding steps, and be accountable for the metric. A pilot owned by a team that does not run the workflow has no path to production because nobody can say yes to going live.
Written, in advance, in two directions: the measured thresholds that justify going to production, and the thresholds below which you stop. For example, a target task-success rate and a maximum acceptable rework rate. Exit criteria turn killing a pilot into a pre-agreed decision rather than an admission of failure, which is the only way pilots get killed on time.
Ready to install the workflow?
Book a free AI Reality Check and build one real thing from your own work, live.