Adoption
Your Copilot licences are unused. Here is the 30-day diagnostic.
Low seat utilisation is a workflow problem wearing a training problem's clothes. A 30-day audit that finds out which, without surveilling your team.


Unused AI licences are usually a workflow problem, not a training problem. MIT's 2025 NANDA study found 95% of enterprise GenAI pilots produced no measurable P&L return, with the cause organisational rather than technical. The fix is to audit which workflows the licence was meant to change.
- MIT NANDA, July 2025: 95% of enterprise GenAI pilots returned no measurable P&L impact.
- MIT also found only about 5% of custom enterprise AI tools ever reach production.
- More than 90% of surveyed firms had employees using personal LLMs for work anyway.
- Seat utilisation measures access. Workflow throughput measures whether the work changed.
- Cancel the seats you cannot attach to a named workflow. Reinvest the money in one workflow.
Unused AI licences are usually a workflow problem, not a training problem. MIT's 2025 NANDA study found 95% of enterprise GenAI pilots produced no measurable P&L return, with the root cause organisational rather than technical. The fix is to audit which workflows the licence was supposed to change, then instrument those workflows, not to run more training.
Key takeaways
- MIT NANDA, July 2025: 95% of enterprise GenAI pilots returned no measurable P&L impact.
- Only about 5% of custom enterprise AI tools reach production, per the same study.
- More than 90% of surveyed firms had employees using personal LLMs for work regardless.
- Utilisation counts access. Throughput counts change. They are not correlated.
What does low licence utilisation actually tell you?
It tells you the tool did not become load-bearing in anyone's day. That is all it tells you. It does not tell you whether people are capable, motivated, or trained. It tells you the work still runs the way it ran before the seats were bought.
This is the finding underneath the headline number. MIT's NANDA study (52 executive interviews, 153 survey responses and over 300 public deployments) found that 95% of enterprise generative AI pilots produced no measurable P&L return (MIT NANDA, July 2025). The cause it identifies is a learning and workflow gap: tools that cannot retain feedback or adapt to organisational context. Not model quality. Not user skill.
The 95% figure has been contested since publication, largely on the grounds that a base of 52 interviews and 153 responses is thin for a claim about the whole enterprise economy. That criticism is fair and we make it ourselves in our analysis of why 95% of GenAI pilots fail. The directional finding (that access was provisioned far ahead of workflow redesign) matches what a utilisation dashboard shows in almost every organisation that runs one.
Why more training does not fix it
Run a training push and utilisation rises. It rises for roughly three to six weeks, then decays toward its previous level. This is the most reliably reproducible result in enterprise enablement and it is also the least useful one.
The reason is structural. A workflow is a sequence of steps, approvals, formats and handoffs. Training changes how a person performs one step. It does not change the approval that comes after that step, the format the next team expects, or the system of record that will not accept a generated draft. The trained user produces their output faster and then waits exactly as long as they waited before.
MIT's own data makes the point sharply from the other direction: more than 90% of the firms it surveyed had employees using personal LLM accounts for work whether or not the enterprise seat existed (AIGL analysis of the MIT report, 2025). The capability was already in the building. What was missing was a process that could absorb it. That distinction is the whole subject of workflow absorption.
The 30-day utilisation audit, step by step
Thirty days is enough to get a defensible answer and short enough that nobody reorganises around it. Run it in four blocks.
Days 1–5: reconstruct the original claim. Find the business case that justified the seats. Written or verbal, someone said this tool would change something. Write down what, in one sentence, per department. If nobody can state it, that is your finding and you can stop.
Days 6–12: pull tenant-level telemetry, not user-level. You want seats provisioned, seats with any activity in 30 days, and seats with activity in 10 or more of the last 30 days. That third number is the only one worth acting on. Do not pull individual prompt logs. Do not report by named user. The moment this becomes a performance conversation, every subsequent input you get is theatre.
Days 13–22: shadow two workflows the licence was supposed to change. Sit with the people who run them. Time the steps. Count the handoffs. Note where the output goes and who has to accept it. Two workflows, four to six sessions each, is enough to see where the time actually sits, and it is almost never where the licence was pointed.
Days 23–30: baseline three numbers and write the recommendation. The recommendation has three buckets: workflows to instrument and pursue, workflows to leave alone, and seats to cancel at renewal. A diagnostic that recommends only expansion is not a diagnostic.
This is the same shape as our two-week AI adoption diagnostic, compressed and run internally. If you have the people to run it yourself, run it yourself.
Which three numbers to instrument first
Pick a workflow. Baseline these three before anyone touches a tool.
| Number | How to capture it | Why it survives scrutiny |
|---|---|---|
| Cycle time per instance | Timestamp on entry, timestamp on final acceptance | It is the only number the sponsor already believes |
| Rework rate | Share of instances returned or corrected downstream | Catches AI output that is fast and wrong |
| Escalation rate | Share of instances that leave the standard path | Detects work being pushed sideways rather than done |
Three numbers, one workflow, measured for two weeks before anything changes. Without the pre-baseline you have no argument in December, because you will be comparing a measured present against a remembered past, and memory is generous. METR's randomised trial is the cleanest demonstration available: experienced developers using early-2025 AI tools were measured as 19% slower on real tasks in mature repositories while estimating afterwards that AI had made them roughly 20% faster (METR, July 2025). METR has since stated that the result is historical and does not necessarily reflect current tools or workflows (METR, February 2026): the tools moved. The gap between what people believed and what was measured did not, and that gap is the reason you baseline.
What to do with the licences you should cancel
Cancel them at renewal, quietly, and say why in one line: no named workflow. Do not run a campaign about it. The point is not to punish a failed rollout; the point is to stop paying for access as a substitute for change.
Then do the arithmetic that makes the case. A mid-market centre carrying a few hundred unused seats is spending a meaningful annual sum on the option to open a chat window. That same budget, pointed at one instrumented workflow with an owner and an evaluation harness, buys a result you can put in front of a CFO. One workflow that provably improved beats four hundred seats that provably existed.
Keep the seats attached to workflows you are actually pursuing. Fund those properly. Everything else is optionality you are not exercising.
What this means for a GCC transformation owner
You are going to be asked for a number this quarter. The temptation is to report utilisation, because it exists, it is easy to pull, and it goes up when you push it. Resist it. A utilisation number that rises while cycle time is flat is a number that will be used against you in twelve months when someone asks what changed.
Report instead on one workflow: its baseline, its current state, and the date the delta will be measurable. That is a smaller claim and a defensible one. It also positions the conversation where the budget actually lives (transformation, not L&D) because you are describing a process outcome rather than a learning outcome.
Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, naming escalating costs, unclear business value and inadequate risk controls (Gartner, 25 June 2025). "Unclear business value" is what a utilisation dashboard produces when the renewal conversation arrives. The audit above is how you arrive at that conversation with something else.
Sources
- MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- AIGL, State of AI in Business 2025: figure breakdown, 2025. https://www.aigl.blog/state-of-ai-in-business-2025/
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
- Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Related reading: why your real competitor is the unused licence · the difference between AI adoption and workflow absorption · what a two-week AI adoption diagnostic produces
Read next
- Your real competitor is the unused licence you already pay forWe are not competing with another training vendor. We compete with the Coursera, Pluralsight and Copilot seats you already bought and no longer use, and a catalogue seat, however good, cannot change a workflow that exists only inside your organisation.
- Workflow absorptionWorkflow absorption measures whether AI has actually changed how work runs (steps redesigned, cycle time reduced, errors cut) as opposed to adoption, which only counts access such as seats and logins.
- Evaluation harnessAn agent evaluation harness is a repeatable test suite that scores an AI agent's outputs against fixed, versioned cases before and after every change, so teams can tell regression from variance.
More from the blog
- Adoption metrics are lying to you: the six numbers that are notLogins, completions and prompt counts measure activity, not outcome. Six metrics that hold up under scrutiny, how to baseline a workflow in a week, and how to instrument without surveilling.
- How to choose the first workflow to automate (and three you should not)The first workflow decides whether the whole programme survives. A four-axis scoring grid, the case for back-office over customer-facing, and the three types to leave alone.
- What the free AI training tier covers, and exactly where it stopsFree AI training is now genuinely good and reaches agent orchestration and MCP. What it structurally cannot do is change your workflow, on your data, in your systems, under your governance.
Frequently asked questions
There is no credible published benchmark, and chasing one is the wrong move. Utilisation counts logins, not changed work. A department where 30% of seats are used inside one redesigned workflow is further ahead than one where 90% of seats open a chat window daily and no process has changed. Measure cycle time on a named workflow instead.
Cancel every seat you cannot attach to a named workflow at renewal, and keep the ones you can. The test is not whether someone opened the tool. It is whether a specific, instrumented process runs faster or with fewer errors because that seat exists. Redirect the recovered budget into instrumenting one workflow properly.
Audit at the workflow level, not the person level. Use aggregate tenant-level telemetry, sampled voluntary work-shadowing with consent, and process metrics that already exist: ticket cycle time, rework rate, escalation counts. Never report individual prompt logs to managers. Surveillance destroys the honest self-reporting the audit depends on.
Training reliably raises utilisation for a few weeks and rarely changes throughput. MIT located the failure in organisational workflow gaps rather than skills gaps. If the process still requires the same approvals, handoffs and formats it did before, a trained user produces the same output slightly faster and the process time is unchanged.
IT owns provisioning; the business owns the workflow, and only the workflow owner can move the number. Utilisation reported by IT alone becomes a procurement metric nobody acts on. Name one accountable owner per workflow, give them the baseline and the target, and report both to the same forum.
Ready to install the workflow?
Book a free AI Reality Check and build one real thing from your own work, live.