Start with a small test, or take on the whole process at once?
95% of enterprise generative AI pilots produce no measurable return, stalling at the production threshold due to unmanaged debt and organizational readiness gaps. This leaves leaders risking significant financial benefits and critical workflow failures.
The question
We have identified one process where AI might help. We can run a small contained test, take on the whole process including its exceptions and handoffs, or do nothing until we see it working somewhere else first. The evidence on pilots that never reach production is not encouraging. What separates a test that teaches us something from one that quietly dies, and when is a narrow pilot the wrong shape for the problem in the first place?
Counsel's position
Integrate AI into the whole process, focusing on critical reasoning tasks with human oversight, to ensure production readiness and measurable impact.
Verdict
The verdict: Integrate AI into the whole process, focusing on critical reasoning tasks with human oversight, to ensure production readiness and measurable impact.
How the criteria decide
2 of 3 criteria resolved on cited evidence. 1 had none either way.
| Criterion | Favours | Evidence |
|---|---|---|
| Chance of reaching everyday use | Take on the whole process | Enterprise AI pilots stall due to sequencing problems, not technology failures Most of the time that I see is that the pilot was built next to reality, not inside reality. As soon as production hits, everything that wasn't in the pilot that you deliberately put out suddenly becomes the work. Enterprise GenAI pilots deliver no measurable profit-and-loss impact As a standalone portal with its own URL and its own login, usage dropped to zero within a month. Resurfaced inside Teams, where the questions were already being asked, it stuck. 95% of enterprise generative AI pilots produce no measurable return It found that 95% of enterprise generative AI pilots produced no measurable profit-and-loss return, and that only 5% of custom enterprise AI tools reached production. |
| What we learn either way | Take on the whole process | Enterprise AI pilots stall due to sequencing problems, not technology failures Most of the time that I see is that the pilot was built next to reality, not inside reality. As soon as production hits, everything that wasn't in the pilot that you deliberately put out suddenly becomes the work. Generative AI pilots fail due to unmanaged production debt Vibes-based assessment is the silent killer of AI projects. Without objective, quantifiable metrics, you cannot safely iterate on your system. |
| Disruption to the people doing the work | No evidence either way |
Enterprise AI pilots stall due to sequencing problems, not technology failures
Successful deployments start with a narrow production slice embedded in a real workflow, rather than a standalone pilot built next to reality.
Generative AI pilots fail due to unmanaged production debt
Moving from a successful demo to a reliable system requires systematically paying down technical, operational, and evaluation debts rather than just optimizing the model.
Enterprise GenAI pilots deliver no measurable profit-and-loss impact
AI initiatives succeed when they are treated as business workflow redesigns championed by business owners, rather than standalone IT deployments.
95% of enterprise generative AI pilots produce no measurable return
Pilots stall at the production threshold due to organizational readiness gaps—specifically unclear ownership, missing monitoring budgets, and unresolved compliance questions—rather than technology failures.
Human decision gates reduce critical AI workflow failures from
Structuring AI workflows with explicit human oversight at critical junctures and restricting models to reasoning tasks prevents unreliable outputs from advancing.
- (Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable
Read another verdict
- Our competitors advertise AI and we don't — match them, or hold the line?
- Our people already put client files into ChatGPT — ban it, frame it, or supply a tool?
- Our most experienced person retires in two years — how do we keep what they know?
- We can't hire the experienced people we need — automate, train up, or outsource?
- Slow our EU AI Act prep now the deadline's moved to 2027?
- Use AI to flatten middle management this year?
- Let an AI agent act on its own — or keep a human in the loop?
- Invest in pre-build costing or post-deployment ROI tracking?