Start with a small test, or take on the whole process at once?

95% of enterprise generative AI pilots produce no measurable return, stalling at the production threshold due to unmanaged debt and organizational readiness gaps. This leaves leaders risking significant financial benefits and critical workflow failures.

· Counsel verdict · AIssential

The question

We have identified one process where AI might help. We can run a small contained test, take on the whole process including its exceptions and handoffs, or do nothing until we see it working somewhere else first. The evidence on pilots that never reach production is not encouraging. What separates a test that teaches us something from one that quietly dies, and when is a narrow pilot the wrong shape for the problem in the first place?

Counsel's position

Integrate AI into the whole process, focusing on critical reasoning tasks with human oversight, to ensure production readiness and measurable impact.

Verdict

The verdict: Integrate AI into the whole process, focusing on critical reasoning tasks with human oversight, to ensure production readiness and measurable impact.

How the criteria decide

2 of 3 criteria resolved on cited evidence. 1 had none either way.

CriterionFavoursEvidence
Chance of reaching everyday useTake on the whole process

Enterprise AI pilots stall due to sequencing problems, not technology failures

Most of the time that I see is that the pilot was built next to reality, not inside reality. As soon as production hits, everything that wasn't in the pilot that you deliberately put out suddenly becomes the work.

The AI in Business Podcast

Enterprise GenAI pilots deliver no measurable profit-and-loss impact

As a standalone portal with its own URL and its own login, usage dropped to zero within a month. Resurfaced inside Teams, where the questions were already being asked, it stuck.

Data Engineering on Medium

95% of enterprise generative AI pilots produce no measurable return

It found that 95% of enterprise generative AI pilots produced no measurable profit-and-loss return, and that only 5% of custom enterprise AI tools reached production.

The AI Journal

What we learn either wayTake on the whole process

Enterprise AI pilots stall due to sequencing problems, not technology failures

Most of the time that I see is that the pilot was built next to reality, not inside reality. As soon as production hits, everything that wasn't in the pilot that you deliberately put out suddenly becomes the work.

The AI in Business Podcast

Generative AI pilots fail due to unmanaged production debt

Vibes-based assessment is the silent killer of AI projects. Without objective, quantifiable metrics, you cannot safely iterate on your system.

Towards Data Science

Disruption to the people doing the workNo evidence either way

Enterprise AI pilots stall due to sequencing problems, not technology failures

Successful deployments start with a narrow production slice embedded in a real workflow, rather than a standalone pilot built next to reality.

Generative AI pilots fail due to unmanaged production debt

Moving from a successful demo to a reliable system requires systematically paying down technical, operational, and evaluation debts rather than just optimizing the model.

Enterprise GenAI pilots deliver no measurable profit-and-loss impact

AI initiatives succeed when they are treated as business workflow redesigns championed by business owners, rather than standalone IT deployments.

95% of enterprise generative AI pilots produce no measurable return

Pilots stall at the production threshold due to organizational readiness gaps—specifically unclear ownership, missing monitoring budgets, and unresolved compliance questions—rather than technology failures.

Human decision gates reduce critical AI workflow failures from

Structuring AI workflows with explicit human oversight at critical junctures and restricting models to reasoning tasks prevents unreliable outputs from advancing.

Read another verdict

Get Counsel for your own decisions →