How do we measure the return on an AI workflow — and what baseline is honest?
Eighty-seven percent of leaders credit AI output entirely to humans, yet organizations see no measurable return on GenAI investments. Without redesigning workflows and measuring output, leaders risk failing to extract value from their AI initiatives.
The question
We need to report the return on an AI-enabled workflow to our board. What should we measure, what baseline is defensible, how do we avoid crediting AI for gains it did not cause, and how long should the measurement run before we make a scale-or-stop decision?
Counsel's position
Implement a phased, output-driven measurement framework tracking leading, operational, and financial metrics against a robust counterfactual baseline.
Verdict
The verdict: Implement a phased, output-driven measurement framework tracking leading, operational, and financial metrics against a robust counterfactual baseline.
How the criteria decide
3 of 4 criteria resolved on cited evidence. 1 had none either way.
| Criterion | Favours | Evidence |
|---|---|---|
| metrics to measure | Both equally | Top performers are more likely to redesign and measure workflows We suggest tracking leading indicators weekly, operational metrics monthly, and financial outcomes quarterly. AI to ROI - By Ray Rike and Peter Buchanan Organizations see no measurable return on GenAI investments Approve one more AI pilot that cannot remember the work, handle exceptions, or answer to a single business owner, and the cost leaves the software line and shows up as backlog, rework, and renewal theater. |
| defensible baseline | No evidence either way | |
| attribution of gains | Both equally | Eighty-seven percent of leaders credit AI output entirely to humans Eighty-seven percent of leaders admitted AI output is sometimes or always credited entirely to the human employee. |
| measurement duration | Both equally | AI pilots fail when measured by fixed durations instead of outputs Our theory on why so many pilots fail is that companies tend to pick an AI tool and a pilot duration and qualitatively check in with users at the end of that time. |
Eighty-seven percent of leaders credit AI output entirely to humans
Categorizing AI as software rather than labor hides the cost of human supervision and inflates perceived ROI.
Top performers are more likely to redesign and measure workflows
Establishing baselines before deployment and tracking leading indicators weekly prevents waiting for quarterly financial results to identify failing projects.
AI pilots fail when measured by fixed durations instead of outputs
Reversing the standard pilot process—choosing the output first and varying the AI tools until the dial moves—provides a concrete metric for board reporting.
Agent unit costs multiply rather than fall as deployments scale
Measuring the "cost per successful outcome"—including failed attempts and human escalations—reveals whether an accurate agent is actually losing money.
Organizations see no measurable return on GenAI investments
Pilots that lack memory, exception handling, and a single business owner fail to produce P&L impact.
Read another verdict
- Start with a small test, or take on the whole process at once?
- Our competitors advertise AI and we don't — match them, or hold the line?
- Our people already put client files into ChatGPT — ban it, frame it, or supply a tool?
- Our most experienced person retires in two years — how do we keep what they know?
- We can't hire the experienced people we need — automate, train up, or outsource?
- Slow our EU AI Act prep now the deadline's moved to 2027?
- Use AI to flatten middle management this year?
- Let an AI agent act on its own — or keep a human in the loop?