How do we measure the return on an AI workflow — and what baseline is honest?

Eighty-seven percent of leaders credit AI output entirely to humans, yet organizations see no measurable return on GenAI investments. Without redesigning workflows and measuring output, leaders risk failing to extract value from their AI initiatives.

· Counsel verdict · AIssential

The question

We need to report the return on an AI-enabled workflow to our board. What should we measure, what baseline is defensible, how do we avoid crediting AI for gains it did not cause, and how long should the measurement run before we make a scale-or-stop decision?

Counsel's position

Implement a phased, output-driven measurement framework tracking leading, operational, and financial metrics against a robust counterfactual baseline.

Verdict

The verdict: Implement a phased, output-driven measurement framework tracking leading, operational, and financial metrics against a robust counterfactual baseline.

How the criteria decide

3 of 4 criteria resolved on cited evidence. 1 had none either way.

CriterionFavoursEvidence
metrics to measureBoth equally

Top performers are more likely to redesign and measure workflows

We suggest tracking leading indicators weekly, operational metrics monthly, and financial outcomes quarterly.

AI to ROI - By Ray Rike and Peter Buchanan

Organizations see no measurable return on GenAI investments

Approve one more AI pilot that cannot remember the work, handle exceptions, or answer to a single business owner, and the cost leaves the software line and shows up as backlog, rework, and renewal theater.

Artificial Intelligence on Medium

defensible baselineNo evidence either way
attribution of gainsBoth equally

Eighty-seven percent of leaders credit AI output entirely to humans

Eighty-seven percent of leaders admitted AI output is sometimes or always credited entirely to the human employee.

CIO

measurement durationBoth equally

AI pilots fail when measured by fixed durations instead of outputs

Our theory on why so many pilots fail is that companies tend to pick an AI tool and a pilot duration and qualitatively check in with users at the end of that time.

Artificial intelligence - Crunchbase News

Eighty-seven percent of leaders credit AI output entirely to humans

Categorizing AI as software rather than labor hides the cost of human supervision and inflates perceived ROI.

Top performers are more likely to redesign and measure workflows

Establishing baselines before deployment and tracking leading indicators weekly prevents waiting for quarterly financial results to identify failing projects.

AI pilots fail when measured by fixed durations instead of outputs

Reversing the standard pilot process—choosing the output first and varying the AI tools until the dial moves—provides a concrete metric for board reporting.

Agent unit costs multiply rather than fall as deployments scale

Measuring the "cost per successful outcome"—including failed attempts and human escalations—reveals whether an accurate agent is actually losing money.

Organizations see no measurable return on GenAI investments

Pilots that lack memory, exception handling, and a single business owner fail to produce P&L impact.

Read another verdict

Get Counsel for your own decisions →