August 24, 2026AIssential

Where the AI Failure Numbers Come From

A provenance register for the statistics everyone quotes

TL;DR — Key Takeaways
  • If you need one defensible paragraph for a board paper, it is at the top of this page, with sources.
  • Five statistics carry most of the public argument about AI failure. Three have no study behind them at all, one is a forecast that expired unscored, and one is a real document that says something much narrower than its headline.
  • The most-quoted of all — "MIT: 95% of GenAI pilots fail" — comes from a self-labelled v0.1 preliminary paper with no bibliography, built on 52 interviews and 153 leaders recruited at conferences. Its own figures contradict each other in two places, and the article that made it famous misreported the method.
  • We then measured how far it travelled: in our corpus of 81,319 AI articles the MIT figure appears in 54, and four of those question it.
  • The executive surveys everyone quotes describe the firms that answer AI surveys. Official statistics on a mandatory frame put adoption near one in five: 18% of French firms with 10+ employees (INSEE, n=11,000), 19.95% across the EU27 (Eurostat).
  • What has actually been measured is less dramatic and more useful: roughly 20–30% of enterprise AI initiatives reach production and meet expectations, outright failure runs about 20%, and a large middle sits in partial success.
  • We got one wrong ourselves and corrected it: we quoted "about 6% attribute more than 5% of EBIT to AI" for months. That figure is not in the McKinsey report.
  • This page is a living register. Every primary source is named, with its sample size and method. Corrections are welcome and will be dated.

You are about to put one of five numbers into a board paper, a steering committee deck or a business case. Each of them is repeated so widely that using one feels safe.

Three of the five have no study behind them at all. One is a forecast that expired without ever being scored. The fifth is a real document that says something considerably narrower than its headline — and the version most people quote misdescribes even its method.

Nobody loses their job over a bad statistic. But if you are asking for budget, "95% of pilots fail" arms the person who wants to say no — and it came from nowhere.

This page exists so that does not happen. For each statistic: where it actually comes from, what it actually says, and what to use in its place — with sample sizes, methods and dates, so the replacement survives the same scrutiny. It is dated, it is corrected in public, and the primary sources are named so you can check us as easily as we checked them.


If you need one defensible paragraph, use this one

Lift it as it stands. Every figure below is from a survey with a disclosed method and sample, and each source is named in full further down.

Adoption is broad among the organisations that answer AI surveys and thin across the economy as a whole. 88% of respondents report regular AI use in at least one business function, only about a third are scaling it, and of the 39% who attribute any EBIT impact to AI most place that impact below 5% (McKinsey, The State of AI in 2025, 1,993 participants across 105 nations). Set that against official statistics on a mandatory frame rather than a respondent panel: 18% of French firms with ten or more employees used at least one AI technology in 2025 (INSEE, enquête TIC, n=11,000), and 19.95% across the EU27 (Eurostat, isoc_eb_ai). Both are true, and the gap between them is a population difference, not a contradiction — the executive panels describe the firms that answer AI surveys, the statistical agencies describe every firm. Independent surveys put the outcome distribution in the same range: 28% of use cases fully succeed and 20% fail outright (Gartner, n=782, infrastructure and operations respondents), while 42% of companies now scrap most of their AI initiatives, up from 17% a year earlier (S&P Global, n=1,006). The pattern across measured studies is that roughly 20–30% of initiatives reach production and meet expectations, about 20% fail outright, and the majority sit in partial success. The binding constraint reported is not technology but attribution: fewer than one in five organisations track well-defined KPIs for their generative-AI work (McKinsey).

And if someone quotes "95% of AI pilots fail" in the meeting, the shortest accurate reply is: that figure comes from a self-labelled v0.1 preliminary paper with no bibliography, based on 52 interviews and 153 leaders recruited at conferences, and it defines success as an impact somebody remarked on. The measured surveys put outright failure nearer 20%.

You do not need the rest of this page to use those two paragraphs. The rest is the evidence that they hold.


"87% of data science projects never reach production"

What it is: a sponsored article published on VentureBeat in July 2019, reporting a remark made on stage at VentureBeat Transform. The speaker, IBM's Deborah Leff, said — and the phrasing matters — "I think CIO Dive Magazine says that only 13%…".

A verbal recollection of a trade-magazine claim, inside sponsored content. There is no study, no dataset, no method, no sample. The number has nonetheless underpinned a decade of MLOps marketing.

Use instead: Gartner, fielded Q4 2023, n=644 — 48% of AI projects reach production, with a median of eight months from prototype to production.


"85% of AI projects fail"

What it is: a Gartner forecast from 2018, which predicted that through 2022, 85% of AI projects would deliver erroneous outcomes due to bias in data, algorithms or the teams managing them.

Three things happened to it. "Erroneous outcomes" became "failure". The forecast window closed in 2022 and was never scored. And a prediction about output quality became, in retelling, a measured failure rate.

Use instead: Gartner's I&O survey, April 2026, n=782 — 28% of AI use cases fully succeed, 20% fail outright, leaving a ~52% partial-success middle. ⚠️ Infrastructure & Operations respondents only; do not generalise it to the whole enterprise.


"RAND: more than 80% of AI projects fail"

What it is: RAND report RR-A2680-1 (2024) does contain the sentence. It also contains the footnote attached to it — footnote 13 — which points to Kahn, Jeremy, "Want Your Company's A.I. Project to Succeed? …", Fortune, 26 July 2022: a trade-press interview with a vendor CEO.

RAND's own hedge, "by some estimates", is dropped by essentially everyone who quotes it. What RAND actually did was interview 65 engineers and data scientists to understand why projects fail. Its headline finding is worth more than the number people take from it: the most-cited root cause, named by 84% of interviewees, is misunderstanding or miscommunicating the problem the AI is meant to solve. RAND's own caveat: most interviewees were non-managerial, so the result may skew toward blaming leadership.

Use instead: S&P Global (451 Research), fielded October–November 2024, n=1,006 — 42% of companies now scrap most of their AI initiatives, up from 17% the year before.


"MIT: 95% of GenAI pilots fail"

This one is a real document, and it deserves the longest entry, because it is the most quoted and the most misread.

The source is MIT NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025). Its method: 52 structured interviews, plus 153 senior leaders surveyed "across four major industry conferences", plus a review of 300 public initiatives.

What the paper says about itself: it is labelled "Preliminary Findings", version 0.1. It carries no bibliography of any kind. Its disclaimer states that the views expressed do not reflect the positions of any affiliated employers — so it is not an institutional MIT finding. And Project NANDA commercialises agentic-AI infrastructure while prescribing "learning-capable systems" as the remedy.

How "success" is defined, verbatim: implementations "users or executives have remarked as causing a marked and sustained productivity and/or P&L impact." A remark, not a measurement.

Its own figures disagree with each other. The executive summary says "only 2 of 8 major sectors show meaningful structural change"; the body says, three times, "seven of nine sectors show little structural change"; the exhibit lists eight. One section heading says 50% of GenAI budgets go to sales and marketing; the body of the same section says approximately 70%. The $30–40 billion enterprise-investment figure is uncited, there being no references section.

And the version that travelled the world was wrong about the method. Fortune, 18 August 2025, described it as "150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments." The report says 52 interviews and 153 leaders surveyed. That inflated version is the one most people are repeating.

Finally, the 95% does not say what the headline says. The report's claim is that 95% of organizations are getting zero return, and the 5% applies to custom, task-specific tools — general-purpose LLM tools reached 40% in the same document.


"10-20-70: 70% of AI's value comes from people and process"

What it is: a heuristic from a 2019 TED@BCG talk by Sylvain Duranton, posted by BCG's own account in September 2019, and restated verbatim in a 2025 press release. It was never empirically derived, and it is not an output of the 1,803-executive survey it now appears alongside. The frequently seen corroboration — "MIT Sloan found 70% too" — traces to nothing we could find, and appears to be an artefact of SEO content.

The underlying claim, that the organisation matters more than the algorithm, is well supported. The ratio is not.

Use instead: Bloom, Brynjolfsson, Foster, Jarmin, Patnaik, Saporta-Eksten & Van Reenen, American Economic Review 2019, 32,000 US plants via Census MOPS data — structured management practice explains roughly 20% of the productivity spread, about twice what IT explains, and 40% of that variation occurs within the same firm. Same company, same technology, same budget, different practice, different output. That is the rigorous version of the claim 10-20-70 gestures at.


What has actually been measured

Each of these has a disclosed method and a sample you can check.

FindingSourcen
88% use AI in at least one function; only one third are scaling; 39% attribute any level of EBIT impact, and most of those put it under 5%McKinsey, The State of AI in 2025, fielded 25 Jun – 29 Jul 20251,993 participants, 105 nations
28% of use cases fully succeed; 20% fail outrightGartner I&O survey, Apr 2026 ⚠️ I&O only782
42% now scrap most AI initiatives, up from 17%; 46% of PoCs scrappedS&P Global / 451 Research, fielded Oct–Nov 20241,006
25% of AI initiatives delivered expected ROI; 16% scaledIBM IBV CEO Study 20252,000 CEOs
74% hope to grow revenue from AI; 20% already doDeloitte State of AI 20263,235
48% of AI projects reach production; 8 months prototype → productionGartner, fielded Q4 2023644

And one tier above all of those — official statistics, mandatory frames, no self-selection:

FindingSourcen
18% of French firms (10+ employees) used at least one AI technology in 2025, up 8 points on 2024; usage tripled between 2023 and 2025 in firms under 250. Adoption tracks size almost perfectly — 15% at 10–49, 31% at 50–249, 58% at 250+ — which the authors attribute to the fixed costs of adoptionINSEE, Insee Première n° 2120, enquête TIC-entreprises 202511,000 enterprises in France
19.95% across the EU27 on the same definition; France 18.16%, Germany 25.97%. By size, France sits at the euro-area average in the 50–249 band (30.8 vs 30.4)Eurostat, dataset isoc_eb_ai, indicator E_AI_TANY, 2025EU-wide official statistics

These two matter more than anything above them, and not because the numbers are lower. They are the only entries here drawn from a frame that includes the firms who would never answer a survey about AI. Every other row in this section describes people who chose to respond.

Read together — and this next sentence is our synthesis across those surveys, not a published figure: roughly 20–30% of enterprise AI initiatives reach production and meet their ROI expectations, outright failure runs about 20%, and the large remainder is partial success.

Less dramatic than 95%. Also more useful, because it points somewhere: the gap is not adoption, it is attribution.


How far the bad numbers actually travel

Everything above is about where five statistics came from. This section is about where they went, and it is the one part of this page we could measure rather than trace.

We run an ingestion pipeline over AI coverage. At the time of writing it holds 81,319 articles. We searched all of them for the figures in this register:

FigureArticles containing it
"MIT: 95% of GenAI pilots fail"54
"RAND: 80% of AI projects fail"11
"30–50% of RPA projects fail"9
"Gartner: 85% of AI projects fail"6
"87% never reach production"3
"73% of AI agent projects fail"2

Of the 54 articles carrying the MIT figure, four contain any language questioning it. The other fifty pass it along as established fact — in podcast summaries, in vendor posts, in a funding announcement, and in sponsored content.

How this was measured, and its limits. We matched patterns against each article's title, summary and generated takeaway, not its full body, so every count here is a floor, not a census. "Contains it" is not the same as "asserts it": we separated the two with a keyword check for sceptical framing and then read a sample by hand to confirm the split. And this is one corpus of AI-focused English coverage, not the internet. Treat the table as an order of magnitude, and the 54-versus-4 ratio as the finding.

That ratio is the reason this page exists. A statistic with no study behind it does not get weaker as it spreads. It gets more credible, because the fiftieth person to repeat it is repeating something they have now seen forty-nine times.


One of them was ours

For months we quoted "about 6% of companies attribute more than 5% of EBIT to AI", crediting McKinsey's State of AI. It appeared in our site copy, in our own posts, and in comments where we were correcting other people's numbers.

On 8 August 2026 we opened the PDF instead of the search results. The figure is not in the report. The >5% threshold appears only inside the compound definition of "AI high performers" — respondents reporting more than 5% of EBIT and "significant value" attributable to AI — and the report never states what share of respondents that is. What it does say is that 39% attribute any level of EBIT impact, and that most of those put it below 5%.

Our number was plausible, probably close, and unsourced. That last word is the one that matters, because it is the same failure we are documenting above. Site copy was corrected the same day, in thirteen places, English and French.

And a second one, found while writing this page. Our own strategy document carried "this is why 40% of AI agent projects get abandoned (Gartner)". Gartner's actual sentence is a forecast: "over 40% of agentic AI projects will be canceled by the end of 2027" (press release, 25 June 2025), and the only data underneath it is a January 2025 poll of 3,412 webinar attendees. We had put a prediction in the present tense — rule three on the checklist below, which we wrote — and worse, we had made it carry a cause Gartner never states: its stated reasons are escalating costs, unclear business value and inadequate risk controls, not the operating model we were arguing for. The sentence is gone. We now cite the measured Gartner survey instead (n=782): "ROI from AI is not driven by the sophistication of the model, but by how well the technology is integrated, governed, and aligned with real operational needs."

Two in one page, both ours, both found by opening the primary. That is the honest rate at which this happens to people who are trying.


How to check a number yourself, in four minutes

  1. Find the sentence in the primary document, not in an article about it. If the primary cannot be found, that is your answer.
  2. Read the footnote. The RAND number survives everything except its own footnote.
  3. Check whether it is a forecast. Forecasts are written in the future tense and rarely scored afterwards. Two of the five above are forecasts wearing the clothes of measurements.
  4. Check what "success" or "failure" was defined as. "Remarked as causing impact" and "measured against a pre-period baseline" are not the same claim, and the difference is usually the whole story.

What we could not verify

Kept here deliberately, because a register that only lists its wins is a marketing page.

  • "By 2028, one in four enterprise software purchases will be made by AI agents with no human in the loop." Could not be traced to any primary Gartner release.
  • "Gartner: 41%/42% of prototypes reach production." Appears only in a client-gated document.
  • "42% show zero ROI; the profitable 58% set KPIs first." A vendor blog, no survey, no method, no sample. It is also the claim we would most like to be true, which is exactly why we are not using it.
  • "82% of banks don't measure ROI on any technology investment." Attributed to a 2025 industry survey, reachable only via a secondary blog.
  • "30–50% of RPA projects fail." EY Financial Services Ireland, 2016, verbatim in context: "we are often called upon to help companies when their first attempt failed… we have seen as many as 30 to 50% of initial RPA projects fail." No sample, no denominator — and the population described is EY's own remediation caseload.


Corrections

This page is a living register. If you can show that something here is wrong — including a primary source we have misread — send it and we will change it, dated, in public. That is the only way a page like this is worth anything.

Last updated: 24 August 2026.

Make the AI decision you can defend.

Try AIssential for free →
Last updated: August 24, 2026