Key takeaways
- In S&P Global Market Intelligence’s 2025 survey, the share of companies abandoning most of their AI initiatives before production rose from 17% to 42% in a year.
- Gartner’s 30% and 40% figures are forecasts about abandonment, not measured outcomes, and they name cost, unclear value and weak risk controls as the causes.
- RAND’s interviews with 65 practitioners trace most failures to a misunderstood problem, missing data and weak infrastructure, not to model quality.
- A pilot that ships has a named owner, a measured baseline, one process, a costed integration and a stated risk limit before the first prompt is written.
In S&P Global Market Intelligence’s 2025 survey of 1,006 mid-level and senior IT and line-of-business professionals in North America and Europe, the share of companies abandoning the majority of their AI initiatives before they reach production rose from 17% to 42% year over year. Respondents also reported that, on average, 46% of projects are scrapped between proof of concept and broad adoption.1
That is the number to put next to any pilot budget. A pilot is a bet with a fixed cost and an uncertain payoff, and a cancelled pilot leaves behind the invoice, the lost months and a sceptical finance team. The question for a COO or CFO is not whether AI works in a demo. It is why so many demos never become a process that someone runs every day, and what you can check before spending the money.

What the forecasts and surveys say
Gartner said in July 2024 that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, inadequate risk controls, escalating costs or unclear business value. In the same release it put the cost of the GenAI deployment approaches it examined at $5 million to $20 million.2
Unfortunately, there is no one size fits all with GenAI, and costs aren’t as predictable as other technologies.
In June 2025 Gartner made a similar call for agents: over 40% of agentic AI projects will be cancelled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. It also described “agent washing”, the relabelling of chatbots, robotic process automation and AI assistants as agentic, and estimated that only about 130 of the thousands of agentic AI vendors are real. In a January 2025 poll of 3,412 webinar attendees, 19% said their organisation had made significant investments in agentic AI and 42% had made conservative ones.3
Figure 1
Abandonment figures, by source and what each one measures
These sources measure different things. S&P reports what surveyed companies say they did. Gartner publishes predictions about project cancellation. RAND, in a 2024 report, notes that by some estimates more than 80% of AI projects fail, twice the rate of IT projects without AI. That number is a cited estimate, and RAND did not measure it.4 Treat the range as a warning about direction, not a precise failure rate.
Why pilots stall
RAND interviewed 65 data scientists and engineers, each with at least five years of experience building AI and machine-learning models, and named five leading root causes: stakeholders misunderstand or miscommunicate the problem, so models are optimised for the wrong metrics or do not fit the business workflow; the organisation lacks the data to train an effective model; the team chases the latest technology instead of a real user problem; the infrastructure to manage data and deploy models is inadequate; and the technology is pointed at problems too hard for AI.4
Read next to Gartner’s causes (cost, value, risk controls, data quality), the pattern is consistent. Our reading, which is judgement and not a finding from those studies, is that it shows up in a pilot as five checkable gaps:
- No owner. The pilot sits with an innovation team, while the process belongs to an operations leader who never agreed to change how the work is done.
- No baseline. Nobody measured handling time, error rate or cost per case before the pilot, so no result can be called a saving.
- The wrong process. The pilot targets a visible task instead of one with volume, rules and a measurable cost, which is RAND’s point about problems that do not fit the workflow.4
- Integration costed late. The model works, and then connecting it to the ERP, the document store and the approval chain turns out to cost more than the first estimate.
- Risk left unstated. Nobody has written down which errors are tolerable, who reviews them and what happens when the system is wrong, which is the Gartner “inadequate risk controls” cause in practice.2

What the pilots that ship do differently
RAND’s recommendations describe the other side. Technical staff need to understand the project purpose and domain context. Leaders should choose enduring problems and be ready to commit a product team to a specific problem for at least a year. Successful projects focus on the problem, not the technology, and invest up front in infrastructure. Leaders also need technical experts to judge whether AI can solve the problem at all.4

S&P’s survey adds a cost of getting this wrong. Companies with higher project failure rates were more prone to encounter resistance from customers and employees and had a higher level of concern about reputational damage.1 The method below is ours, built from those findings and from the gaps above.
Figure 2
A pre-pilot checklist (our method)
- 01
Name the owner
One operations leader who owns the process, the budget line and the decision to switch the old way off.
- 02
Measure the baseline
Volume, handling time, error rate and cost per case, taken over a normal period before any tool is chosen.
- 03
Pick one process
High volume, written rules, a measurable cost and an end-to-end path from input to booked result.
- 04
Cost the integration
Price the connections to existing systems and the data clean-up before the model, and include them in payback.
- 05
Set the risk limit
Define which errors are tolerable, who reviews them and the point at which the pilot stops.
What the evidence does not settle
There is also a cost to the opposite policy. If every pilot must have a signed baseline and a costed integration, some useful experiments will not start. The sensible limit is a small, time-boxed test of a cheap idea, with the full checklist applied before anything is wired into a live process.
What to do next
- List the AI pilots currently running and, for each, write the name of the operations owner and the measured baseline. Any pilot missing either is a candidate to pause.
- Rank candidate processes by volume and by cost per case, and pick the one where the work is rule-based and the result can be booked end to end.
- Get a written estimate of integration and data preparation before approving the build, and compare payback with and without it.
- Agree the tolerable error types and the review step with risk and compliance before the first live run.
- Set a stop date and a stop rule. If the baseline has not moved by then, close the pilot and record why.
If you want a first estimate of where AI could cut cost in your own workflows, the free pre-audit is a short questionnaire, and our AI readiness benchmark covers the organisational side of the same question.
Newmind Partners
Find out where AI would pay in your workflows
Newmind Partners designs and builds AI workflows that cut operating cost. Start with the free pre-audit for a first estimate, or run a Feasibility audit for a scored report on one workflow.
Sources
- S&P Global Market Intelligence, “AI experiences rapid adoption, but with mixed outcomes” (Voice of the Enterprise: AI & Machine Learning, Use Cases 2025), 1,006 respondents in North America and Europe. spglobal.com
- Gartner, “Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025” (29 July 2024). gartner.com
- Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027” (25 June 2025), poll of 3,412 webinar attendees in January 2025. gartner.com
- RAND Corporation, “The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed” (Ryseff, De Bruhl, Newberry, 13 August 2024), 65 interviews. rand.org




