Back to all articles

Strategy 7 min read

Most companies cannot show AI ROI because they never measured a baseline

CEO surveys, an MIT report and a Danish labour-market study point the same way: gains are easier to report than to prove. Here is how to set a baseline before you build.

Two colleagues planning at a whiteboard covered in sticky notes
Photo Walls.io / Pexels

Key takeaways

  • In IBM’s 2025 survey of 2,000 CEOs, 25% said their AI initiatives had delivered the expected return over the last few years, and 16% had scaled them enterprise-wide.
  • In a Danish study of about 25,000 workers, adopters reported time savings of 2.8% of work hours, yet earnings and recorded hours showed no effect larger than 1%.
  • The much-quoted MIT figure that 95% of organizations see no return rests on a small sample and a narrow definition of value, so it shows how hard measurement is without settling whether AI pays.
  • A baseline of time per task, cost per unit, error rate and cycle time, taken before the build, is the only way to separate a real gain from a reported one.

IBM surveyed 2,000 chief executives in 33 countries between February and April 2025. Only 25% said their AI initiatives had delivered the expected return on investment over the last few years, and 16% said they had scaled an initiative enterprise-wide.1 In the same release, 68% of the CEOs said their organization has clear metrics to measure innovation ROI effectively.1 Read together, the two numbers describe a gap: most executives believe they can measure returns, and most say the returns have not arrived.

The 68% is a self-assessment about innovation in general, not a check of what the metrics measure, so it should not be read as proof that baselines exist. In practice many AI projects start with a demo, a licence and a hopeful estimate, and a finance team that is asked a year later whether it paid off has nothing to compare against. The cost of that omission is concrete: a project that cannot show a return is hard to renew, hard to expand and easy to cut.

What the headline numbers can and cannot tell you

The most quoted figure in this debate comes from the MIT NANDA report “The GenAI Divide: State of AI in Business 2025”. As reported by Campus Technology in August 2025, it says that 95% of organizations are seeing no business return and that just 5% of integrated AI pilots are extracting millions in value.2 The research ran from January to June 2025 and combined a systematic review of more than 300 publicly disclosed AI initiatives, 52 structured interviews and 153 survey responses from senior leaders at four major conferences.2 The full report was not openly available for this article, so every statement here about its content is attributed to that secondary account.

The report has drawn criticism on method. Thomas Wieberneit, writing in CustomerThink in November 2025, argues that the sample is small, at about 150 survey responses and 52 structured interviews, and that the report looks solely at custom, task-specific agentic AI implementations rather than tools such as ChatGPT or Copilot.3 He also points out that it counts tools that mainly raise individual productivity as not delivering value, because they do not show up in profit and loss.3 These are one analyst’s views, but they identify the real issue. Whether a company reports 5% or 25% success depends on how value is defined, and a definition chosen after the fact will always favour whoever chose it.

So the headline numbers agree on one point only: many organizations cannot show a return. They do not agree on how large the problem is, and none of them says which projects failed because the technology was weak and which failed because nobody wrote down what success looked like. A baseline is what lets you tell the two apart.

A hand holding a pen over printed financial charts
Headline surveys report what leaders believe; they do not measure your process. Photo Kindel Media / Pexels

Reported productivity and measured outcomes are different things

The strongest evidence on the difference between the two comes from Denmark. Anders Humlum and Emilie Vestergaard linked two survey rounds, from late 2023 and 2024, covering 11 exposed occupations, about 25,000 workers and 7,000 workplaces, to administrative records of earnings and hours.4 Workers who used chatbots reported average time savings of 2.8% of work hours, and 6% to 7% on the days they used them.4 The records show something else: the confidence intervals rule out effects larger than 1% on either earnings or recorded hours.4 The NBER version of the paper, revised in March 2026, states its bound as no effect larger than 2% two years after the launch of ChatGPT.5

The authors also note that randomized trials in the same occupations have documented gains often exceeding 15%.4 The gap between 15% in trials, 2.8% in self-reports and a figure indistinguishable from zero in payroll records has several possible causes. Trials measure a defined task under supervision, while real jobs include the tasks that did not get faster. Saved minutes may go into other work that nobody records. And 43% of the workers in the study were explicitly encouraged by their employer to use chatbots, while 21% were allowed to, so adoption is uneven.4 The study does not settle which explanation dominates, and it covers wages and hours in a national labour market, not the cost per unit inside one department.

For an executive the lesson is practical. A survey asking staff whether the tool saves time measures perception. A baseline measures the task. If the two disagree, the baseline is the one a finance team can audit.

How to set a baseline before you build

A baseline is a measurement of the current process, taken before any change, using the same method you will use afterwards. It needs a unit of work that a person can define in a sentence, such as an invoice processed, a ticket resolved or a contract reviewed. It also needs enough volume to be stable: a week of data is often distorted by one unusual customer, and a quarter is better.

Figure 1

A baseline in four measurements (a method, not sourced data)

  1. 01

    Time per task

    Sample real cases and time the whole task from start to finish, including waiting and rework, not only the time a person spends typing.

  2. 02

    Cost per unit

    Divide the full cost of the process, labour plus software plus overhead, by the units completed in the same period. Write the definition down.

  3. 03

    Error rate

    Have a second person re-check a sample of finished cases and record the share that are wrong or reworked.

  4. 04

    Cycle time

    Record the elapsed time from request to completion, since a task can get faster while the queue around it gets longer.

The four measures protect each other. Time per task alone rewards a tool that is fast and wrong. Error rate alone rewards a tool that is accurate and slow. Cost per unit is the number finance cares about, and it is the one most likely to be undefined, because the labour, software and overhead belong to different budgets.

A coiled yellow measuring tape on a blue background
Taken together, the four measures stop a tool from improving one number at the expense of another. Photo Ann H / Pexels

Counterfactuals and a worked example

A before-and-after comparison is not enough, because things other than the tool change between the two periods: volumes, staffing, seasonality. A counterfactual answers the question of what would have happened without the tool. The simplest form is a control group: run the new workflow for some teams or some case types, keep the old one for comparable others, and compare the same four measures over the same weeks. When that is not possible, a stable pre-period of several months gives a trend to compare against, though it is weaker.

Risks and criticism of measuring first

Baselines have costs of their own. Measuring takes staff time before any benefit appears, and it can delay a project that would have worked. Workers who know they are being timed change their behaviour. Metrics also narrow attention: a team judged on handling time may stop doing the unmeasured work that kept customers satisfied. A baseline of the wrong thing is worse than none, because it produces confident numbers about the wrong question.

There is also a case against waiting. Some of the value of a tool shows up as quality, as capacity for work that was previously dropped, or as staff retention, and these resist a unit-cost measure. The Danish study measures wages and hours, not quality of output, so a null result there does not mean no benefit anywhere. The honest position is that a baseline reduces self-deception without removing judgment, and that a pilot should state in advance which results would make the team stop.

A laptop showing data analytics charts in a bright room
Some value only shows after launch, so the baseline should be quick rather than exhaustive. Photo Lukas Blazek / Pexels

What to do next

  • Pick one workflow and define its unit of work in a sentence that two people would apply the same way.
  • Measure time per task, cost per unit, error rate and cycle time for at least one full quarter, and write down the definition of each cost line.
  • Set a stop rule before the pilot: the result, on those four measures, that would end the project.
  • Arrange a comparison group or a stable pre-period so that you can estimate what would have happened without the tool.
  • Ask users what they think the tool saves, then compare that answer with the measured figure and keep both.
  • If you want a first estimate of where AI could cut cost in your workflows, the free pre-audit is a short questionnaire, and the AI readiness benchmark shows where organizations stand on readiness.

Newmind Partners

Find out where AI would pay in your workflows

Newmind Partners designs and builds AI workflows that cut operating cost. Start with the free pre-audit for a first estimate, or run a Feasibility audit for a scored report on one workflow.

Free pre-audit

Sources

  1. IBM Newsroom, “IBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles” (6 May 2025), survey of 2,000 CEOs in 33 countries and 24 industries, February to April 2025: 25% of AI initiatives delivered expected ROI; 16% scaled enterprise-wide; 68% say their organization has clear metrics to measure innovation ROI effectively. newsroom.ibm.com
  2. David Ramel, Campus Technology, “MIT Report: Most Organizations See No Business Return on Gen AI Investments” (26 August 2025), reporting the MIT NANDA study “The GenAI Divide: State of AI in Business 2025”: 95% see no business return; research January to June 2025, over 300 initiatives reviewed, 52 interviews, 153 survey responses. campustechnology.com
  3. Thomas Wieberneit, CustomerThink, “The Great GenAI Divide: Debunking the Myth of 95% Failure” (6 November 2025): critique of the MIT NANDA report’s sample size and definition of value. Author’s views. customerthink.com
  4. Anders Humlum and Emilie Vestergaard, “Large Language Models, Small Labor Market Effects”, Becker Friedman Institute working paper 2025-56 (9 May 2025): two survey rounds, 11 occupations, about 25,000 workers and 7,000 workplaces in Denmark; self-reported time savings 2.8% of work hours; no effect larger than 1% on earnings or recorded hours; randomized trials often exceeding 15%. bfi.uchicago.edu
  5. Anders Humlum and Emilie Vestergaard, “Still Waters, Rapid Currents”, NBER Working Paper 33777 (issued May 2025, revised March 2026): null effects on earnings and recorded hours, ruling out effects larger than 2% two years after the launch of ChatGPT. nber.org
Let’s talk

See where AI would cut your costs.

Take the free pre-audit: a short questionnaire that returns a first estimate of where AI could cut cost in your workflows. Prefer a scored report on one workflow? Start a Feasibility audit.

Free pre-audit