Back to all articles

People 7 min read

What controlled studies say about AI and productivity: it is uneven

Four studies, four different answers: faster writing, slower coding, worse judgement outside AI’s reach, and no visible effect on Danish pay. What that means for your rollout.

A software developer working at a dual-monitor setup
Photo Zayed Hossain / Pexels

Key takeaways

  • In a controlled experiment with 758 BCG consultants, AI users did better on tasks inside AI’s capability and 19 percentage points worse on one task outside it.
  • Experienced open-source developers took 19% longer with AI tools in a randomised trial while believing AI had made them 20% faster.
  • Danish administrative data rule out earnings or recorded-hours effects larger than 2% two years after ChatGPT’s launch.
  • The studies point to task fit and measurement as the deciding variables, which makes a tool rollout without task selection a weak bet.

In a 2025 randomised trial, 16 experienced developers worked through 246 real issues in their own open-source repositories, with AI tools allowed on some issues and banned on others. When AI was allowed, they took 19% longer. Before the study they had expected AI to make them 24% faster, and afterwards they still believed it had made them 20% faster.1

That gap between measured and felt productivity matters to anyone signing off a licence budget. Executives hear enthusiastic reports from their teams and a stream of vendor claims, and the controlled evidence is less tidy than either. This article walks through four of the better-known studies, says what each one can and cannot support, and ends with a way to decide where AI deserves money in your own workflows.

Inside and outside the frontier: the BCG experiment

The most cited field experiment involved 758 strategy consultants at Boston Consulting Group, about 7% of the firm’s individual-contributor consultants. They were randomly assigned to no AI, GPT-4, or GPT-4 with a prompt-engineering overview.2

On 18 realistic consulting tasks chosen to sit inside what the model could do, consultants with AI completed 12.2% more tasks on average, finished 25.1% faster, and produced results rated more than 40% higher in quality than the control group. Those below the average performance threshold improved by 43% against their own baseline scores; those above it improved by 17%.2

The authors also designed one task outside the frontier, where the model’s plausible-looking answer was wrong. Consultants without AI got it right 84.5% of the time. The two AI groups scored 60% and 70%, an average drop of 19 percentage points.2 The authors call the boundary between these two regimes a “jagged technological frontier”: tasks that look similar in difficulty can fall on either side of it, and nothing on the screen tells the user which side they are on.2

Limits matter here. The tasks were built by the researchers, the outside-the-frontier test was a single task, and the participants were consultants working on isolated problems. It is evidence about a mechanism, not a forecast for your finance department.

For a buyer, the outside-the-frontier result is the more useful half of the paper. A workflow that produces confident, fluent output will be trusted in proportion to how fluent it is, and the experiment shows what happens when that trust is misplaced. Gains of 25.1% in speed on suitable tasks and a loss of 19 percentage points in accuracy on an unsuitable one can coexist inside the same team, on the same day, with the same tool.2

Three professionals working together with laptops and documents
In the consulting experiment, results depended on whether the task sat inside what the model could do. Photo Kindel Media / Pexels

Writing tasks: faster, better, and more even

Noy and Zhang gave 444 college-educated professionals occupation-specific writing tasks, such as press releases and short reports, in a preregistered online experiment, and randomly exposed half of them to ChatGPT. In the working-paper version, time taken fell by 0.8 standard deviations and output quality rose by 0.4 standard deviations.3

The same paper reports that the tool narrowed the gap between workers, because lower-ability participants benefited more, and that it mostly substituted for effort rather than adding skill. Workers shifted their time towards idea generation and editing and away from rough drafting.3

The caveats are the task and the setting. These were short, self-contained writing assignments done online by paid participants, not months of drafting inside a regulated firm with approval chains. Standard-deviation effects also say little about money until someone attaches a cost to an hour and a defect.

There is also a question of who captures the benefit. If a tool lifts weaker writers towards the level of stronger ones, a manager may see more uniform output rather than a lower headcount need. Whether that is worth paying for depends on what a poor first draft currently costs you in review time, and neither paper answers that for your organisation.

Experienced developers: slower, and sure they were faster

The METR trial is the awkward one. Its developers were experienced contributors to large open-source projects, averaging more than 22,000 stars and over 1 million lines of code, which they had worked on for years. Issues averaged about two hours. Developers could choose their tools, mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet.1

The result, again, was 19% longer with AI allowed.1 The sample is small at 16 people, the tools are of early 2025, and METR’s own page now states that the results are out of date and points to newer data published in February 2026 on late-2025 tools.1 The durable finding is the perception gap. People who sat through the slowdown still reported a speed-up, so self-reported gains from a team survey are a weak basis for a business case.

Why would skilled people be slower with a capable tool? The study does not prove a mechanism, and we will not pretend it does. A plausible reading is that on code a person already knows well, the time spent writing a prompt, reading the output and correcting it can exceed the time spent typing the fix. If so, the same tool might help a newcomer to the codebase and hinder a veteran. The study was not designed to test that.

Hands typing on a laptop keyboard
Experienced developers took longer with AI allowed, and believed they had been faster. Photo ClickerHappy / Pexels

The labour-market view: small effects on pay and hours

Humlum and Vestergaard linked large-scale surveys on chatbot adoption in Denmark to administrative labour-market records. Using difference-in-differences estimates at worker and workplace level, they find precise null effects on earnings and on recorded hours, and they rule out effects larger than 2% two years after ChatGPT’s launch.4

Two readings are open. Time saved may not show up in pay or hours because it is absorbed into other work or never captured by the employer. Or adoption has been shallow in practice. The data cannot separate the two, and the study does not measure output per hour inside a company.

For an executive this is a useful brake on the most optimistic forecasts, and an equally useful brake on the most pessimistic ones. A null in national records two years in does not show that individual tasks are not getting faster. It shows that, in this country and period, the saving has not yet moved wages or hours in a way the data can detect.

Figure 1

Four studies, four different results

Setting and scopeMeasured result
BCG consultants, tasks inside AI’s capability (758 people) Randomised field experiment 12.2% more tasks, 25.1% faster, more than 40% higher quality
BCG consultants, one task outside it Randomised field experiment 19 percentage points less likely to be correct
Professional writing tasks (444 people) Preregistered online experiment 0.8 SD less time, 0.4 SD higher quality
Experienced open-source developers (16 people, 246 issues) Randomised trial, early-2025 tools 19% longer with AI allowed
Danish workers, administrative records Difference-in-differences No effect on earnings or recorded hours larger than 2%
Source: Dell’Acqua et al.2; Noy & Zhang3; METR1; Humlum & Vestergaard4. Different tasks, populations and metrics: the rows are not comparable as a ranking.

What the pattern suggests for a rollout

No single study settles the question, and these four do not measure the same thing. Read together, though, they point the same way. Where a task sat inside the model’s competence and was well specified, as in the BCG tasks and the writing assignments, results were positive. Where the work was deep, context-heavy and already done by experts, as in the METR trial, they were negative. Where the model’s answer looked right and was not, people trusted it too readily.

That reading is ours, not a conclusion any of these papers states. It does suggest that the unit of analysis is the task and the workflow around it, with a human checking step placed where errors are likely. The BCG authors observed that some consultants split work between themselves and the AI while others worked with it continuously, and that these were distinct patterns of successful use.2 How the work is organised affects the outcome, which a licence count cannot capture.

A second implication concerns how you judge success. Self-reported time saved is the cheapest metric to collect and, on the METR evidence, among the least reliable.1 Counting licences activated is cheaper still and says nothing about outcomes. Metrics that survive scrutiny are slower to gather: elapsed time per finished item, defects found downstream, and the share of AI output that is rewritten before use.

Colleagues brainstorming with colourful sticky notes on a wall
Treat the first deployment as an experiment with a hypothesis and a point at which to stop. Photo Vitaly Gariev / Pexels

What to do next

Treat the first deployment as an experiment with a hypothesis and a stop rule. A practical sequence follows. Our benchmark on knowledge work offers a way to compare your own task mix against reference data.

Figure 2

A method for deciding where AI earns its place

  1. 01

    List tasks, not roles

    Pick five to ten recurring tasks with a clear output and a clear definition of correct.

  2. 02

    Sort by checkability

    Prefer tasks where a wrong answer is cheap to spot, and put a human check on those where it is not.

  3. 03

    Measure a baseline

    Record time per item, error rate and rework before anyone touches a tool.

  4. 04

    Run a controlled trial

    Give half the team the tool, or alternate weeks, and compare the same metrics rather than opinions.

  5. 05

    Redesign the workflow

    Change who drafts, who reviews and where the check sits, and re-measure before extending the licences.

If you want a first estimate of which of your workflows are candidates, the free pre-audit is a short questionnaire that returns one. Whatever you use, insist on a number measured on your own work before you scale a rollout, because the published evidence says gains are real in some places, absent in others, and hard to feel from the inside.

Newmind Partners

Find out where AI would pay in your workflows

Newmind Partners designs and builds AI workflows that cut operating cost. Start with the free pre-audit for a first estimate, or run a Feasibility audit for a scored report on one workflow.

Free pre-audit

Sources

  1. METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (10 July 2025): 16 developers, 246 issues. metr.org
  2. Dell’Acqua et al., “Navigating the Jagged Technological Frontier”, Harvard Business School Working Paper 24-013 (2023): 758 consultants at Boston Consulting Group. mitsloan.mit.edu
  3. Noy & Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence” (working paper, 2 March 2023): 444 college-educated professionals. economics.mit.edu
  4. Humlum & Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI”, NBER Working Paper 33777 (issued May 2025, revised March 2026): Denmark, surveys linked to administrative records. nber.org
Let’s talk

See where AI would cut your costs.

Take the free pre-audit: a short questionnaire that returns a first estimate of where AI could cut cost in your workflows. Prefer a scored report on one workflow? Start a Feasibility audit.

Free pre-audit