Key takeaways
- In a study of 5,179 support agents at one software company, an AI assistant raised issues resolved per hour by 14% on average and by 34% for novice and low-skilled agents.
- The same study found minimal impact on experienced, highly skilled agents and cautions that quality may fall for the most skilled.
- Klarna reported in February 2024 that its assistant did the work of 700 full-time agents; in May 2025 its chief executive said investing in the quality of human support is the way forward.
- In Gartner’s October 2025 survey of 321 service leaders, 20% had cut agent staffing because of AI, so the evidence for wholesale replacement is thin.
In February 2024 Klarna announced that its AI assistant had handled 2.3 million customer service conversations in its first month, about two-thirds of the company’s customer service chats, and was doing work equivalent to 700 full-time agents.1 The company also reported that resolution time had fallen from 11 minutes to under 2 minutes, that repeat inquiries were down 25%, and that the assistant would improve profit by an estimated $40 million in 2024.1 All of these are Klarna’s own figures, taken from a press release covering the assistant’s first month.
In May 2025 the story had moved on. According to Entrepreneur, reporting on a Bloomberg interview with chief executive Sebastian Siemiatkowski, Klarna was recruiting remote customer service staff again. Hiring had been paused about a year earlier, and headcount had fallen 22% to 3,500, mostly through attrition.2 The article describes the chatbots as cheaper but producing lower-quality service.
Really, investing in the quality of human support is the way of the future for us
The two announcements bracket a decision that many finance and operations leaders now face: whether AI should sit beside the support agent or take the agent’s place. The strongest independent evidence concerns the first option. For the second, what exists is mostly company-reported or drawn from surveys of executives, and it points toward caution.
Two designs with different economics
An assistant sits inside the agent’s screen. It reads the conversation as it happens and suggests replies or points to relevant material, and the agent decides what to send. A replacement system talks to the customer directly and hands the case to a person only when it fails or when the customer asks. The first changes how fast and how well a person works. The second changes how many people are needed.
The distinction matters for the arithmetic. An assistant raises the output of people you already employ, so the gain appears as spare capacity, and it becomes a saving only if you reduce hiring, hours or headcount. A replacement system takes a share of contacts out of the human queue entirely, so the saving per contact is larger, and so is the exposure when the system gets a case wrong. A weak suggestion gets edited by an agent. A weak answer from an autonomous bot goes straight to the customer.

What the measured evidence shows
The most cited measurement of the assistant model comes from Erik Brynjolfsson, Danielle Li and Lindsey Raymond. They studied the staggered introduction of a generative AI conversational assistant, using data on 5,179 customer support agents at one Fortune 500 enterprise software company that serves small businesses in the US.3 The work is chat-based technical support, and most of the agents (89%) are outside the US, mainly in the Philippines.3 Because the tool reached agents at different times, the authors could compare agents who had it with agents who did not.
Productivity, measured as issues resolved per hour, rose 14% on average.3 The gain was uneven: 34% for novice and low-skilled workers and minimal for experienced and highly skilled workers.3 Agents with two months of tenure who used the tool performed as well as untreated agents with more than six months of tenure.3 The paper also reports a rise in customer sentiment of 0.18 points, about half a standard deviation, and lower attrition among newer workers: about 10 percentage points, a 40% decrease from a 25% baseline in that group.3
Figure 1
Change in issues resolved per hour with an AI assistant
The same paper carries a warning that is easy to skip. It cautions that AI assistance may lower the quality of conversations handled by the most skilled agents.3 An assistant that lifts the average while nudging down the best people is still a good trade in many operations, but it is a trade, and it should be measured as one.
Read the scope before quoting the headline. This is one firm, one product and one chat-based support operation, with an average chat of about 40 minutes and a baseline of 2.6 resolutions per hour.3 A different channel, such as short voice calls, or a different company may produce a different number. The study measures assistance. It does not test whether the agents could be removed.
What happened when a company went further
Klarna’s 2024 figures describe volume and speed over the assistant’s first month, as reported by the company itself.1 The 2025 reporting describes a company that had paused hiring, let headcount fall, and then decided that human support quality needed investment.2 One company’s reversal does not measure how many agents AI can replace. Firms with a high share of simple, repetitive requests may be able to automate more than Klarna chose to. What the episode shows is that a headline about contacts handled and a headline about service quality can diverge.
Surveys point the same way, though they record leaders’ reports and expectations rather than measured outcomes. In June 2025 Gartner predicted that by 2027, 50% of organizations that expected to significantly reduce their customer service workforce will abandon those plans.4 That is a forecast. The press release also cites a March 2025 poll of 163 service leaders, in which 95% plan to keep human agents.4
In a later Gartner survey of 321 customer service leaders in October 2025, 20% reported having reduced agent staffing because of AI, and 55% reported stable staffing with higher volumes.5 Both are self-reported. They suggest that, so far, AI in service has more often absorbed growth than shrunk teams.

What the capacity is worth
Finance teams will want a cost per contact. This article does not use one, because none of the sources opened for it gives a figure whose scope fits a general claim. For published cost-per-contact ranges, see Newmind’s customer service cost benchmark. What can be done with the sourced numbers is a capacity calculation, with the inputs labelled.
Two further points. Novices gain most, so the benefit concentrates in the teams with the highest turnover, and the study found lower attrition among newer workers, from a 25% baseline.3 Each avoided departure saves recruiting and training time, which a capacity calculation does not capture. And an assistant that brings a two-month agent to the level of an agent with more than six months shortens the ramp-up period for every new hire.3
Risks and limits
- Single-site evidence. The headline study covers one company and one channel. Treat its percentages as a hypothesis to test in your own operation.
- Experienced agents. The measured gain for highly skilled agents was minimal, and quality may fall for them. A rollout that treats all agents alike can waste licences or erode your best people.
- Self-reported results. Klarna’s 2024 figures and Gartner’s surveys describe what companies say. Neither is an audited measurement of cost or satisfaction.
- Reversal risk. Cutting headcount ahead of evidence is expensive to undo: Klarna’s own account in 2025 involved paused hiring, falling headcount and renewed recruitment.
There is also a labour effect. An assistant that lets two-month agents match six-month agents compresses the value of experience, which has consequences for pay, retention and the career ladder. The study reports lower attrition among newer workers, but one firm over a short window is not a settled answer on wages or morale.

What to do next
- Measure a baseline first: issues resolved per hour, repeat contacts and customer sentiment, split by agent tenure.
- Roll an assistant out in stages so some teams act as a comparison group, as the study’s design allowed.
- Report results by tenure band, expecting the largest gains among new agents and checking quality among the most experienced.
- Treat any fully automated flow as a separate pilot, with an easy route to a person and tracking of escalations and repeat contacts.
- Decide in writing how freed capacity will be used, whether through slower hiring, attrition or absorbing growth, before claiming a saving.
Newmind’s free pre-audit is a short questionnaire that returns a first estimate of where AI could cut cost in your workflows. Whatever the pilot shows, report it the way the research above does: one firm, one channel, one period, with the limits stated.
Newmind Partners
Find out where AI would pay in your workflows
Newmind Partners designs and builds AI workflows that cut operating cost. Start with the free pre-audit for a first estimate, or run a Feasibility audit for a scored report on one workflow.
Sources
- Klarna, “Klarna AI assistant handles two-thirds of customer service chats in its first month” (27 February 2024), press release. klarna.com
- Entrepreneur, “Klarna CEO reverses course by hiring more humans, not AI” (9 May 2025), reporting a Bloomberg interview of 8 May 2025. entrepreneur.com
- Brynjolfsson, Li & Raymond, “Generative AI at Work”, NBER Working Paper 31161 (April 2023, revised November 2023); data on 5,179 customer support agents at one company. nber.org
- Gartner, press release of 10 June 2025 on customer service workforce plans; includes a March 2025 poll of 163 service leaders and a forecast for 2027. gartner.com
- Gartner, press release of 2 December 2025 on AI-driven headcount reduction; survey of 321 customer service leaders, October 2025. gartner.com




