The lesson from Klarna isn’t that AI in customer service doesn’t work. In 2024 Klarna said its AI assistant performed work equivalent to roughly 700 full-time roles and cut average resolution time from eleven minutes to under two. That is not the same as firing 700 people. In 2025 the CEO also acknowledged that cost had weighed too heavily in how support was organised, and the company reinvested in human support for complex cases. The precise lesson is hybrid service, not a complete reversal.
That’s not a defeat, it’s a correction. And it’s exactly the kind of correction you can learn from without putting 700 jobs on the line yourself. Here’s Renforza’s rundown of what happened, according to whom, and what it means for your first agent.
What did Klarna actually do?
Klarna, the Swedish payments company, put an AI assistant to work in customer service and was strikingly open about it. According to Klarna’s own announcement, the assistant handled 2.3 million conversations in its first month, two thirds of all chats. The company described that capacity as equivalent to the work of roughly 700 full-time employees; average resolution time fell from eleven minutes to under two, according to the same source.
On paper that’s a dream case. Faster, cheaper, scalable. It’s also exactly the case that ends up in sales decks, usually without the sequel. And the sequel is the interesting part.
Why did Klarna bring human support back in?
Because speed is not sufficient for every conversation. CEO Sebastian Siemiatkowski said in 2025 that an excessive focus on cost had resulted in lower quality. Klarna therefore began testing a smaller model for in-house, flexible human support so customers could still reach a person. This was neither a return to the old staffing level nor an abandonment of AI.
The model became hybrid: the AI assistant handles high volume while humans remain available for the moments that matter. Klarna’s 2025 annual report filed with the SEC says the assistant handled 80 percent of customer-service chats that year, with no decline in overall consumer satisfaction according to the company. That qualifies the popular story of a total failure.
Why that model stuck is easy to explain. An agent is excellent at patterns and useless at exceptions it doesn’t recognize as exceptions. An angry customer reads to a language model as text, not as a relationship at risk. Where that goes structurally wrong is covered in more depth in our piece on what an AI agent can’t do.
Why isn’t this an argument against AI?
Because the mistake wasn’t in the technology, it was in the ambition. “Replace the department” is a different assignment than “catch the routine work.” The first forces the agent into terrain where it’s structurally weak: escalations, emotions, exceptions. The second lets it do what it’s good at and keeps humans free for the rest.
That aligns with a broader signal. A field report from Project NANDA at MIT said 95 percent of the organisations studied saw no rapid GenAI P&L impact. External partnerships outperformed internal builds in the implementations studied. That is neither a universal failure probability nor causal proof, but it does underline the importance of scope, integration and relevant experience.
Klarna automated at scale and publicly corrected the human layer. That gives you a useful lesson without forcing the case into a simplistic success-or-failure story.
What does this mean for your first agent?
Three things, concretely.
Don’t start with customer contact. The temptation is strong, because that’s where the volume is and the savings look spectacular. But it’s also exactly the spot with escalations, emotions, and liability. Air Canada had to honor a discount its chatbot invented, Forbes reported, and the German Higher Regional Court of Hamm ruled that a company is directly liable for what its chatbot says. Start with something boring and reversible instead.
Design the human checkpoint from day one. Klarna later reinvested in it. You can build it in right away. Decide up front which cases the agent handles itself and which it hands to a human, and make that handoff cheap rather than absent.
Measure quality, not just speed. Handling time from eleven to two minutes is easy to count and therefore tempting. Quality on difficult conversations is harder to count and therefore easy to ignore. Put that yardstick next to the first from the start.
For employers, the summary is sober: the question isn’t whether you’ll still need people, but exactly where. Routine work can go to the agent, the moments with money, emotions, and signatures stay human. How to draw that line is also covered on our page for employers.
The lesson in one sentence
Klarna showed that an AI assistant can handle an enormous amount of routine work while reachable human support still has value. The hybrid model is not a weak compromise but a deliberate design.
Renforza places agents as the third candidate on your shortlist, and that’s exactly why we don’t steer toward the biggest leap but toward the process an agent can genuinely handle. In an agents intake we settle together on where the agent catches the routine work and where the human checkpoint belongs. That way you don’t have to pay for Klarna’s lesson yourself.
Sources
- Klarna: first-month results for its AI assistant, 27 February 2024
- Klarna Group, Annual Report 2025, filed with the SEC in 2026
- Library of Congress on the OLG Hamm chatbot-liability ruling, 9 June 2026


