A 2025 field report from Project NANDA at MIT said that 95 percent of the organisations studied saw no rapid, measurable GenAI impact on their profit and loss statement. That figure is directional, not a universal failure probability for every AI project. The report draws on interviews, a survey and an analysis of public implementations; it is not a randomised or peer-reviewed effect study.
That number gets quoted everywhere these days, usually as a club: sceptics use it to prove AI is hype, vendors use it to prove you need their help. What sits behind the number is more useful, because it points to better project design.
What does the report actually say?
Three things. One: within the group studied, few integrated GenAI initiatives quickly reached demonstrable P&L impact. Two: the report emphasises data quality, workflow integration and systems that learn too little from feedback. Three: external partnerships outperformed internal builds in the implementations studied. The frequently quoted ratio of roughly 67 versus 33 percent comes from that report and should be read as an observed pattern, not a guaranteed success rate.
Note what the report does not prove. It does not prove that 95 percent of every AI pilot everywhere fails, nor that external guidance itself causes success. The narrower signal is more useful: a demo becomes business value only when data, workflow, ownership and measurement line up.
So why do those pilots fail?
Three recurring causes from the report, translated into scenes you will recognise.
Data quality. The pilot is supposed to draft quotes from customer data, but the customer data lives in three systems that contradict each other. The agent produces beautifully formatted nonsense, trust evaporates within two weeks, and the pilot dies quietly. Data cleanup is unglamorous prep work, and it gets skipped because the demo worked fine without it. The demo ran on clean sample data.
Workflow integration. The tool exists, but nobody has to use it. It lives in a separate tab, the old process keeps running, and after the kickoff, usage drops to zero. A pilot without an agreed-upon place in the process isn’t a pilot. It’s a subscription.
Tools that don’t learn. The user corrects the same mistake for the twentieth time and the tool does nothing with it. At that point the tool isn’t a colleague. It’s an intern who never remembers anything, and people rightly stop investing in the relationship.
Why is this actually good news for smaller companies?
Because many failure causes are organisational, and small organisations can have an advantage there. A large enterprise often has many systems and approval layers. A smaller company can assign one owner and scope one process more quickly. That makes small, measurable starts easier, although it guarantees nothing.
Figures from Statistics Netherlands show how much room remains: in 2025, 29.8 percent of Dutch SMEs with 10 to 249 employees used AI technology, versus 66.2 percent of large companies. Smaller companies starting now can use lessons from earlier implementations.
How do you improve the odds of demonstrable results?
Four rules, derived straight from the failure causes:
- Pick one process, not a vision. Not “bringing AI into the organization” but “the incoming invoices.” How to pick that process is covered in our piece on which process to hand to an AI agent first.
- Measure in euros or hours, from day one. Agree up front what success means: this many hours of processing time saved, this many fewer errors. If the result can’t land on a P&L or a timesheet, it isn’t a result.
- Clean the data first. Boring, but often decisive. Budget the preparation explicitly.
- Bring in relevant experience if you have not done this before. It need not be a large consulting engagement; the point is having someone at the table who recognises the technical and organisational pitfalls.
And if the pilot fails anyway?
Then you want it to have been small enough to learn from rather than be buried by. A well-designed pilot that fails tells you exactly where your process or your data isn’t ready. That’s annoying but useful information, and considerably cheaper than learning the same lesson at production scale.
This is also the lens Renforza applies to agents: a placement with a yardstick, not an experiment on a drip feed. In an intake, we settle on one process, the measurable outcome and the point where we would stop. That way every agent starts with the question many pilots skip: what will tell us in eight weeks that this is working?
Sources and scope
- The GenAI Divide: State of AI in Business 2025, a field report from Project NANDA at MIT
- Statistics Netherlands: use of AI technology by company size in 2025, published 16 March 2026
The 95 percent figure is an outcome produced by the NANDA report’s dataset and definitions. Do not use it as a universal probability for every type of AI project.


