Spend billions of dollars on something and you’d expect a receipt. However, most of the enterprise generative AI pilots examined did not produce one.

A preliminary report from MIT’s Project NANDA, published in 2025, found that roughly 95 percent of integrated enterprise AI pilots showed no measurable effect on profit and loss. Only about 5 percent were extracting millions in value. The authors wrote that “The 95% failure rate for enterprise AI solutions represents the clearest manifestation of the GenAI Divide.”

Note that this is one report, not settled evidence, and it has not been peer-reviewed. It presents preliminary findings gathered between January and June 2025 from a review of more than 300 publicly disclosed AI initiatives, interviews with representatives from 52 organizations, and survey responses from 153 senior leaders. Its figures were partly based on interviews and self-reported outcomes, and the sample might not represent every industry or region. 

Why everyone blamed the wrong thing

When a pilot flops, the instinct is to blame the tool. Is the model smart enough? Is it making things up? Do we need the newer, bigger one? That’s the comfortable question, because it turns a management problem into a shopping problem, and shopping problems get solved by writing another check.

The MIT researchers point somewhere less flattering. Their diagnosis was that “The core barrier to scaling is not infrastructure, regulation, or talent. It is learning.” That’s the report’s read, and I’d treat it as the main barrier the researchers landed on, not the only one that could matter. It reframes the whole thing. The bottleneck they describe sits between the software and the workflow, not simply inside the underlying model.

Workflows, memory, and the learning gap

So what is a “learning gap” in practice? The report describes tools that get dropped into a workflow and then stall because “Most GenAI systems do not retain feedback, adapt to context, or improve over time.” A generic chatbot is impressive in a demo and forgettable by the second week if it doesn’t remember what your team told it yesterday or bend to how your operation actually runs.

Two patterns in particular seem to sink projects. The first is where the money goes. In the report’s directional estimate, about half of generative AI budgets went to sales and marketing, while some of the clearest returns showed up in dull back-office automation.

Companies spent where the excitement was, not necessarily where the payoff was.

The second is the urge to build. Lead author Aditya Challapally saw a pattern of enterprises trying to develop their own tools. Homegrown tools tended to underperform. In the report’s interview sample, external partnerships involving customized, learning-capable tools reached deployment about 67 percent of the time, compared with about 33 percent for internally built tools. The authors cautioned that these were self-reported outcomes and that the relationship did not prove buying caused the better results.

There’s a quieter finding underneath all this that I keep coming back to. Even inside companies whose official pilots flopped, employees kept using AI anyway. The report describes a “shadow AI economy”, with workers from over 90 percent of surveyed companies reporting regular use of personal AI tools for work alongside the failed rollouts. Demand was never the problem, it seems. The workers had already voted with their browser tabs. What the organizations couldn’t do was fold that behavior into how work officially gets done.

What the 5 percent did differently

The interesting group is the small one. Challapally noted that “Some large companies’ pilots and younger startups are really excelling with generative AI”. Note his “some,” because this is a subset, not a rule about who wins.

His explanation for the winners is narrow. It isn’t scale, budget, or the fanciest model. Challapally put it this way: “It’s because they pick one pain point, execute well, and partner smartly with companies who use their tools.” That’s his read, not a proven formula, but the shape of it is telling. One problem, solved properly, with a partner who knows the field. Not a platform, not a company-wide transformation program. A single stuck task, unstuck.

“Is the AI good enough?” likely gets you a vendor demo and a bigger invoice. The sharper question is probably whether the organization is actually built to absorb the thing. Can the tool learn from your people? Is it aimed at a real bottleneck instead of the flashiest department? Has anyone redesigned the work around it, rather than bolting it onto the old way?

That seems to be where the 95 percent and the 5 percent part ways. The underlying models were often capable. Most enterprise systems and workflows were not ready to turn that capability into measurable value.