Does AI actually deliver ROI yet? Not yet, at scale; redesign first.
Across three years and one hundred sources, the evidence converges: AI creates real value at the task level, loses nearly four of every ten saved hours to rework77, and rarely reaches the P&L: except where the work itself was redesigned first.
The loudest "AI pays" numbers come from companies that sell AI. But the most-quoted number on the other side is just as weak: MIT's "95% of pilots fail" rests on a few hundred interviews and survey responses, a sample its own coverage describes two different ways, and its authors call the figure directional rather than audited52.
The case for “yes” is real too. PwC found productivity growing almost four times faster in the industries most exposed to AI.46 The OECD calls 2025 an early signal.92 And the best number cuts both ways: when NBER found 80% of organizations seeing no impact82, that came from what roughly 6,000 executives reported about their own firms. It is the same class of evidence we discount when it points the other way.
"ROI follows work redesign, not tool deployment, and the investment plan should reflect that."
AI investment does not translate into organizational value in one move. It has to survive four stages to reach the P&L, and most of it doesn't. (For now.)
Turning an AI investment into value is a sequence, not a transfer. Four stages, and value is lost at three of them.
Two things are trying to reach your P&L: the time AI saves and the work people can suddenly do that they could not do before. They get stuck in different places. (Select a stage to open it.)
Of 100 units of value, how much reaches each stage
The task
Hours saved and quality gained on a discrete piece of workThis part works. Support agents resolve 15% more issues an hour.38 Writing takes 40% less time.4 Legal tasks improved 50 to 130% in a trial with law students.32
People now do work outside their own job. Almost 17% of what they ask AI for belongs to somebody else's role.99
Nothing is lost yet.
This stage is settled, which is why every vendor demo lives here. The time really is saved. Support agents resolve 15% more issues an hour. The least experienced improve in both speed and quality; the most experienced gain a little speed and lose a little quality. Writing takes 40% less time. Legal tasks improved 50 to 130% in a trial with law students, and quality went up, not down. Doctors write up notes faster and burn out less.48
One study cuts against it. Experienced developers were 19% slower with AI, and thought they were 20% faster.49 The time saved is real. Asking people how much time they saved is not.
Ask: did we measure this, or is this based on perception?
After fixing
What survives rework, verification and redirectionAbout 37% of the time saved goes straight back into fixing what AI wrote. Only 14% of people finish ahead.77
Somebody still has to check work done outside their field, usually the specialist you were trying to skip. Nobody has given them that job.
To a colleague.
The person using AI books the saving. The person who receives their work pays for it. 41% of workers got AI-written work they had to untangle, about two hours each time.63 The saving is easy to see and the cost is spread out, so your reports show a gain your P&L never gets.
There is a second reason. Some of that untangling is work done by somebody outside their own field, using a tool that removed the friction that used to stop them. It arrives looking finished and sounding certain. That is harder to fix than a question would have been.
Ask: when AI gets it wrong, whose day pays for it?
The books
Booked financial impact an executive can point toTwenty people save fifteen minutes a day. That is a full-time job. It arrives as twenty slightly easier days, and nobody can put that in a report.
Trying a task is not the same as finishing it well. Nobody has said who signs it off.
Into slack.
This is arithmetic before it is culture. To use scattered spare minutes you have to gather them, and that means changing what jobs cover and how work moves between people. Two-thirds of employees are given no guidance at all on what to do with the time AI frees up, and more than half do not redirect it to anything strategic (BCG 2026, n=11,749 across 14 markets).96 Where reinvestment was measured directly, 72% of sales teams put almost none of their 4.8 saved hours a week back to work.66
So: 12% of CEOs see both lower costs and higher revenue.78 Over 80% of organizations see nothing at all.82 This is the biggest loss of the four, and none of it is a technology problem.
One more loss worth naming. In the call-centre study, the weakest agents improved most and nearly caught the best. Some operators concluded they no longer needed to hire the best. But the AI was passing on what the best agents knew. Cut them and you have banked a saving by destroying the thing that produced it.
Ask: what did the freed-up hours actually produce, and who decided that?
The market
Excess return to shareholders of the organizations that adoptedWhatever is left has to beat what you paid for the computing, in a market where your competitors bought the same tool.
To your customers, through price. And to the company that sold you the tool.
So far the money has gone to the companies selling AI, not the ones using it.80 Shares in AI buyers have simply tracked the market. Morgan Stanley disagrees and reports better margins for adopters, but that is a broker forecast, not a result.84
Anything your competitors can also buy is not an advantage. It is the price of staying in the game. What you own is your ability to turn what AI produces into money, which is stages two and three. Nobody sells that, and nobody can copy it from your press release.
There is a newer problem underneath this. Organizations are spending tens of millions on AI usage, and nobody can predict what a given job will cost, including the AI itself. Projects stall or blow the budget halfway through. You cannot calculate a return when you do not know the bill.
Ask: if every competitor deploys the same tool, what remains ours?
Details and footnotes 3
Do not multiply these numbers together. The figures on the bars come from different studies with different samples. The 63 beside after fixing is one study’s rework share, not 63% of the gains on the stage above. The bars show roughly which stages lose the most, and nothing more precise than that.
This is why the research looks like it disagrees with itself. The legal study measured stage one.32 The survey of 6,000 executives measured stage three.82 They are counting different things.
The argument against all this: every handoff is also a wait, and waiting usually takes far longer than the work itself. A marketer who no longer waits two days for a developer has compressed cycle time even though no hour appeared on any timesheet, and that is a real P&L item invisible to hour-counting. Crossover also runs higher in small workspaces, 18.9% at 2–5 seats against 16.3% at 101 or more99, which suggests boundaries get crossed when no specialist is available. Large organizations keep their specialists, so the coordination gain is structurally hardest to capture in exactly the organizations reading this.
When people say AI creates value they mean at least six different things. The column below rates how good the evidence is on each, not how well AI performs. Strong evidence can still show a disappointing result, and on several of these it does. The first evidence sweep behind this brief hunted hours and money, and on that basis four of these channels looked barely researched. They are not. A second sweep, through the research on how organizations actually behave, found what the first one was not built to see.
| Channel | Evidence base | Why |
|---|---|---|
| Hours | Strong | Controlled trials, payroll records, representative firm panels. The best-evidenced channel in the literature. |
| Scope | Emerging | One large-scale usage study, vendor-authored, counting messages rather than outcomes. |
| Coordination | Emerging | A randomized field experiment (INSEAD, Jan 2026, 316 employees across 42 teams)76 found grounded AI raised collaboration and knowledge-network density, with specialists gaining centrality and generalists gaining throughput.One site, one working paper. |
| Quality | Emerging at task level, thin at P&L | Abundant at task level and now with two system-level datasets, both negative: delivery stability down 7.2% (DORA) and duplicated code up eightfold (GitClear)2430. Colonoscopy detection fell 6.0 points after routine AI exposure.51 Quality net of AI-introduced error remains unmeasured. |
| Innovation | Strong | Peer-reviewed work shows AI raising individual creativity while compressing collective diversity (Doshi & Hauser, PNAS replication)1454, and LLM research ideas rated more novel at proposal, then falling behind human ideas once executed. |
| Morale | Strong, and split by method | The split is measurement against perception, not pole against pole. Population-panel evidence finds no sizeable wellbeing harm36; the review literature on algorithmic management finds real autonomy and surveillance costs41; and the one measured interpersonal cost is clear, with 42% viewing colleagues who send AI slop as less trustworthy63. |
Details and footnotes 2
A brief claiming “AI does not deliver value” while measuring only hours would be overclaiming. The defensible statement is narrower: on the channel with the best evidence, value is real at stage one and rarely survives to stage three.
Why the verdict speaks only about money. The question this brief asks is whether AI investment reaches the P&L, so the verdict is scoped to that. It is not a claim that nothing happens elsewhere. Innovation and morale both have strong evidence bases, and what they show is the same shape as everything else here.
Individual gains, collective costs. AI raises one worker’s satisfaction and lowers the trust of the colleague receiving their output. It lifts the weakest performers toward the strongest, then removes the reason to keep employing the strongest, who were the source of what the model learned. Coordination is the one channel where the signs currently agree, with individual centrality gains coinciding with denser networks overall. Everywhere else, the stage-one case and the stage-three case are measuring different levels of the same organization, and a business case built on the first without accounting for the second is not a business case.
Four chairs in the C-suite, four versions of the question
"I told the board AI would transform our cost base. Where is it?"
56% of 4,454 CEOs report no significant financial benefit to date.78 The pressure is now investor-facing.
"Can I attribute any P&L movement to our AI spend, and book the savings?"
Attribution is the core problem: individually saved time leaks out of the organization before it can be booked.
"Is this a tooling problem or a work-design problem, and am I about to be handed a RIF target justified by unproven gains?"
Fewer than 1% of last year's layoffs traced to actual AI productivity gains79 (Gartner).
"Are we cutting ahead of proven returns, and what's our rehiring exposure?"
Gartner warns organizations that cut on AI's promise are already rehiring for roles they eliminated.