In short: The pilot-to-production gap is closing fast, but the pilot-to-profit gap is not - because most businesses measure whether the AI works instead of whether the P&L moved, and those are different questions.
The pilot worked. The demo landed. Someone said “this is going to change how we operate” and meant it. Six months on, the tool is live, people use it, and nobody is unhappy.
And your operating costs are exactly where they were.
This is the most common AI conversation we have, and it is almost never a conversation about a failed project. The technology usually works. The pilot-to-production gap is genuinely closing. The pilot-to-profit gap is not, and they are two entirely different problems with two entirely different fixes.
Here is what the 2026 data actually shows, the three costs that quietly eat the saving, and what the businesses with numbers on the board did differently.
1. Shipping is no longer the hard part. The proof is in the numbers.
For two years the received wisdom was that AI projects die before production. That was true, and it is becoming less true fast.
Lenovo’s CIO Playbook 2026, with research from IDC (published 27 January 2026, 3,120 IT and business decision-makers), found that 46% of AI proof-of-concepts have already progressed into production. Compare that with the figure everyone quoted through 2025, from the same research pairing a year earlier: for every 33 proof-of-concepts launched, four made it to widescale deployment. Roughly one in eight.
So the engineering problem is being solved. Which makes the next set of numbers harder to explain away.
PwC’s 2026 Global CEO Survey is the largest sample available on this question - 4,454 CEOs across 95 countries, fielded from September to November 2025. Its finding: only 12% of CEOs say AI has delivered both cost and revenue benefits. 56% report no significant financial benefit to date.
Deloitte’s 2026 State of AI in the Enterprise (3,235 leaders, 24 countries) frames the same gap from the other direction: 74% of organizations hope to grow revenue through AI, while just 20% already are.
Nobody in these surveys is reporting that the AI does not work. They are reporting that they cannot find the money.
2. Why a working pilot does not become a saving
The single most revealing statistic we found this year is not about technology at all. BCG surveyed 152 CEOs at companies with revenues over $500m (published 22 July 2026) and found that nearly nine in ten say they are now seeing cost or revenue benefits from AI in targeted areas - and that only 14% have clearly defined the P&L impact for all their AI initiatives. More than half named linking AI to the P&L as a key barrier: a 42-point gap between the problem and the practice.
Read those two together. Benefits are visible in targeted areas. P&L impact is defined almost nowhere. That is not a technology failure, it is a measurement failure, and it has a predictable shape. The habits that close that gap are not complicated, and we list them in 27 AI rules for business owners.
BCG’s earlier AI Radar (January 2025, 1,803 executives) found 60% of companies were not defining or monitoring any financial KPIs for AI value creation at all. If you never set the number, the saving cannot show up - because “saving” is a comparison, and there is nothing to compare against.
The trap almost everyone falls into
A pilot is judged on whether the AI produced good output. A business case is judged on whether a cost line went down. Those are different tests, and passing the first tells you almost nothing about the second. Most AI pilots are never actually tested against the business case, because nobody wrote one down before starting. One of the quietest examples is inference cost, where prompt caching can look enabled and save nothing at all.
3. The perception gap is real, and it is measurable
Here is the finding every owner should know before trusting a productivity estimate from their own team.
METR ran a controlled study of 16 experienced open-source developers across 246 real tasks (published 10 July 2025). Developers predicted AI tools would make them 24% faster. In fact, they took 19% longer when allowed to use AI. And even after finishing, they still believed they had been sped up by about 20%.
Sit with that. Not only was the productivity gain absent, the people doing the work could not tell. They were slower and felt faster.
This is a small study on experienced developers working on mature codebases, and we would not stretch it into a claim about all AI work. But it is enough to retire one very common practice: asking your team whether the AI is saving time and writing the answer into a business case. Self-reported time savings are the weakest evidence in this entire field, and they are what most AI ROI claims rest on.

4. The three costs nobody budgets
When a working pilot fails to move the P&L, the money has usually gone to one of three places.
The verification tax
Somebody now checks the output. If that person is a senior member of staff spot-checking everything, you have added a cost that scales with usage while removing one that did not. This is the cost most often left out entirely, because the reviewing happens informally and never gets logged as AI cost.
The plumbing, which is most of the work
The canonical reference here is a decade old and still correct. Google researchers wrote in Hidden Technical Debt in Machine Learning Systems (NIPS, 2015) that “only a small fraction of real-world ML systems is composed of the ML code”, with the surrounding infrastructure “vast and complex”. Their opening observation applies exactly to AI pilots: developing and deploying is relatively fast and cheap, and maintaining is difficult and expensive.
The 2026 numbers bear this out. In a survey of 372 organizations on AI cost management, the top sources of unexpected spend were data platforms (56%) and network access costs (52%). Large language models ranked fifth. Four in five organizations missed their AI cost forecasts by 25% or more. Some of that is a tier problem rather than a plumbing problem, which is why it pays to be deliberate about whether to build, buy or just use a chat tool for each piece of work.
The work that came back
The saving is often real and then quietly reversed. Commonwealth Bank of Australia told about 45 staff in July 2025 that their roles were redundant after deploying a voice bot it said would cut roughly 2,000 calls a week. The union disputed the volumes, noting staff were being offered overtime and team leaders were answering calls. On 21 August 2025 the bank reversed the decision and apologised, stating its initial assessment “did not adequately consider all relevant business considerations and the roles were not redundant”.
A saving that has to be given back was never a saving. It was a forecast.
The strongest objection to all of this
“Not everything valuable shows up in the P&L. Faster answers, better decisions and less drudgery are real even if I cannot invoice them.” Completely true, and we would never argue a business should only fund what it can count. But be honest about which one you bought. If you approved AI spend on a cost-saving case, judge it on cost. If you approved it to improve quality of work, say so out loud, stop calling it an efficiency program, and measure quality instead. The damage comes from funding it as one and defending it as the other.
5. Where the money actually came from, in the cases that worked
MIT’s State of AI in Business 2025 is best known for a statistic we would ask you to treat sceptically (see Myth vs Facts below). Its more useful section is the one nobody quotes, on where returns genuinely appeared: elimination of outsourcing contracts worth $2m to $10m a year in customer service and document processing, and around a 30% reduction in external creative and agency spend.
And the crucial detail: those gains “came without material workforce reduction”. The report is explicit that ROI emerged from reduced external spend - dropping BPO contracts, cutting agency fees, replacing expensive consultants - rather than from cutting staff.
That matches what we see. External spend is contractual, visible and cancellable, so a reduction lands in the accounts immediately. Internal time saved diffuses into other work and shows up nowhere. If your AI business case depends on staff time savings, you are relying on the one category that is hardest to measure and easiest to lose.
A real number from our own deployment
In the D2C deployment where we cut operational costs 68%, the target was set before we built anything: roughly £140,000 a year of staff time across three functions, doing work that turned out to be about 85% rule-following. Support resolution went from 11 minutes to under 90 seconds and the build paid back in under five months - and those figures are net of the system’s own operating cost, including inference. That last clause is the whole discipline. A saving quoted gross of what the AI costs to run is not a saving.
The other consistent marker is workflow redesign rather than workflow automation. BCG’s July 2026 CEO research found high performers were roughly seven times more likely to redesign workflows and reshape the business end-to-end, rather than bolting AI onto the process as it stood. Automating a bad process faster mostly produces the same result sooner.
Myth vs Facts
Myth: “95% of AI projects fail - everyone knows that.”
Fact: That figure comes from one July 2025 report labeled preliminary findings, version 0.1, based on 52 interviews and 153 leaders surveyed at conferences. It says 95% of organizations saw zero return, not that 95% of pilots failed - and by its own funnel, of organizations that actually ran a pilot, roughly a quarter succeeded. The PDF is no longer served from MIT’s own domain. Use PwC’s 56% from 4,454 CEOs instead.
Myth: “Our pilot proved the business case.”
Fact: A pilot proves feasibility. Only 14% of CEOs have clearly defined the P&L impact of all their AI initiatives, and 60% were not tracking any financial KPI for AI value at all. If no baseline was recorded before the pilot, the business case remains unproven whatever the pilot showed.
Myth: “The team says it saves them hours, so it is working.”
Fact: In METR’s controlled study, developers were 19% slower with AI tools while believing they were 20% faster. Perceived time savings are not evidence, and they point the wrong way often enough to be dangerous in a business case.
Myth: “We need to get the model or the prompts right.”
Fact: Model spend is rarely the problem. Data platforms and networking were the top sources of unexpected AI cost across 372 organizations, with LLMs fifth. The expensive part is the plumbing and the checking, exactly as ML researchers described in 2015. Most of that plumbing cost is the same short list of data problems we hit on almost every build.
Working pilot, no savings: where to look
| What you are seeing | What it usually means | What to do |
|---|---|---|
| Everyone likes it, no cost line moved | No baseline was ever recorded | Measure the old process now; you cannot back-date a baseline |
| Time saved but headcount unchanged | Saved time diffused into other work | Retarget at external spend: agencies, outsourcing, contractors |
| A senior person checks all output | You added a verification cost as usage grew | Reduce review by scope, not by hope: narrow what it does unsupervised |
| Bill higher than forecast | Data movement and integration, not tokens | Audit platform and egress costs before optimizing prompts |
| Same process, now with AI in it | Automation without redesign | Redesign the workflow; high performers are ~7x more likely to |
| Savings claimed, then reversed | The forecast was mistaken for a result | Hold the reduction for a full quarter before booking it |
What this means if you are running AI in your business
The uncomfortable part is that most of this has to happen before the pilot, and almost nobody does it, because a pilot is exciting and a baseline is admin.
If you have not started at all, the barrier is usually not what you think it is. And if your pilot is already live and unmeasured, you have not lost - but you do have to do the boring thing now. Pick the single process the AI touches most, measure how it runs today, and set a target with a currency symbol and a date attached. It is the same discipline that separates the projects that reach production from the ones that stall, which we covered in why agentic AI projects fail.
- Write down what the current process costs before you change it, in money, not in impressions
- Name the cost line that should fall, and who owns that line
- Quote every saving net of what the AI costs to run, inference included
- Look for external spend to cut before internal time to save
- Count the review time as a cost of the AI, because that is what it is
- Redesign the workflow rather than inserting AI into the existing one
- Hold any claimed reduction for a full quarter before treating it as banked
One thing worth knowing about the wider mood: BCG’s AI Radar 2026, covering 2,360 executives including 640 CEOs, found 94% plan to keep investing in AI even if it does not drive immediate returns. That is a lot of patience. It is also exactly the condition in which unmeasured spending survives for years, so being the business that can actually show its numbers is becoming a genuine advantage rather than good hygiene.
Can you prove your AI is paying for itself?
Tick each one that is true today.
- We did not record what the process cost before we introduced AI
- No specific cost line was named as the one that should fall
- Our savings estimate comes from asking the team how much time they save
- We do not subtract the AI’s running cost from the saving we quote
- Nobody counts the time spent checking AI output as a cost
- The underlying process is unchanged, just with AI inside it
- We could not say today what any one AI tool costs us per completed task
If this were your deployment, here is our first move
We would not touch the AI. We would spend the first week on a cost attribution exercise: where the operational budget goes, which of that work is genuinely rule-following, and which external contracts exist because that work has to happen. That analysis is what makes a target real, and it is the step that got skipped in almost every stalled deployment we have been called into.
Then we would pick one process, set a number with a date, and instrument it before changing anything. Unglamorous, and it is the difference between a tool people like and a line on the P&L that moved.
Your pilot working was never the finish line. It was the point at which the actual work started. If you have something live and cannot yet prove what it earns, talk to us and we will help you find the number - including if the honest answer is that there is not one yet.