Skip to main content
A marketing executive reviewing an initiative board in a glass office before a budget review.

Your AI Budget Is Lying

AI spending is rising faster than the evidence needed to defend it. Five 2026 datasets point to the same gap: adoption is visible, but return is still difficult to prove.

By Dellon S.11 min read

The AI budget is not the proof. The proof is a measurable change in an outcome that mattered before the tool arrived, with enough context for another person to test the claim. If nobody else can reproduce the comparison, the number is a claim, not a result.

A CMO can spend heavily, report time saved, and still be unable to answer the CFO's basic question: what changed, compared with what, at what full cost? Phantom ROI starts when that missing evidence is replaced by a plausible narrative.

The AI audit did not arrive as a single meeting. It arrived as a collection of numbers that refuse to agree with one another.

The audit question is narrower than “did the team use AI?” It is: which decision changed, what baseline moved, and what evidence would have existed if the tool had never been introduced?

A 2026 Writer survey found that 59% of companies spend at least $1 million a year on AI, while only 29% report significant returns. That gap is too large to dismiss as a normal learning curve. It says many companies have moved money, tools, and expectations into production before moving the measurement system with them.

The pressure is visible in marketing. Supermetrics found that 80% of marketers feel pressure to adopt AI, but only 6% say it is fully embedded in their workflows. Gartner's 2026 CMO Spend Survey put the average share of marketing budgets allocated to AI at 15.3%, while only 30% of CMOs said their organizations were ready to scale AI initiatives. Spend is no longer the signal of intent. It is the thing that now needs an explanation.

This is not an argument against AI. It is an argument against treating adoption as a result. A workflow can be popular and still fail to improve revenue, quality, speed, risk, or cost. The CMO who can separate those cases will have a stronger position than the CMO with the most tools.

Five datasets, one shape

The studies use different samples and definitions, so they should not be combined into one precise benchmark. They do, however, describe the same operating pattern: spend and experimentation are widespread; embedded capability and proven financial impact are not.

29%

Writer

report significant returns from AI

6%

Supermetrics

say AI is fully embedded in workflow

30%

Gartner

are ready to scale AI initiatives

39%

McKinsey

report enterprise-level EBIT impact

5%

MIT / Fortune

of pilots reach measurable P&L impact

The figures are not interchangeable. Writer's survey measures reported significant returns, McKinsey's State of AI asks about enterprise-level EBIT impact, and the MIT/Fortune study examines pilots. The useful comparison is directional. A large number of organizations are buying and testing AI; a much smaller group can show where the value appeared and how they know it was caused by the change.

That is the gap the CMO inherits. Finance does not need another statement that the market is moving quickly. It needs a portfolio view that distinguishes a controlled experiment, a useful capability investment, a scaled operating improvement, and a tool that has become an expensive habit.

The portfolio view should also show concentration. If one vendor, one channel, or one operator accounts for most of the reported return, the result has a different risk profile from a gain distributed across repeatable workflows. Concentration does not invalidate the return. It tells you what must be protected, documented, and tested before the result is treated as durable.

Marketing and finance colleagues reviewing an AI initiative portfolio board and a stack of budget papers.
The useful AI review is a joint audit of initiatives, evidence, and decisions. It is not a tour of the latest tools.

How ROI becomes phantom

Phantom ROI usually begins with a real improvement. A team uses an assistant to draft faster, an analyst automates a recurring report, or a media operator finds a useful pattern sooner. The problem appears when that local gain is promoted into a business return without defining what it displaced or what it changed.

There are four common substitutions. Time saved becomes value without an agreed use for the time. More output becomes performance without a quality threshold. A platform's attributed revenue becomes incremental revenue without a counterfactual. And a cost reduction becomes savings without counting implementation, review, integration, training, and risk.

Once the first number enters a budget deck, it gains a second life. The next period compares the new claim with the previous claim instead of comparing the work with a baseline. A plausible estimate becomes a trend. The trend becomes a reason to renew. The renewed budget makes the estimate look validated.

01

Money in

Licenses, APIs, consultants, integration, training, and review enter the budget.

02

Savings claimed

Anecdotal hours or platform attribution become the first return estimate.

03

Dashboard made

The estimate is repeated without a baseline, quality gate, or comparable period.

04

Budget renewed

The repeated number is treated as proof, and the loop starts again.

The missing gate: define the outcome, baseline, counterfactual, quality threshold, and full cost before the result is reported.

Set the baseline first

Record the current cycle time, error rate, throughput, approval rate, or cost per unit before the AI workflow changes. If the baseline is a range, preserve the range and its sample. A single unusually slow week creates an attractive but false improvement when it is used as the comparison.

Choose a measurement window that matches the work. A creative workflow may need several campaign cycles. A reporting workflow may need four weekly closes. The window should be long enough to include normal exceptions, not only the day when the new tool looked fastest.

Translate time into capacity

Time saved is not cash saved unless the capacity is actually released or redeployed. Ask what the team did with the recovered hours. Did it remove contractor spend, shorten a launch, increase tested volume, improve review, or simply create room for work that was already overdue?

Keep the time measure and the business measure separate. It is reasonable to report both. It is not reasonable to multiply estimated hours by a salary rate and call the result revenue without showing the decision that turned capacity into value.

Make confidence visible

Every return claim should carry a confidence level and a reason. High confidence might mean a stable baseline, a controlled comparison, and a result that clears a pre-set threshold. Medium confidence might mean a credible operational gain with limited downstream data. Low confidence means the signal is useful for the next test but not ready for a renewal case.

Confidence is not a disclaimer added at the end. It changes the decision. A low-confidence result earns another instrumented period. A high-confidence result can earn scale. The same number should not trigger both decisions.

A dark dashboard visual with a declining red graph representing a plausible but ungrounded AI return story.
A clean chart can make an untested assumption look like a measured result. The missing fields matter more than the polish.

The super-user illusion

Writer also found that super-users across marketing, sales, HR, and support represented roughly 40% of staff and saved 4.5 times more time than laggards. That is an important signal, but it is easy to report it incorrectly. It does not mean 40% of marketing teams are super-users. It describes the share of staff across four functions in the study.

More importantly, the number is not a scaling plan. A super-user can be evidence that a workflow is possible. It is not evidence that the workflow is transferable, compliant, measurable, or economically useful at the team level. The person may be carrying hidden work: prompt design, data cleaning, review, exception handling, and judgment that the dashboard does not count.

Scaling begins when the team can describe the workflow without relying on the person who discovered it. What input is allowed? What does the system produce? What review is mandatory? Which errors stop the process? What outcome moves? Who owns the metric after launch? Until those answers are written down, the company has a talented operator, not a repeatable capability.

Test portability deliberately. Have a second operator run the workflow with the same inputs, then compare the time, quality, and exception log. If the result depends on private context or undocumented judgment, the next investment is documentation and training, not another license.

What the 29% do differently

The high-return group is not defined by a better slogan. It is defined by the discipline around the work. The practices below make a return claim inspectable.

01

Start with the outcome

Name the business change before naming the tool. It can be a cost, speed, quality, revenue, risk, or capacity measure. If the outcome cannot be written in one sentence, the pilot is not ready for a return claim.

02

Instrument the workflow

Record the baseline, intervention, quality threshold, review time, and full cost. Keep the unit of analysis small enough to compare. A portfolio can hold different units, but each initiative needs a legible measurement contract.

03

Scale or stop on purpose

Set the decision rule before the result arrives. Scale when the outcome clears the threshold at an acceptable cost. Instrument again when confidence is low. Stop when the work consumes attention without earning another period.

That discipline also changes what counts as a useful early result. A pilot does not need to claim revenue impact on day one. It can prove that a workflow is faster at the same quality, that review risk is understood, or that a data dependency blocks scale. An honest non-result is more valuable than a false positive because it changes the next decision.

McKinsey's State of AI research found that 39% of respondents reported enterprise-level EBIT impact from AI, while its analysis also describes how unevenly that impact is distributed. The lesson is not that the number is a target every team can borrow. It is that impact must be located. Which workflow, which business unit, which period, and which evidence made the claim credible?

The same measurement pressure appears outside the headline surveys. Forbes' analysis of marketing measurement describes the strain on deterministic attribution as AI changes the path to action. Prosus' research on AI agents points to the other half of the problem: a small number of systems can carry disproportionate impact, so a blended portfolio average can hide the workflows that need instrumentation. The measurement contract should identify those systems before the result is reported.

Surviving the CFO conversation

Build the answer before the audit is scheduled. Five lines are enough to expose a weak claim and make a strong one portable.

EY and the AIUC-1 consortium found that 64% of companies with more than $1 billion in revenue attributed more than $1 million in 2025 losses to AI failures. That is a different measurement problem from marketing ROI, but it reinforces the same operating point: AI portfolios need ownership, evidence, and a decision path that includes downside.

Bring finance into the measurement contract before the review, not only into the meeting. Finance can challenge cost categories, timing, and the difference between an accounting saving and an avoided expense. Marketing can explain the workflow, quality tradeoffs, and customer impact. The joint record is stronger because each team can test the other team's preferred assumption.

The goal is not to force every experiment into a neat payback period. It is to stop a portfolio from hiding unlike things under one label. Discovery work can be funded as discovery. Capability work can be funded as capability. A scaled improvement can be funded against a return. The failure is calling all three proven ROI.

If the result is not visible yet, say that plainly. A measured pilot with no material lift can still answer whether the idea deserves more time, a different design, or a stop. The audit becomes easier when uncertainty is named before someone has to discover it in a spreadsheet. Review that uncertainty monthly, not only during budget season.

01

Outcome

What changed in the business, and why did it matter?

02

Baseline

What was true before the workflow, and what is the comparison?

03

Result

What moved, by how much, over which period, at what quality?

04

Full cost

What did licenses, people, integration, review, and risk actually cost?

05

Decision

Scale, instrument, change, or stop. What happens next and who owns it?

FAQs

What is phantom ROI in AI marketing?+

Phantom ROI is a return claim that sounds plausible but cannot be tied to a defined outcome, a baseline, a comparable period or counterfactual, and the full cost of producing it. It is not necessarily fabricated. It is ungrounded, so a budget conversation cannot distinguish durable improvement from a good story.

Why are AI marketing pilots hard to measure?+

AI work often changes several variables at once: tools, workflow, team behavior, creative volume, media mix, and data access. If the team does not define the outcome and baseline before launch, later reporting becomes a reconstruction from partial platform data and self-reported time savings.

What should a CMO measure first?+

Start with one business outcome and its operational driver. For example, connect a shorter campaign production cycle to the number of tested variants, qualified pipeline, conversion rate, or cost per approved asset. Record the baseline, full cost, measurement window, and decision rule before the workflow changes.

Does the 29% figure mean the rest of AI spending is wasted?+

No. The Writer finding is a signal about reported significant returns, not a verdict on every initiative. Some work is still in discovery, some creates capability that has not reached the P&L, and some should be stopped. The useful question is whether the company can separate those cases with evidence.

How can a marketing team prepare for an AI spend audit?+

Create a portfolio register with the owner, business outcome, baseline, current result, full cost, confidence level, and next decision for every material initiative. Review it with finance before the formal audit. The register should make it easy to scale proven work, instrument ambiguous work, and stop work that cannot earn another period.

An executive leaving a glass office into a rainy city evening after a budget review.

AI spend becomes credible when the return can survive a change in owner, a new quarter, and a skeptical question.

The audit is not asking for optimism. It is asking what earned another period.