The AI audit did not arrive as a single meeting. It arrived as a collection of numbers that refuse to agree with one another.
The audit question is narrower than “did the team use AI?” It is: which decision changed, what baseline moved, and what evidence would have existed if the tool had never been introduced?
A 2026 Writer survey found that 59% of companies spend at least $1 million a year on AI, while only 29% report significant returns. That gap is too large to dismiss as a normal learning curve. It says many companies have moved money, tools, and expectations into production before moving the measurement system with them.
The pressure is visible in marketing. Supermetrics found that 80% of marketers feel pressure to adopt AI, but only 6% say it is fully embedded in their workflows. Gartner's 2026 CMO Spend Survey put the average share of marketing budgets allocated to AI at 15.3%, while only 30% of CMOs said their organizations were ready to scale AI initiatives. Spend is no longer the signal of intent. It is the thing that now needs an explanation.
This is not an argument against AI. It is an argument against treating adoption as a result. A workflow can be popular and still fail to improve revenue, quality, speed, risk, or cost. The CMO who can separate those cases will have a stronger position than the CMO with the most tools.
Five datasets, one shape
The studies use different samples and definitions, so they should not be combined into one precise benchmark. They do, however, describe the same operating pattern: spend and experimentation are widespread; embedded capability and proven financial impact are not.
29%
Writer
report significant returns from AI
6%
Supermetrics
say AI is fully embedded in workflow
30%
Gartner
are ready to scale AI initiatives
39%
McKinsey
report enterprise-level EBIT impact
5%
MIT / Fortune
of pilots reach measurable P&L impact
The figures are not interchangeable. Writer's survey measures reported significant returns, McKinsey's State of AI asks about enterprise-level EBIT impact, and the MIT/Fortune study examines pilots. The useful comparison is directional. A large number of organizations are buying and testing AI; a much smaller group can show where the value appeared and how they know it was caused by the change.
That is the gap the CMO inherits. Finance does not need another statement that the market is moving quickly. It needs a portfolio view that distinguishes a controlled experiment, a useful capability investment, a scaled operating improvement, and a tool that has become an expensive habit.
The portfolio view should also show concentration. If one vendor, one channel, or one operator accounts for most of the reported return, the result has a different risk profile from a gain distributed across repeatable workflows. Concentration does not invalidate the return. It tells you what must be protected, documented, and tested before the result is treated as durable.

How ROI becomes phantom
Phantom ROI usually begins with a real improvement. A team uses an assistant to draft faster, an analyst automates a recurring report, or a media operator finds a useful pattern sooner. The problem appears when that local gain is promoted into a business return without defining what it displaced or what it changed.
There are four common substitutions. Time saved becomes value without an agreed use for the time. More output becomes performance without a quality threshold. A platform's attributed revenue becomes incremental revenue without a counterfactual. And a cost reduction becomes savings without counting implementation, review, integration, training, and risk.
Once the first number enters a budget deck, it gains a second life. The next period compares the new claim with the previous claim instead of comparing the work with a baseline. A plausible estimate becomes a trend. The trend becomes a reason to renew. The renewed budget makes the estimate look validated.
01
Money in
Licenses, APIs, consultants, integration, training, and review enter the budget.
02
Savings claimed
Anecdotal hours or platform attribution become the first return estimate.
03
Dashboard made
The estimate is repeated without a baseline, quality gate, or comparable period.
04
Budget renewed
The repeated number is treated as proof, and the loop starts again.
Set the baseline first
Record the current cycle time, error rate, throughput, approval rate, or cost per unit before the AI workflow changes. If the baseline is a range, preserve the range and its sample. A single unusually slow week creates an attractive but false improvement when it is used as the comparison.
Choose a measurement window that matches the work. A creative workflow may need several campaign cycles. A reporting workflow may need four weekly closes. The window should be long enough to include normal exceptions, not only the day when the new tool looked fastest.
Translate time into capacity
Time saved is not cash saved unless the capacity is actually released or redeployed. Ask what the team did with the recovered hours. Did it remove contractor spend, shorten a launch, increase tested volume, improve review, or simply create room for work that was already overdue?
Keep the time measure and the business measure separate. It is reasonable to report both. It is not reasonable to multiply estimated hours by a salary rate and call the result revenue without showing the decision that turned capacity into value.
Make confidence visible
Every return claim should carry a confidence level and a reason. High confidence might mean a stable baseline, a controlled comparison, and a result that clears a pre-set threshold. Medium confidence might mean a credible operational gain with limited downstream data. Low confidence means the signal is useful for the next test but not ready for a renewal case.
Confidence is not a disclaimer added at the end. It changes the decision. A low-confidence result earns another instrumented period. A high-confidence result can earn scale. The same number should not trigger both decisions.

The super-user illusion
Writer also found that super-users across marketing, sales, HR, and support represented roughly 40% of staff and saved 4.5 times more time than laggards. That is an important signal, but it is easy to report it incorrectly. It does not mean 40% of marketing teams are super-users. It describes the share of staff across four functions in the study.
More importantly, the number is not a scaling plan. A super-user can be evidence that a workflow is possible. It is not evidence that the workflow is transferable, compliant, measurable, or economically useful at the team level. The person may be carrying hidden work: prompt design, data cleaning, review, exception handling, and judgment that the dashboard does not count.
Scaling begins when the team can describe the workflow without relying on the person who discovered it. What input is allowed? What does the system produce? What review is mandatory? Which errors stop the process? What outcome moves? Who owns the metric after launch? Until those answers are written down, the company has a talented operator, not a repeatable capability.
Test portability deliberately. Have a second operator run the workflow with the same inputs, then compare the time, quality, and exception log. If the result depends on private context or undocumented judgment, the next investment is documentation and training, not another license.
What the 29% do differently
The high-return group is not defined by a better slogan. It is defined by the discipline around the work. The practices below make a return claim inspectable.
01
Start with the outcome
Name the business change before naming the tool. It can be a cost, speed, quality, revenue, risk, or capacity measure. If the outcome cannot be written in one sentence, the pilot is not ready for a return claim.
02
Instrument the workflow
Record the baseline, intervention, quality threshold, review time, and full cost. Keep the unit of analysis small enough to compare. A portfolio can hold different units, but each initiative needs a legible measurement contract.
03
Scale or stop on purpose
Set the decision rule before the result arrives. Scale when the outcome clears the threshold at an acceptable cost. Instrument again when confidence is low. Stop when the work consumes attention without earning another period.
That discipline also changes what counts as a useful early result. A pilot does not need to claim revenue impact on day one. It can prove that a workflow is faster at the same quality, that review risk is understood, or that a data dependency blocks scale. An honest non-result is more valuable than a false positive because it changes the next decision.
McKinsey's State of AI research found that 39% of respondents reported enterprise-level EBIT impact from AI, while its analysis also describes how unevenly that impact is distributed. The lesson is not that the number is a target every team can borrow. It is that impact must be located. Which workflow, which business unit, which period, and which evidence made the claim credible?
The same measurement pressure appears outside the headline surveys. Forbes' analysis of marketing measurement describes the strain on deterministic attribution as AI changes the path to action. Prosus' research on AI agents points to the other half of the problem: a small number of systems can carry disproportionate impact, so a blended portfolio average can hide the workflows that need instrumentation. The measurement contract should identify those systems before the result is reported.
Surviving the CFO conversation
Build the answer before the audit is scheduled. Five lines are enough to expose a weak claim and make a strong one portable.
EY and the AIUC-1 consortium found that 64% of companies with more than $1 billion in revenue attributed more than $1 million in 2025 losses to AI failures. That is a different measurement problem from marketing ROI, but it reinforces the same operating point: AI portfolios need ownership, evidence, and a decision path that includes downside.
Bring finance into the measurement contract before the review, not only into the meeting. Finance can challenge cost categories, timing, and the difference between an accounting saving and an avoided expense. Marketing can explain the workflow, quality tradeoffs, and customer impact. The joint record is stronger because each team can test the other team's preferred assumption.
The goal is not to force every experiment into a neat payback period. It is to stop a portfolio from hiding unlike things under one label. Discovery work can be funded as discovery. Capability work can be funded as capability. A scaled improvement can be funded against a return. The failure is calling all three proven ROI.
If the result is not visible yet, say that plainly. A measured pilot with no material lift can still answer whether the idea deserves more time, a different design, or a stop. The audit becomes easier when uncertainty is named before someone has to discover it in a spreadsheet. Review that uncertainty monthly, not only during budget season.
01
Outcome
What changed in the business, and why did it matter?
02
Baseline
What was true before the workflow, and what is the comparison?
03
Result
What moved, by how much, over which period, at what quality?
04
Full cost
What did licenses, people, integration, review, and risk actually cost?
05
Decision
Scale, instrument, change, or stop. What happens next and who owns it?

