Skip to main content
The ROI Accounting Trap: Why AI Costs Stay Visible But Benefits Don't
June 13, 2026·8 min read

The ROI Accounting Trap: Why AI Costs Stay Visible But Benefits Don't

Uber burned $30M in four months on AI. When asked if it was worth it, the COO said 'it's very hard to draw a line' between cost and features. That's not a measurement problem-it's a fundamental accounting gap in how companies value AI.

DS
Dellon S.

Digital Marketing

Uber's COO admitted it last month: the company burned through its entire 2026 AI budget in four months. When asked if the investment was worth it, he said something that should terrify every CMO: "It's very hard to draw a line between rising AI costs and useful features for customers."

Not "hard to measure." Hard to draw a line.

That's not a measurement problem. That's an accounting problem. And it's about to get worse.

The paradox is simple: AI vendors have made inference costs hypervisible. You see every token, every API call, every dollar spent. But the value of that inference-the actual business impact-remains invisible. This creates a structural gap where billions are being spent on AI while companies have literally no way to prove ROI. Unlike other tech investments, this gap only widens the more you spend.

The Cost Transparency Illusion

Here's what happened at Uber. Finance knew exactly what they spent. Token costs are transparent. API logs are auditable. Every inference shows up in billing with the precision of a utility meter: per-request cost, timestamp, model, token count, latency.

It feels scientific. It feels real.

But CMOs don't care about token counts. They care about:

  • Did this feature reduce customer churn?
  • Did it increase engagement time?
  • Did it improve product-market fit?
  • Did it help us compete against rivals shipping the same AI feature?

Those numbers stayed hidden.

Comparison of visible costs vs invisible benefits in AI spending
Every cost is tracked down to the token. Every benefit stays theoretical.

Inference costs exploded at Uber because Claude, ChatGPT, and Mixtral got smarter, not cheaper. Longer reasoning chains cost tokens. Bigger context windows cost tokens. Better hallucination-avoidance requires redundant reasoning that costs tokens.

So the engineering team's answer to "What did we get for $30M?" was: "Better code generation accuracy. Fewer incorrect API calls. Faster response times."

True. Measurable to an engineer. Meaningless to a CMO. Because the product roadmap assumed these improvements would drive adoption and engagement. They didn't. The company still had the same churn. The same engagement plateau. The same market share problem. Just a $30M invoice that's hard to justify.

The cost is real. The benefit stays theoretical.

Why Marketing Can't Measure AI Output

The problem runs deeper than just bad analytics.

When you buy a database, you know the tradeoff: $X per month for Y GB of storage, with Z millisecond query latency. You measure the benefit: product uptime improves, user experience improves, revenue correlation exists. The two live in different silos, but they're connectable.

When you deploy Claude as the backbone of customer support, the equation breaks:

  • Cost: 1M tokens/month times $0.003 equals $3K. Crystal clear.
  • Benefit: Did it reduce support ticket volume? You'd need to control for seasonality, pricing changes, competitor promotions. Did it improve CSAT? You'd need to control for ticket complexity, agent tenure, product quality. Did it impact churn? You'd need to control for everything.

Marketing teams don't have the experimental rigor to run these controls. Engineering teams building AI pipelines aren't incentivized to quantify marketing impact. So the cost sits on one side of the ledger (auditable, growing, undeniable) and the benefit sits on the other (hoped-for, theoretical, unmeasured).

By the time a CMO asks "Was this worth it?", the budget is spent. The vendor relationship is locked. The infrastructure is woven into the product. And the honest answer is: "We'll find out next quarter." Spoiler: next quarter, you still won't know.

This is different from other tech spending because other tech has external benchmarks. Unsure if your new CRM is working? Compare your sales cycle to competitors' public benchmarks. Unsure about your email platform? Check industry CSAT averages.

But AI? There are no benchmarks for "good inference." No public data on "Does Claude support improve NPS?" No industry standard for "ROI per token spent." So every company is flying blind, hoping their AI investment pays off while competitors do the same thing in the dark.

The Vendor's Incentive Structure

Here's what makes this worse: vendors are structurally incentivized to hide ROI measurement.

A Claude API call costs $0.0003. Sounds free. Run 100K inferences a day? Still cheap, you might think. But run them in production for millions of users, millions of times daily, with more aggressive prompting and fine-tuning, and suddenly you're burning $500K/quarter with nothing to show except "better model outputs."

Vendors know this. So they don't talk about cost-per-inference. They talk about "model capability uplift" and "reasoning depth" and "reduced hallucinations." Those are inputs to ROI, not outputs.

Marketing manager checking AI billing dashboard with frustration
The moment a CMO realizes token costs are exploding but business metrics aren't moving.

A CMO paying attention hears: "We're using the best available model. We're getting higher-quality outputs." But there's no mechanism to measure whether that quality actually matters to the business.

Meanwhile, the finance team sees: "AI spend is up 300% YoY." The CEO asks why. Nobody can connect it to revenue, user retention, or competitive advantage. By then, four months of budget are burned. Switching costs are high. The vendor is embedded.

The Lock-In Shadow

This is where the game gets strategic.

The longer this gap persists (measurable costs versus unmeasurable benefits), the more locked in you become.

Why? Switching becomes a cost argument, not a benefit argument.

If you're on Claude and someone says "Switch to a cheaper model," the response is: "We've optimized every prompt for Claude's reasoning patterns. Our infrastructure is built around Claude's latency characteristics. The migration cost is $500K and three months."

But here's the thing: you never had a benefit measurement for Claude in the first place. You just have a cost. And now you have switching costs too.

This is vendor lock-in that doesn't feel like lock-in. It's not lock-in by contract. It's lock-in by opacity. The vendor wins not by being obviously superior, but by making the cost of staying marginally lower than the cost of leaving. And they do this by ensuring you never measure whether the model is actually helping your business.

Case Study: The Content Generation Trap

Take content generation, probably the #1 use case for Claude and ChatGPT in marketing.

A typical enterprise pays $50K-150K/month for Claude API calls used by content teams. The pitch is simple: "Generate 100 blog posts a month instead of 20. Five times the productivity."

Sounds great. Sounds measurable.

But what actually happens:

  1. Publish 100 AI-generated posts instead of 20 human posts.
  2. Measure traffic to the new posts.
  3. Compare to traffic on human-written posts from two years ago.
  4. Celebrate 40% more traffic.
  5. Don't control for: SEO algorithm changes, competitive landscape shifts, topic seasonality, link-building improvements, user behavior evolution.

Three months later, organic traffic plateaus. The content team says: "We need better prompting." Finance says: "We're still spending $150K/month." The team tries better prompting, spends $200K/month, and traffic stays flat.

At this point, you've got $600K in sunk AI costs and no way to measure whether the content strategy actually worked. You're locked in because your entire content pipeline is built around AI-generated drafts. Switching to human-only content would require rebuilding the workflow and admitting the AI experiment failed.

So you keep spending. You keep hoping. You keep optimizing Claude prompts instead of measuring actual ROI.

The vendor wins. The CMO loses. And nobody can prove it because the accounting was never there to begin with.

What Actually Needs to Happen

  1. Stop conflating inference cost with business impact. Your Claude bill is $50K this month. Fact. But did it increase revenue? Did it reduce churn? Did it improve product metrics that matter? These are separate questions with separate answers. Don't assume they're connected.

  2. Run marketing experiments, not just deployments. If you're using AI for customer support, set up a real A/B test. Route half your support volume through AI-assisted agents; route the other half to standard agents. Measure: resolution time, CSAT, cost-per-ticket, follow-up rate. Measure for 90 days. Then you have real data.

  3. Separate the infrastructure cost from the feature value. The cost of inference is an engineering budget item. The value of "better support" is a business outcome. They live in different accounting categories. Treat them that way.

  4. Question the model assumptions. When a vendor says "This new model has 30% better reasoning," ask: (a) better at what specifically? (b) how much of that improvement matters to my use case? (c) am I paying for capability I don't use? (d) what's the switching cost if I need something different in six months?

  5. Build cost forecasting into capacity planning. If your inference volume grows 2x YoY and your cost grows 8x YoY, that's a problem. Track this slope. Set alerts at 50% budget burn. Don't find out in month five that you've spent the whole year's allocation.

  6. Create an AI ROI scorecard. For every major AI deployment: cost, business metric targeted, baseline value of that metric, measured impact after 90 days, confidence level in the measurement. If you can't fill it out, you shouldn't be spending on that AI.

The Bottom Line

Uber's COO's honest quote is going to echo through earnings calls for the next two years.

Because the same structural problem is everywhere. Marketing teams are deploying AI everywhere (personalization, content generation, attribution, ad copy, recommendation engines) but almost none of them have built the measurement infrastructure to prove ROI.

The costs are real, auditable, and growing. The benefits are theoretical, hidden, and unmeasured. And by the time anyone asks "Is this worth it?", the switching costs are too high. The vendor is too embedded. The infrastructure is too deep.

The question shifts from "Is this worth it?" to "How do I optimize what I'm already committed to?"

That's not a technology problem. That's a business problem. And it starts with numbers that refuse to add up.