Skip to main content
AI Advertising Measurement Is Mostly Guesswork Right Now
August 3, 2026·8 min read

AI Advertising Measurement Is Mostly Guesswork Right Now

AI advertising measurement is moving from clicks to brand perception, but teams still cannot prove whether AI visibility changes sentiment, demand, or revenue.

DS
Dellon S.

Digital Marketing

AI AdvertisingMarketing MeasurementBrand StrategyGenerative AI

AI advertising measurement still feels like guesswork

AI advertising measurement has a new problem: the ad may be doing its job before anyone clicks, and the brand may be changing inside an answer nobody on the media team can fully inspect.

That is the logic behind Japan Airlines working with Jellyfish to study how its brand appears across AI models, as Ad Age reports. The airline reportedly shows up strongly in Google Gemini and more weakly in ChatGPT, then plans to track whether advertising changes that gap and shifts sentiment. It is a smart test. It is also a warning that the old dashboard is about to become even less useful.

A marketing analyst studies ambiguous AI brand signals on a dark wall of data.

The metric is moving upstream

For years, digital advertising measurement worked backward from an action. An impression became a visit, a visit became a conversion, and the conversion could be assigned to a campaign with varying degrees of confidence.

AI systems interrupt that chain. A person can ask an assistant which airline is best for a trip, receive a recommendation, and never visit the airline's website until the booking is already decided. The meaningful exposure happened inside a generated answer. The visible click, if it arrives at all, is late evidence.

That is why the idea of share of model matters. It asks a different question from share of search. Not “How often do we rank?” but “When a relevant person asks a relevant question, how often does the model mention us, describe us accurately, and place us in the right competitive frame?”

The problem is that model visibility is not a stable impression. The answer can change with wording, location, account history, current events, retrieval sources, and the model itself. A brand can gain mentions while losing the qualities that made those mentions valuable.

The industry has already seen a version of this measurement confusion in generative search. My earlier piece on the AI search measurement crisis made the basic point: a citation count can rise while commercial visibility quietly falls. AI advertising adds paid influence to the same fog.

A marketer compares sentiment notes, campaign results, and abstract conversational AI results at a dark desk.

Exposure is not persuasion

Japan Airlines' experiment is interesting because it treats model presence as something that can be changed, observed, and connected to sentiment. Most brand teams will want the same thing. They will ask whether paid exposure in an AI environment makes people more likely to trust, prefer, or buy from the brand.

Those are not the same measurement.

A model can mention a company more often because its campaign generated fresh content. That may improve recall without improving preference. It can also improve preference for the wrong reason, such as a temporary news cycle or a highly repeated claim that does not survive human scrutiny.

The reverse is just as uncomfortable. A brand can become less visible in an assistant's answer while its actual business remains healthy. The model may simply be favoring a different source set, a new product comparison format, or a competitor that has published more machine-readable material.

This is where marketers need to separate four layers:

  • Presence: Did the model mention the brand?
  • Framing: What did the model say the brand is good at?
  • Confidence: Did the answer sound certain, qualified, or skeptical?
  • Action: Did the exposure change a measurable business behavior?

Most early “AI visibility” reports collapse all four into a single score. That is convenient, and mostly useless. A brand that appears in 40 percent of answers but is described as expensive, unreliable, or irrelevant has not won the metric.

The more honest dashboard would show the answer text, the prompt family, the competing brands, the source citations, and the change over time. It would also show uncertainty. If a platform cannot reproduce the result across a controlled prompt set, the number is a signal, not a KPI.

Paid influence makes the fog thicker

The arrival of advertising inside conversational products raises the stakes because paid and organic influence can sit next to each other in the same answer journey. A user may see an ad, ask a follow-up question, and then receive an answer shaped by the information available to the system. The journey feels organic even when the first nudge was not.

That creates a measurement problem familiar to anyone who has worked in brand media, but with fewer clean edges. Traditional view-through attribution was already controversial because it treated exposure as evidence of causation. In an AI interface, the exposure can alter the user's next question, which alters the answer, which alters the eventual action. The path is interactive instead of linear.

The industry response will probably be another layer of modeled attribution. More panels, more lift studies, more “influence scores.” Some of that will be useful. A lot of it will be beautifully presented speculation.

The useful test is simpler: can the team define an intervention and compare it with a credible control? That fits Google's people-first content guidance, which favors useful evidence over polished claims. If paid activity changes how a model describes a brand, do similar prompts in an untreated market change less? If sentiment improves, does branded search, direct traffic, qualified leads, or bookings move afterward? If nothing downstream moves, the team should be willing to call the result awareness rather than performance.

That discipline matters because AI platforms are eager to sell the idea that every answer is a measurable funnel stage. It is not. Sometimes an answer is just an answer, and sometimes it is the beginning of a decision that disappears into a private conversation.

A candid phone photo of a small marketing team reviewing an AI campaign late at night.

The old dashboard is losing context

Clicks had weaknesses, but they came with useful context. Teams could see the ad, the landing page, the audience, the timestamp, and the eventual event. They could argue about attribution while looking at roughly the same evidence.

AI exposure breaks that shared view. The user sees one answer. The brand sees a sampled output. The platform sees a much larger interaction graph. The measurement vendor sees whatever it can collect through prompts, panels, browser data, or integrations.

Those are four different realities.

The result is a temptation to treat model output as a media placement. That framing is too narrow. AI answers are closer to a mixture of media, recommendation, customer service, and reputation. The correct owner may not be the media team. It may be brand, SEO, product marketing, communications, and customer experience working from the same evidence.

That organizational problem is already visible in the spread of AI agents across marketing operations. In the entry-level marketing squeeze, I argued that automation changes who gets to touch the work. Measurement is next. A junior analyst may not be asked to build a report from scratch if the platform can generate one, but someone still has to decide whether the report describes reality.

The answer is not to abandon measurement. It is to measure the parts that can survive inspection:

  • Run a fixed prompt library and publish the exact prompts internally.
  • Track mention, framing, accuracy, sentiment, and citations separately.
  • Compare model outputs with brand search, direct demand, qualified pipeline, and customer research.
  • Record model, region, language, date, and account conditions for every observation.
  • Use holdouts or matched markets before calling paid exposure causal.

That list sounds less glamorous than a single AI visibility score. Good. The job is not to make uncertainty look expensive. The job is to reduce it.

A candid smartphone photo of a strategist checking how an AI assistant describes a brand in a coffee shop.

What Japan Airlines is really testing

The visible story is that an airline wants stronger representation in ChatGPT. The deeper story is that brands are beginning to treat model behavior as part of their public reputation.

That is a meaningful shift. A bad product review can be answered. A weak search result can be improved. A model's compressed description of a brand may be repeated thousands of times without a clear place to correct it.

Japan Airlines is not just buying exposure. It is testing whether paid communication can influence the information environment that shapes future answers. That is much closer to reputation management than standard media buying.

It also exposes the limit of the current vocabulary. “AI advertising” makes the work sound like a new placement. “Model reputation” is closer to the actual problem, but it forces teams to admit that they do not control the interface, the answer, or the user's next question.

The brands that handle this well will not chase every mention. They will decide what they want models to understand, build credible evidence around those claims, and measure whether the understanding holds across prompts and platforms. They will also keep a human review loop, because a fluent wrong answer is still wrong.

My recent analysis of vendor lock-in in agentic marketing made a related argument: the most dangerous dependency is not a tool that stops working. It is a tool that keeps working while quietly changing the terms of the decision.

AI advertising measurement has that same risk. The reports will get cleaner before the evidence gets better.

The honest number is often a range

There is no reason to wait for perfect measurement. Brands can start with controlled prompts, documented baselines, and a clear distinction between exposure and impact. They can ask whether the model is accurate before celebrating that it is favorable.

But they should resist the pressure to turn every model mention into a precise return-on-ad-spend calculation. That is not sophistication. It is old attribution anxiety wearing a new interface.

The next generation of brand measurement will probably combine model audits, customer research, search behavior, and incrementality testing. It may even produce useful numbers. For now, the honest result is often a range, a confidence level, and a transcript that lets another person disagree with you.

That may feel less impressive than a dashboard. It will be much more valuable when the CMO asks the only question that matters: did the market change, or did the report?