The first time a brand team sees an AI visibility report, someone usually asks the wrong question: “What is our score?”
The better question is harder. What does the model say about us, when it has to choose?
That distinction matters because AI search visibility is no longer a fringe SEO concern. People are asking ChatGPT, Gemini, Google AI Overviews, and shopping agents to compare products, explain categories, and recommend what to buy. The answer is often the new shelf, the new shortlist, and sometimes the entire decision.
On August 3, the Interactive Advertising Bureau released Measuring Visibility in the AI Era, its first attempt to give the market a shared measurement language. The timing is right. The market has more than 20 vendors selling AI visibility tools, according to the IAB, yet those tools can return different answers for the same brand because their prompts, models, sampling, and scoring systems differ.
That is not a small reporting problem. It means a marketing team can spend a quarter optimizing for a number that another provider would calculate differently.
The shift is also showing up in Google’s own guidance for optimizing content for generative AI search, which keeps the emphasis on useful, crawlable, people-first content rather than a secret AI ranking trick. Marketing Dive’s coverage makes the commercial pressure plain: brands that cannot track AI discovery will struggle to tell whether their visibility is real or merely reported.

AI search visibility is not one number
The IAB’s framework organizes visibility around four measures: presence, prominence, portrayal, and persuasion. The vocabulary is useful because it separates four things that most dashboards flatten into one percentage.
Presence asks whether the brand appears at all. Prominence asks where it appears and how much attention it receives. Portrayal asks whether the answer describes the brand accurately and in the way the business wants. Persuasion asks whether the answer gives a person a reason to keep considering it.
A brand can score well on the first measure and fail the other three.
It can be mentioned in a category answer, buried under six competitors, described with an outdated product detail, and never recommended. The dashboard says visible. The buyer experiences irrelevance.
That is why the familiar language of rankings starts to break down. A Google result has a position. An AI answer has a narrative. One brand may be cited first in one response, recommended second in another, and omitted entirely when the user adds a constraint such as budget, location, or safety.
My earlier analysis of the AI search measurement crisis made the same point from a different angle: if teams only count mentions, they are measuring exposure without meaning.
The real metric is decision influence
Marketing measurement has always preferred what is easy to count. Impressions are easier than memory. Clicks are easier than trust. A rank is easier than a buyer’s changing mental model.
AI search makes that weakness visible because the interface compresses research and recommendation into one response. The user may never visit the brand site. They may not even remember which source supplied the fact. They simply leave with a shortlist that feels reasonable.
That means the most valuable question is not “How often did the model mention us?” It is “Did the model move the user toward or away from us?”
A practical measurement stack should therefore track at least four layers:
- Eligibility: Does the brand appear for the prompts that define the category?
- Context: Is it surfaced for the audiences, use cases, and constraints that matter?
- Accuracy: Are the facts, prices, availability, and differentiators correct?
- Choice: Does the answer make the brand easier to select than its alternatives?
The first layer is visibility. The last layer is value.
This is also where the distinction between directional and decision-grade measurement becomes important. The IAB framework says providers should disclose the quality and rigor of their data, including stability and reproducibility. A report based on a small set of prompts can still be useful, but it should not be dressed up as a precise market share number.
The honest label might be “directional signal.” That is less exciting than a 78.4 visibility score, but far more useful to a CMO deciding where to put budget.
The answer can be visible and wrong
A brand’s most dangerous AI-search problem may not be invisibility. It may be a confident, persuasive error.
A model can misstate a product’s ingredients, confuse two similarly named companies, repeat an old review, or describe a service as available in a market where it is not. The response can still look polished. It can still cite sources. It can still send a customer to a competitor.

That is why portrayal deserves its own operational owner. SEO teams can improve crawlability and source coverage, but they cannot treat factual accuracy as a metadata task. Product marketing, customer support, legal, and local operators all have a role in keeping the public record legible.
The work is not glamorous. It looks like checking whether the same product name means the same thing across pages, fixing stale comparison copy, clarifying service areas, and publishing evidence that a machine can understand without guessing.
The brands that win here will not necessarily be the ones producing the most AI-written content. They will be the ones with the cleanest, most consistent, most defensible body of information for models to retrieve and combine.
That connects directly to the brand narrative problem created by LLM hallucinations. If the model is filling gaps in your story, it is not neutral. It is making editorial decisions on your behalf.
Prominence is a strategic choice
Most teams will want to optimize for presence first because omission feels like the emergency. It is not always the emergency.
A brand can be present in the wrong conversations. A luxury product that appears in every “cheapest option” answer may be technically visible while quietly damaging its positioning. A regulated company that gets described as a casual alternative may gain impressions and lose trust.
Prominence should be evaluated against the brand’s actual strategy. The goal is not maximum mention volume. The goal is the right kind of mention in the moments where the brand can credibly win.
That requires prompt portfolios built around decisions, not keywords. Test questions should include category entry points, comparison questions, objections, use cases, local availability, replacement scenarios, and post-purchase support. Then segment results by audience and intent.
A weekly prompt like “What are the best project management tools?” tells you very little. A portfolio that includes “What should a five-person agency use if client approvals are the bottleneck?” tells you whether the brand is understood in a useful context.
The difference is the same one I wrote about in agentic shopping and traditional marketing. Once the interface starts making the shortlist, broad awareness becomes less valuable than being legible at the point of selection.
Build an evidence loop, not a scorecard
AI visibility should live in a loop with four steps.
Listen. Run a stable set of prompts across the models and interfaces that your customers actually use. Keep the prompts versioned. Do not quietly change the test every week and call the result a trend.
Diagnose. Separate presence, prominence, portrayal, and persuasion. Record the exact answer, citations, competitor mentions, and factual errors. A screenshot is often more valuable than a score.
Repair. Fix the source material first. Improve product pages, comparison pages, documentation, profiles, retailer data, reviews, and third-party references. Do not begin with a flood of generic articles written for a machine.
Re-test. Measure stability over time and across repeated runs. If a brand disappears after a small wording change, that is not necessarily a failure. It may reveal that the category association is weak or the evidence base is thin.

This loop also gives agencies a better client conversation. Instead of promising to “increase AI visibility,” they can explain which layer is broken and what evidence would improve it. That is a much more adult service than selling a magic GEO score.
The uncomfortable new standard
The IAB framework is not a perfect answer. It is a starting point, and the industry will argue over definitions, sampling, and whether persuasion can ever be measured reliably from the outside.
That argument is healthy. Measurement should be challenged before it becomes budget allocation.
The more immediate mistake would be waiting for a universal score. There will not be one clean number that makes AI discovery behave like a search results page. The interfaces are different, the models change, the prompts are personal, and the answer itself is part of the product experience.
Brands need a better standard anyway. They need to know whether they are present, whether they are understood, whether they are represented accurately, and whether the answer gives a buyer a credible reason to choose them.
That is the uncomfortable shift. AI search visibility is not a ranking metric. It is a test of whether the market can explain your value when you are not in the room.
