← Back to blog

The IAB's new AI visibility framework sets a measurement standard no one yet audits

The IAB's new AI visibility framework sets a measurement standard no one yet audits

On August 3rd the IAB published Measuring Visibility in the AI Era, a 36-page framework that is, as far as anyone can tell, the first attempt to standardize how a brand's presence inside the answers built by ChatGPT, Google Gemini, Microsoft Copilot, Perplexity and Claude gets measured. That distinction comes first because it changes everything that follows: this measures organic visibility, not advertising placed inside those answers. On ChatGPT ads we already wrote before, back when OpenAI's own reporting was the only read available. This note is about the other problem: how to know whether your brand shows up, where it shows up, and whether any of that moves something, when nobody yet audits the people selling you the number.

From showing up to being chosen: the chain the framework proposes

The framework organizes visibility inside AI answers into four causal layers, the 4 P's: Presence, Prominence, Portrayal and Persuasion. Presence measures how often the brand appears across a set of relevant prompts. Prominence measures where it appears within the answer, whether it's the first source cited or a passing mention. Portrayal measures how the brand gets described, how accurately and with what sentiment. Persuasion measures whether that citation pushes traffic or an action. The chain breaks at any link: showing up often doesn't matter if the description is wrong, and an accurate description doesn't guarantee a click either. Earning a good portrayal is ground we already covered in the note on Answer Engine Marketing: optimizing to get cited is a different job from optimizing to get clicked.

Exploratory, directional, decision-grade: where the 50-query floor fits

The framework distinguishes three tiers of measurement quality, not two. Below 50 queries per program, the result is exploratory. Past that threshold it becomes directional: useful for spotting a trend, whether the brand is gaining or losing presence month over month, but the IAB itself rules it out for budget decisions. The third tier, decision-grade, additionally requires meeting nine documented dimensions, the list worth bringing into a meeting with a vendor:

  • Query volume
  • Sample size
  • Prompt-type coverage
  • Testing cadence
  • Reproducibility
  • Data validation
  • Methodology documentation
  • Platform coverage
  • Cross-platform aggregation

The 50-query threshold applies per program, not per market. A multi-market program aiming for decision-grade status needs separate measurement per market with adapted queries, but the document doesn't quantify a minimum for each one. The most common mistake is treating an exploratory or directional measurement as if it were solid enough to move budget, the same logic we worked through in the note on how PostHog turns a release into a measurable question: a metric can be real and still fall short of decision-grade.

No certification means the burden of proof sits with the buyer

The framework certifies nothing yet. The document states, literally, that it "establishes no certification and evaluates no individual provider." Meeting the nine criteria is, for now, a claim the vendor makes about itself: no third party audits it. The document presents itself as the foundation for a future IAB certification program, one that would formalize the three tiers and verify claims independently, but that program doesn't exist yet. It's a possibility, not an announcement.

Faced with that lack of auditing, the framework offers what is probably its strongest line: where a provider cannot or will not disclose against a required item, that absence itself is to be treated as a signal. The standard doesn't chase vendors. It shifts the burden of proof onto the buyer: the question a provider dodges is already, in the document's own terms, part of the answer.

Before August 3rd, a CMO shown an "AI visibility score" wasn't short on the ability to ask what was behind it. What was missing was a shared standard for judging the answer. That bar exists now, even if it still depends on the provider using it honestly.

The underlying problem is concrete: only 16% of brands measure their AI visibility systematically today, and there are already more than 20 companies selling tools to measure it, each with a different methodology and, in most cases, an undisclosed one. Two vendors can audit the same brand in the same week and return numbers that don't match, with nobody able to explain why. It's also worth saying that the IAB is funded by dues from its member companies, and some of them sell the very measurement tools the framework is now asking to be transparent. That doesn't invalidate the technical criteria, but it's worth keeping in mind while reading them.

Adapt the queries to the market, don't translate them: the part left for us to read between the lines

The framework asks that a program covering multiple markets adapt its query design to each market's language and cultural frame; that's the document's own wording. In our reading, that rules out simply translating a set built in English: adapting implies redesigning, even though the framework doesn't put it in those exact terms. This matters for anyone measuring brand across more than one country in the region: a vendor demonstrating that its methodology reaches decision-grade in the United States says nothing about whether it reaches that bar in Mexico, and even less about whether that proof was replicated in Portuguese for Brazil. The document also doesn't quantify a minimum number of queries per market, a gap worth flagging to any regional vendor. We recommend asking for concrete evidence that the Spanish-language queries were designed for that market, not adapted from a set built for another language.

This matters most for teams already spending budget against an AI visibility score, or about to start: there, the nine criteria are the list to bring before signing anything. For a brand that doesn't yet show up on the radar of these answers, the framework matters less than understanding first how much of its category's search volume has already moved from Google to an AI assistant. On that, we already wrote when we quantified, using the Bocconi study, how much AI search erodes traditional search: it's the data point that explains why visibility inside an answer is starting to carry as much weight as ranking in a search result.

The framework still doesn't solve the problem that prompted this note: there's no certification and no external auditor, and the document itself admits as much. What does change is what to do in the meantime. With no third party verifying anything, the only real lever is the one the framework itself offers: ask for the nine criteria in writing and treat any silence as part of the answer. A vendor that answers all nine points can still be wrong, but one that can't answer them has already said, without saying it, most of what needed to be known.

Sources