AIO Library

The New Marketing Scoreboard

As answers get assembled before the click, the metrics that survive are the ones platforms actually publish, and most of them measure appearance rather than outcome.

ReferenceAI Optimization2026-08-05

Evidence: emerging

Platform documentation from Google, Microsoft, and OpenAI confirms that appearance and citation reporting now exists for AI surfaces, and independent panel data from Pew Research Center documents reduced clicking when AI summaries appear, but no platform publishes complete click or outcome data for these surfaces, so any full scoreboard remains partially uninstrumented.

Why the old scoreboard stopped reporting the whole game

Digital marketing measurement has rested for roughly two decades on a single chain of events: an impression produces a click, the click produces a session, and the session produces something countable on the destination. Nearly every widely used dashboard, attribution model, and reporting template inherits that chain. Rankings mattered because they predicted clicks. Clicks mattered because they delivered sessions. Sessions mattered because they were where measurement actually happened.

When an answer is assembled on the results page, or inside an assistant, the click becomes optional rather than necessary. The chain does not degrade gracefully at that point. It breaks at the first link, and every metric downstream of it inherits the break.

The clearest public evidence for that break comes from a Pew Research Center analysis published on 22 July 2025. It followed 900 US adults who consented to share browsing data, covering 68,879 Google searches conducted during March 2025, of which 12,593 produced an AI summary. In visits where an AI summary appeared, users clicked a traditional search result in 8 percent of visits. In visits without a summary, they did so in 15 percent. Links inside the summary itself were clicked in 1 percent of visits.

Those are findings from one panel, one platform, one country, and one month. They do not describe how AI systems work, and they do not explain user motivation. What they do establish, with real measurement rather than inference, is that the click can no longer be assumed to be the unit of interest. Google's own documentation offers a different reading of the same surface, stating that when people click from results pages carrying AI Overviews, those clicks are higher quality in the sense that users are more likely to spend more time on the site. Both statements can hold at once: fewer clicks, longer sessions. AIOFacts treats neither as settled, and notes that the two parties measuring have different interests in the answer.

What the platforms now report, and what they withhold

The most consequential development for measurement is not a study. It is that two of the largest search platforms began publishing first-party appearance data for their AI surfaces, and that both stopped short of publishing a complete outcome number.

Google introduced Search Generative AI performance reports in Search Console on 3 June 2026, giving site owners a dedicated view of how often their URLs appeared in generative AI features across Search and Discover. The report covers impressions broken out by page, country, device, and date. It does not include click data. Asked about that omission, a Google spokesperson said the company was continuing to work with website owners to understand what insights would be most helpful and would introduce additional metrics over time. No timeline has been published.

Microsoft moved earlier and along a different axis. AI Performance in Bing Webmaster Tools entered public preview on 10 February 2026, reporting total citations, average cited pages, page-level citation activity, and grounding queries, defined by Microsoft as the key phrases the AI used when retrieving content that was referenced in AI-generated answers. Microsoft states in the same announcement that grounding query data represents a sample of overall citation activity, and that the metrics reflect citation frequency rather than page importance, ranking, or placement. A June 2026 update added Intents, Topics, Citation Share, and Compare, where Citation Share is defined as the percentage of citations attributed to a site out of all citations shown across all sites for the same grounding query.

OpenAI documents its crawlers but not its publisher-facing outcomes. Its crawler overview separates OAI-SearchBot, used to surface websites in ChatGPT's search features, from GPTBot, used to crawl content that may be used in training foundation models, from ChatGPT-User, which handles user-initiated visits and is documented as not crawling the web automatically. That separation is operationally useful, because it lets a publisher distinguish retrieval access from training access in server logs. It is not a reporting surface, and OpenAI publishes no equivalent of the Bing citation dashboard.

The asymmetry is itself the finding. Appearance is now reported by more than one major platform. Outcome is reported by none of them for these surfaces. A scoreboard built in 2026 is built on partial instrumentation by design, not by oversight, and any measurement practice that pretends otherwise is describing a system it cannot see.

The crawl-to-refer ratio, and what it does not measure

One network-scale metric entered the vocabulary quickly and is frequently misread. Cloudflare published the crawl-to-refer ratio on 1 July 2025, authored by David Belson and Sam Rhea. It is calculated as the total number of HTML content requests from user agents associated with a given search or AI platform, divided by the total number of HTML requests whose Referer header contained a hostname associated with that platform.

In the week of 19 to 26 June 2025, Cloudflare reported an Anthropic ratio of 70,900 to 1 and a Mistral ratio of 0.1 to 1, the latter reflecting roughly ten times more referrals than crawls. Cloudflare states its own limitation directly: traffic referred by Claude's native app does not include a Referer header, the same is suspected of other native applications, and therefore the calculations may overstate the respective ratios by an amount that is unclear.

The ratio measures the terms of exchange between a platform and the open web at network scale, in a specific week, using a signal that native applications are known to omit. It does not measure how any individual entity is represented, whether that representation is accurate, or whether the exchange is commercially good or bad for a given business. A high ratio is not evidence that a platform describes you poorly. A low ratio is not evidence that it describes you well. Quoting either as a marketing performance figure moves a network statistic into a job it was never built to do.

Four things that can actually be counted

Stripping out what cannot be evidenced leaves a short list. Each item below has real instrumentation behind it, and each has a stated ceiling on what it can tell you.

The list is ordered by how strongly the measurement can be supported, not by how interesting the number looks in a report.

  • Retrievability. Whether the systems that fetch content for answers can reach it, parse it, and render it. This is measured in your own server logs against documented user agents, and it is the only item on this list you control outright and can verify without a third party.
  • Appearance and citation frequency. How often URLs surface in generative features. Reported by Google as impressions with no click data, and by Bing as citations and citation share with an explicit sampling caveat. These are platform-reported numbers, which makes them the strongest external evidence available and also ties their definitions to platform decisions you do not control.
  • Referral behaviour and what follows it. Sessions that arrive with an assistant or AI surface as the referrer, and what those sessions do. First-party, but structurally incomplete: Cloudflare documents that native application traffic can arrive without a Referer header at all, so an absent referral is not proof of an absent referral.
  • Representation accuracy. Whether the description a system gives of an entity is correct, current, and complete. No platform reports this. Every figure in this category is a sample statistic produced by whoever measured it.

Representation accuracy has no dashboard

The fourth item is the one most businesses actually care about, and it is the one with the least infrastructure behind it. The practical question is not how often a page was cited. It is whether the account a system gives of a company, its products, its pricing, its service area, and its credentials matches reality.

Generated answers are not stable objects. The same question can return different text across sessions, phrasings, regions, devices, model versions, and account states. That is not a defect to be engineered around, it is a property of the systems, and it has a direct methodological consequence: any representation accuracy figure is a sample statistic and inherits every property of its sampling design. The prompt set, the number of repetitions, the region, the date, the platform, the model version where it is disclosed, and whether personalisation or memory was active all change the result.

AIOFacts position: a representation figure quoted without its method is not a measurement, it is an assertion with a decimal point attached. Where the method cannot be published, the honest output is the observation and its date, not a score. Unknown is a valid finding. Unknown is never the same as zero, and an unmeasured attribute must never be rendered as a failing one.

Volatility, versions, and why numbers should never be overwritten

Platform behaviour varies and changes without notice, and the reporting surfaces themselves are moving. Bing expanded its AI Performance report from the February 2026 preview to a substantially different set of capabilities by 16 June 2026, and Microsoft states plainly in that update that citation patterns can shift due to changes in user behaviour, evolving models, freshness signals, partner refresh cycles, and broader changes across the web itself. Google's report launched with a regional rollout and an open commitment to add metrics later.

Two practices follow, and they are the difference between a measurement record and a rolling guess. First, every recorded figure carries its date, its platform, and the version of whatever engine produced it. Second, a re-measurement is a new observation, appended, never an update that overwrites the old one. A score that is silently replaced destroys the only thing that made it useful, which is the ability to see movement and to know whether the movement was in the world or in the instrument.

Cross-platform comparability is weak and should be stated as such. An impression in Google's generative AI report and a citation in Bing's AI Performance report are not the same event, are not counted the same way, and are not defined by the same party. They can be tracked side by side. They must never be summed into a single total, and a composite built by averaging them is arithmetic performed on incompatible units.

What a defensible scoreboard looks like

This section is an AIOFacts proposal, not a description of industry practice and not a standard. It is offered because the alternative currently on offer, a single opaque visibility score, fails the first question anyone should ask of a number, which is how it was produced.

A defensible scoreboard has five properties. It is sourced, meaning every figure traces to a platform report or a first-party log. It is dated, because a figure without a date describes a moment nobody can identify. It is versioned, so a change in the measuring instrument is distinguishable from a change in the world. It is bounded, meaning the claim attached to each number does not exceed what the number can support. And it is reproducible, meaning someone else following the stated method would arrive at a comparable result within a stated variance.

Ordering the four countable items by evidential strength gives the reporting hierarchy: platform-reported appearance data first, because it comes from the party that observed the event; first-party server logs and referral data second, because you observed them yourself and can audit the collection; sampled representation testing third, and only where the method is published alongside the result; third-party visibility scores last, and only where their methodology is disclosed in enough detail to be criticised.

The failure mode worth naming is the composite. A single number blending platform impressions, sampled prompt tests, and a proprietary weighting is attractive precisely because it hides the joins. It will move, and nobody will be able to say why.

What the evidence does not support

Three conclusions circulate widely and are not carried by the material cited here. The first is that clicks are obsolete. The Pew data documents a reduction in one context, not an elimination, and Google's documentation makes a competing claim about the quality of the clicks that remain. Reduced is not gone, and a metric under pressure still needs measuring.

The second is that citations function as a replacement currency for links, with the accumulation logic transferred intact. Microsoft states directly that its citation metrics reflect citation frequency rather than page importance, ranking, or placement. That is the platform reporting the number telling users what it does not mean, and it is worth taking at face value.

The third is any account of internal weighting: which factors a system prioritises, how it decides between sources, or what it rewards. Nobody outside a model's operator can observe that, the operators have not published it, and a scoreboard that claims to measure it is measuring something else and mislabelling the result. Everything reliable in this piece describes observable inputs and platform-reported outputs. The mechanism between them remains closed.

Key points

  • Pew Research Center measured 8 percent clicking on a traditional result when an AI summary appeared against 15 percent when none did, across 68,879 Google searches by 900 US adults in March 2025.
  • Google's Search Console generative AI performance report, launched 3 June 2026, provides impressions by page, country, device, and date, and explicitly excludes click data.
  • Bing Webmaster Tools AI Performance, in public preview since 10 February 2026, reports citations, grounding queries, and citation share, with Microsoft stating the grounding data is a sample and that the metrics reflect frequency, not importance or ranking.
  • Cloudflare's crawl-to-refer ratio measures the terms of exchange at network scale in a given week, and Cloudflare itself warns the figures may overstate the ratios because native app referrals carry no Referer header.
  • An impression in Google's report and a citation in Bing's are different events counted by different parties, and must never be summed into a combined score.
  • Every recorded figure should carry its date, platform, and engine version, and a re-measurement should be appended as a new observation rather than overwriting the previous one.

What this page cannot establish

  • No platform publishes click or conversion data for its AI surfaces, so the commercial outcome of an impression or citation cannot currently be measured from platform reporting.
  • The true volume of referral traffic from assistants is not knowable from Referer headers alone, because Cloudflare documents that native application traffic can arrive without one, and the size of that gap is unquantified.
  • Whether the Pew findings from March 2025 still describe current behaviour is unestablished: the platforms, the interfaces, and the link presentation have all changed since, and no equivalent panel study covering the current period has been identified here.
  • There is no published, independently validated method for measuring representation accuracy across AI systems, so figures in that category cannot be compared between vendors.

Sources

What supports this page

  1. Google users are less likely to click on links when an AI summary appears in the results
    Pew Research Center · dataset · accessed 2026-08-05
  2. The crawl before the fall of referrals: understanding AI's impact on content providers
    Cloudflare · dataset · accessed 2026-08-05
  3. Introducing Search Generative AI performance reports in Search Console
    Google Search Central · platform-documentation · accessed 2026-08-05
  4. AI Features and Your Website
    Google Search Central · platform-documentation · accessed 2026-08-05
  5. Introducing AI Performance in Bing Webmaster Tools Public Preview
    Microsoft Bing Webmaster Blog · platform-documentation · accessed 2026-08-05
  6. New AI Visibility Insights in Bing Webmaster Tools: Intents, Topics, Citation Share, Compare
    Microsoft Bing Search Blog · platform-documentation · accessed 2026-08-05
  7. Overview of OpenAI Crawlers
    OpenAI · platform-documentation · accessed 2026-08-05
  8. Google Search Console AI performance reports and controls to block your content in AI responses
    Search Engine Land · reporting · accessed 2026-08-05

Questions

Common questions

Should traffic and rankings be removed from reporting?

The available evidence does not support removing them. Pew documented a reduction in clicking when AI summaries appeared, not an absence of clicking, and Google's documentation argues that the remaining clicks may involve longer time on site. The defensible change is to stop treating traffic as the complete picture and to add appearance and retrievability measures alongside it, each labelled with what it can and cannot show.

Which single metric best measures AI visibility?

There is no single metric that can carry that weight, and any product offering one is combining incompatible units. Platform-reported appearance data, first-party crawler and referral logs, and sampled representation testing measure different things with different reliability. AIOFacts proposes reporting them separately with their sources and dates rather than blending them into a composite whose movements cannot be explained.

Is citation count a reliable measure of authority?

Microsoft addresses this directly in its own documentation, stating that the AI Performance metrics reflect citation frequency and not page importance, ranking, or placement. Citation counts are useful as evidence that content is being retrieved and referenced on a given platform. Treating them as a transferable authority score goes beyond what the platform reporting them says they mean.

Why does a measurement need a version number attached?

Because both the platforms and the measurement tools change, sometimes within months. Bing revised its AI reporting substantially between February and June 2026, and Microsoft notes that citation patterns shift with evolving models and refresh cycles. Without a recorded version and date, a change in a score cannot be distinguished from a change in the instrument that produced it, which makes any trend line unreadable.

What can be measured without relying on any platform report?

Retrievability, using your own server logs against documented user agents such as OAI-SearchBot, GPTBot, and ChatGPT-User, whose separate purposes OpenAI publishes. This tells you whether retrieval systems can reach and parse your content, which is a precondition for everything else on the scoreboard. It is the one measurement you fully own and can audit end to end.

One term, still unsettled, documented in the open.

Read the AIOFacts working definition, versioned and sourced, then see how the terminology is actually used in the wild.

The AIO Definition AIO Truth