AIO Library
Running an AI Visibility Audit
A repeatable observation method for seeing how AI-powered systems currently access, represent, and cite an entity, and an honest account of what such an audit can and cannot establish.
Evidence: supported
The measurement surfaces described here are documented by the platforms that operate them and by published measurement research, but the audit protocol assembled from those surfaces is an AIOFacts proposal, and no audit can observe how any AI system reaches its output.
What an AI visibility audit is, and what it can honestly claim
An AI visibility audit is a structured, repeatable observation of how AI-powered systems currently access, represent, and cite an entity. It is a record of observed outputs and observed access, taken at a stated moment, using a stated method. It is not a diagnosis of how any system works internally, and nothing in this piece should be read as describing the internals of a model. No one outside a model's operator can see those, so an audit is limited to what is publicly observable from the outside plus what the platforms choose to report to you about yourself.
The reason to run one now is that a growing share of discovery is mediated by summarized answers rather than by a list of links, which changes what visibility even means. Pew Research Center, analyzing the browsing activity of roughly 900 U.S. adults who agreed to share it during March 2025, reported that users clicked a traditional search result link on 8% of visits to pages that carried an AI summary, compared with 15% of visits to pages without one, and that only 1% of visits to pages with an AI summary produced a click on a source cited inside that summary. Those figures describe one platform, one country, and one month. They do not describe every AI surface. They are enough, though, to explain why counting sessions alone no longer tells a business whether AI systems can find and represent it.
An audit answers a narrower and more answerable question: given a defined set of questions a real customer might ask, what do these specific systems currently say, what do they currently cite, and can they currently reach the material that would let them answer accurately.
The measurement problem: the same question twice
The first thing an audit must account for is that a single observation proves very little. Ask the same assistant the same question twice and the answer can move. A 2026 preprint by Dmitrij Żatuchin, published on arXiv and not peer reviewed, examines this directly for brand mentions. Working from a fully crossed corpus of 12,933 model responses covering 20 Central and Eastern European brands, 8 query languages, and 3 models, it separates the movement in a measured brand outcome into distinct sources: resampling the same prompt, paraphrasing the prompt, changing the model, and changing the query language. Its reported decomposition attributes a substantial share of variance to plain resampling and a substantial share to query language, which means a single screenshot of a single answer carries little information about a brand.
The practical consequence is a design requirement, not a caveat. If a measurement moves on its own, then a before-and-after comparison built on one observation each cannot distinguish a real change from noise. The same preprint notes that practice has converged on repeating each prompt a small fixed number of times, commonly five, and reports that beyond roughly that point, adding further repeats improves precision far less than adding models or languages does. AIOFacts treats that as a reasonable starting design rather than a settled standard: the useful takeaway is that breadth across systems and phrasings buys more than depth on a single phrasing.
An audit that does not state its sample size, its repeat count, and its date is not repeatable, and an observation that is not repeatable cannot be used to claim improvement later.
First-party reporting: what platforms now tell you about yourself
Two of the largest operators now report some version of AI-surface performance directly to site owners, which makes part of an audit first-party rather than inferred. In June 2026, Google introduced Search Generative AI performance reports in Search Console, described as dedicated views of impressions within generative AI features on Search, such as AI Overviews and AI Mode, as well as generative AI features in Discover. The documented breakdowns are impressions, pages, countries, devices, and dates, with hourly, daily, weekly, and monthly granularity. Google described the reports as rolling out to a subset of sites.
The shape of that report matters as much as its existence. It reports impressions, not clicks, and it does not surface a click-through rate for AI features specifically. That means it can tell you whether your URLs appeared inside these features and roughly how often, and it cannot tell you what happened next. Search Engine Journal reported in February 2026 that Bing Webmaster Tools added an AI performance view showing how often a site's pages were cited in Copilot and AI-generated answers, including page-level citation data and the grounding queries Copilot generates internally when retrieving web content.
Both surfaces should be read as partial by construction. Each covers only its own operator's products. Neither covers assistants that do not publish webmaster reporting at all. An audit that uses them should record which report it read, on what date, and for which property, because the reports themselves are new and their coverage has been changing.
- Record impressions in generative AI features from Search Console, by page and by date, noting that click data for those features is not provided.
- Record citation counts and grounding queries from Bing Webmaster Tools where available.
- Note explicitly which assistants your reporting does not cover, so absence of data is never read as absence of visibility.
Access and crawler behavior: the part you fully control
The most verifiable layer of an audit is whether AI-related crawlers can reach your material at all, because this is measurable from your own logs and configuration rather than inferred from someone else's output. OpenAI documents three distinct crawlers with different roles: GPTBot, which it describes as used to crawl content for training foundation models, OAI-SearchBot, which it describes as used to surface and link to sites in ChatGPT search and not used to collect training data, and ChatGPT-User, which it describes as user-triggered fetching on behalf of someone in a conversation. Because these are separate user agents, a robots.txt policy can treat them differently, and OpenAI notes that changes to robots.txt can take on the order of a day to be reflected in its search behavior.
Google's own optimization guidance for generative AI features states that its generative AI models use publicly accessible, crawlable content, and that ensuring content is crawlable is a prerequisite for appearing in those features. The same documentation describes retrieval-augmented generation, where responses are grounded in retrieved pages, and query fan-out, where a single user query produces a set of concurrent related queries that fetch additional results. Google's documentation is explicit that not every query triggers fan-out.
Cloudflare has published a further measurement worth adopting: the crawl-to-refer ratio, calculated by dividing HTML page requests from a platform's crawler user agents by HTML page requests whose Referer header points at that platform. Whatever the number turns out to be for a given site, computing it from your own logs converts a vague worry into an observation you own. An audit should therefore include a log-derived table of which AI-related user agents requested which paths, how often, and with what response codes, alongside a plain reading of the current robots.txt and any WAF or bot-management rules that may be blocking requests the site owner did not intend to block.
The prompt panel: sampling assistant answers repeatably
The observational core of the audit is a fixed panel of prompts, run on a fixed schedule, with the results logged rather than screenshotted. AIOFacts proposes building the panel from questions with commercial reality behind them, in three groups: direct entity questions such as who a company is and what it does, category questions where the entity might or might not be named, and comparison or suitability questions where a recommendation is actually being requested. Each prompt should exist in at least two or three paraphrases, because paraphrase is one of the documented sources of movement in measured brand outcomes.
For each run, log the system and any visible model version, the date and time, the geography and account state used, the full response text, and every source the response cited with its URL. Citations are the most durable artifact an audit produces, because a cited URL is checkable by anyone later, whereas a summarized claim is not. Repeat each prompt a fixed number of times per run and keep the number constant across runs, since changing the sampling design between runs makes two audits incomparable no matter how careful each one was individually.
Two limits deserve to be written into the report rather than discovered by a reader. Platform behavior varies, and results can differ by account, location, language, and product surface, so a panel run from one account in one country describes that vantage point and not the world. And an assistant naming a business is an observation about that response, not evidence of how any system ranks or evaluates anything.
The readable record: what an AI system could verify about you
The final layer looks at the material itself, on the reasoning that a system can only represent what it can access and reconcile. This part of the audit is a documentation review rather than a prediction. It asks whether the entity is described consistently wherever it appears, whether structured data is present and valid against the Schema.org vocabulary, whether the entity's own site states plainly what it is, where it operates, and who is responsible for it, and whether independent third-party sources exist that say the same things.
Google's guidance for generative AI features is worth quoting for its restraint: it states that from Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and that creating content people find unique, compelling, and useful is likely to influence presence in generative AI search more than other suggestions. That is a platform describing its own products, and AIOFacts records it as such rather than generalizing it to other systems.
The AIOFacts working framework organizes this review into seven signals: Entity Clarity, Trust Signals, Citation Authority, Structured Data, Semantic Consistency, Content Accessibility, and Brand Recognition. This is a proposal for structuring an audit so that findings land somewhere consistent and can be compared across quarters. It is not a description of what any AI system weighs, and it should never be presented as one.
- Validate structured data against the Schema.org vocabulary and record errors rather than summarizing them.
- Compare the entity's name, address, contact details, and description across every place they appear publicly, and list disagreements.
- List the independent sources that describe the entity, with URLs, and note where none exist.
Cadence, scoring, and the conclusions an audit cannot support
A defensible cadence is a full baseline audit, then repeats at a fixed interval, with the method version recorded each time. Quarterly is a reasonable default for most organizations, because the platform surfaces themselves change on roughly that timescale and because a shorter interval mostly measures the variance described earlier. If the method changes, the change should be dated and the prior series should not be silently restated.
A composite score can be useful internally as a way to track your own work, provided it is labeled as your own instrument. It is not a reading of an external reality and it is not comparable to anyone else's number. The most common failure in this category is presenting an internally computed visibility score as though a platform had issued it.
Three conclusions an audit cannot support, however carefully it is run: that a change on the site caused a change in an AI answer, because the observation is uncontrolled and the systems change independently; that a system prefers, rewards, or ranks anything, because that describes internals nobody outside the operator can see; and that results generalize from the assistants sampled to those not sampled. An audit that states its findings and stops short of these is more useful than one that overreaches, because the first can be built on and the second cannot.
Where the evidence runs out
The reporting that exists is new, partial, and asymmetric. Google's generative AI performance reports launched in June 2026 to a subset of sites and report impressions without click data. Bing's AI performance view, as reported in February 2026, reports citations rather than outcomes. Assistants without webmaster reporting are observable only through prompt sampling, which is exactly the layer with the most documented variance. This is not a complete picture and no method described here makes it one.
There is also no established external standard for what an AI visibility audit must contain. The terminology in this field remains unsettled, and different practitioners use AIO, GEO, AEO, and LLMO for overlapping and sometimes different scopes. AIOFacts treats those narrower terms as narrower, which is a position, not a fact. What can be said with confidence is that an audit which records its method, its date, its sample, and its sources can be checked and repeated by someone else, and an audit that does not cannot.
Key points
- A single observation of a single AI answer is not evidence. Published measurement work indicates answers move on resampling and on paraphrase, so a repeatable audit fixes its prompt set, its repeat count, and its schedule before it draws any comparison.
- Two first-party reports now exist: Google's Search Generative AI performance reports in Search Console, introduced June 2026 with impressions but not clicks or CTR, and Bing Webmaster Tools AI performance data covering Copilot citations and grounding queries.
- Crawler access is the most verifiable layer. OpenAI documents GPTBot, OAI-SearchBot, and ChatGPT-User as separate user agents with different purposes, so robots.txt policy and server logs give you evidence you own outright.
- Cloudflare's crawl-to-refer ratio, crawler requests divided by referred requests for the same platform, can be computed from your own logs and turns a general concern into a specific measurement.
- Log cited URLs, not screenshots. A cited URL is checkable by a third party months later, which is what makes an audit auditable.
- State plainly what the audit did not cover. Assistants without webmaster reporting and outside the prompt panel are unmeasured, and unmeasured must never be reported as absent.
What this page cannot establish
- Whether anything found in an audit causes a change in how any AI system represents an entity. The observation is uncontrolled, and the systems change on their own schedule.
- How representative a prompt panel is. There is no published basis for choosing a panel size that generalizes to real user queries across assistants, geographies, and languages.
- What proportion of AI-mediated exposure any first-party report captures. Google's and Microsoft's reports cover their own products only, and other assistants publish no comparable data to site owners.
- Whether the variance structure reported in the arXiv preprint, which studied 20 Central and Eastern European brands across 8 languages and 3 models, holds for other markets, languages, and model families.
Sources
What supports this page
- Introducing Search Generative AI performance reports in Search Console
Google Search Central Blog · platform-documentation · accessed 2026-07-27 - AI Features and Your Website
Google Search Central · platform-documentation · accessed 2026-07-27 - Optimizing your website for generative AI features on Google Search
Google Search Central · platform-documentation · accessed 2026-07-27 - Overview of OpenAI Crawlers
OpenAI · platform-documentation · accessed 2026-07-27 - Google users are less likely to click on links when an AI summary appears in the results
Pew Research Center · dataset · accessed 2026-07-27 - Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers
arXiv (preprint, not peer reviewed) · expert-analysis · accessed 2026-07-27 - The crawl before the fall of referrals: understanding AI's impact on content providers
Cloudflare · dataset · accessed 2026-07-27 - Bing Webmaster Tools Adds AI Citation Performance Data
Search Engine Journal · reporting · accessed 2026-07-27
Questions
Common questions
How often should an AI visibility audit be repeated?
A fixed interval matters more than a short one, and quarterly is a reasonable default. Published work on answer variance indicates that repeated observations of the same prompt differ on their own, so audits run too frequently mostly measure noise. Whatever interval is chosen, the method version and date should be recorded each time so the series stays comparable.
Can an audit tell me why an AI assistant did or did not mention my business?
No. An audit records what a system output on a given date from a given vantage point. The reasoning behind that output is internal to the operator and is not publicly observable. Any explanation offered for a specific mention or omission is an inference, and it should be labeled as one.
Do the new platform reports replace prompt sampling?
They complement it and cover different ground. Google's Search Console generative AI reports and Bing's AI performance data are first-party and precise about their own surfaces, but Google's report provides impressions without click or CTR data for AI features, and neither covers assistants outside its own products. Prompt sampling is the only route to those other systems, and it carries the variance problem described above.
Is blocking AI crawlers a reasonable audit outcome?
It is a decision with a documented trade-off rather than a technical detail. OpenAI documents separate user agents for training, search, and user-triggered fetching, so a site can permit one and refuse another. Google's optimization guidance states that its generative AI features draw on publicly accessible, crawlable content, so a block applied broadly can remove a site from surfaces the owner intended to appear in. The audit's job is to report what is currently blocked and what that currently means, not to make the choice.
What makes an audit repeatable rather than just thorough?
Four recorded things: the exact prompt set including paraphrases, the number of repeats per prompt, the date and vantage point of each run, and the full cited URLs returned. Any audit missing one of these can be read but cannot be reproduced or compared to a later run. Thoroughness without those is a snapshot, not a measurement.
One term, still unsettled, documented in the open.
Read the AIOFacts working definition, versioned and sourced, then see how the terminology is actually used in the wild.