AIO Library

Proof Over Promises: Why Evidence Behaves Differently From Claims in AI Recommendation

AI assistants increasingly answer by retrieving documents and grounding responses in them, which makes a checkable claim and an unsupported promise behave differently in ways that are partly documented and partly unknown.

ReferenceAI Optimization2026-08-04

Evidence: supported

Platform documentation from Google and Microsoft confirms that AI answer surfaces retrieve and ground responses in web documents, and multiple independent studies document how often the resulting citations fail verification. What no source establishes is that evidence-bearing content is ranked, retrieved, or cited more often than assertion-bearing content: no operator publishes a weighting, so the practical advantage described here is an inference from documented mechanism, not a measured outcome.

The distinction, stated plainly

A promise is a statement about quality that the reader is asked to accept. A proof is a statement that carries the means of its own checking: a number with a method behind it, a date, a named source, a document someone else can open. Both are text. Both can sit on the same page in the same typeface. The difference only becomes visible when something tries to verify them.

That difference matters more now than it did five years ago, because a growing share of answers about a business, a product, or a topic are assembled by systems that retrieve documents first and write second. Google describes retrieval-augmented generation, which it also calls grounding, as a technique used to improve the quality, accuracy, and freshness of AI responses by relying on its core Search ranking systems to retrieve relevant, up-to-date web pages. Microsoft describes an equivalent step for Copilot, exposing the internal search phrases it calls grounding queries in its Bing Webmaster Tools AI Performance preview, announced 10 February 2026.

AIOFacts does not claim to know how any model weighs what it retrieves. What is publicly observable is narrower and still useful: retrieval happens against documents, and a document that contains a checkable statement gives every downstream check something to land on. A document that contains only an assertion does not.

What the platforms actually document, and what they decline to say

Google's published guidance on generative AI features is unusually direct about the absence of a special lever. It states that structured data isn't required for generative AI search, and there's no special schema.org markup you need to add, while still recommending structured data as part of general practice. Its documentation on AI features states that to be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, and that existing preview controls including nosnippet, data-nosnippet, and max-snippet continue to govern what can appear. The same guidance describes a query fan-out technique, issuing multiple related searches across subtopics.

On the question of what content earns that placement, Google's own emphasis is on substance rather than markup: creating content that people find unique, compelling, and useful will likely influence your website's presence in generative AI search in the long run more than any of the other suggestions, and it names commodity content, described as often based on common knowledge, as the thing that adds little. Google also states that from Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.

That last sentence is worth reading carefully rather than as a settlement of terminology. AIOFacts treats AIO as a broader practice than search visibility because it includes systems that are not search engines at all, and treats GEO and AEO as narrower framings. That is an AIOFacts position, and Google's framing is a different one. Neither is an industry agreement, and the vocabulary in this area remains unsettled.

Citation is not verification, and the gap is measured

The strongest available argument for evidence over assertion is not that evidence is rewarded. It is that the citation layer currently sitting between a document and a reader is demonstrably unreliable, which raises the value of anything that can be independently confirmed.

The Tow Center for Digital Journalism, in research by Klaudia Jaźwińska and Aisvarya Chandrasekar published 6 March 2025, ran 1,600 queries across eight generative search tools, using ten articles from each of twenty news publishers and asking each tool to identify the headline, publisher, date, and URL from a direct excerpt. Collectively the tools answered incorrectly on more than 60 percent of queries. Error rates varied widely by tool, from 37 percent for Perplexity to 94 percent for Grok-3. The finding most relevant here is about confidence rather than accuracy: ChatGPT incorrectly identified 134 articles yet signalled uncertainty only 15 times across 200 responses and never declined to answer.

Later work separates the two things that a citation appears to promise. The 2026 preprint Cited but Not Verified evaluated inline citations from fourteen models and reported link validity above 94 percent and topical relevance above 80 percent for leading systems, while factual accuracy of those same citations ranged from 39 to 77 percent. A separate 2026 preprint on reference hallucinations, working across 53,090 URLs on one benchmark and 168,021 on another, reported that 3 to 13 percent of citation URLs had no record in the Wayback Machine and likely never existed, with 5 to 18 percent non-resolving overall.

Read together, these describe a specific failure: a citation that looks correct, resolves correctly, and is about roughly the right subject, while not actually supporting the sentence attached to it. A reader who follows the link finds nothing. A reader who does not follow it finds nothing wrong.

The honest complication: assertion is not always penalised

A reference article that only assembled supporting evidence would be doing the thing it warns against. There is measured evidence pointing the other way, and it should be stated.

The Phare benchmark, published by Giskard on 30 April 2025, tested how leading models handle misinformation and debunking. It found that when a controversial claim was presented with high confidence, using framings such as claiming certainty or invoking an authority, debunking performance dropped by as much as 15 percent compared with neutral framing. It separately found that system instructions emphasising conciseness produced up to a 20 percent drop in hallucination resistance in extreme cases. The same work noted that models ranking highest on preference-based leaderboards are not necessarily the most resistant to hallucination.

The implication is uncomfortable and worth holding. Confident, unsourced assertion is not reliably filtered out, and in some conditions confident framing measurably reduces the pushback a false claim receives. So the case for proof over promises is not that unsupported claims fail. It is that unsupported claims cannot be confirmed by anything outside the page that makes them, which leaves them dependent on retrieval conditions nobody outside a model operator can see or predict.

The anatomy of a checkable claim

If the practical distinction is checkability, it is worth being specific about what makes a sentence checkable. None of the following is a ranking technique, and AIOFacts makes no claim that any of them causes citation. They are the properties that allow a human reader, a fact-checking pass, or a retrieval system carrying a source document to confirm or refute a statement rather than only repeat it.

  • A named subject rather than a category: which product, which programme, which jurisdiction.
  • A quantity with its method attached: what was measured, over what period, by whom, using what sample.
  • A date on the claim itself, not only on the page, so that a reader can tell whether the statement has expired.
  • An external reference with a resolvable URL, so the claim can be checked somewhere other than the page asserting it.
  • An explicit statement of scope and exception, including where the claim does not hold.
  • A stated evidence status, so that a proposal, an observation, and a confirmed fact are not presented at the same weight.

Identity and corroboration: who is making the claim

A verifiable claim still requires a resolvable claimant. Schema.org documents sameAs as an entirely general purpose disambiguation mechanism that can be used on any entity in any schema.org description, typically pointing to an external reference page that unambiguously identifies the item. That is a published standard and its purpose is not in dispute.

What is in dispute, or more precisely unestablished, is the effect. Practitioner writing frequently asserts that entity identifiers increase how often a site is cited by AI answers. AIOFacts records that as unverified: it conflicts with Google's own statement that no special markup is required, and no operator has published data on the point. The defensible reason to make an entity resolvable is that ambiguity is a real failure mode with observable consequences, including the misattribution documented in the Tow Center work, where tools frequently cited syndicated or republished versions instead of the originating publisher.

The same logic applies to corroboration. A claim that appears only on the site that benefits from it can be retrieved but not cross-checked. A claim that also appears in a filing, a dataset, a standards document, a court record, or independent reporting can be. That is not a distribution tactic. It is the difference between one instance of a statement and a statement with a second, independent instance behind it.

Provenance: proof about the artifact rather than the assertion

A parallel line of work addresses a narrower question: not whether a claim is true, but whether a file is what it says it is and came from where it says it came from. The C2PA technical specification, currently published at version 2.4, defines Content Credentials as tamper-evident, cryptographically signed provenance metadata attached to media assets, describing their origin and edit history.

This is not a solution to claim verification and should not be described as one. Provenance metadata can attest that an image was captured by a particular device and altered in particular ways. It cannot attest that a sentence about market performance is accurate. It is included here because it illustrates the general direction the infrastructure has taken over several years: from asserted authenticity toward attested authenticity, with a machine-readable artifact standing behind the assertion. Whether that direction extends to textual claims at scale is not established, and adoption levels for text specifically are not something AIOFacts can verify from primary sources.

What follows in practice, and what does not

The measurement situation is improving slightly. Microsoft's AI Performance preview reports total citations, average cited pages per day, grounding queries, and page-level citation activity across Copilot, AI-generated summaries in Bing, and select partner integrations. That is a first-party count of citation events on one platform family. It is not a count across all assistants, it does not explain selection, and platform behaviour varies. Anyone reporting a single number for AI visibility is aggregating across systems that publish no comparable metric.

Meanwhile the document-level version of the same problem has a partial engineering answer. Anthropic documents a Citations feature that returns references to the specific passages of a supplied source document used to generate a response, with cited text extracted directly from the provided document. That constrains a model to source material a developer supplied. It says nothing about open web retrieval, and it is not a claim about how any consumer assistant selects public pages.

The practical conclusion AIOFacts is willing to state is modest. Publishing evidence rather than assertion does not guarantee retrieval, citation, or recommendation, and anyone selling that guarantee is selling something no operator has confirmed. What it does is make a claim survivable: checkable by a reader, confirmable against a second source, and correctable when it becomes wrong. Given a citation layer that resolves links correctly while supporting the underlying statement between 39 and 77 percent of the time, being the source that holds up under checking is a durable position even where its retrieval advantage is unproven.

Key points

  • Google documents grounding, also called retrieval-augmented generation, as the mechanism behind its AI features, and states that eligibility requires only that a page be indexed and eligible for a snippet, with no special markup.
  • Tow Center research across 1,600 queries and eight tools found incorrect answers on more than 60 percent of citation queries, with per-tool error rates from 37 to 94 percent.
  • A 2026 evaluation found citation link validity above 94 percent while factual support for the attached claim ran between 39 and 77 percent: a link that resolves is not a claim that checks out.
  • Separate 2026 work found 3 to 13 percent of citation URLs had no record of ever existing, which means fabricated references pass surface-level citation checks.
  • Confident framing is not automatically penalised. The Phare benchmark measured debunking performance dropping by up to 15 percent when a claim was asserted with high confidence, so the case for evidence rests on checkability rather than on assertion being filtered out.
  • Microsoft's Bing Webmaster Tools AI Performance preview is one of the few first-party citation counts available, covering Copilot and Bing AI summaries only. No cross-platform equivalent exists.

What this page cannot establish

  • Whether evidence-bearing content is retrieved or cited more often than assertion-bearing content. No AI platform publishes a selection weighting, and the studies cited here measure citation accuracy, not selection.
  • Whether entity identifiers such as schema.org sameAs links affect inclusion in AI answers. The standard's disambiguation purpose is documented; any effect on citation frequency is not, and Google states no special markup is required.
  • Whether the citation accuracy rates reported in 2025 and 2026 still hold. These systems change frequently and the studies are point-in-time measurements against specific model versions.
  • How the sycophancy effect measured by Phare interacts with retrieval. It was measured on user-supplied framing, and it is not established that the same effect applies to confident assertions found in retrieved documents.

Sources

What supports this page

  1. Google's Guide to Optimizing for Generative AI Features on Google Search
    Google Search Central · platform-documentation · accessed 2026-08-04
  2. AI Features and Your Website
    Google Search Central · platform-documentation · accessed 2026-08-04
  3. Introducing AI Performance in Bing Webmaster Tools (Public Preview)
    Microsoft Bing Webmaster Blog · platform-documentation · accessed 2026-08-04
  4. AI Search Has a Citation Problem
    Tow Center for Digital Journalism, Columbia Journalism Review · expert-analysis · accessed 2026-08-04
  5. Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
    arXiv preprint · expert-analysis · accessed 2026-08-04
  6. Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
    arXiv preprint · expert-analysis · accessed 2026-08-04
  7. Good Answers Are Not Necessarily Factual Answers: An Analysis of Hallucination in Leading LLMs (Phare benchmark)
    Giskard · expert-analysis · accessed 2026-08-04
  8. Citations
    Anthropic · platform-documentation · accessed 2026-08-04
  9. sameAs
    Schema.org · published-standard · accessed 2026-08-04
  10. Content Credentials: C2PA Technical Specification, Version 2.4
    Coalition for Content Provenance and Authenticity (C2PA) · published-standard · accessed 2026-08-04

Questions

Common questions

Does publishing sourced evidence make an AI assistant more likely to cite a page?

That is not established and AIOFacts does not claim it. No AI platform publishes how sources are selected, and the research cited here measures citation accuracy rather than selection. What evidence does reliably provide is checkability: a claim with a date, a method, and an external reference can be confirmed or refuted by a reader or a downstream system, and an unsupported claim cannot.

If AI systems cite sources, does that mean the citation supports the claim?

Frequently not. A 2026 evaluation of fourteen models found link validity above 94 percent and topical relevance above 80 percent, while factual support for the attached statement ranged from 39 to 77 percent. Separate work found that 3 to 13 percent of citation URLs appear never to have existed. A resolving link and a supported claim are different properties and should be checked separately.

Is structured data required to appear in AI answers?

Google states directly that structured data isn't required for generative AI search and that there's no special schema.org markup you need to add, while still recommending it as part of general practice for rich results. Platform behaviour varies and other systems document their own requirements differently. Treat any claim that a particular markup guarantees inclusion in AI answers as unverified.

Why does this article include evidence that weakens its own argument?

Because the Phare benchmark findings on confident framing genuinely complicate the case, and omitting them would make the piece an argument rather than a reference. The honest position is that unsupported assertion is not reliably filtered out by current systems, and that the durable advantage of evidence is its survivability under checking, not a guaranteed retrieval benefit.

How should a claim be marked when its evidence is incomplete?

State the evidence status alongside the claim rather than presenting everything at uniform confidence. AIOFacts uses labels including confirmed, supported, emerging, disputed, unverified, and position, and treats unknown as a valid published output. A claim carrying an accurate status label is more useful than a claim asserted with false certainty, and it is correctable later without retracting the page.

One term, still unsettled, documented in the open.

Read the AIOFacts working definition, versioned and sourced, then see how the terminology is actually used in the wild.

The AIO Definition AIO Truth