AIO Library
Case Studies as Machine-Readable Evidence
A case study written for persuasion and a case study written to be checked are different documents, and the second one is the one a retrieval system can quote without inventing the qualifiers.
Evidence: position
The vocabularies, crawler roles, and structured data policies described here are documented by schema.org, W3C, Google, OpenAI, and Anthropic; the claim that structuring case study outcomes this way improves how AI systems represent them is an AIOFacts position, and Google's own guidance states that no new markup is required to appear in its AI features.
Why the case study is a hard document for a machine
A case study is the most evidence-dense page most organizations publish, and it is usually the least structured. It contains a named subject, a starting condition, an intervention, a time period, a measured outcome, and an attribution of cause. Every one of those is a distinct factual component with its own reliability. In practice they arrive fused into prose that was written to persuade a reader, where the number is a headline, the time period is implied, the baseline is absent, and the causal claim is carried by the word 'after'.
That was survivable when the document's only job was to be read by a person who could see the whole page and apply their own discount rate to it. It is a different problem when passages are retrieved and quoted in isolation. A sentence lifted out of a case study carries whatever qualifiers were inside that sentence and none of the qualifiers that were three paragraphs away. The qualifier that stayed behind is the part that made the claim honest.
This piece is about a narrow question: what can be done to an outcome claim, at the level of text and markup, so that a system fetching the page can identify what was measured, about whom, over what period, by what method, and who is asserting it. It is not about making claims more likely to be repeated. It is about making them checkable when they are.
The documented layer: crawlers, retrieval, and sentence-level citation
Very little is publicly known about how any AI system chooses what to cite. What is documented is the plumbing around that decision, and the plumbing is worth understanding before anything else.
OpenAI documents three separate user agents with three separate jobs: GPTBot, described as collecting content that could contribute to training its models; OAI-SearchBot, described as surfacing websites in ChatGPT's search features, where OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers; and ChatGPT-User, described as handling user-initiated fetches, which OpenAI notes is not used to crawl the web automatically. Anthropic documents a parallel split: ClaudeBot for content that may contribute to model training, Claude-User for retrieval prompted by an individual's query, and Claude-SearchBot, described as navigating the web to improve search result quality. Each can be addressed separately in robots.txt.
The consequence for evidence is structural rather than clever. Access and citation are governed by different agents, so a page can be excluded from a search surface while remaining reachable in other ways, and the decision about which agents may fetch a case study is a decision about whether its evidence can be retrieved at all.
Anthropic's Citations documentation describes a mechanism at the other end of the pipe: when documents are supplied with citations enabled, Claude can cite individual sentences, chaining consecutive sentences for longer passages, with PDF text extracted and chunked into sentences. That is documentation of one product's behavior, not a general law. But it names the unit that matters in this context. If the citable unit can be a sentence, then a sentence is the smallest thing that has to be true on its own.
What platform documentation does not claim about markup
It would be convenient to say that structured data causes citation. The available platform documentation does not support that, and one publisher contradicts it directly.
Google's guidance on AI features states that there are no additional requirements to appear in AI Overviews or AI Mode, and, in its own words, that 'You don't need to create new machine readable files, AI text files, or markup to appear in these features.' The same page directs site owners to existing preview controls, naming nosnippet, data-nosnippet, max-snippet, and noindex, as the mechanism for limiting inclusion.
This is the single most important honesty constraint on the topic. Structured data on a case study is not documented by Google as an entry ticket to its AI features. What structured data does is reduce the amount of inference required to determine what a page asserts, which is a parsing benefit and an accuracy benefit, and which may or may not influence any given system. AIOFacts treats the parsing benefit as real and the visibility benefit as unestablished, and keeps those two statements apart on purpose.
Google's structured data general guidelines add a constraint that applies whatever the benefit turns out to be: markup must be a true representation of the page content, content that is not visible to readers should not be marked up, and structured data must not be used to deceive or mislead. A case study whose JSON-LD asserts a figure the visible page does not state is not a stronger case study. It is a policy violation with a machine-readable audit trail.
Vocabulary that already exists for evidence
There is no schema.org type called CaseStudy, and inventing one solves nothing, because a vocabulary is only useful to the extent that consumers already understand it. What exists instead is a set of general purpose terms that decompose an outcome claim into its parts.
Schema.org's Observation type specifies observations about an entity at a particular time, and defines properties including measuredProperty, observationAbout, observationPeriod, marginOfError, measurementMethod, and measurementQualifier, with StatisticalVariable available to name the metric being measured. That is close to a formal skeleton of a case study result: what was measured, about whom, over what window, with what uncertainty, by what method. The honest caveat is adoption. Schema.org marks Observation as a term in its 'new' area, explicitly pending implementation feedback and adoption, and reports low usage across domains. Publishing it is a bet on a vocabulary that is standardized but not widely consumed.
Schema.org's Claim and ClaimReview separate an assertion from an assessment of it. Google documents ClaimReview as markup for pages that review a claim made by others, and the 'by others' is doing real work: the type is built for third party assessment, not for self-certification. An organization cannot fact-check its own case study into a higher tier of evidence by marking it up.
W3C's PROV-O, a Recommendation published in 2013, provides the vocabulary for the part nothing else covers: lineage. Its properties, including wasDerivedFrom, wasGeneratedBy, wasAttributedTo, and wasQuotedFrom, describe how one entity came from another and who is responsible for it. A reported figure that can point at the dataset it was derived from, and at the agent that produced it, is a different object from a figure that appears with no ancestry.
- Observation, StatisticalVariable, measurementMethod, marginOfError: the shape of a measured result
- Claim and ClaimReview: separating an assertion from an independent assessment of it
- Dataset and its distributions: the underlying numbers, where they can be released
- PROV-O wasDerivedFrom and wasAttributedTo: where a figure came from and who stands behind it
The self-serving problem, stated plainly
A case study is an entity publishing praise of itself. That is not a criticism, it is a description, and there is documented precedent for treating it as a distinct category.
Google's review snippet documentation defines self-serving reviews as reviews about an entity placed on that entity's own site, whether in its own markup or through an embedded third party widget, and states that pages using LocalBusiness or other Organization structured data are not eligible for the review star feature when the reviewed entity controls the reviews about itself. The rule is scoped to review snippets and should not be stretched into a general law about AI systems. What it demonstrates is that at least one major platform found the self-published testimonial problem serious enough to write an eligibility rule about it.
The implication for case studies is not that self-published outcome claims are worthless. It is that the discount is structural and cannot be removed by better writing. What can change is whether the claim is independently checkable: whether the client is named and can be reached, whether the metric came from a system the reader can also see, whether a third party has published anything that corroborates it. Corroboration is the only variable under the publisher's influence that actually moves this.
Practice: what a checkable case study contains
The practice that follows is mostly editorial rather than technical, and it starts from the assumption that any single sentence may travel alone.
Write outcome sentences as self-contained units. 'Organic sessions rose from 4,100 to 7,900 per month between March and September 2025, measured in Google Analytics 4 for the client's primary domain' survives extraction. 'Traffic nearly doubled' does not, because the baseline, the window, the metric definition, and the measurement instrument are all somewhere else on the page. The first sentence is longer and less quotable in a marketing sense, which is the tradeoff being made deliberately.
Then keep the markup subordinate to the visible text. Under Google's structured data general guidelines, markup describes what is on the page; it is not a second, more optimistic version of the page. Anything asserted in JSON-LD should be readable in the HTML body by a person.
- Name the subject entity with a resolvable identifier where possible, not just a trade name
- State the measurement window with explicit start and end dates, not a duration
- State the baseline as a number, because a percentage change without a baseline cannot be verified or falsified
- Name the measurement instrument and the account or property the figure came from
- State the method, including whether anything was controlled for, and say so plainly when nothing was
- Give uncertainty where it exists, using marginOfError or equivalent language, rather than presenting a point estimate as exact
- Link to any independent corroboration, and mark its absence as absence rather than filling the space
- Record what changed and when, so that a figure quoted a year later can be dated by whoever quotes it
Failure modes worth naming
The first is causal overreach, and it is endemic. A case study observes one subject, over one period, with no control group, while many things changed at once. The correct claim is that a set of actions was taken and a set of outcomes was observed in the same window. Structuring that claim well makes the overreach more visible, not less, because a machine-readable observation has a slot for method and the slot is empty.
The second is markup that outruns the page. Adding an Observation with a marginOfError to a figure whose margin was never calculated produces a precise-looking artifact built on nothing. Formal structure raises the apparent reliability of a claim without changing its actual reliability, and that gap is a real hazard of this entire practice rather than an argument against it.
The third is the assumption that any of this is measurable in outcomes. Even where a system does surface a case study, publishers generally cannot see which passage was used or why. Attribution reporting for AI assistant surfaces is limited and varies by platform, so a claim that structuring evidence produced a citation is usually not verifiable from the publisher's side.
The AIOFacts position
AIOFacts proposes that outcome claims be published as structured, dated, sourced observations, with the causal claim stated separately and more weakly than the measurement, and with markup that never asserts more than the visible page. The proposal rests on two things that are documented: the vocabularies exist and are standardized, and retrieval systems can quote at passage or sentence granularity. It does not rest on any claim about how a model weighs anything, because that is not publicly known.
The honest summary is that this is defensive work. It reduces the chance of a claim being repeated without its qualifiers, it makes an assertion falsifiable by a reader who wants to check it, and it produces an internal record of what was actually measured. Whether it increases the frequency of citation is unestablished, and Google's own guidance is that no new markup is required for its AI features. A practice that is worth doing only if it lifts visibility is not worth doing here. A practice that makes published claims true in a way that survives extraction is worth doing regardless.
Key points
- A case study fuses six separable components: subject, baseline, intervention, period, measurement, and causal attribution. Retrieval can separate them whether or not the publisher did.
- Anthropic documents that Claude can cite individual sentences and chain consecutive ones, which makes the self-contained sentence the practical unit of a checkable claim.
- OpenAI and Anthropic each document three distinct crawler roles for training, user-initiated fetch, and search. Access and citation are governed separately, so robots.txt decisions determine whether evidence is reachable at all.
- Google states directly that no new machine readable files or markup are needed for its AI features, so structured data on a case study should be justified by parsing accuracy and honesty, not by an assumed visibility gain.
- Schema.org's Observation supplies measuredProperty, observationPeriod, marginOfError, and measurementMethod, which is close to the skeleton of an outcome claim, but schema.org marks it as a new term with limited adoption.
- Google's structured data general guidelines require markup to be a true representation of visible page content, which makes an unsupported marginOfError or an invisible figure a policy problem rather than an optimization.
What this page cannot establish
- Whether any AI system uses schema.org structured data as an input when selecting or attributing sources. No major platform documents this, and Google states that no new markup is required to appear in its AI features.
- Whether Observation, StatisticalVariable, or PROV-O markup is consumed by any consumer-facing AI assistant. Schema.org marks Observation as a new term pending adoption, and no platform documents parsing it.
- How self-published outcome claims are treated relative to third party corroboration inside AI assistants. Google's self-serving rule is documented for review snippets specifically, and cannot be generalized beyond that surface.
- Whether structuring a case study changes citation frequency. Publishers generally cannot observe which passage a system used, so the effect is not measurable from the publishing side.
Sources
What supports this page
- AI features and your website
Google Search Central · platform-documentation · accessed 2026-08-10 - Structured data general guidelines
Google Search Central · platform-documentation · accessed 2026-08-10 - Review snippet (Review, AggregateRating) structured data
Google Search Central · platform-documentation · accessed 2026-08-10 - Fact Check (ClaimReview) structured data
Google Search Central · platform-documentation · accessed 2026-08-10 - Observation
Schema.org · published-standard · accessed 2026-08-10 - ClaimReview
Schema.org · published-standard · accessed 2026-08-10 - PROV-O: The PROV Ontology
W3C · published-standard · accessed 2026-08-10 - Bots: GPTBot, OAI-SearchBot and ChatGPT-User
OpenAI · platform-documentation · accessed 2026-08-10 - Does Anthropic crawl data from the web, and how can site owners block the crawler?
Anthropic · platform-documentation · accessed 2026-08-10 - Citations
Anthropic · platform-documentation · accessed 2026-08-10
Questions
Common questions
Is there a schema.org type for a case study?
No. There is no CaseStudy type in schema.org, and creating a custom one would not help, because consumers only interpret vocabulary they already understand. The workable approach is to decompose the document using existing terms: Article or Report for the page, Observation and StatisticalVariable for the measured result, Organization for the subject, and PROV-O for lineage.
Will marking up a case study get it cited by ChatGPT or Claude?
That is not established, and no platform documents structured data as a citation input. Google states explicitly that no new markup is required to appear in its AI features. Markup can reduce the inference required to determine what a page asserts, which is a parsing and accuracy benefit; treating it as a visibility mechanism goes beyond what any publisher has documented.
Can ClaimReview be used to fact-check our own case study?
It should not be. Google documents ClaimReview as markup for a page that reviews a claim made by others, and schema.org defines it as a fact-checking review of a claim made or reported in some creative work. Self-certification through ClaimReview inverts the purpose of the type and is likely to be treated as misleading markup under Google's structured data general guidelines.
Why does the self-serving problem matter if AI systems are not search engines?
It matters because the underlying asymmetry does not depend on the platform: an entity publishing praise of itself controls the evidence. Google's review snippet rule, which makes pages ineligible for star features when the reviewed entity controls the reviews about itself, shows that at least one platform wrote an eligibility rule about it. AIOFacts does not extend that rule to AI assistants, because nothing documents that it applies there.
What is the single highest-value change to make to an existing case study?
State the baseline as a number and the measurement window as explicit dates, in the same sentence as the outcome. A percentage change with no baseline and no window cannot be verified or falsified by anyone, and it loses nothing further when quoted in isolation because there was nothing there to lose.
One term, still unsettled, documented in the open.
Read the AIOFacts working definition, versioned and sourced, then see how the terminology is actually used in the wild.