Structured data and schema for AI: a practical guide
Search “schema markup for AI” and you will find a lot of confident numbers. Pages claiming schema makes you 36% more likely to appear in AI summaries. Claims that FAQ markup triples citations. Assertions that structured data is the foundation of generative engine optimization.
Try to trace any of those figures to a source.
We spent a while doing exactly that, and what follows is what we found. It is less exciting than the marketing and considerably more useful, because it tells you where to spend your time.
The honest answer, up front
Here is the state of the evidence as it actually stands.
One AI platform has explicitly confirmed schema helps. Microsoft. Fabrice Canel, principal product manager at Bing, stated at SMX Munich in March 2025 that schema markup helps Microsoft’s language models understand content for Copilot. That is a direct, on-record confirmation from the vendor.
Google has confirmed something broader and vaguer. In April 2025 Google confirmed structured data gives an advantage in search results, and its May 2026 guidance named structured data as a supporting signal for AI features. Useful, but not the same as saying schema drives AI citations.
Three major platforms have said nothing at all. OpenAI, Anthropic and Perplexity have made no public statement on whether they preserve or use schema when crawling.
The one published correlation study found nothing. An analysis of schema coverage against LLM citation rates across OpenAI, Gemini and Perplexity found no correlation. Sites with comprehensive markup performed roughly the same as sites with almost none.
There are no peer-reviewed studies. Zero, as of early 2026.
So why does this page still recommend implementing schema?
Because “no measured correlation with citations” is not the same as “no value.” Schema does several things that are demonstrably real, just not the thing it is usually sold on. The rest of this page is about those things, and about spending your implementation time where it actually pays.
One caveat worth stating plainly: most published research on this subject comes from companies selling schema tools, companies selling link tools, or companies competing with both. Read the findings, then check who benefits from them. Including ours.
Two consumers, one markup
This is the mental model that makes everything else make sense, and it is missing from most guides.
Your structured data has two entirely different audiences, and they use it differently.
<script type="application/ld+json">
{ "@type": "Organization", … }
</script>| Search engines | AI models | |
|---|---|---|
| What they do with it | Run it through a validator and, if it qualifies, render a rich result | Read it as text alongside everything else on the page |
| Do they need it valid? | Yes, strictly. Invalid markup produces nothing | Less so. They are reading, not validating |
| What it earns you | Stars, FAQs, breadcrumbs, product cards in the SERP | Clearer understanding of what your page asserts |
| Is it guaranteed to be used? | No, eligibility is not entitlement | No, and mostly unconfirmed |
| Where it must live | In the HTML | In the HTML, and it must survive without JavaScript |
The practical consequence: a rich result being deprecated does not mean the markup stopped being useful. Google turning off a SERP feature says nothing about whether a language model reading your page finds the JSON-LD helpful. Those are two different consumers making two different decisions.
Hold onto that. It explains the single most common mistake in this whole area, which comes up in the deprecation section below.
What each platform has actually said
Three of the five biggest AI platforms have simply never addressed it.
What we can say from testing rather than documentation: pages with clean structured markup tend to be cited at higher rates than dense unformatted prose. But that observation is confounded, and badly. Sites that implement schema properly also tend to have better information architecture, clearer headings, named authors and stronger technical foundations. Attributing the citation difference to the JSON-LD specifically is a leap the data does not support.
Anyone who tells you otherwise has not separated the variables, because separating them is genuinely hard and nobody has published an attempt.
How schema could be helping
If there is no measured citation correlation, what is the mechanism people are arguing for? There are three plausible ones, and it is worth knowing which are speculative.
Training pipelines typically retain document metadata, including JSON-LD from the HTML head. So schema plausibly contributes to what a model learns about your organization during training.
Logical, undocumented by any vendor
RAG systems fetch and read pages at query time. Schema sitting in that HTML is available to be read.
Whether any specific system weights it is unconfirmed outside Microsoft’s statement
This is the strongest of the three and the one worth building for.
Better supported: models extract facts more accurately from structured data
Research does show that language models extract information more accurately from structured data than from unstructured text. That is a different claim from “schema drives citations”, and it is better supported.
Put concretely: your page says you are “a leading provider serving clients nationwide.” A model reading that has to infer what you are, where you are and what you do. Your Organization schema, with a proper address, service list and sameAs links to your verified profiles, says the same thing in a form that requires no inference at all.
That is what schema is genuinely good at. It is not a ranking trick. It is removing ambiguity about facts you want repeated correctly.
Which, on a platform where you are trying to be described accurately without a listing to claim, is worth more than it sounds.
Which schema types actually matter
Most sites over-implement breadth and under-implement depth. Ten types marked up thinly is worse than four marked up properly.
Implement if they describe what you actually publish. Do not add them because a list said to.
The Review and AggregateRating warning
Self-serving review markup, where a business marks up reviews about itself on its own site, is against Google’s guidelines and risks a manual action. Mark up reviews of products you sell. Do not mark up reviews of your own business on your own service pages. This is the most common way a well-intentioned schema implementation causes actual damage.
The deprecation trap
Here is where the two-consumers model earns its keep.
Google has deprecated several rich result types over the past few years. FAQPage rich results were restricted in August 2023 to a small set of government and health authorities. HowTo rich results followed in September 2023.
- Aug 2023FAQPage rich results restrictedTo a small set of government and health authorities
- Sept 2023HowTo rich results deprecatedFollowed a month later
- Mar 2025Microsoft confirms schema helps CopilotFabrice Canel, SMX Munich
- Apr 2025Google: structured data gives an advantageIn search results
- May 2026Google names it a supporting signalFor AI features
Most guides draw the obvious conclusion and tell you to remove the markup.
That conclusion does not follow.
Google stopped rendering a SERP feature. That tells you nothing about whether a language model reading your HTML finds a clean FAQPage block useful. Those are different systems making different decisions about the same markup.
| Rich result status | Worth keeping for AI? | |
|---|---|---|
| FAQPage | Deprecated for most sites, Aug 2023 | Yes. It structures question and answer pairs, which is exactly the shape retrieval favours |
| HowTo | Deprecated, Sept 2023 | Yes, for genuine step-by-step content |
| Article | Active | Yes, both |
| Product | Active | Yes, both |
| BreadcrumbList | Active | Yes, both |
The honest framing: keeping deprecated markup costs you a few kilobytes and some maintenance. The AI benefit is plausible and unproven. That trade is worth making, and you should know it is a judgment call rather than a certainty.
What you should not do is keep a bloated HowTo block on a page that is not a how-to, purely because a checklist told you to. Markup that misdescribes your content is worse than no markup at all, for both consumers.
The entity layer that does the actual work
If you take one section away from this page, take this one. This is where structured data stops being a formality and starts doing something no amount of copywriting can.
Verified profiles the model can use to confirm who you are.
Three mechanisms matter.
@id: give every entity a permanent name
Without @id, each page’s schema is an island. With it, you are building a connected graph that says “this Organization on the about page is the same Organization publishing this article.”
"@id": "https://example.com/#organization"Use a consistent URI pattern sitewide and never change it. This is the single highest-leverage structured data decision most sites never make.
sameAs: connect your entity to the rest of the web
sameAs links your entity to its verified presence elsewhere. This is what lets a model reading your site connect it to your LinkedIn, your Wikidata entry, your Crunchbase profile and your Companies House record.
"sameAs": [
"https://www.linkedin.com/company/example",
"https://www.crunchbase.com/organization/example",
"https://www.wikidata.org/wiki/Q00000000",
"https://x.com/example"
]For entity clarity, which is the one thing every AI platform demonstrably cares about, this array does more work than any other line of schema on your site.
@graph: put it all in one block
Rather than scattering separate JSON-LD blocks, use a single @graph array with internal references. Cleaner to maintain, easier to parse, and it makes the relationships explicit rather than implied.
A working example
A complete @graph block for a service business page. Adapt the values, keep the structure.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Organization",
"@id": "https://example.com/#organization",
"name": "Example Company",
"url": "https://example.com/",
"logo": {
"@type": "ImageObject",
"url": "https://example.com/logo.png",
"width": 600,
"height": 60
},
"description": "One canonical sentence describing the business. Use this exact wording everywhere.",
"address": {
"@type": "PostalAddress",
"streetAddress": "123 Example Street",
"addressLocality": "Example City",
"addressRegion": "Example Region",
"postalCode": "00000",
"addressCountry": "CA"
},
"contactPoint": {
"@type": "ContactPoint",
"telephone": "+1-555-555-0100",
"contactType": "customer service",
"areaServed": "CA",
"availableLanguage": "English"
},
"sameAs": [
"https://www.linkedin.com/company/example",
"https://www.crunchbase.com/organization/example",
"https://x.com/example"
]
},
{
"@type": "WebSite",
"@id": "https://example.com/#website",
"url": "https://example.com/",
"name": "Example Company",
"publisher": { "@id": "https://example.com/#organization" }
},
{
"@type": "Person",
"@id": "https://example.com/#jane-doe",
"name": "Jane Doe",
"jobTitle": "Head of Strategy",
"worksFor": { "@id": "https://example.com/#organization" },
"url": "https://example.com/team/jane-doe/",
"sameAs": ["https://www.linkedin.com/in/janedoe"]
},
{
"@type": "Service",
"@id": "https://example.com/services/consulting/#service",
"name": "Strategy Consulting",
"serviceType": "Business consulting",
"provider": { "@id": "https://example.com/#organization" },
"areaServed": {
"@type": "AdministrativeArea",
"name": "Example Region"
},
"description": "What the service is, in one clear sentence."
},
{
"@type": "BreadcrumbList",
"itemListElement": [
{ "@type": "ListItem", "position": 1, "name": "Home", "item": "https://example.com/" },
{ "@type": "ListItem", "position": 2, "name": "Services", "item": "https://example.com/services/" },
{ "@type": "ListItem", "position": 3, "name": "Consulting", "item": "https://example.com/services/consulting/" }
]
}
]
}
</script>Note what is happening structurally: one Organization node, defined once, referenced by everything else through @id. The Person node links to the Organization via worksFor. The Service names the Organization as provider. Nothing is repeated, and every relationship is explicit.
That is a small knowledge graph, and it is the thing worth building.
Getting it wrong
| Mistake | Why it hurts | Fix |
|---|---|---|
| Markup that contradicts the page | Both consumers penalize mismatch. Google treats it as spam | Schema describes what is visibly on the page. Always |
| Self-serving review markup | Guideline violation, manual action risk | Mark up reviews of products, never of your own business on your own site |
| JSON-LD injected by JavaScript | Most AI crawlers do not execute JS, so it is invisible to them | Server-side render it into the HTML, as every site we design does |
| No @id anywhere | Every page’s schema is an isolated island, no graph forms | Consistent URI pattern sitewide |
| Empty or placeholder sameAs | Wastes the highest-value entity signal available | Link only to profiles you actually control and maintain |
| Ten types, all thin | Breadth without depth helps nobody | Four types done properly beats ten done badly |
| Copy-pasted from a generator, never checked | Wrong types, invented properties, stale syntax | Validate, then read what it actually says |
| NAP that does not match the Google Business Profile | Actively damages local trust signals | Byte-for-byte consistency, including punctuation |
| Set and forget | Prices, hours, staff and services all change | Review quarterly alongside content |
The mismatch row deserves emphasis. Schema asserting facts your page does not support is the one failure mode that makes things actively worse rather than merely neutral.
If your markup says you are open until 9pm and your page says 5pm, you have made a model less certain about you, not more.
Testing and validation
Three tools, three jobs.
Schema.org Validator checks whether your markup is valid against the vocabulary. Use this first. It catches invented properties and structural errors.
Google Rich Results Test checks whether Google considers you eligible for a rich result. Narrower than the vocabulary, since Google supports a subset. Eligibility is not entitlement.
Search Console’s structured data reports show what Google actually found across your site at scale, which is where you spot the template that broke three months ago.
One test none of these perform: whether an AI model can see your markup at all. For that, fetch your page with JavaScript disabled and confirm the JSON-LD is present in the raw HTML. If it is not, everything above is academic. The same check, and what else blocks AI crawlers, is covered in our crawler access guide.
Schema: Priority order
A check, not a build. First in importance if it fails.
In order, highest return first.
- One Organization node with a complete sameAs array. Sitewide, with a stable @id. If you do nothing else, do this. It is the entity foundation everything else rests on.
- Author entities. Person schema on real author pages, with credentials and sameAs to their professional profiles, linked from every Article via author. This is E-E-A-T rendered machine-readable, and almost nobody does it properly.
- Article schema on everything you publish. Author, datePublished, dateModified, publisher. Cheap, universally supported.
- Consolidate into a single @graph with @id references replacing repetition.
- Type-specific markup where genuinely relevant. Service, Product, LocalBusiness, Event. Match reality, do not pad.
- Verify it survives without JavaScript. Last, because it is a check rather than a build, and first in importance if it fails.
What deliberately is not on this list: chasing every schema type in the vocabulary. The returns fall off a cliff after the entity layer, and the time is better spent on the things the evidence actually supports, which is clear content structure, named authors and third-party corroboration. That wider work is what AI optimization covers.
Frequently asked questions
Does schema markup help with AI search?
Which schema types do AI models actually parse?
Should I keep FAQPage schema now that Google deprecated the rich result?
Is JSON-LD better than microdata for AI?
Will schema markup get me cited by ChatGPT?
Does schema need to be in the HTML or can JavaScript inject it?
How much schema is too much?
Can I mark up reviews of my own business?
Want this checked properly?
Structured data is one of about a dozen things we audit when a business is not appearing in AI answers. On its own it is rarely the cause. As part of a wider entity problem, it frequently is.