Structured data and schema for AI: a practical guide

Search “schema markup for AI” and you will find a lot of confident numbers. Pages claiming schema makes you 36% more likely to appear in AI summaries. Claims that FAQ markup triples citations. Assertions that structured data is the foundation of generative engine optimization.

Try to trace any of those figures to a source.

We spent a while doing exactly that, and what follows is what we found. It is less exciting than the marketing and considerably more useful, because it tells you where to spend your time.

Last updated · · Part of our plain-English AI search guides
01The honest answer

The honest answer, up front

Here is the state of the evidence as it actually stands.

What the platforms have said
Microsoft
Google
OpenAI
Anthropic
Perplexity
ConfirmedPartialNo public statement
1
published correlation study. It found no correlation
0
peer-reviewed studies, as of early 2026

One AI platform has explicitly confirmed schema helps. Microsoft. Fabrice Canel, principal product manager at Bing, stated at SMX Munich in March 2025 that schema markup helps Microsoft’s language models understand content for Copilot. That is a direct, on-record confirmation from the vendor.

Google has confirmed something broader and vaguer. In April 2025 Google confirmed structured data gives an advantage in search results, and its May 2026 guidance named structured data as a supporting signal for AI features. Useful, but not the same as saying schema drives AI citations.

Three major platforms have said nothing at all. OpenAI, Anthropic and Perplexity have made no public statement on whether they preserve or use schema when crawling.

The one published correlation study found nothing. An analysis of schema coverage against LLM citation rates across OpenAI, Gemini and Perplexity found no correlation. Sites with comprehensive markup performed roughly the same as sites with almost none.

There are no peer-reviewed studies. Zero, as of early 2026.

So why does this page still recommend implementing schema?

Because “no measured correlation with citations” is not the same as “no value.” Schema does several things that are demonstrably real, just not the thing it is usually sold on. The rest of this page is about those things, and about spending your implementation time where it actually pays.

One caveat worth stating plainly: most published research on this subject comes from companies selling schema tools, companies selling link tools, or companies competing with both. Read the findings, then check who benefits from them. Including ours.

02Two consumers

Two consumers, one markup

This is the mental model that makes everything else make sense, and it is missing from most guides.

Your structured data has two entirely different audiences, and they use it differently.

One markupIn the HTML
<script type="application/ld+json">
  { "@type": "Organization", … }
</script>
Consumer
Search engines
Validate
Strictly. Invalid markup produces nothing
Rich result
Stars, FAQs, breadcrumbs, product cards
Consumer
AI models
Read as text
Alongside everything else on the page
Understanding
A clearer picture of what the page asserts
A rich result being deprecated says nothing about the second consumer
Search enginesAI models
What they do with itRun it through a validator and, if it qualifies, render a rich resultRead it as text alongside everything else on the page
Do they need it valid?Yes, strictly. Invalid markup produces nothingLess so. They are reading, not validating
What it earns youStars, FAQs, breadcrumbs, product cards in the SERPClearer understanding of what your page asserts
Is it guaranteed to be used?No, eligibility is not entitlementNo, and mostly unconfirmed
Where it must liveIn the HTMLIn the HTML, and it must survive without JavaScript

The practical consequence: a rich result being deprecated does not mean the markup stopped being useful. Google turning off a SERP feature says nothing about whether a language model reading your page finds the JSON-LD helpful. Those are two different consumers making two different decisions.

Hold onto that. It explains the single most common mistake in this whole area, which comes up in the deprecation section below.

03What platforms said

What each platform has actually said

Microsoft CopilotConfirmed
Schema helps Bing’s LLMs understand content
Fabrice Canel, SMX Munich, March 2025
GoogleConfirmed for search, partial for AI
Structured data gives an advantage in search; a supporting signal for AI features
Google, April 2025 and May 2026 guidance
OpenAI / ChatGPTUnknown
No public statement
—
Anthropic / ClaudeUnknown
No public statement
—
PerplexityUnknown
No public statement
—
Confidence as each vendor has stated it. A dash means there is nothing to cite

Three of the five biggest AI platforms have simply never addressed it.

What we can say from testing rather than documentation: pages with clean structured markup tend to be cited at higher rates than dense unformatted prose. But that observation is confounded, and badly. Sites that implement schema properly also tend to have better information architecture, clearer headings, named authors and stronger technical foundations. Attributing the citation difference to the JSON-LD specifically is a leap the data does not support.

Anyone who tells you otherwise has not separated the variables, because separating them is genuinely hard and nobody has published an attempt.

04How it could help

How schema could be helping

If there is no measured citation correlation, what is the mechanism people are arguing for? There are three plausible ones, and it is worth knowing which are speculative.

01
Training ingestion

Training pipelines typically retain document metadata, including JSON-LD from the HTML head. So schema plausibly contributes to what a model learns about your organization during training.

Support

Logical, undocumented by any vendor

02
Retrieval-time parsing

RAG systems fetch and read pages at query time. Schema sitting in that HTML is available to be read.

Support

Whether any specific system weights it is unconfirmed outside Microsoft’s statement

03
Entity disambiguation

This is the strongest of the three and the one worth building for.

Support

Better supported: models extract facts more accurately from structured data

Only the third is backed by research, and it is a different claim from “schema drives citations”

Research does show that language models extract information more accurately from structured data than from unstructured text. That is a different claim from “schema drives citations”, and it is better supported.

Put concretely: your page says you are “a leading provider serving clients nationwide.” A model reading that has to infer what you are, where you are and what you do. Your Organization schema, with a proper address, service list and sameAs links to your verified profiles, says the same thing in a form that requires no inference at all.

That is what schema is genuinely good at. It is not a ranking trick. It is removing ambiguity about facts you want repeated correctly.

Which, on a platform where you are trying to be described accurately without a listing to claim, is worth more than it sounds.

05Types that matter

Which schema types actually matter

Most sites over-implement breadth and under-implement depth. Ten types marked up thinly is worse than four marked up properly.

Tier 1Implement these on every site
Organization
Sitewide, usually homepage
The core entity node. Everything else references it. Carries sameAs, address, logo, contact
WebSite
Homepage
Establishes the site entity, supports sitelinks search box
BreadcrumbList
Every page
Makes hierarchy explicit, still earns rich results
Article or BlogPosting
Every content page
Carries author, dates and publisher, which is E-E-A-T made machine-readable
Person
Author pages, team pages
The author entity layer. Underused and genuinely valuable
Tier 2Implement where relevant
LocalBusiness or a subtype
Local businesses
Geo, hours, service area. Must match your Google Business Profile exactly
Service
Service pages
States what you sell in machine-readable form
Product
E-commerce
Price, availability, identifiers. Increasingly relevant to AI shopping surfaces
FAQPage
Q&A sections
See the deprecation section, the reasoning has changed
Event
Events
Still earns rich results, still well supported
Review / AggregateRating
Where genuinely earned
Read the warning below before touching this
Tier 3Situational
RecipeJobPostingCourseSoftwareApplicationVideoObjectHowToDatasetMedicalEntity

Implement if they describe what you actually publish. Do not add them because a list said to.

The Review and AggregateRating warning

Self-serving review markup, where a business marks up reviews about itself on its own site, is against Google’s guidelines and risks a manual action. Mark up reviews of products you sell. Do not mark up reviews of your own business on your own service pages. This is the most common way a well-intentioned schema implementation causes actual damage.

06The deprecation trap

The deprecation trap

Here is where the two-consumers model earns its keep.

Google has deprecated several rich result types over the past few years. FAQPage rich results were restricted in August 2023 to a small set of government and health authorities. HowTo rich results followed in September 2023.

Rich result changeStatement about AI
  1. Aug 2023
    FAQPage rich results restricted
    To a small set of government and health authorities
  2. Sept 2023
    HowTo rich results deprecated
    Followed a month later
  3. Mar 2025
    Microsoft confirms schema helps Copilot
    Fabrice Canel, SMX Munich
  4. Apr 2025
    Google: structured data gives an advantage
    In search results
  5. May 2026
    Google names it a supporting signal
    For AI features
Google turned off SERP features. It said nothing about what models do with the markup

Most guides draw the obvious conclusion and tell you to remove the markup.

That conclusion does not follow.

Google stopped rendering a SERP feature. That tells you nothing about whether a language model reading your HTML finds a clean FAQPage block useful. Those are different systems making different decisions about the same markup.

Rich result statusWorth keeping for AI?
FAQPageDeprecated for most sites, Aug 2023Yes. It structures question and answer pairs, which is exactly the shape retrieval favours
HowToDeprecated, Sept 2023Yes, for genuine step-by-step content
ArticleActiveYes, both
ProductActiveYes, both
BreadcrumbListActiveYes, both

The honest framing: keeping deprecated markup costs you a few kilobytes and some maintenance. The AI benefit is plausible and unproven. That trade is worth making, and you should know it is a judgment call rather than a certainty.

What you should not do is keep a bloated HowTo block on a page that is not a how-to, purely because a checklist told you to. Markup that misdescribes your content is worse than no markup at all, for both consumers.

07The entity layer

The entity layer that does the actual work

If you take one section away from this page, take this one. This is where structured data stops being a formality and starts doing something no amount of copywriting can.

Inside your site · @id references
WebSitepublisher →
PersonworksFor →
Articlepublisher →
Serviceprovider →
The entity
Organization
"@id": "https://example.com/#organization"
Outside your site · sameAs
← LinkedIn← Wikidata← Crunchbase← Companies House← X

Verified profiles the model can use to confirm who you are.

Defined once, referenced by everything. Nothing repeated, every relationship explicit

Three mechanisms matter.

@id: give every entity a permanent name

Without @id, each page’s schema is an island. With it, you are building a connected graph that says “this Organization on the about page is the same Organization publishing this article.”

JSON-LD · @id
"@id": "https://example.com/#organization"

Use a consistent URI pattern sitewide and never change it. This is the single highest-leverage structured data decision most sites never make.

sameAs: connect your entity to the rest of the web

sameAs links your entity to its verified presence elsewhere. This is what lets a model reading your site connect it to your LinkedIn, your Wikidata entry, your Crunchbase profile and your Companies House record.

JSON-LD · sameAs
"sameAs": [
  "https://www.linkedin.com/company/example",
  "https://www.crunchbase.com/organization/example",
  "https://www.wikidata.org/wiki/Q00000000",
  "https://x.com/example"
]

For entity clarity, which is the one thing every AI platform demonstrably cares about, this array does more work than any other line of schema on your site.

@graph: put it all in one block

Rather than scattering separate JSON-LD blocks, use a single @graph array with internal references. Cleaner to maintain, easier to parse, and it makes the relationships explicit rather than implied.

08A working example

A working example

A complete @graph block for a service business page. Adapt the values, keep the structure.

JSON-LD · complete @graph
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Organization",
      "@id": "https://example.com/#organization",
      "name": "Example Company",
      "url": "https://example.com/",
      "logo": {
        "@type": "ImageObject",
        "url": "https://example.com/logo.png",
        "width": 600,
        "height": 60
      },
      "description": "One canonical sentence describing the business. Use this exact wording everywhere.",
      "address": {
        "@type": "PostalAddress",
        "streetAddress": "123 Example Street",
        "addressLocality": "Example City",
        "addressRegion": "Example Region",
        "postalCode": "00000",
        "addressCountry": "CA"
      },
      "contactPoint": {
        "@type": "ContactPoint",
        "telephone": "+1-555-555-0100",
        "contactType": "customer service",
        "areaServed": "CA",
        "availableLanguage": "English"
      },
      "sameAs": [
        "https://www.linkedin.com/company/example",
        "https://www.crunchbase.com/organization/example",
        "https://x.com/example"
      ]
    },
    {
      "@type": "WebSite",
      "@id": "https://example.com/#website",
      "url": "https://example.com/",
      "name": "Example Company",
      "publisher": { "@id": "https://example.com/#organization" }
    },
    {
      "@type": "Person",
      "@id": "https://example.com/#jane-doe",
      "name": "Jane Doe",
      "jobTitle": "Head of Strategy",
      "worksFor": { "@id": "https://example.com/#organization" },
      "url": "https://example.com/team/jane-doe/",
      "sameAs": ["https://www.linkedin.com/in/janedoe"]
    },
    {
      "@type": "Service",
      "@id": "https://example.com/services/consulting/#service",
      "name": "Strategy Consulting",
      "serviceType": "Business consulting",
      "provider": { "@id": "https://example.com/#organization" },
      "areaServed": {
        "@type": "AdministrativeArea",
        "name": "Example Region"
      },
      "description": "What the service is, in one clear sentence."
    },
    {
      "@type": "BreadcrumbList",
      "itemListElement": [
        { "@type": "ListItem", "position": 1, "name": "Home", "item": "https://example.com/" },
        { "@type": "ListItem", "position": 2, "name": "Services", "item": "https://example.com/services/" },
        { "@type": "ListItem", "position": 3, "name": "Consulting", "item": "https://example.com/services/consulting/" }
      ]
    }
  ]
}
</script>
Every value is a placeholder. Replace them with your own

Note what is happening structurally: one Organization node, defined once, referenced by everything else through @id. The Person node links to the Organization via worksFor. The Service names the Organization as provider. Nothing is repeated, and every relationship is explicit.

That is a small knowledge graph, and it is the thing worth building.

09Getting it wrong

Getting it wrong

MistakeWhy it hurtsFix
Markup that contradicts the pageBoth consumers penalize mismatch. Google treats it as spamSchema describes what is visibly on the page. Always
Self-serving review markupGuideline violation, manual action riskMark up reviews of products, never of your own business on your own site
JSON-LD injected by JavaScriptMost AI crawlers do not execute JS, so it is invisible to themServer-side render it into the HTML, as every site we design does
No @id anywhereEvery page’s schema is an isolated island, no graph formsConsistent URI pattern sitewide
Empty or placeholder sameAsWastes the highest-value entity signal availableLink only to profiles you actually control and maintain
Ten types, all thinBreadth without depth helps nobodyFour types done properly beats ten done badly
Copy-pasted from a generator, never checkedWrong types, invented properties, stale syntaxValidate, then read what it actually says
NAP that does not match the Google Business ProfileActively damages local trust signalsByte-for-byte consistency, including punctuation
Set and forgetPrices, hours, staff and services all changeReview quarterly alongside content

The mismatch row deserves emphasis. Schema asserting facts your page does not support is the one failure mode that makes things actively worse rather than merely neutral.

If your markup says you are open until 9pm and your page says 5pm, you have made a model less certain about you, not more.

10Testing

Testing and validation

Three tools, three jobs.

Schema.org Validator checks whether your markup is valid against the vocabulary. Use this first. It catches invented properties and structural errors.

Google Rich Results Test checks whether Google considers you eligible for a rich result. Narrower than the vocabulary, since Google supports a subset. Eligibility is not entitlement.

Search Console’s structured data reports show what Google actually found across your site at scale, which is where you spot the template that broke three months ago.

One test none of these perform: whether an AI model can see your markup at all. For that, fetch your page with JavaScript disabled and confirm the JSON-LD is present in the raw HTML. If it is not, everything above is academic. The same check, and what else blocks AI crawlers, is covered in our crawler access guide.

11Priority order

Schema: Priority order

Lower return
05Type-specific markup where relevant
04One @graph with @id references
03Article schema on everything
02Author entities
01Organization with a complete sameAs
Highest return
06 · Gate
Verify it survives without JavaScript

A check, not a build. First in importance if it fails.

In order, highest return first.

  1. One Organization node with a complete sameAs array. Sitewide, with a stable @id. If you do nothing else, do this. It is the entity foundation everything else rests on.
  2. Author entities. Person schema on real author pages, with credentials and sameAs to their professional profiles, linked from every Article via author. This is E-E-A-T rendered machine-readable, and almost nobody does it properly.
  3. Article schema on everything you publish. Author, datePublished, dateModified, publisher. Cheap, universally supported.
  4. Consolidate into a single @graph with @id references replacing repetition.
  5. Type-specific markup where genuinely relevant. Service, Product, LocalBusiness, Event. Match reality, do not pad.
  6. Verify it survives without JavaScript. Last, because it is a check rather than a build, and first in importance if it fails.

What deliberately is not on this list: chasing every schema type in the vocabulary. The returns fall off a cliff after the entity layer, and the time is better spent on the things the evidence actually supports, which is clear content structure, named authors and third-party corroboration. That wider work is what AI optimization covers.

12FAQ

Frequently asked questions

Does schema markup help with AI search?

Partially, and less than commonly claimed. Microsoft has confirmed schema helps Bing’s language models understand content for Copilot. Google has confirmed structured data gives an advantage in search and named it a supporting signal for AI features. OpenAI, Anthropic and Perplexity have said nothing publicly, and one published analysis found no correlation between schema coverage and citation rates across OpenAI, Gemini and Perplexity. Implement it as infrastructure, not as a citation lever.

Which schema types do AI models actually parse?

No AI vendor publishes a list. What is defensible from testing and vendor statements: Organization, Person, Article, Product, Service and FAQPage carry the most useful entity and factual information. The @id and sameAs properties do more work for entity clarity than any individual type.

Should I keep FAQPage schema now that Google deprecated the rich result?

Yes, if the page genuinely contains questions and answers. Google deprecated a SERP feature, which says nothing about whether a language model reading your HTML finds structured question and answer pairs useful. The cost of keeping it is negligible.

Is JSON-LD better than microdata for AI?

Yes, for practical reasons. JSON-LD sits in a single block that is trivially separable from page content, which makes it easier for any parser to handle. Microdata and RDFa are interleaved with markup and harder to extract cleanly.

Will schema markup get me cited by ChatGPT?

There is no evidence it will on its own. The published correlation analysis found none. What schema does reliably is reduce ambiguity about who you are and what you do, which matters on every platform. Citations are driven by content quality, structure and third-party corroboration, which is what our ChatGPT SEO guide covers.

Does schema need to be in the HTML or can JavaScript inject it?

It must be in the served HTML. Most AI crawlers do not execute JavaScript, so client-side injected JSON-LD is invisible to them even when Google can see it. Our guide to which AI bots to allow covers what they can and cannot read.

How much schema is too much?

When it stops describing what is actually on the page. Marking up content you do not have is the one mistake that makes things worse rather than neutral.

Can I mark up reviews of my own business?

Not on your own site. Self-serving review markup violates Google’s guidelines and risks a manual action. Mark up reviews of products you sell, not of your own services.
Free AI visibility audit

Want this checked properly?

Structured data is one of about a dozen things we audit when a business is not appearing in AI answers. On its own it is rarely the cause. As part of a wider entity problem, it frequently is.