How to measure AI visibility: a complete methodology
Ask ChatGPT the same question twice and you may get two different sets of businesses named. Ask it tomorrow and you may get a third.
That single fact breaks every measurement habit the search industry spent twenty years building. There is no position to check, no tracked keyword to look up, no number that stays still long enough to put in a report.
What there is instead is a distribution, and a well-understood way to estimate one. This page is the methodology: how to build a prompt set, how many observations you genuinely need, how to avoid contaminating your own data, and how to report numbers that survive someone checking your work.
Fair warning on one finding below: by these standards, a lot of published AI visibility research, including a lot of agency reporting, does not have enough data to support what it claims.
Last updated · Part of our knowledge base
Why this is not rank tracking
| SEO rank tracking | AI visibility | |
|---|---|---|
| Same query, same day | Same result, give or take personalization | Frequently different results |
| What you are measuring | A position | A probability |
| One check tells you | The answer | Almost nothing |
| Method | Look it up | Sample it, repeatedly |
| Stability | Changes when the index changes | Changes between runs, for no external reason |
The scale of the instability is not marginal. Analysis of AI Overviews found the cited content shifting roughly 70% of the time on repeat queries. Not over months. On a re-run.
So a single check is not a small sample. It is one draw from a distribution you have not characterized, and treating it as a result is how agencies end up reporting improvements that are noise and missing declines that are real. A free tool like our AI visibility checker is useful exactly as that: a first snapshot, not a measurement.
Everything that follows exists to solve that problem properly.
Three different things, and most reports confuse them
These get used interchangeably. They are not the same, they move independently, and conflating them produces reports that mislead in both directions.
Does the model say our name?
Does it link us as a source?
Did anyone click through?
| Measure | The question it answers | How you capture it | The trap |
|---|---|---|---|
| Named rate | Does the model say our name? | Prompt testing | Invisible in analytics. This is the primary metric |
| Cited rate | Does it link us as a source? | Prompt testing, logging citations | Much rarer than being named |
| Referral traffic | Did anyone click through? | GA4 | Systematically understates influence |
The gap between the first two is large and platform-specific. ChatGPT mentions brands roughly 3.2 times more often than it cites them with a link. Perplexity, built around attribution, cites on nearly every answer, averaging around 22 sources per response, against roughly six for Claude and seven for Copilot.
Which produces a reporting trap worth naming. A brand with a strong named rate and thin referral traffic is not failing. It is being recommended in conversations that never touch its analytics. Report only on GA4 sessions and you will conclude the program is not working while it is working.
The reverse trap is rarer but real: healthy referral numbers from a handful of prompts can mask absence across the rest of the category.
Measure all three. Lead with named rate.
Building the prompt set
Your prompt set is the instrument. A biased instrument produces precise measurements of the wrong thing, and the most common bias is flattery: people build prompt sets from queries where they already do well.
Composition
A usable set runs 150 to 300 prompts, distributed roughly like this:
| Category | Share | Example shape |
|---|---|---|
| Category discovery | 30% | “Best [service] in [city]”, “Top [category] providers” |
| Problem-led | 25% | “How do I fix [problem the customer has]” |
| Comparison | 20% | “[Competitor] alternatives”, “[A] vs [B]” |
| Qualification | 15% | “How much does [service] cost”, “How to choose a [provider]” |
| Branded | 10% | “What does [your brand] do”, “Is [your brand] any good” |
That last row is small and disproportionately important. Branded prompts measure accuracy rather than presence, and they are where you discover the model has your pricing wrong, your service list outdated, or your location in the wrong country.
Four rules
- Write them as a customer would type them. Not as keywords. “Best accountant in Vancouver for a small construction business” is a real prompt. “vancouver accountant” is a keyword someone pasted.
- Include prompts you expect to lose. A set built only on your strengths measures your confidence, not your visibility.
- Cover the sub-questions. On surfaces using query fan-out, the engine invents its own sub-queries. Your set should include the questions it will generate, not only the ones your customer types.
- Freeze it. Once baselined, the set does not change. Adding prompts mid-program makes period comparison invalid, and it is the easiest way to accidentally manufacture an improvement. Version the set, and if you must revise it, start a new baseline and say so in the report.
How much data you actually need
This is the part nobody publishes, and it determines whether anything else you do is meaningful.
You are estimating a proportion: the share of prompt-runs where your brand is named. That is a standard statistical problem with a standard answer.
The margin of error on your named rate
At 95% confidence, assuming a true named rate around 30%:
| Observations | Margin of error | What that means |
|---|---|---|
| 20 | ±20.1 points | A measured 30% could truly be 10% or 50% |
| 50 | ±12.7 points | Still too wide to act on |
| 100 | ±9.0 points | Directionally useful, not reportable as precise |
| 200 | ±6.4 points | Usable |
| 300 | ±5.2 points | Solid |
| 500 | ±4.0 points | Strong |
| 1,000 | ±2.8 points | Research grade |
Read the first row again. Twenty manual spot checks produce an estimate with a twenty point margin of error. You could measure 30% when the truth is 10%, and report a catastrophe as a modest result, or the reverse.
This is why “I asked ChatGPT a few times and we didn’t come up” is not a finding. It is also why a great deal of published AI visibility content, built on small manual samples, should be read with more suspicion than it usually gets.
Getting to a usable number
Observations are prompts multiplied by runs, per platform:
| Prompt set | Runs | Observations per platform |
|---|---|---|
| 100 | 1 | 100 |
| 150 | 3 | 450 |
| 200 | 3 | 600 |
| 300 | 3 | 900 |
150 prompts run three times gives you 450 observations and a margin of error around ±4.3 points. That is a defensible number, and it is the minimum we would put in front of a client.
One statistical note worth knowing: the margin is widest when the true rate sits near 50% and narrows toward the extremes. If your named rate is 5% or 90%, you need less data to pin it down than the table suggests. Most brands starting out are well below 30%, which works slightly in your favour at baseline and against you as you improve.
Detecting change is harder than measuring a level
Here is where most reporting quietly breaks.
Measuring where you stand is one problem. Proving you moved is a harder one, because now two noisy estimates have to be far enough apart that the difference cannot be explained by sampling.
Sample size needed per period, at 95% confidence with 80% power:
| The change you want to prove | Observations needed per period |
|---|---|
| 30% to 45% (15 points) | 160 |
| 30% to 40% (10 points) | 354 |
| 10% to 20% (10 points) | 197 |
| 30% to 35% (5 points) | 1,374 |
| 30% to 32% (2 points) | 8,391 |
That bottom row is the important one. Small improvements are effectively unprovable at normal sampling volumes. A two point shift needs over eight thousand observations per period to distinguish from noise.
Three practical consequences.
- Do not report small movements as results. If your measurement carries ±5 points and the number moved 3, nothing happened that you can demonstrate.
- Set expectations around double-digit shifts. Below ten points, you are usually reporting noise with a narrative attached.
- Use trend lines, not point comparisons. Six months of monthly measurement showing consistent upward drift is more persuasive than any two individual months, and it is more honest, because it does not depend on a single comparison surviving scrutiny.
There is a commercial temptation here that is worth naming. Small sample sizes produce big swings, and big swings make impressive slides.
Resisting that is the difference between a measurement practice and a sales tool.
Running it cleanly
Your own setup can contaminate the result before you record a single number.
Memory and chat history personalize answers to you
Earlier prompts in the same thread steer later ones
Local results vary by IP
Personalization learns what you look at
Index freshness varies
Different systems with different retrieval
That last source deserves its own paragraph, because most tooling depends on it.
The API is not the product. Testing through an API may not reflect what a person using the consumer app actually sees. Retrieval behaviour, system prompts and grounding can all differ. Some tools measure via API because it is cheap and scalable, then report the result as consumer visibility.
That is not necessarily wrong, but it is a methodological choice, and it should be disclosed. When evaluating any AI visibility tool, ask how it collects data. If the answer is vague, treat the numbers as directional.
The principle: whatever you choose, hold it constant.
Changing collection method between periods invalidates the comparison more thoroughly than any sample size problem.
What to record
A binary named or not-named throws away most of the signal. Capture this per observation:
| Field | Values | Why it matters |
|---|---|---|
| Named | Yes / No | The primary metric |
| Cited | Yes / No, with URL | Separate from naming, and rarer |
| Position | 1st, 2nd, 3rd, later | Being named first carries more weight than being named |
| Sentiment | Positive / Neutral / Negative / Caveated | “X is solid but expensive” is not a clean win |
| Accuracy | Correct / Partially wrong / Wrong | Branded prompts especially |
| Competitors named | List | Your actual competitive set, often not who you assumed |
| Sources cited | Domain list | The most actionable field on this list |
| Search triggered | Yes / No | Whether the model retrieved or answered from training |
Two of these do more work than the rest.
Sources cited turns measurement into a task list. Once you know which twelve domains a model pulls from when answering questions in your category, you stop theorizing about authority and start working a target list.
Search triggered tells you which problem you have. If the model retrieved and still did not name you, that is a content and corroboration problem you can fix in weeks. If it answered from training knowledge without searching, that is an entity problem measured across model releases.
A weighted score, if you need one number
Executives want a single figure. If you build one, publish the formula:
Visibility Score = (Named rate × 50)
+ (Cited rate × 25)
+ (First-position rate × 15)
+ (Positive sentiment rate × 10)The weights are a judgment, not a discovered truth. Say so in the footnote. A composite score with an undisclosed formula is the same genre of thing as an invented ranking factor.
Tooling
| Tool | Best for | Rough entry price |
|---|---|---|
| Profound | Enterprise depth, largest datasets | Enterprise |
| Semrush AI Visibility Toolkit | Breadth of model coverage, solid infrastructure | ~$99/mo per domain |
| Ahrefs Brand Radar | Teams already on Ahrefs | Add-on |
| Peec AI | Mid-market, clean interface | Mid |
| Otterly.AI | Straightforward tracking | Low to mid |
| Scrunch | Bot crawl visibility alongside citations | Mid |
| LLM Pulse | Cheapest credible entry point | ~€49/mo |
| SE Visible | Mid-market alternative | ~$99/mo |
Three things to check before buying any of them.
- Collection method. API or consumer interface? This is the question most vendors answer vaguely and it determines what you are actually measuring.
- Runs per prompt. A tool checking each prompt once a week is giving you a sample size of one per period. The margin of error table applies to tools exactly as it applies to manual testing.
- Export. If you cannot get raw observations out, you cannot verify the numbers or compute your own confidence intervals.
Run two tools if the budget allows. One for depth, one as a sanity check. When they disagree materially, that disagreement is information about methodology, and it is worth understanding before you report either number.
Manual testing still has a place, and it is not measurement. It is for reading the actual answers: how you are described, what the tone is, which competitor is getting the better write-up. Tools give you the distribution. Reading twenty answers yourself gives you the texture, and the texture is usually where the insight lives.
Metrics that matter, and vanity ones
The blended average deserves a warning. Studies consistently find 80 to 91% of cited URLs appearing on only one engine, and one analysis found 62% brand disagreement across ChatGPT, AI Mode and AI Overviews. Even Google’s own two surfaces cite the same URLs only about 13.7% of the time while reaching similar conclusions 86% of the time.
Averaging across platforms with that little overlap produces a number that describes no platform anyone actually uses.
Always report per platform.
The reporting framework
What a defensible monthly report contains, in order:
“34% (±4.3)”. The interval is not pedantry, it is what stops a 3 point wobble being sold as progress.
Six rows. No average.
A line, not two bars.
Your named rate against the top three competitors on the identical prompt set.
Which domains the models drew on, and what changed since last period. This is the section clients find most useful, because it is the only one that names specific actions.
Anything the models say about you that is wrong, flagged with the source producing the error.
Prompt set version, number of runs, collection method, date range. Boring, and the thing that makes everything above verifiable.
That last item is the tell. A report without a method note cannot be checked, and a number that cannot be checked is a claim rather than a measurement.
Seven ways this goes wrong
| The failure | Why it matters | The fix |
|---|---|---|
| Sample too small | ±20 points at n=20, so any conclusion is unsupported | 450+ observations per platform |
| Prompt set built on strengths | Measures confidence, not visibility | Include prompts you expect to lose |
| Set changed mid-program | Period comparison becomes invalid | Freeze and version it |
| Testing while logged in | Personalization contaminates every observation | Clean sessions, always |
| Blending platforms into one average | Hides platform-specific reality behind a meaningless mean | Report per platform |
| Reporting small movements as results | Noise presented as progress | Only report movement exceeding your margin of error |
| Referral traffic as the sole metric | Misses most of the influence, since naming outpaces citing | Named rate first, traffic third |
Frequently asked questions
How do I measure AI visibility?
How many prompts do I need?
Why do AI answers change every time I ask?
Can I just check manually?
What is the difference between being named and being cited?
Should I use a tool or build my own tracking?
How often should I measure?
How long before I see change?
Can I compare my AI visibility to a competitor’s?
Want this run properly?
We baseline a prompt set before any work starts, re-run it monthly across six platforms, and report named rate with confidence intervals, per platform, against your competitors.
The method note is in every report. You can check our working.