How to measure AI visibility: a complete methodology

Ask ChatGPT the same question twice and you may get two different sets of businesses named. Ask it tomorrow and you may get a third.

That single fact breaks every measurement habit the search industry spent twenty years building. There is no position to check, no tracked keyword to look up, no number that stays still long enough to put in a report.

What there is instead is a distribution, and a well-understood way to estimate one. This page is the methodology: how to build a prompt set, how many observations you genuinely need, how to avoid contaminating your own data, and how to report numbers that survive someone checking your work.

Fair warning on one finding below: by these standards, a lot of published AI visibility research, including a lot of agency reporting, does not have enough data to support what it claims.

Last updated · Part of our knowledge base

01Not rank tracking

Why this is not rank tracking

SEO rank trackingDeterministic
Run 1Position #3
Run 2Position #3
Run 3Position #3
One check tells you
The answer
AI visibilityProbabilistic
Run 1Brand ABrand BYou
Run 2Brand CBrand ABrand D
Run 3YouBrand E
One check tells you
Almost nothing. It is one draw from a distribution
Illustrative runs. On repeat AI Overview queries, cited content shifted roughly 70% of the time
SEO rank trackingAI visibility
Same query, same daySame result, give or take personalizationFrequently different results
What you are measuringA positionA probability
One check tells youThe answerAlmost nothing
MethodLook it upSample it, repeatedly
StabilityChanges when the index changesChanges between runs, for no external reason

The scale of the instability is not marginal. Analysis of AI Overviews found the cited content shifting roughly 70% of the time on repeat queries. Not over months. On a re-run.

So a single check is not a small sample. It is one draw from a distribution you have not characterized, and treating it as a result is how agencies end up reporting improvements that are noise and missing declines that are real. A free tool like our AI visibility checker is useful exactly as that: a first snapshot, not a measurement.

Everything that follows exists to solve that problem properly.

02Three measures

Three different things, and most reports confuse them

These get used interchangeably. They are not the same, they move independently, and conflating them produces reports that mislead in both directions.

01Lead with this
Named rate

Does the model say our name?

Captured by
Prompt testing
Shows in your analytics
No
02
Cited rate

Does it link us as a source?

Captured by
Prompt testing, logging citations
Shows in your analytics
No
03
Referral traffic

Did anyone click through?

Captured by
GA4
Shows in your analytics
Yes
ChatGPT: how often a brand is named vs linked
Named
3.2×
Cited
1×
MeasureThe question it answersHow you capture itThe trap
Named rateDoes the model say our name?Prompt testingInvisible in analytics. This is the primary metric
Cited rateDoes it link us as a source?Prompt testing, logging citationsMuch rarer than being named
Referral trafficDid anyone click through?GA4Systematically understates influence

The gap between the first two is large and platform-specific. ChatGPT mentions brands roughly 3.2 times more often than it cites them with a link. Perplexity, built around attribution, cites on nearly every answer, averaging around 22 sources per response, against roughly six for Claude and seven for Copilot.

Which produces a reporting trap worth naming. A brand with a strong named rate and thin referral traffic is not failing. It is being recommended in conversations that never touch its analytics. Report only on GA4 sessions and you will conclude the program is not working while it is working.

The reverse trap is rarer but real: healthy referral numbers from a handful of prompts can mask absence across the rest of the category.

Measure all three. Lead with named rate.

03The prompt set

Building the prompt set

Your prompt set is the instrument. A biased instrument produces precise measurements of the wrong thing, and the most common bias is flattery: people build prompt sets from queries where they already do well.

A prompt set of 150 to 300Share of prompts
Category discovery
Problem-led
Comparison
Qualification
Branded
Branded prompts measure accuracy, not presence: wrong pricing, stale services, the wrong country

Composition

A usable set runs 150 to 300 prompts, distributed roughly like this:

CategoryShareExample shape
Category discovery30%“Best [service] in [city]”, “Top [category] providers”
Problem-led25%“How do I fix [problem the customer has]”
Comparison20%“[Competitor] alternatives”, “[A] vs [B]”
Qualification15%“How much does [service] cost”, “How to choose a [provider]”
Branded10%“What does [your brand] do”, “Is [your brand] any good”

That last row is small and disproportionately important. Branded prompts measure accuracy rather than presence, and they are where you discover the model has your pricing wrong, your service list outdated, or your location in the wrong country.

Four rules

  • Write them as a customer would type them. Not as keywords. “Best accountant in Vancouver for a small construction business” is a real prompt. “vancouver accountant” is a keyword someone pasted.
  • Include prompts you expect to lose. A set built only on your strengths measures your confidence, not your visibility.
  • Cover the sub-questions. On surfaces using query fan-out, the engine invents its own sub-queries. Your set should include the questions it will generate, not only the ones your customer types.
  • Freeze it. Once baselined, the set does not change. Adding prompts mid-program makes period comparison invalid, and it is the easiest way to accidentally manufacture an improvement. Version the set, and if you must revise it, start a new baseline and say so in the report.
04How much data

How much data you actually need

This is the part nobody publishes, and it determines whether anything else you do is meaningful.

You are estimating a proportion: the share of prompt-runs where your brand is named. That is a standard statistical problem with a standard answer.

Margin of error, 95% confidence
True rate 50%True rate 30%True rate 10%
USABLE±0±5±10±15±20±25±301020501002005001,000OBSERVATIONS (LOG SCALE)±20.1±12.7±9.0±6.4±5.2±4.0±2.8450 obs · ±4.2
Computed: 1.96 × √(p(1−p)/n). The margin is widest at 50% and narrows toward the extremes

The margin of error on your named rate

At 95% confidence, assuming a true named rate around 30%:

ObservationsMargin of errorWhat that means
20±20.1 pointsA measured 30% could truly be 10% or 50%
50±12.7 pointsStill too wide to act on
100±9.0 pointsDirectionally useful, not reportable as precise
200±6.4 pointsUsable
300±5.2 pointsSolid
500±4.0 pointsStrong
1,000±2.8 pointsResearch grade

Read the first row again. Twenty manual spot checks produce an estimate with a twenty point margin of error. You could measure 30% when the truth is 10%, and report a catastrophe as a modest result, or the reverse.

This is why “I asked ChatGPT a few times and we didn’t come up” is not a finding. It is also why a great deal of published AI visibility content, built on small manual samples, should be read with more suspicion than it usually gets.

Getting to a usable number

Observations are prompts multiplied by runs, per platform:

Prompt setRunsObservations per platform
1001100
1503450
2003600
3003900

150 prompts run three times gives you 450 observations and a margin of error around ±4.3 points. That is a defensible number, and it is the minimum we would put in front of a client.

One statistical note worth knowing: the margin is widest when the true rate sits near 50% and narrows toward the extremes. If your named rate is 5% or 90%, you need less data to pin it down than the table suggests. Most brands starting out are well below 30%, which works slightly in your favour at baseline and against you as you improve.

05Detecting change

Detecting change is harder than measuring a level

Here is where most reporting quietly breaks.

Measuring where you stand is one problem. Proving you moved is a harder one, because now two noisy estimates have to be far enough apart that the difference cannot be explained by sampling.

Named rate, two periods, 95% intervals
Period 1Period 2
30% → 35% 450 obs per periodOverlap: could be noise
30% ±4.235% ±4.4
30% → 45% 450 obs per periodClear: a real change
30% ±4.245% ±4.6
30% → 35% 1,374 obs per periodClear: a real change
30% ±2.435% ±2.5
15%25%35%45%55%
Computed intervals. A 5 point move needs about 1,374 observations per period before the two stop overlapping

Sample size needed per period, at 95% confidence with 80% power:

The change you want to proveObservations needed per period
30% to 45% (15 points)160
30% to 40% (10 points)354
10% to 20% (10 points)197
30% to 35% (5 points)1,374
30% to 32% (2 points)8,391

That bottom row is the important one. Small improvements are effectively unprovable at normal sampling volumes. A two point shift needs over eight thousand observations per period to distinguish from noise.

Three practical consequences.

  • Do not report small movements as results. If your measurement carries ±5 points and the number moved 3, nothing happened that you can demonstrate.
  • Set expectations around double-digit shifts. Below ten points, you are usually reporting noise with a narrative attached.
  • Use trend lines, not point comparisons. Six months of monthly measurement showing consistent upward drift is more persuasive than any two individual months, and it is more honest, because it does not depend on a single comparison surviving scrutiny.

There is a commercial temptation here that is worth naming. Small sample sizes produce big swings, and big swings make impressive slides.

Resisting that is the difference between a measurement practice and a sales tool.

06Running it cleanly

Running it cleanly

Your own setup can contaminate the result before you record a single number.

01
Logged-in accounts

Memory and chat history personalize answers to you

FIXClean, logged-out sessions every time
02
Conversation memory

Earlier prompts in the same thread steer later ones

FIXOne prompt per session, no follow-ups
03
Location

Local results vary by IP

FIXFix location deliberately, document it
04
Your own browsing

Personalization learns what you look at

FIXNever test from the machine you research on
05
Time of day

Index freshness varies

FIXRandomize or fix the schedule, then keep it
06Matters most
API vs consumer product

Different systems with different retrieval

FIXChoose one, disclose it, hold it constant

That last source deserves its own paragraph, because most tooling depends on it.

The API is not the product. Testing through an API may not reflect what a person using the consumer app actually sees. Retrieval behaviour, system prompts and grounding can all differ. Some tools measure via API because it is cheap and scalable, then report the result as consumer visibility.

That is not necessarily wrong, but it is a methodological choice, and it should be disclosed. When evaluating any AI visibility tool, ask how it collects data. If the answer is vague, treat the numbers as directional.

The principle: whatever you choose, hold it constant.

Changing collection method between periods invalidates the comparison more thoroughly than any sample size problem.

07What to record

What to record

A binary named or not-named throws away most of the signal. Capture this per observation:

Observation #0147Sample record
Platform: ChatGPTPrompt: P-032Run: 2 of 3
NamedYes
CitedYes · example.com/services
Position2nd
SentimentCaveated: “solid but expensive”
AccuracyPartially wrong: outdated pricing
Competitors namedBrand A, Brand C
Sources cited4 domains: directory, review site, listicle, own site
Search triggeredYes
Highlighted: the two fields that do the most work
FieldValuesWhy it matters
NamedYes / NoThe primary metric
CitedYes / No, with URLSeparate from naming, and rarer
Position1st, 2nd, 3rd, laterBeing named first carries more weight than being named
SentimentPositive / Neutral / Negative / Caveated“X is solid but expensive” is not a clean win
AccuracyCorrect / Partially wrong / WrongBranded prompts especially
Competitors namedListYour actual competitive set, often not who you assumed
Sources citedDomain listThe most actionable field on this list
Search triggeredYes / NoWhether the model retrieved or answered from training

Two of these do more work than the rest.

Sources cited turns measurement into a task list. Once you know which twelve domains a model pulls from when answering questions in your category, you stop theorizing about authority and start working a target list.

Search triggered tells you which problem you have. If the model retrieved and still did not name you, that is a content and corroboration problem you can fix in weeks. If it answered from training knowledge without searching, that is an entity problem measured across model releases.

A weighted score, if you need one number

Executives want a single figure. If you build one, publish the formula:

Visibility Score = (Named rate × 50)
                 + (Cited rate × 25)
                 + (First-position rate × 15)
                 + (Positive sentiment rate × 10)

The weights are a judgment, not a discovered truth. Say so in the footnote. A composite score with an undisclosed formula is the same genre of thing as an invented ranking factor.

08Tooling

Tooling

ToolBest forRough entry price
ProfoundEnterprise depth, largest datasetsEnterprise
Semrush AI Visibility ToolkitBreadth of model coverage, solid infrastructure~$99/mo per domain
Ahrefs Brand RadarTeams already on AhrefsAdd-on
Peec AIMid-market, clean interfaceMid
Otterly.AIStraightforward trackingLow to mid
ScrunchBot crawl visibility alongside citationsMid
LLM PulseCheapest credible entry point~€49/mo
SE VisibleMid-market alternative~$99/mo

Three things to check before buying any of them.

  • Collection method. API or consumer interface? This is the question most vendors answer vaguely and it determines what you are actually measuring.
  • Runs per prompt. A tool checking each prompt once a week is giving you a sample size of one per period. The margin of error table applies to tools exactly as it applies to manual testing.
  • Export. If you cannot get raw observations out, you cannot verify the numbers or compute your own confidence intervals.

Run two tools if the budget allows. One for depth, one as a sanity check. When they disagree materially, that disagreement is information about methodology, and it is worth understanding before you report either number.

Manual testing still has a place, and it is not measurement. It is for reading the actual answers: how you are described, what the tone is, which competitor is getting the better write-up. Tools give you the distribution. Reading twenty answers yourself gives you the texture, and the texture is usually where the insight lives.

09Metrics and vanity

Metrics that matter, and vanity ones

Report these
Be careful with these
Named rate, with confidence interval
A single composite “AI score” with no published formula
Named rate by platform
A blended cross-platform average, which hides everything
Competitor named rate, same prompt set
Cherry-picked screenshots of good answers
Citation source list, and changes to it
Total mentions with no denominator
Sentiment and accuracy on branded prompts
Referral traffic alone
Trend over three or more periods
Month-on-month comparison of small movements
Crawler access status
“Impressions” invented by extrapolation

The blended average deserves a warning. Studies consistently find 80 to 91% of cited URLs appearing on only one engine, and one analysis found 62% brand disagreement across ChatGPT, AI Mode and AI Overviews. Even Google’s own two surfaces cite the same URLs only about 13.7% of the time while reaching similar conclusions 86% of the time.

Averaging across platforms with that little overlap produces a number that describes no platform anyone actually uses.

Always report per platform.

10The report

The reporting framework

What a defensible monthly report contains, in order:

Monthly AI visibility reportIn this order
01
Headline named rate, with its confidence interval

“34% (±4.3)”. The interval is not pedantry, it is what stops a 3 point wobble being sold as progress.

02
Per-platform breakdown

Six rows. No average.

03
Trend, minimum three periods

A line, not two bars.

04
Competitive position

Your named rate against the top three competitors on the identical prompt set.

05
Citation sources

Which domains the models drew on, and what changed since last period. This is the section clients find most useful, because it is the only one that names specific actions.

06
Accuracy and sentiment

Anything the models say about you that is wrong, flagged with the source producing the error.

07
Method note

Prompt set version, number of runs, collection method, date range. Boring, and the thing that makes everything above verifiable.

That last item is the tell. A report without a method note cannot be checked, and a number that cannot be checked is a claim rather than a measurement.

Seven ways this goes wrong

The failureWhy it mattersThe fix
Sample too small±20 points at n=20, so any conclusion is unsupported450+ observations per platform
Prompt set built on strengthsMeasures confidence, not visibilityInclude prompts you expect to lose
Set changed mid-programPeriod comparison becomes invalidFreeze and version it
Testing while logged inPersonalization contaminates every observationClean sessions, always
Blending platforms into one averageHides platform-specific reality behind a meaningless meanReport per platform
Reporting small movements as resultsNoise presented as progressOnly report movement exceeding your margin of error
Referral traffic as the sole metricMisses most of the influence, since naming outpaces citingNamed rate first, traffic third
11FAQ

Frequently asked questions

How do I measure AI visibility?

Build a fixed prompt set of 150 to 300 buying questions, run it repeatedly across each AI platform in clean logged-out sessions, and record whether your brand is named, cited, in what position, with what sentiment, and which sources the model used. Report the named rate per platform with a confidence interval, and track the trend across at least three periods.

How many prompts do I need?

150 to 300, run at least three times each. That produces 450 to 900 observations per platform, giving a margin of error of roughly 4 to 5 points at 95% confidence. Fewer than 100 observations leaves a margin too wide to support any conclusion.

Why do AI answers change every time I ask?

These systems are non-deterministic and retrieval is re-run per query. Analysis of AI Overviews found cited content shifting around 70% on repeat queries. This is why a single check is not a measurement and why sampling repeatedly is the only valid approach.

Can I just check manually?

For texture, yes, and it is worth doing. For measurement, no. Twenty manual checks carry a margin of error of about 20 points, which is wide enough that the result cannot distinguish a serious problem from a decent position. The AI visibility checker runs a full prompt set instead.

What is the difference between being named and being cited?

Named means the model says your brand name. Cited means it links you as a source. Being named is far more common: ChatGPT mentions brands roughly 3.2 times more often than it cites them. Named rate is the primary metric because it captures influence that never reaches your analytics.

Should I use a tool or build my own tracking?

A tool, for most businesses. Check three things before buying: whether it collects via API or consumer interface, how many runs it performs per prompt per period, and whether you can export raw observations. A tool running each prompt once a week has the same sample size problem as manual checking.

How often should I measure?

Monthly for reporting, weekly for priority prompts if the budget allows. More frequent measurement mostly buys you earlier warning rather than better precision, since precision comes from observations per period rather than periods per year.

How long before I see change?

Crawler access problems can resolve in days. Entity and citation source work typically produces measurable movement in two to four months. Remember that you need roughly 350 observations per period to prove a 10 point shift, so build the measurement capacity before you need to demonstrate the result.

Can I compare my AI visibility to a competitor’s?

Yes, and it is one of the more useful things you can do. Run the identical prompt set and record every brand named in every answer. That gives you competitor named rates on exactly the same instrument, which is a cleaner comparison than anything available in traditional SEO.
Free AI visibility audit

Want this run properly?

We baseline a prompt set before any work starts, re-run it monthly across six platforms, and report named rate with confidence intervals, per platform, against your competitors.

The method note is in every report. You can check our working.