How we measure AI visibility.
Every number on this site — in a case study, an article, or a client report — comes from the rules on this page. It is written to be argued with: if a figure here does not follow these rules, that is a defect and we want to hear about it.
Version 2026-08-28.1
The metrics, and what each one is not
These names mean the same thing everywhere on this site. When a page shows a figure, it names which of these metrics it is; where a case study reports a metric other than AI Visibility, the case study says so rather than relabelling it.
AI Visibility
Also called: mention rate
The share of relevant answers, across repeated runs of a fixed prompt panel, in which the assistant states the business name at all.
- Denominator:
- relevant answers returned for the panel in the window
- Does not mean:
- that the business was recommended — being named in a list of many is still a mention
Recommendation Rate
The share of relevant answers in which the business is put forward as an option a buyer should consider, rather than merely referenced.
- Denominator:
- relevant answers returned for the panel in the window
- Does not mean:
- that the buyer saw the answer, clicked, or converted
Share of Voice
The business's share of all brand names appearing across the relevant answers — how much of the named field belongs to it.
- Denominator:
- all brand mentions across those answers, the business included
- Does not mean:
- the same thing as AI Visibility — a small field can produce a high share of voice at a low mention rate
Citation Rate
The share of relevant answers that display at least one source URL on a domain the business owns.
- Denominator:
- relevant answers returned for the panel in the window
- Does not mean:
- that the cited page caused the answer, or that uncited sources were unused — and it is not comparable across providers that differ in how often they display sources at all
Citation Share
Of every source URL displayed across the relevant answers, the proportion on domains the business owns.
- Denominator:
- all source URLs displayed across those answers
- Does not mean:
- a measure of cause — it maps where the observed public evidence sits, not why a brand was selected
Source-Type Share
The same citation counts grouped by the kind of source — owned pages, review platforms, directories, editorial, community — so the gap points at a type of work.
- Denominator:
- all source URLs displayed across those answers
- Does not mean:
- a ranking of which source type an engine prefers in general
Brand-Accuracy Rate
The share of relevant answers describing the business whose stated facts — services, area, hours, ownership, claims — are correct.
- Denominator:
- relevant answers that described the business at all
- Does not mean:
- that the remaining answers were hostile — most inaccuracy is stale, not adversarial
What actually gets measured
- Prompt panel
- A fixed set of buyer questions written the way a customer would ask them, agreed with the client before the baseline run. It covers the categories a buyer moves through — discovery ("who does X near me"), comparison, qualification, and problem-first phrasing — and it does not contain the brand name unless the prompt is deliberately a brand-knowledge probe. Once a baseline exists, the panel is frozen: adding or rewording a prompt starts a new series.
- Cell — the unit of measurement
- One prompt × one provider × one answering mode × one market. Every number on this site aggregates cells, never raw responses, because a provider that returns 40 responses is not 40 times more informative than one that returns one. "Mode" matters as much as provider: the same model with retrieval enabled and disabled is two different measurement conditions and is never pooled.
- Repeats
- Each cell is run repeatedly within the measurement window rather than once. Repeats are what turn a single answer into an estimate; they are not independent trials (see uncertainty below). Reporting thresholds are stated below.
- Providers
- ChatGPT, Claude, Perplexity, and Google AI Overview by default; additional surfaces by agreement. Provider and model version are recorded with every run, because a provider update is a change to the measurement instrument, not to the business.
- Market
- The location and language the prompts are issued for. Local-intent answers differ sharply by market, so a market is part of the cell identity and results are never averaged across markets without saying so.
- Window
- A named start and end date. Every published figure belongs to a window, and a window is what a later window is compared against.
What counts, and what is excluded
Most disagreements about an AI-visibility number are really disagreements about the denominator. Ours is fixed in advance.
Relevant answers are the denominator
A response counts toward a metric only if the assistant actually answered the buyer question. Refusals, "I cannot browse", topic redirects, and answers about a different category are excluded from the denominator rather than counted as a miss — counting them as misses would let a provider outage look like a visibility drop.
Failed runs are recorded, not silently dropped
Transport errors, rate limits, and truncated responses are logged with a reason and excluded. A cell whose failure rate is high enough to bias the estimate is reported as incomplete instead of being reported as a number.
Named is named — not implied
A brand counts as named when the answer states the business name. Category descriptions that fit the business without naming it do not count. Near-miss variants (misspellings, an old trading name) are counted, and flagged separately as an entity-consistency finding.
Mention is not recommendation
Being named in passing and being put forward as an option are separate metrics, coded separately. A rising mention rate with a flat recommendation rate is a real and common pattern, and collapsing the two hides it.
How we decide a change is real
This is the part most AI-visibility reporting gets wrong, including some earlier copy on this site. The correction matters more than the convenience.
- 1
An observed range is not a confidence interval
The lowest and highest values seen across runs describe what happened in this sample. They are not an interval estimate for the underlying rate, and they get wider — not narrower — the more you measure. Where this site shows a spread, it says which of the two it is.
- 2
Repeated answers are not independent coin flips
Runs of the same prompt on the same provider share the prompt, the retrieval pool, and the model version, so their outcomes are correlated. Treating each response as an independent Bernoulli trial and shrinking the interval by √N understates the true uncertainty, sometimes by a lot. The effective sample size is closer to the number of prompt-provider cells than to the number of responses.
- 3
Intervals come from a cluster bootstrap
Uncertainty is estimated by resampling whole prompts (with their repeats attached) rather than individual responses, so the correlation inside a cell is carried into the interval instead of being assumed away. Reported at 95% unless stated otherwise.
- 4
Period-over-period changes are paired
A month-to-month difference is computed on the same prompt panel measured twice, and the interval is bootstrapped on the paired per-prompt differences. Comparing two independently-estimated percentages would throw away the pairing that makes a small real change detectable.
- 5
Minimum reporting thresholds
A per-provider figure is reported when the panel has at least 15 relevant prompts and at least 5 valid repeats per cell. Below that the result is shown as directional and labelled as such. A change is described as a change only when the paired interval excludes zero; otherwise it is reported as "no detectable change", which is a finding, not a failure.
When two numbers may be compared
Two AI-visibility figures are comparable only when both were produced under:
- the same prompt panel, word for word
- the same providers and answering modes
- the same market and language
- the same competitor set
- the same relevance and coding rules
- a stated measurement window for each side
Which is why cross-vendor scores rarely compare
A score from another tool was almost certainly produced on a different panel, a different provider mix, and a different set of coding rules. Two such numbers can differ by a factor of several without either being wrong. Ask any vendor — including us — for the six items on the left before treating their number as a benchmark. Our article on what a good AI visibility score is works through the consequences.
What this method cannot tell you
A method is only credible if its boundaries are published alongside it. These are ours.
It does not establish why a brand was chosen
Retrieval, internal model knowledge, ranking, synthesis, and the citations a product displays are separate steps that differ by provider and mode. A citation shown next to an answer is evidence associated with that answer, not proof that the cited page caused the recommendation.
It does not measure model internals
We observe the product surface a buyer sees. We do not have access to weights, retrieval indexes, or ranking signals, and we do not infer them from output patterns.
It does not attribute revenue
Visibility measurement counts answers, not leads. Where a client attributes leads to AI-assisted discovery, that is client-reported and labelled as such.
It does not generalise across categories
A panel describes one category in one market. Numbers from one engagement are not a benchmark for another, and we do not present them as one.
It cannot separate our work from everything else
An engagement runs in a live market. Competitors act, providers update, seasons turn. A before/after difference on a fixed panel is a measured association over that window, not an isolated treatment effect. Only a held-out comparison — matched competitors or untouched topics measured on the same panel — narrows that gap, and where we have one, we say so.
Exploratory analyses we have not yet published a study design for — such as the cross-engine signal rankings in our provider-comparison article — are labelled as exploratory on the page that shows them and are not covered by this methodology.
Changing this page
A rule change starts a new series
If a definition, a denominator rule, or a threshold changes, figures produced before and after are not pooled. The version stamp at the top of this page is what a report cites.
Corrections are published, not quietly edited
How we handle a wrong number after publication — and who reviews research before it ships — is set out in our editorial and evidence policy.
Revision history
- 2026-08-28.1
- Citation Rate's denominator corrected to all relevant answers, matching AI Visibility and Recommendation Rate. It previously read "relevant answers that displayed any sources at all", which contradicted this page's own definition of the metric and inflated the rate for providers whose interfaces rarely display sources. A Citation Rate produced under 2026-08-28 is not comparable with one produced under this version.
- 2026-08-28
- First publication of the measurement contract.
See the method run on your category
A free preview runs a short panel against your domain. The full engagement runs the design above.
Check my AI visibility