There is no universal “good” AI visibility score—and one scan cannot establish your visibility. A 70 is not automatically strong, and a 15 is not automatically weak. In practical terms, a good score places you at or above the matched peer median on priority non-branded prompts, closes the gap to the leader over time, and maintains accurate, repeatable results.
A defensible answer requires six fixed boundaries: the buyer-prompt panel, AI products and modes, market and language, measurement window, scoring method, and weighting contract. Predeclare how prompts, intent clusters, products, markets, and repeated runs contribute to an aggregate. Then compare the brand with a frozen peer set, repeat the runs, and track the gap across comparable windows.
A run is one answer from one AI product for one fixed question. An iteration repeats that prompt/product setup so variation is visible. A measurement window groups the repeated runs used to compare one period with another.
Good AI visibility means winning a useful share of relevant recommendations versus a defined peer set—with accurate facts and results stable enough to reproduce.
That definition deliberately separates visibility from business impact. A score does not prove traffic, leads, or revenue. It tells you what happened in a measured panel and where to investigate.
First ask what the score measures
“AI visibility score” is not standardized. One platform may report simple mention rate. Another may weight answer position, citations, sentiment, or performance relative to competitors. Two dashboards can both show 50 while describing different outcomes.
Define the rules before scoring: which prompts count as recommendation opportunities, what language counts as an endorsement, how brand aliases are matched, what qualifies as an attributable citation, and what constitutes a material factual error. Apply those rules consistently to every brand and report any responses that required manual review.
| Metric | Transparent formula | Question answered |
|---|---|---|
| Mention rate | Valid answers substantively naming the brand ÷ all valid answers | How often was the brand present? |
| Recommendation rate | Valid answers to prompts designated recommendation-eligible before collection that recommend the brand ÷ all valid answers to those prompts | How often was it actually recommended? |
| First-position rate among ordered recommendations | Valid ordered recommendation answers placing the brand first ÷ all valid ordered recommendation answers | How often did it lead a shortlist? |
| Tracked-brand share of voice | Brand mentions ÷ mentions of all tracked brands | What portion of observed competitive attention did it receive? |
| Attributed citation rate | Citation-enabled answers containing a citation mapped to a material claim about the brand ÷ valid citation-enabled answers containing such a claim | How often was a material brand claim directly supported? |
| Citation share | Relevant citations to the brand's domain ÷ all relevant citations under a predeclared relevance rule | What share of observed source support did it own? |
| Conditional factual-accuracy rate | Adjudicated appearances with all material checkable claims verified ÷ appearances containing at least one material checkable claim that could be adjudicated | Was verifiable visibility correct? Report unverified appearances separately. |
Separate branded and non-branded prompts
“What does Acme do?” tests whether an AI system can resolve Acme. “Best emergency plumber near me” tests competitive discovery. Blending both can make a brand look visible because it answers questions that already contain the brand name.
Also publish provider-level results. A blended 30% can hide 60% visibility in one product and zero in another.
Compare every brand under identical conditions
A standalone score says only how one brand performed in one test. A competitive benchmark asks a more useful question: when the same buyer asks the same questions in the same AI products, how often are we selected versus the alternatives?
Choose the comparison peer group before calculating results. Include businesses or substitutes that serve the same buyer need, geography, and category—not merely the competitors your SEO team already tracks. If the scan reveals additional recurring AI-shortlist brands, treat them as exploratory and add them to the next frozen benchmark window.
Compare gaps, not just ranks
“We rank second” can hide a tiny difference or a structural gap. For every priority metric, calculate both:
- Gap to leader: target rate minus the highest competitor rate.
- Gap to peer median: target rate minus the median rate among predeclared competitors, excluding the target.
- Prompt wins: priority prompts where the target beats or ties the leader across repeated runs.
- Provider gaps: the same comparison inside ChatGPT, Gemini, Perplexity, Claude, or another measured surface—not only as a blended total.
| Comparison | What to inspect | What the gap can reveal |
|---|---|---|
| Overall peer group | Mention, recommendation, first-choice, citation, and accuracy rates | Whether the brand is absent, present, or genuinely selected |
| Prompt cluster | Category, problem, comparison, trust, price, urgency, and location | Where competitors own the buyer journey |
| AI product | Provider-level rates and repeated-run ranges | Whether a blended score hides a provider-specific weakness |
| Recommendation evidence | Reasons, claims, and recurring source domains supporting each winner | Which evidence difference is worth testing next |
| Accuracy | Incorrect services, locations, prices, credentials, or brand identity | Whether higher visibility is helping or spreading bad facts |
For implementation details, use the AI visibility audit checklist. To identify the right cohort, see why AI competitors may differ from Google competitors.
A worked benchmark without invented industry averages
The fictional example uses 40 fixed non-branded buyer prompts, three AI products, and three runs per prompt: 360 scheduled responses in one shared panel. After nine provider failures, 351 responses remained valid. Every tracked brand was scored against those same 351 responses. The 300-answer recommendation denominator was determined from prompt eligibility before examining brand outcomes.
| Brand | Mention | Recommendation | First position | Conditional accuracy |
|---|---|---|---|---|
| Northstar | 119/351 (34%) | 72/300 (24%) | 33/300 (11%) | 115/119 (97%) |
| Target brand | 95/351 (27%) | 45/300 (15%) | 12/300 (4%) | 78/95 (82%) |
| Harbor | 77/351 (22%) | 42/300 (14%) | 15/300 (5%) | 72/77 (94%) |
| Cascade | 42/351 (12%) | 21/300 (7%) | 6/300 (2%) | 40/42 (95%) |
The target is not simply “27% visible.” It is second in mention rate, nearly tied with Harbor in recommendations, behind Harbor on first position, and weak on conditional factual accuracy. The action is not “raise the score.” It is to investigate which high-intent prompts produce mentions without endorsement and which facts are wrong.
Report the numerator, denominator, failures, and results by provider and prompt cluster. These estimates describe the tested panel—not every possible buyer question. If you publish uncertainty ranges, account for repeated answers from the same prompts rather than treating every generated response as fully independent.
A scan is a snapshot; visibility is a pattern
A single scan captures one set of generated answers at one moment. It cannot tell you whether the result is typical, whether a competitor usually wins, or whether a change you shipped improved anything.
AI answers vary because model versions, retrieval systems, available sources, location context, product modes, and variable generation can change. Even with the same prompt, one run may mention the brand and the next may omit it. That is why a serious benchmark needs two time scales:
What to keep fixed—and what to log when it changes
| Keep fixed when possible | Log for every window | Why it matters |
|---|---|---|
| Prompt panel and intent labels | Added, removed, or rewritten prompts | Prompt changes can move the score without any market change |
| Competitor cohort and aliases | New entrants, exits, mergers, rebrands | Share of voice changes when the denominator changes |
| Products, modes, language, and location | Model/version when exposed, browsing and account state | Provider changes can create a break that makes before-and-after results hard to compare |
| Run count and scoring rules | Failures, unavailable citations, classification changes | Different denominators make adjacent windows incomparable |
| Priority prompt clusters | Site changes, profile updates, PR, review, and content releases | Creates hypotheses about movement without claiming causation |
For the full repetition methodology, see why one AI visibility scan is not enough.
Read the delta together with the competitor gap
- Your observed rate rises and the leader gap shrinks: a smaller measured gap, subject to cluster-aware uncertainty and replication across matched windows.
- Your observed rate rises but competitors rise faster: absolute rate increased while relative position weakened.
- Your observed rate is flat while peers fall: relative position rose even though the dashboard number did not.
- Every brand jumps at once: suspect a model, provider, prompt, or scoring change before crediting your campaign.
- Only one product moves: inspect that provider's answers and source mix instead of declaring a market-wide win.
Four ways to read the benchmark
Set targets against gaps, not arbitrary numbers
A useful target names the peer group, metric, segment, period, and quality guardrail. For example:
Over the next eight weeks, close half the recommendation-rate gap to the peer-group leader for five priority non-branded purchase prompts, while keeping factual accuracy above the current baseline.
That target can be measured. “Reach an AI visibility score of 70” cannot be interpreted unless the scoring system and baseline remain unchanged.
- Use the peer median, calculated from predeclared competitors excluding the target, to judge whether the brand has entered the competitive cluster.
- Use the leader gap to quantify the next attainable measured difference.
- Use prompt clusters to keep observed movement tied to real buyer intent.
- Use run-level distributions and cluster-aware uncertainty to distinguish observed movement from ordinary within-panel variation; do not rely on the repeated-run range alone.
- Use accuracy as a guardrail so more visibility does not mean more misinformation.
Do not compare scores across vendors unless formulas, prompts, products, geography, run rules, and denominators match. For planning work after the baseline, see how to set AI visibility targets you can budget for and why one scan is not enough.
Run a directional preview
Spot possible gaps worth validating—not a stable benchmark. The free check is one quick sample; a defensible comparison requires repeated runs of a frozen brand and competitor peer group over time.
For an example of the evidence behind a full analysis, see the published redacted benchmark.
Run the free directional checkDiscuss a repeated-run benchmarkFrequently asked questions
What is a good AI visibility score?
There is no universal threshold. A useful result is competitive for a fixed non-branded prompt panel and matched peer group, accurate when the brand appears, and reproducible across repeated runs.
Is an AI visibility score of 50 good?
Not enough information. You need the formula, denominator, prompts, provider mix, cohort, and repeated-run variation. A 50 in one tool may not equal a 50 in another.
How should I benchmark against competitors?
Use identical prompts, products, markets, run counts, brand matching, and outcome definitions for every brand. Compare mentions, recommendations, first choice, tracked-brand share of voice, citations, accuracy, and stability separately.
How often should I measure?
Repeat runs for the baseline, then re-run the same panel consistently and after meaningful changes. Document model, product, prompt, and source changes before attributing movement to your work.
Why is one AI visibility scan not enough?
One scan is a snapshot of variable answers at one moment. It cannot show recurrence, competitor movement, or whether a later change improved your relative position. Use repeated runs within each window and repeat the matched competitor panel over time.