Measurement · Competitive benchmarking

What is a good AI visibility score?

A score is only useful when its formula is clear and every brand is measured against the same buyer questions, AI products, market, and time window.

Example dashboard52Not interpretable yet

Which prompts? Which providers? Mention or recommendation? Against which competitors? How stable were repeated runs?

The short answer

There is no universal “good” AI visibility score—and one scan cannot establish your visibility. A 70 is not automatically strong, and a 15 is not automatically weak. In practical terms, a good score places you at or above the matched peer median on priority non-branded prompts, closes the gap to the leader over time, and maintains accurate, repeatable results.

A defensible answer requires six fixed boundaries: the buyer-prompt panel, AI products and modes, market and language, measurement window, scoring method, and weighting contract. Predeclare how prompts, intent clusters, products, markets, and repeated runs contribute to an aggregate. Then compare the brand with a frozen peer set, repeat the runs, and track the gap across comparable windows.

A run is one answer from one AI product for one fixed question. An iteration repeats that prompt/product setup so variation is visible. A measurement window groups the repeated runs used to compare one period with another.

Good AI visibility means winning a useful share of relevant recommendations versus a defined peer set—with accurate facts and results stable enough to reproduce.

That definition deliberately separates visibility from business impact. A score does not prove traffic, leads, or revenue. It tells you what happened in a measured panel and where to investigate.

First ask what the score measures

“AI visibility score” is not standardized. One platform may report simple mention rate. Another may weight answer position, citations, sentiment, or performance relative to competitors. Two dashboards can both show 50 while describing different outcomes.

Define the rules before scoring: which prompts count as recommendation opportunities, what language counts as an endorsement, how brand aliases are matched, what qualifies as an attributable citation, and what constitutes a material factual error. Apply those rules consistently to every brand and report any responses that required manual review.

Keep these outcomes separate before creating any summary score.
MetricTransparent formulaQuestion answered
Mention rateValid answers substantively naming the brand ÷ all valid answersHow often was the brand present?
Recommendation rateValid answers to prompts designated recommendation-eligible before collection that recommend the brand ÷ all valid answers to those promptsHow often was it actually recommended?
First-position rate among ordered recommendationsValid ordered recommendation answers placing the brand first ÷ all valid ordered recommendation answersHow often did it lead a shortlist?
Tracked-brand share of voiceBrand mentions ÷ mentions of all tracked brandsWhat portion of observed competitive attention did it receive?
Attributed citation rateCitation-enabled answers containing a citation mapped to a material claim about the brand ÷ valid citation-enabled answers containing such a claimHow often was a material brand claim directly supported?
Citation shareRelevant citations to the brand's domain ÷ all relevant citations under a predeclared relevance ruleWhat share of observed source support did it own?
Conditional factual-accuracy rateAdjudicated appearances with all material checkable claims verified ÷ appearances containing at least one material checkable claim that could be adjudicatedWas verifiable visibility correct? Report unverified appearances separately.
Citation unavailable is not zero. If a product or mode does not expose citations, mark that field N/A. Failed responses are not negative outcomes: report failures separately and calculate rates on valid responses while preserving fixed prompt/product weights—or mark the aggregate incomparable when missing cells cannot be handled consistently.

Separate branded and non-branded prompts

“What does Acme do?” tests whether an AI system can resolve Acme. “Best emergency plumber near me” tests competitive discovery. Blending both can make a brand look visible because it answers questions that already contain the brand name.

Also publish provider-level results. A blended 30% can hide 60% visibility in one product and zero in another.

Compare every brand under identical conditions

A standalone score says only how one brand performed in one test. A competitive benchmark asks a more useful question: when the same buyer asks the same questions in the same AI products, how often are we selected versus the alternatives?

Choose the comparison peer group before calculating results. Include businesses or substitutes that serve the same buyer need, geography, and category—not merely the competitors your SEO team already tracks. If the scan reveals additional recurring AI-shortlist brands, treat them as exploratory and add them to the next frozen benchmark window.

Framework moving from an uninterpretable score to a matched competitor benchmark using fixed prompts, products, repeated runs, and separate outcome metrics
A score becomes interpretable only after the test panel and metric rules are fixed. The resulting status is relative to that peer group, not an industry-wide grade.
Freeze the buyer-question panelCover category, problem, comparison, trust, price, urgency, and location intent. Keep branded prompts in a separate segment.
Freeze products and contextRecord AI product, mode, model when available, language, location, account state, browsing, and date.
Run the whole peer group togetherUse the same prompts, providers, dates, and run count for the target and competitors. Otherwise the comparison is not apples to apples.
Apply one brand-matching ruleDocument aliases, parent brands, domains, misspellings, subsidiaries, and ambiguous names before scoring.
Repeat comparable runsOne generated answer can vary even under identical conditions. Repeat the panel so recurring winners separate from one-off appearances.
Preserve the evidence safelyKeep prompt, answer, provider, date, citations, ordering, and classification. Redact personal or confidential details before publishing examples.

Compare gaps, not just ranks

“We rank second” can hide a tiny difference or a structural gap. For every priority metric, calculate both:

  • Gap to leader: target rate minus the highest competitor rate.
  • Gap to peer median: target rate minus the median rate among predeclared competitors, excluding the target.
  • Prompt wins: priority prompts where the target beats or ties the leader across repeated runs.
  • Provider gaps: the same comparison inside ChatGPT, Gemini, Perplexity, Claude, or another measured surface—not only as a blended total.
Why one blended leaderboard is not enough.
ComparisonWhat to inspectWhat the gap can reveal
Overall peer groupMention, recommendation, first-choice, citation, and accuracy ratesWhether the brand is absent, present, or genuinely selected
Prompt clusterCategory, problem, comparison, trust, price, urgency, and locationWhere competitors own the buyer journey
AI productProvider-level rates and repeated-run rangesWhether a blended score hides a provider-specific weakness
Recommendation evidenceReasons, claims, and recurring source domains supporting each winnerWhich evidence difference is worth testing next
AccuracyIncorrect services, locations, prices, credentials, or brand identityWhether higher visibility is helping or spreading bad facts
Do not compare your score from July with a competitor's score from August. Models, retrieval, sources, and market conditions may have changed. Competitors belong in the same measurement window.

For implementation details, use the AI visibility audit checklist. To identify the right cohort, see why AI competitors may differ from Google competitors.

A worked benchmark without invented industry averages

The fictional example uses 40 fixed non-branded buyer prompts, three AI products, and three runs per prompt: 360 scheduled responses in one shared panel. After nine provider failures, 351 responses remained valid. Every tracked brand was scored against those same 351 responses. The 300-answer recommendation denominator was determined from prompt eligibility before examining brand outcomes.

Illustrative data only; not a Plastorium customer result or industry benchmark.
BrandMentionRecommendationFirst positionConditional accuracy
Northstar119/351 (34%)72/300 (24%)33/300 (11%)115/119 (97%)
Target brand95/351 (27%)45/300 (15%)12/300 (4%)78/95 (82%)
Harbor77/351 (22%)42/300 (14%)15/300 (5%)72/77 (94%)
Cascade42/351 (12%)21/300 (7%)6/300 (2%)40/42 (95%)

The target is not simply “27% visible.” It is second in mention rate, nearly tied with Harbor in recommendations, behind Harbor on first position, and weak on conditional factual accuracy. The action is not “raise the score.” It is to investigate which high-intent prompts produce mentions without endorsement and which facts are wrong.

Report the numerator, denominator, failures, and results by provider and prompt cluster. These estimates describe the tested panel—not every possible buyer question. If you publish uncertainty ranges, account for repeated answers from the same prompts rather than treating every generated response as fully independent.

A scan is a snapshot; visibility is a pattern

A single scan captures one set of generated answers at one moment. It cannot tell you whether the result is typical, whether a competitor usually wins, or whether a change you shipped improved anything.

AI answers vary because model versions, retrieval systems, available sources, location context, product modes, and variable generation can change. Even with the same prompt, one run may mention the brand and the next may omit it. That is why a serious benchmark needs two time scales:

Repeated runs inside each windowRun the fixed panel multiple times per product. This estimates recurrence and prevents one unusually favorable answer from becoming the baseline.
Repeated windows over timeRe-run the same panel on a consistent cadence. This shows whether the target is closing the competitor gap, losing ground, or merely bouncing inside normal variation.
Illustrative line chart comparing target, competitor leader, and peer median recommendation rates across four repeated measurement windows
Illustrative values only. The observed leader gap was smaller in Window 4 under the same measurement rules, but the single late increase does not by itself establish a durable trend or campaign effect.

What to keep fixed—and what to log when it changes

A time series is comparable only when its documented test setup is visible.
Keep fixed when possibleLog for every windowWhy it matters
Prompt panel and intent labelsAdded, removed, or rewritten promptsPrompt changes can move the score without any market change
Competitor cohort and aliasesNew entrants, exits, mergers, rebrandsShare of voice changes when the denominator changes
Products, modes, language, and locationModel/version when exposed, browsing and account stateProvider changes can create a break that makes before-and-after results hard to compare
Run count and scoring rulesFailures, unavailable citations, classification changesDifferent denominators make adjacent windows incomparable
Priority prompt clustersSite changes, profile updates, PR, review, and content releasesCreates hypotheses about movement without claiming causation

For the full repetition methodology, see why one AI visibility scan is not enough.

Read the delta together with the competitor gap

  • Your observed rate rises and the leader gap shrinks: a smaller measured gap, subject to cluster-aware uncertainty and replication across matched windows.
  • Your observed rate rises but competitors rise faster: absolute rate increased while relative position weakened.
  • Your observed rate is flat while peers fall: relative position rose even though the dashboard number did not.
  • Every brand jumps at once: suspect a model, provider, prompt, or scoring change before crediting your campaign.
  • Only one product moves: inspect that provider's answers and source mix instead of declaring a market-wide win.
Do not optimize for a smooth chart. Preserve methodology changes as visible breaks in the series. If the prompt panel or scoring formula changes materially, establish a new baseline rather than splicing incompatible numbers together.

Four ways to read the benchmark

Low absolute rate, peer-group leaderThe category may be hard to enterProtect the lead, then test whether the panel contains the buyer questions that matter commercially.
High rate, behind the leaderVisibility exists; competitive selection is weakCompare recommendation reasons and evidence by prompt cluster rather than chasing more generic mentions.
Strong mentions, weak citationsRecognition exceeds attributable supportInspect source patterns. Do not assume every uncited mention is bad or that every citation caused the answer.
Growth with poor stabilityOne run may be flattering youRepeat the same panel. Stability measures reproducibility, not quality: a brand can also be stably absent or stably misdescribed.
A practical status framework: call a result uninterpretable when methodology is missing; behind when it trails the matched peer group on priority prompts; competitive when it sits within the peer cluster; strong when it reaches the peer group's top tier across priority metrics; and leading when it repeatedly outperforms the cohort while remaining accurate. These are Plastorium decision labels—not universal score bands.

Set targets against gaps, not arbitrary numbers

A useful target names the peer group, metric, segment, period, and quality guardrail. For example:

Over the next eight weeks, close half the recommendation-rate gap to the peer-group leader for five priority non-branded purchase prompts, while keeping factual accuracy above the current baseline.

That target can be measured. “Reach an AI visibility score of 70” cannot be interpreted unless the scoring system and baseline remain unchanged.

  • Use the peer median, calculated from predeclared competitors excluding the target, to judge whether the brand has entered the competitive cluster.
  • Use the leader gap to quantify the next attainable measured difference.
  • Use prompt clusters to keep observed movement tied to real buyer intent.
  • Use run-level distributions and cluster-aware uncertainty to distinguish observed movement from ordinary within-panel variation; do not rely on the repeated-run range alone.
  • Use accuracy as a guardrail so more visibility does not mean more misinformation.

Do not compare scores across vendors unless formulas, prompts, products, geography, run rules, and denominators match. For planning work after the baseline, see how to set AI visibility targets you can budget for and why one scan is not enough.

Run a directional preview

Spot possible gaps worth validating—not a stable benchmark. The free check is one quick sample; a defensible comparison requires repeated runs of a frozen brand and competitor peer group over time.

For an example of the evidence behind a full analysis, see the published redacted benchmark.

Run the free directional checkDiscuss a repeated-run benchmark

Frequently asked questions

What is a good AI visibility score?

There is no universal threshold. A useful result is competitive for a fixed non-branded prompt panel and matched peer group, accurate when the brand appears, and reproducible across repeated runs.

Is an AI visibility score of 50 good?

Not enough information. You need the formula, denominator, prompts, provider mix, cohort, and repeated-run variation. A 50 in one tool may not equal a 50 in another.

How should I benchmark against competitors?

Use identical prompts, products, markets, run counts, brand matching, and outcome definitions for every brand. Compare mentions, recommendations, first choice, tracked-brand share of voice, citations, accuracy, and stability separately.

How often should I measure?

Repeat runs for the baseline, then re-run the same panel consistently and after meaningful changes. Document model, product, prompt, and source changes before attributing movement to your work.

Why is one AI visibility scan not enough?

One scan is a snapshot of variable answers at one moment. It cannot show recurrence, competitor movement, or whether a later change improved your relative position. Use repeated runs within each window and repeat the matched competitor panel over time.