AI visibility · Category diagnosis

How to diagnose AI visibility in your category Beyond generic SEO signals.

Measure candidate variables against defined buyer questions, providers, and markets. Do not turn an observed association into a universal ranking-factor claim.

Plastorium Research · August 4, 2026 · 13 min read

Short answer: diagnose the category, not a mythical universal factor

SEO signals can matter, but neither an SEO metric nor a displayed citation is a universal AI ranking factor. Treat each as a candidate variable, observed in a defined panel, and test only useful, truthful changes.

AI visibility is not one outcome. A brand can rank in Google, be mentioned, be recommended, be cited, or be described accurately. Those can all move differently. Saying something was “observed in this defined panel” is a finding. Saying “this caused the recommendation” is a much bigger claim. It needs causal evidence that a normal visibility scan does not supply.

This is distinct from What Actually Drives AI Visibility? , the Reddit/monthly-driver piece. This article is a category-specific measurement method, not a source or ranking-factor claim.

Separate the outcomes before you explain them

Measure What it records Useful evidence Does not prove
Different measures answer different questions; none establishes a causal mechanism.
SEO rank Search position under specified conditions Query, date, location, device AI mention or recommendation
AI mention Brand appears in answer Answer text and run Best-fit selection
Recommendation Brand is presented as a choice Coding rule, framing, alternatives Why it was selected
Citation Displayed source associated with answer URL, claim context, provider/mode That the citation caused selection
Accuracy Material fact is correct Primary-record check That accuracy causes recommendation

Read What Is AI Citation Share? for source definitions and You Rank on Google, but AI Still Skips You for the SEO/AI split.

Define the category, prompt cluster, and buyer evidence map

Before you measure anything, write down the basics. Specify the offer, market, buyer role, buying stage, comparable alternatives, provider/mode, and test window.

Then decide in advance which buyer questions you will track. Group them into clusters, for example:

  • Discovery — first-look questions from a new buyer
  • Shortlist — narrowing down a set of options
  • Comparison — weighing this brand against alternatives
  • Risk validation — checking for dealbreakers
  • Urgent availability — can this brand help right now

For each cluster, map the buyer facts that actually matter, such as:

  • Integrations
  • Security
  • Implementation
  • Service area
  • Hours
  • Qualifications
  • Price policy

Lock the exact prompt wording, outcome rules, run count, location/language, and competitor cohort before you start. Keep unavailable citations as N/A, not zero. This prevents after-the-fact storytelling and makes it possible to compare results later.

A buyer evidence map is not a signal-manufacturing checklist. Reviews, security pages, citations, and SEO measures can be candidate variables. Their presence does not establish that they cause recommendations.

Five candidate-variable families—not ranking factors

1. Search and site accessibility Query rank, indexation, crawlable text, relevance, technical errors, and structured facts.
2. Entity and factual consistency Name, offer, geography, qualifications, and contact facts across primary and eligible records.
3. Owned proof and documentation Product, implementation, policy, security, pricing, service, or qualification pages.
4. Independent evidence context Legitimate editorial coverage, profiles, reviews, partnerships, and references; never infer they cause selection.
5. Prompt, provider, and market fit Intent, urgency, buyer role, region, language, freshness, provider/mode, and competitor set—often confounders.
Keep uncertainty Failed runs, absent citations, and inaccessible sources are data-quality information, not negative signals.

B2B SaaS and urgent local service can show different patterns

Dimension B2B SaaS Urgent local service Implication
Diagnostic priorities differ because the buyer question and evidence differ.
Question Which workflow tool fits a regulated team? Who can repair a furnace tonight near me? Intent is a confounder.
Evidence Integration, security, implementation, support Service area, hours, availability, qualifications Map facts first.
Safe test Clarify verified buyer-critical documentation Correct verified eligible operating facts Improve information quality, not signals.

Measurement design and controls

Control Hold or record Why
Controls reduce ambiguity, not confounding to zero.
Prompt version Exact wording, intent, priority Wording changes outputs.
Provider context Product, mode, time, citation availability Provider change can explain movement.
Market context Location, language, account/device state Locality changes candidates.
Repeated runs Scheduled, valid, failed, excluded One answer is not stable evidence.
Change log Owned/source/provider events Reveals competing explanations.

Fictional illustrative two-cohort example

Everything below is fictional: not client data, a benchmark, or causal proof. Northstar Workflow, a fictional B2B SaaS company, tests 24 prompts: 216 scheduled / 208 valid . Harbor City Heating, a fictional urgent local service, tests 18 prompts: 162 scheduled / 156 valid . Results stay within each fixed market/provider panel; they are not pooled.

Cohort Observed association Candidate variable worth testing Cannot conclude
Observed patterns in fictional defined panels, not benchmarks or causal proof.
Northstar Workflow Integration/security answer accuracy varied more than non-branded SEO rank. Clearer verified implementation and security documentation. Documentation, security pages, or rank cause recommendations.
Harbor City Heating Availability/service-area accuracy varied in local-intent answers and by provider context. Correct verified operating facts on owned pages and eligible profiles. Profiles, reviews, citations, or local SEO cause recommendations.

The only defensible summary: different evidence gaps were associated with different outcomes in fictional, defined panels. Samples are not benchmarks; confounding remains plausible.

Use an ethical test plan and remeasure

Input Question Evidence
Prioritize learning value and buyer value, not promised recommendation lift.
Buyer materiality Does the fact affect selection or risk? Prompt priority and evidence map
Observed association Does it recur across valid runs? Run-level table, not anecdotes
Actionability Can a truthful fact be improved legitimately? Owner, policy, primary record
Confounding risk Could other changes explain it? Comparator and change log
Learning value Will the result change a decision? Success, failure, stop rules

Some tactics are off-limits, no matter how tempting. Do not:

  • Manufacture reviews
  • Buy undisclosed endorsements
  • Create fake citations
  • Manipulate online communities

Instead, remeasure the same panel after doing real work. When you report results, include valid runs, N/A citations, provider context, and the change log. Say “associated with” and “worth testing.” Never say “caused.” See Why One AI Visibility Scan Is Not Enough and Where Should You Publish Next? .

Code the answer before you look for an explanation

Teams often call every brand appearance a “win.” That turns a mixed answer into a flattering metric, which makes later diagnosis impossible. Write the coding rules before you collect a panel. Apply them to every valid run, and keep the raw answer beside the coded result. A person who did not write the hypothesis should be able to inspect the rule and reach the same label.

Outcome Code as yes when Do not code as yes when Record beside the code
Example coding rules. Adapt the definitions before the first run; do not edit them after seeing a favorable answer.
Mention The brand is explicitly named in the answer body or a clearly enumerated list. It appears only in a cited URL, navigation element, or an ambiguous string. Exact excerpt, answer position, prompt and run ID.
Recommendation The answer presents the brand as a suitable option for the stated buyer need, with affirmative wording or inclusion in a shortlist. The brand is merely described, rejected, or named as background. Selection wording, alternatives, order and hedging.
Displayed citation A visible source link or source label is attached to the answer and can be retained with its URL and nearby claim. The product provides no visible citations or the source cannot be resolved. URL, source type, claim context and availability status.
Accuracy A material fact in the answer matches a dated primary record. The fact is vague, unverifiable, stale, or only looks plausible. Fact checked, primary record, reviewer and date.

Keep a separate unknown state. If an answer says “may offer 24/7 service,” treat that as a mention. It is not automatically an accurate availability claim. If a provider shows a domain with no clear link to the sentence, record the citation as displayed. Leave claim support marked unknown. These distinctions protect both the score and the buyer from false certainty.

Plan for valid runs, exclusions, and uncertainty

A scheduled run is not necessarily evidence. A response can fail in several ways:

  • Fail outright
  • Return a safety refusal
  • Ignore the requested market
  • Truncate before a list
  • Expose no citations
  • Be so malformed it cannot be coded

Define exclusions in advance. Keep every excluded run in the export, with its reason. Never quietly remove a result because it weakens a narrative.

Predeclare the panel List prompts, priority tier, providers or modes, market settings, intended repetition, competitors, and the first and last collection dates. Do not add only the prompts where the brand performs well.
Count the denominator Report scheduled runs, completed runs, valid coded runs, exclusions by reason, and citation-unavailable runs. Mention rate out of 18 valid runs means something different from 18 scheduled runs.
Preserve run-level evidence Store prompt version, timestamp, provider context, answer text, citations, coding result, and reviewer. Aggregates without source rows cannot be audited.
Do not overfit a small panel Small samples can surface a question; they cannot establish a durable category rule. Treat a large swing in a handful of valid runs as a reason to inspect and repeat.

There is no magic sample size that turns an AI visibility scan into a causal study. The right panel size depends on the number of buyer intents, markets, products, and decisions at stake. What matters is that the panel is large enough, and repeated enough, to support a decision. The team also needs to state its uncertainty honestly. If only eight runs were valid for a high-intent cluster, say so. Do not dress it up with percentages that imply more certainty than you have.

Look for competing explanations before funding a “driver”

Category comparisons are especially easy to get wrong. Recommended companies can differ in many ways at once, so a pattern you notice may not have the cause you assume. A competitor might simply have:

  • More reviews
  • Clearer pages
  • Longer market history
  • Stronger direct demand
  • A better-known founder
  • Different availability
  • Sources that changed during the collection window

An observed relationship does not isolate any one of those differences.

Alternative explanation What it can distort How to record it
Questions to ask before treating a pattern as a priority.
Provider or retrieval change Many brands move together, citations appear or disappear, answer format changes. Provider/mode, timestamps, release notes when available, cohort movement.
Prompt or market mismatch A local brand is judged on a non-local prompt, or a SaaS product is compared across different buyer roles. Intent label, location, language, buyer role, eligibility rules.
Real-world business change New service, acquisition, outage, reputation event, certification, pricing or availability change. Dated change log and primary-source verification.
Source freshness or access A page was updated, removed, blocked, paywalled, or newly indexed during the window. URL checks, publish/update date where visible, access status.
Selection and survivorship bias Only visible competitors or successful prompts were included in the cohort. Discovery method, inclusion criteria, missing competitors and exclusions.

A useful counterfactual question is: What else changed, or what else differs, that could plausibly explain this result? If the answer is “several things,” fund the most buyer-useful, reversible improvement. Label it a test, not a proven lever. This is more honest than claiming that a backlink, DA score, review count, or citation created a recommendation.

DA and backlinks belong in the evidence file only as candidate SEO measurements. They can line up with other useful things, like discoverability or editorial coverage. But DA is a proprietary third-party metric, not a universal AI selection factor, and neither is any other single measure. A high-DA site with weak claim fit may be less useful to a buyer than an accurate specialist, local, or product-specific source. Do not buy links, fabricate endorsements, or treat authority scores as a shortcut around evidence quality.

A practical 30-day diagnostic cycle

The goal of a cycle is not to “hack” an answer. It is to make one material buyer fact more useful, measure a defined outcome, and decide whether the next action still deserves attention.

  1. Days 1–3: frame the decision. Name the buyer segment, question cluster, market, products, competitor cohort, outcome definitions, and baseline window. Identify the decision that would change if the evidence moves.
  2. Days 4–7: collect and code the baseline. Run the fixed panel with planned repetitions. Keep track of scheduled, valid, excluded, and citation-unavailable counts. Verify the key answer facts against primary records.
  3. Days 8–11: inspect the evidence gap. Compare accurate owned facts, eligible profiles, documentation, independent context, and competitor framing. List alternative explanations before proposing work.
  4. Days 12–20: ship one truthful, buyer-useful improvement. For example, clarify a verified implementation requirement, correct an eligible service-area profile, or publish a supportable comparison. Keep the owner, approval, URL, date, and policy constraints in the log.
  5. Days 21–30: re-check and decide. Re-run the same priority cluster where feasible. Compare outcomes with the change log and cohort, and report the uncertainty honestly. Then continue, adjust, or stop based on what you learned, not on one lucky answer.

Some work needs longer than 30 days to be discoverable or to change buyer evidence. That is fine. The re-check can confirm publication and factual availability before it claims any AI outcome. Do not force a conclusion simply because a reporting calendar ended.

Keep a decision log that a buyer can audit

A dashboard score alone cannot explain why a team spent money or why a recommendation changed. Instead, keep the smallest useful record: one row per hypothesis, with one linked folder of evidence. A future reviewer should be able to see the claim, the baseline, the actual change, and the uncertainty. They should not have to reverse-engineer a presentation to find it.

Field Example Why it matters
A minimum decision-log row. The expected outcome is a measurement question, not a promised lift.
Buyer question and outcome “Can a regulated operations team implement X?”; recommendation and accuracy Keeps work tied to a commercial decision.
Observed evidence Run IDs, answer excerpts, source availability, valid/excluded counts Prevents a screenshot from becoming the whole case.
Candidate gap and alternatives Verified integration explanation unclear; provider format also changed Preserves competing explanations.
Action and guardrail Publish reviewed integration guide; no unverified claims or paid endorsements Defines a useful, compliant action.
Owner, date, and re-check Product marketing; published May 12; re-check June 12 Makes accountability possible.
Decision after re-check Continue, revise, or stop; include uncertainty and next question Avoids repeating activity without learning.

What the two fictional cohorts would do next

For Northstar Workflow, the buying committee needs implementation, integration, and security facts accurate enough to survive a risk-validation question. The right next action is not “get more backlinks.” It is to verify the product documentation against primary records. It also means making the implementation boundaries clear, and letting legal and security owners approve claims before publication. The re-check should ask two things: do the answers remain accurate, and is the product framed as an eligible option? It should not ask whether a documentation page “caused” a selection.

For Harbor City Heating, an urgent buyer needs a real service area, live availability, qualification, and safe contact path. The team should correct the facts it controls and use legitimate platform correction processes. It should not create fake local coverage, filter out negative reviews, or claim emergency service it cannot deliver. The next panel should measure factual accuracy and shortlist inclusion separately. It should also record any provider or weather-driven changes that could explain the result.

Start with a category-specific visibility baseline

Run a free directional baseline. For a defined buyer-question panel, evidence map, and measurement plan, book an analysis. Neither is a benchmark or a guarantee.

Check my AI visibility Talk to a strategist

Frequently asked questions

Are SEO signals AI ranking factors?

No universal conclusion follows from an observed association. SEO measures can be useful candidate variables, but AI visibility should be measured in a defined category, prompt panel, provider, and market before deciding what is worth testing.

What should a category-specific AI visibility panel include?

Include predeclared buyer questions by intent, relevant competitors, providers or modes, market context, repeated runs, and separate outcomes for mention, recommendation, citation, and answer accuracy.

Why might B2B SaaS and urgent local services show different patterns?

Their buyers ask different questions and need different evidence. Neither pattern proves what causes a recommendation.

How often should I remeasure AI visibility?

Remeasure a fixed panel on a regular cadence and after meaningful changes, preserving prompts, settings, valid-run counts, and a change log.