Open resource · · Version 0.1

How to measure AI recommendation visibility

A working protocol for finding out whether AI assistants actually recommend you, and for measuring it in a way you can repeat, check, and disagree with. Free to use.

Most claims about AI visibility rest on a single screenshot. This is the alternative: fixed queries decided before you look at any answers, repetitions, explicit scoring definitions, and limitations stated up front. It is a pilot, not a validated standard, and it is published so you can run it yourself rather than take anyone's word for where you stand.


Why single checks mislead

Ask an assistant once who the best provider in your category is, and you have one draw from a stochastic system whose retrieval changes between sessions and between products. Ask it three times in fresh conversations and you will often get three different sets of names. Any honest measurement therefore has to fix its questions in advance, repeat them, record the raw outputs, and report the spread rather than a single headline number.

The protocol

  1. Fix the query set before looking at any answers. Decide the prompts first so results cannot be reverse-engineered into a flattering picture. Keep broad category questions in the set alongside your specialist ones. Removing the queries you lose is how a benchmark becomes marketing.
  2. Run each prompt in a fresh conversation. Disable personalization and memory where the product supports it. Account history contaminates the result.
  3. Repeat at least three times per prompt, per engine, per surface. Report the variation you observe.
  4. Record the conditions. Engine, exact product and model where available, timestamp, language, location setting, whether web search was active, and account or memory state.
  5. Treat surfaces as distinct. A consumer app, an API call, and a search-enabled mode retrieve differently. Do not pool them into one number.
  6. Save every raw response. A score without the underlying record is an assertion.

What to record, one row per response

Run identifier, prompt identifier, repetition number, timestamp, engine, model, search mode, location setting, personalization state, path to the saved raw response, whether you were mentioned, whether you were recommended, ordinal position only if the answer was genuinely ranked, which of your URLs were cited, which competitors were named, the ownership category of each cited source, any factual errors about you, and any tool or access failure.

Scoring definitions

Handling failures honestly

Exclude blocked or failed calls from the completed-response denominator and report their count separately, keeping them in the raw log. Never score an engine you could not access as zero visibility. An engine that refused to answer is missing data, not a loss.

A starting query set

These starting prompts are grouped by buyer category. Substitute your own market and keep the structure. The point of the grouping is that a single blended visibility score hides the thing you most need to know, which is that you can be strong in one category and absent in another.

Comparing runs over time

Compare unchanged prompt sets under matched conditions, after a documented content change and after crawl has actually been confirmed in the relevant engine's own tooling. Keep the original baseline intact. Record every other change that happened in the same window. Treat observed movement as descriptive: a benchmark of this size cannot establish that a specific edit caused a specific change, and anyone telling you otherwise is selling something.

What this protocol is not

Common questions

How do I measure whether ChatGPT or Perplexity recommends my business?

Fix a set of prompts before you look at any answers, run each one in a fresh conversation with personalization disabled where the product allows it, repeat each prompt at least three times per engine, and save every raw response. Score three things separately: whether you are mentioned, whether you are actually recommended, and whether a source associated with you is cited. Report the counts and the denominator, not a single blended score. Treat consumer apps, API calls, and search-enabled modes as different surfaces, because they retrieve differently.

Why do I get different answers from the same AI assistant each time?

Because these systems are stochastic and their retrieval changes. A single run tells you almost nothing. That is why repetition is part of the protocol and why variation across repetitions should be reported rather than averaged away. Three runs reduce noise; they do not eliminate uncertainty, and they do not prove that any particular change to your website caused an observed difference.

Does a failed site: search mean my website is not indexed?

No. A site: operator returning nothing is weak evidence and is often unreliable, and an HTTP 200 response is not evidence that new content has been crawled. Confirm crawl and index status in the search engine's own webmaster tooling, and treat a single negative search result as inconclusive rather than proof of absence.

What is the difference between being mentioned and being recommended by an AI assistant?

Mention means your name appears anywhere in the response. Recommendation means the assistant actually puts you forward as an option for the buyer's stated need. They are different outcomes with different commercial value, so score them as separate metrics. Where an answer presents an unordered list, do not assign an ordinal rank to it.

What counts as a credible source when an AI assistant cites something about me?

Record source ownership honestly. Your own website is an owned source. A directory profile you submitted yourself is a self-submitted record, not independent validation. Coverage written about you by someone else is independently authored. A persistent identifier such as a DOI establishes a durable record, not peer review and not an endorsement. Keeping these categories separate is what stops a measurement exercise from flattering you.

Use it

This protocol is free to use and adapt, with attribution. If you run it and find it wrong, or find a better way to control for something, I would rather hear it than not. Version 0.1, published . Revisions will be dated and the prior version kept addressable.

Want this run properly on your organization?

Related: AI search visibility service · Sovereign AI Tracker · Insights