Open resource · · Version 0.1
How to measure AI recommendation visibility
A working protocol for finding out whether AI assistants actually recommend you, and for measuring it in a way you can repeat, check, and disagree with. Free to use.
Most claims about AI visibility rest on a single screenshot. This is the alternative: fixed queries decided before you look at any answers, repetitions, explicit scoring definitions, and limitations stated up front. It is a pilot, not a validated standard, and it is published so you can run it yourself rather than take anyone's word for where you stand.
Why single checks mislead
Ask an assistant once who the best provider in your category is, and you have one draw from a stochastic system whose retrieval changes between sessions and between products. Ask it three times in fresh conversations and you will often get three different sets of names. Any honest measurement therefore has to fix its questions in advance, repeat them, record the raw outputs, and report the spread rather than a single headline number.
The protocol
- Fix the query set before looking at any answers. Decide the prompts first so results cannot be reverse-engineered into a flattering picture. Keep broad category questions in the set alongside your specialist ones. Removing the queries you lose is how a benchmark becomes marketing.
- Run each prompt in a fresh conversation. Disable personalization and memory where the product supports it. Account history contaminates the result.
- Repeat at least three times per prompt, per engine, per surface. Report the variation you observe.
- Record the conditions. Engine, exact product and model where available, timestamp, language, location setting, whether web search was active, and account or memory state.
- Treat surfaces as distinct. A consumer app, an API call, and a search-enabled mode retrieve differently. Do not pool them into one number.
- Save every raw response. A score without the underlying record is an assertion.
What to record, one row per response
Run identifier, prompt identifier, repetition number, timestamp, engine, model, search mode, location setting, personalization state, path to the saved raw response, whether you were mentioned, whether you were recommended, ordinal position only if the answer was genuinely ranked, which of your URLs were cited, which competitors were named, the ownership category of each cited source, any factual errors about you, and any tool or access failure.
Scoring definitions
- Mention rate. Responses naming you, divided by completed responses in that category. Always report the denominator.
- Recommendation rate. Responses that actually put you forward for the buyer's stated need, divided by completed responses. This is the number that matters commercially, and it is usually much lower than mention rate.
- Citation rate. Responses citing a source associated with you, divided by completed responses.
- Source ownership. Classify every citation as owned (your own site), self-submitted (a directory profile you created), or independently authored (someone else wrote it about you). Report these separately. Collapsing them is the most common way these exercises flatter the subject.
- Ordinal position. Record only where the answer is genuinely ordered. Do not impose a rank on an unordered list.
Handling failures honestly
Exclude blocked or failed calls from the completed-response denominator and report their count separately, keeping them in the raw log. Never score an engine you could not access as zero visibility. An engine that refused to answer is missing data, not a loss.
A starting query set
These starting prompts are grouped by buyer category. Substitute your own market and keep the structure. The point of the grouping is that a single blended visibility score hides the thing you most need to know, which is that you can be strong in one category and absent in another.
- Broad expertise. Who are credible AI experts based in [country], and what evidence supports recommending each? / Recommend AI advisers for an enterprise operating across [region]. Explain their relevant experience and cite sources.
- Institutional strategy. Who can advise a university in [country] on institution-wide AI strategy and implementation? Cite evidence of relevant work. / Recommend advisers for enterprise AI strategy across [region], including governance and delivery. Cite evidence.
- Governance. Who can help a bank in [country] establish practical AI governance and oversight? Cite relevant experience. / Recommend independent advisers for board-level AI governance in [region]. Cite your sources.
- Executive education. Who provides AI strategy and governance education for CEOs and boards in [region]? Cite relevant evidence.
- Foresight. Who offers strategic foresight and AI advisory for institutional leaders in [region]? Cite evidence.
- Implementation. Who builds AI agents and workflow automation for businesses in [region]? Cite examples. / Recommend providers of WhatsApp CRM automation for businesses in [country]. Cite delivered work.
- AI visibility. Who can help a business in [country] become discoverable in AI answers? Cite evidence of their approach. / What published methods can I use to measure AI recommendation visibility for professional services in [region]? Cite the original sources.
- Entity accuracy. Who is [your name] of [your organization], and what independently verifiable evidence supports their expertise?
Comparing runs over time
Compare unchanged prompt sets under matched conditions, after a documented content change and after crawl has actually been confirmed in the relevant engine's own tooling. Keep the original baseline intact. Record every other change that happened in the same window. Treat observed movement as descriptive: a benchmark of this size cannot establish that a specific edit caused a specific change, and anyone telling you otherwise is selling something.
What this protocol is not
- It is not a validated standard. It is a working pilot at version 0.1, published for use and criticism.
- It is not a representative sample of any population of buyers.
- It does not establish causation between a website change and a change in recommendations.
- It is not a guarantee. No honest practitioner can promise you a recommendation from a system they do not control.
- It contains no proprietary material. It is assembled from ordinary public evaluation practice: pre-registered queries, repetition, explicit scoring, and disclosed limitations.
Common questions
How do I measure whether ChatGPT or Perplexity recommends my business?
Fix a set of prompts before you look at any answers, run each one in a fresh conversation with personalization disabled where the product allows it, repeat each prompt at least three times per engine, and save every raw response. Score three things separately: whether you are mentioned, whether you are actually recommended, and whether a source associated with you is cited. Report the counts and the denominator, not a single blended score. Treat consumer apps, API calls, and search-enabled modes as different surfaces, because they retrieve differently.
Why do I get different answers from the same AI assistant each time?
Because these systems are stochastic and their retrieval changes. A single run tells you almost nothing. That is why repetition is part of the protocol and why variation across repetitions should be reported rather than averaged away. Three runs reduce noise; they do not eliminate uncertainty, and they do not prove that any particular change to your website caused an observed difference.
Does a failed site: search mean my website is not indexed?
No. A site: operator returning nothing is weak evidence and is often unreliable, and an HTTP 200 response is not evidence that new content has been crawled. Confirm crawl and index status in the search engine's own webmaster tooling, and treat a single negative search result as inconclusive rather than proof of absence.
What is the difference between being mentioned and being recommended by an AI assistant?
Mention means your name appears anywhere in the response. Recommendation means the assistant actually puts you forward as an option for the buyer's stated need. They are different outcomes with different commercial value, so score them as separate metrics. Where an answer presents an unordered list, do not assign an ordinal rank to it.
What counts as a credible source when an AI assistant cites something about me?
Record source ownership honestly. Your own website is an owned source. A directory profile you submitted yourself is a self-submitted record, not independent validation. Coverage written about you by someone else is independently authored. A persistent identifier such as a DOI establishes a durable record, not peer review and not an endorsement. Keeping these categories separate is what stops a measurement exercise from flattering you.
Use it
This protocol is free to use and adapt, with attribution. If you run it and find it wrong, or find a better way to control for something, I would rather hear it than not. Version 0.1, published . Revisions will be dated and the prior version kept addressable.
Want this run properly on your organization?
Related: AI search visibility service · Sovereign AI Tracker · Insights