AI visibility audit: a reproducible 50 prompt checklist, scoring rubric and sample audit
The 50 prompt design with its 600 response arithmetic printed, the scoring rubric in full, and the dated samples we actually hold.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
An AI visibility audit is only worth reporting if you can state its sample size as a piece of arithmetic that a reader can check: prompts times engines times runs. The design on this page is 50 prompts times 4 engines times 3 runs, which is 600 responses, and every rate in the report is a fraction over 600 or over a named subset of it. We have not run that 600 response design as of 29 September 2026, and we will not print a number for a run we have not done. The samples we do hold, with their dates, are these. For clawlaw.in in legal technology: 78 prompts times 1 engine times 1 run, which is 78 responses, on 27 July 2026 on ChatGPT, in which the company was named on 21 of 78; and 18 prompts on 18 August 2026 on ChatGPT, in which it was the top source on 1 of 18. For aiknowsus.com in AI visibility tooling: 24 batches and 72 conversations in September 2026 on Perplexity, holding 161 recorded statements that we were not cited, including one batch of 6 questions in which the engine reported it had cited or recommended us in 0 of 6.
That is the whole honest position in one paragraph, and it is the shape every audit report should take. A sample size, the dates, the engine, and the fraction. If a report you are reading cannot give you those four things, it has not measured anything you can act on, whoever produced it.
How it was measured
This section is the protocol. It is written so that you can run it without us, and so that two people running it on the same market in the same week would produce numbers that can be compared.
Step one. Build the 50 prompt set and freeze it. Fifty prompts in five groups of ten, saved to a file with a version number and the date you wrote it. Every rate you later publish must name the version it came from. The five groups are not interchangeable and mixing them without saying so is the most common way an audit flatters its subject.
- Ten category questions. The best or the top tools, services or suppliers in your category, in your country. No brand named.
- Ten problem questions. The problem in the buyer's own words, before they know your category exists.
- Ten comparison questions. Two or three competitors named against each other. Your own brand appears in none of them.
- Ten how to questions. How to do the job by hand, which is where official sources usually win and a commercial product usually should not.
- Ten commercial questions. Price, contract, hidden cost, what a buyer asks at the end.
Step two. Apply the blind rule with no exceptions. Your own brand name appears in none of the 50 prompts and in no system instruction, wrapper or custom instruction around them. This is measured, not stylistic. On 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt.
Step three. Name the four engines and report them separately. A blended score across engines hides the one you are absent from, which is the only one you needed to know about. Record which interface and which account state, because a signed in account with memory switched on is not the same instrument as a fresh one.
Step four. Run each prompt three times in a fresh chat. Identical prompts produce different answers. One response is a draw, not a behaviour. Three is the minimum that lets you say a result is stable rather than lucky.
Step five. Print the denominator arithmetic on the report. 50 prompts times 4 engines is 200 prompt engine pairs. 200 pairs times 3 runs is 600 responses. A rate over 600 responses and a rate over 50 prompts are different numbers and they differ by a factor of twelve. This is arithmetic, not an observation, and it is the arithmetic that decides whether two audits can be compared at all.
Step six. Log these fourteen fields for every response. One row per response, 600 rows. The sheet is the audit. The chart is decoration.
- Prompt identifier and prompt set version.
- The prompt text exactly as sent.
- Group, one of the five above.
- Engine and interface.
- Run number, one, two or three.
- Date and time, with the time zone.
- Country and city, and the language of the prompt.
- Signed in state, and whether memory or personalisation was active.
- Every brand named in the answer, in the order they appeared.
- Every link the answer offered, as a full URL.
- The five label states, below.
- Whether the answer claimed to have searched, recorded as a claim.
- The labeller's name.
- A link to the saved capture, screenshot and text.
Step seven. Score with five states, not one. A single visibility number collapses five different situations that need five different responses from you, and the fifth one is invisible to every tool we know of.
- Named. Your brand appears in the answer text.
- Linked. A page on your own domain appears in the answer's sources or links.
- Top source. Your page is the first or most prominent source the answer leans on.
- Recommended and cited together. Your brand is put forward as an answer and a page on your domain supports it in the same response. This is the state that matters commercially, and it is the one most reports never separate.
- Used without attribution. The argument is yours and the credit is not. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and had dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement. A human has to read a sample to catch this.
Step eight. Record every search claim as a claim. Where an engine tells you whether it ran a live search, save the sentence and mark the field unverified. Two dated reasons. On 27 July 2026, asked 78 blind questions, ChatGPT then confirmed it had run no live web search for any of them, so 0 of 78 responses were built from what it could find rather than from what it remembered. And in three separate batches of the September 2026 aiknowsus.com audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. An engine's account of its own work is evidence of what it said, not of what it did.
Step nine. Have a second person label a sample and publish the agreement rate. Fifty of the 600 rows is enough. If two labellers disagree on more than a handful, your definitions are the problem and no amount of extra data will fix it.
Step ten. Keep every capture, and be willing to publish them. A number whose raw answers cannot be produced is not auditable, and that applies to our numbers as much as to anybody's.
What the numbers were
Here is every measured figure we hold, each with its numerator, its denominator, its engine, its date and its sector. Read the sector labels carefully. A legal technology result is not evidence about a marketing agency, a dental clinic or a manufacturer, and moving it across is exactly the error this page exists to prevent.
- 21 of 78 responses named the company. Legal technology, clawlaw.in, 27 July 2026, ChatGPT, 78 blind prompts, one run each. In the same reading the breadth order was ProVakil on 28 of 78, CLAW on 21 of 78 and Legistify on 18 of 78. That is a ranking of what the model had absorbed, not of what was true that week, because it had searched for none of the 78.
- 1 of 18 responses had the company as the top source. Legal technology, clawlaw.in, 18 August 2026, ChatGPT, 18 blind commercial questions. In that one answer the company's own comparison page was named in the answer as still being the vendor's own editorial page.
- 0 of 6 responses cited or recommended us. AI visibility tooling, aiknowsus.com, September 2026, Perplexity, six blind questions about our own category. The engine audited its own answers afterwards and reported it had cited or recommended us in none of the six. There was no position for us to hold.
- 161 recorded statements that we were not cited, across 24 batches and 72 conversations, September 2026, Perplexity, on our own domain. That is the engines' own self audits of their own answers, counted by a script over the capture files.
- 0 of 6 questions in which any third party review or directory source reached an answer. Legal technology, clawlaw.in, 17 September 2026, Claude. One well known review site did appear in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
What has not been run, as of 29 September 2026. We have not executed the 50 prompt, 4 engine, 3 run design on any domain. So this page publishes no visibility rate over 600 responses, no rate by prompt group, no engine to engine comparison over a common prompt set, and no inter labeller agreement figure. Those cells are empty on purpose. When the run is complete the figures will appear here with the prompt set version, the four engines, the dates and the failure counts attached.
One further result from that September 2026 capture is worth more to an auditor than any of the counts above. In batch after batch, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. The engine refused to rank a market on a single run. Any audit report that does rank a market on a single run is claiming more than the engine itself will claim.
What this cannot tell you
Seven limits. The first three are the ones that get skipped in published audits, and skipping them is usually the whole reason the headline number looks good.
- It measures your prompt set, not your market. Change ten prompts and the rate moves. That is why the set is frozen and versioned, and why a rate without a version number cannot be compared with anything.
- A rate over one engine is not a rate over AI. Report each engine separately or the number means nothing.
- Three runs is enough to notice instability, not enough to produce a confidence interval. Do not attach a margin of error to 600 responses collected on one week in one country.
- Absence in a response is not absence from the engine. The same prompt can name you on run one and not on run two. That is why the denominator is responses rather than prompts.
- The used without attribution state is only partly detectable. A person has to read the answers, and even then you catch the obvious cases. Both of the cases we hold were admitted by the engine when asked, on 17 September 2026, and would not have been visible otherwise.
- An audit is valid for the period it was run in. These systems change without notice, and a September figure is not a December figure.
- Nothing in this method gets you a position in an AI answer, and no auditor, ourselves included, can promise one. An audit tells you where you stand and which questions you are absent from. What you do about it is a separate job.
Sources and change log
All figures on this page come from two recorded programmes. The clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026. And the aiknowsus.com audit of September 2026, 24 batches and 72 conversations, captured outside this repository. Every count was produced by a script over those files rather than from memory.
One figure on this page can be checked from outside today: on 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. That is one page on one engine, and one first appearance is not a median time to first appearance. We publish no median.
Disclosure. We sell an AI visibility product, AI Knows Us. The protocol above is what our product is built to enforce, which is a design claim you can check by using it, not a benchmark result. The counts above are our own, and until the capture files are published you are taking them on trust.
Version. 29 September 2026, first publication. This page is updated when the 600 response run completes, when any figure is corrected, and when a reader shows us an error.
Common questions
Why 50 prompts rather than 20 or 200?
Fifty divides cleanly into five groups of ten, which is the smallest set that lets you report a rate per group rather than one blended number. Below about ten per group a single odd answer swings the group rate visibly. Above 50 the cost rises faster than the insight, because the marginal prompt is usually a rewording of one you already have. If you only have time for one pass, run 20 and say it was 20.
Can I use one engine for the first audit?
Yes, and label it as one engine everywhere, including in the internal deck. The clawlaw.in baseline of 27 July 2026 was ChatGPT only, 78 prompts, one run each, and every sentence written about it says ChatGPT. That is why it is still quotable two months later.
Our rate came out at zero. Is the audit broken?
Probably not. Zero is a common and real starting point, and we have it on our own domain: 0 of 6 in September 2026 on Perplexity, and 161 recorded statements across the wider capture that we were not cited. Zero is where the work starts, and it is a much more useful reading than an inflated number produced by a prompt that named you.
Should I count a mention of my brand with a link to somebody else's page?
Count it as named and not linked, and keep the two states apart. It is a real situation with a real cause. On 17 September 2026 Claude confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, while the site landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
How often should the audit be repeated?
Monthly on the same frozen set for your own tracking, and a new version of the set no more than twice a year. Any change to the prompt set starts a new baseline, and the old numbers are reported beside the new ones rather than joined to them.
Can a tool run this for me?
Parts of it. A tool can issue prompts, repeat them, split by engine and export the rows, which is most of the labour. It cannot write your prompt set, it cannot judge the used without attribution state, and it cannot tell you whether an engine really searched, because the engine's own account of that is unreliable and we have three dated retractions to show it. Demand raw answer exports from any tool you buy, or you cannot audit the audit.
What to do first
Open a spreadsheet and write the ten category questions, in your buyers' words, with your brand name in none of them. Run those ten on two engines, three times each, which is 60 responses, and label them with the five states. You will have a dated baseline you own by the end of the afternoon, and you will know which of the five groups to build out next. Then extend to 50 prompts and four engines when the first 60 rows have told you what you are actually measuring.