AI search visibility audit: a reproducible 30 prompt method and results

Thirty prompts, three engines, three repeats, and every result reported with its denominator.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

An AI visibility audit is 30 blind prompts, run on named engines, three times each, with every answer labelled by a person and every rate reported as a fraction. The headline number is the share of prompt runs in which a page on your own domain was cited. On our own domain the measured figure is 0 of 6 prompt runs, which is 0 per cent: six blind questions about our own category in September 2026, on Perplexity, after which the engine audited its own answers and reported it had not cited or recommended aiknowsus.com in any of the six. Across the wider capture of 24 batches and 72 conversations in the same month, the phrase recording that we were not cited appears 161 times. We have not yet completed the full 30 prompt run on three engines, and the empty cells below stay empty until we do.

Publishing our own zero is deliberate. A vendor page that reports a strong result for itself and no method is the exact kind of page the engines discount, and we have the observation to prove they do it.

What an AI visibility audit does and does not measure

It measures what a set of named engines answered, for a written list of prompts, under recorded conditions, on stated dates. That is the whole scope. Four things it is regularly claimed to measure and does not.

  • It is not a share of voice for your market. You have not sampled the market. You have sampled 30 prompts you chose.
  • It is not a measure of demand. Nothing in an audit tells you how many buyers asked that question.
  • It is not a ranking. Appearing is not the same as being preferred, and the selection rule is not published by any engine.
  • It is not a forecast. Engines change their behaviour without notice, so a result is valid for its dates.

What it is genuinely good for is finding out which of your buyers' questions you have no page for, and which sources are answering them instead of you. That is worth more than a score.

Choose prompts and engines

Thirty prompts, six groups of five, so the set covers the whole decision rather than only the part where your category is already known.

  • Five category prompts. Best, top, leading, recommend three, what to look for.
  • Five comparison prompts naming competitors and never you.
  • Five problem prompts from before the buyer knows the category exists.
  • Five commercial prompts. Price, inclusions, hidden charges, contract terms, comparing quotes.
  • Five trust prompts. How to check a provider is legitimate, what the risks are, who publishes their coverage or pricing, common complaints.
  • Five local or segment prompts carrying a city, a state or a specific buyer type.

Three rules govern the set. Your brand name appears in none of the 30. The file is versioned and dated, and every reported rate names the version. And once frozen it does not change during the study, because a denominator that moves makes every trend meaningless.

On engines: name them, run at least three, and report each separately. A blended score across engines hides which one you are absent from, and they select sources differently. Our own two programmes used ChatGPT, Claude and Perplexity, and the same questions produced different results on each.

On repeats: three runs per prompt per engine. Thirty prompts, three engines, three repeats is 270 answer runs, which is the denominator you will be reporting against. That is arithmetic, and doing it before you start is what stops a plan from quietly becoming unaffordable.

Run and label each result

Nine steps, in order.

One. Fresh session per run. No follow up questions before the answer is captured.

Two. Record the conditions on every row: date, time, engine, model or mode, country and city, language, signed in state, device, and whether memory or personalisation was active.

Three. Capture the answer as text and every citation exactly, plus a screenshot. Keep the raw citation strings in their own column, before any cleaning.

Four. Label with five states, each its own column: brand named in the text, own domain cited, own page is the top source, competitor named, and material used with no attribution. The fifth exists because it happens. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement.

Five. Record every competitor and every source domain named, because this becomes the most useful output of the audit. Across the whole aiknowsus.com capture of September 2026 the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. That told us where to spend, and it was a by product of labelling.

Six. Record the search claim and mark it unverified. In three separate batches of that September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question. And on 27 July 2026 ChatGPT confirmed it had run no live search for any of 78 questions, which means an entire baseline described what the model remembered.

Seven. Keep the brand name out, absolutely. On 27 July 2026 the branded version of that same 78 question run came back ranking the company first on almost every question, while the blind version placed it second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. We discarded the branded run. If you need a branded check, keep it in a separate labelled set that never enters the numbers.

Eight. Have a second person label a sample of at least 30 runs blind, and publish the agreement rate with the results.

Nine. Report every rate as a fraction over the run count, then the percentage if you want one, then the engines, the prompt set version and the date range. In that order, in the same sentence.

Check website discoverability

Half of a first audit's findings are on your own site, and this part costs an hour. Six checks, in the order that has explained the most for us.

  • Fetch your key pages the way a crawler does and read what comes back. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so a crawler received a page with no prices at all, and the prices the assistants quoted had come from an app store listing instead. That single check explains more absences than any prompt analysis.
  • Confirm the pages you expect to be cited are indexed, one by one, and log the date each was confirmed.
  • Find every other public copy of your key facts and check they agree. On 6 August 2026 ChatGPT noticed that the clawlaw.in website and the app store listing carried different plan names and different prices for the same product, and said so in its answer.
  • Check that every number on your site can survive being verified. Also on 6 August 2026, Claude checked a headline figure on the clawlaw.in site against public numbers, found it could not be true, and advised a buyer against the product, after which the company's accurate claims stopped counting for that answer.
  • Source and date every competitor fact you publish. On 17 September 2026 Claude read two clawlaw.in comparison pages stating competitors' prices with no link, no date and no source. The figures were correct when it checked them independently, and it still treated the pages as advocacy and used official sources instead.
  • Publish your coverage as an explicit list rather than as an adjective. On 17 September 2026, on a question about finding every case against a company, Claude ranked two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and it noted that its ordering reflected price transparency and source authority rather than product quality.

Audit results

The full 30 prompt, three engine, three repeat run has not been completed on our own domain, so no 270 run figure is published here. What follows is every measured share we hold, each with its denominator, its engines and its dates.

  • On domain citations: 0 of 6 prompt runs. Perplexity, September 2026, aiknowsus.com, six blind questions in our own category, self audited by the engine afterwards.
  • 161 recorded statements that we were not cited, across 24 batches and 72 conversations, September 2026, in the engines' own self audits of their answers.
  • Top source: 1 of 18 questions. ChatGPT, 18 August 2026, clawlaw.in, eighteen blind commercial questions. The company's own comparison page was named in the answer as still being the vendor's own editorial page.
  • Named: 21 of 78 questions, against 28 of 78 for ProVakil and 18 of 78 for Legistify. ChatGPT, 27 July 2026, blind. Zero live searches across all 78, confirmed by the engine on the same run.
  • One verified citation of a named URL, 14 days after publication. ChatGPT, 6 August 2026, clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india, for a vendor due diligence question in a zone where the same set had named the company nowhere at baseline.
  • Third party reviews reaching an answer: 0 of 6 questions. Claude, 17 September 2026. One well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.

What these results cannot show

They cannot be combined into one score, because they come from different engines, different prompt sets and different months. They cannot be read as a rate for any category, because each is a count over its own small set. The counts taken over our capture files cannot be checked from outside until we publish those files, and only the 6 August 2026 citation is verifiable today. A change after a site change is not proof the site change caused it. And no audit method entitles anybody, including us, to promise you a position in an answer.

Common questions

Why 30 prompts and not 100?

Because 30 across six groups covers the whole buying decision and can be run properly with repeats in a day or two. A hundred prompts run once, with no repeats and no second labeller, is a worse measurement than 30 run properly, and it costs four times as much.

Is 0 of 6 a bad result or a normal starting point?

It is a normal starting point and it is ours. Zero is where most businesses begin, and the useful part is not the zero, it is the list of which sources answered those six questions instead of us.

Can I count a mention inside a listicle the answer cited?

Count it in its own column and do not fold it into your on domain citation rate. You do not control that page and its author can reorder it next month. On 17 September 2026 the one well known review site in the raw results was discarded for being about American products, which is how quickly a third party page can stop helping.

How do I keep the audit comparable over six months?

Freeze the prompt file, keep the engines, the repeat count, the location and the signed in state identical, rerun on the same day each month, and log every site change with its date in the same workbook. If you must change the prompt set, version it and report the two sets separately rather than joining the numbers.

Who should do the labelling?

A person who did not write the pages, with a written rule for each of the five states, and a second person on a sample. Labelling your own pages generously is the easiest mistake in this work and it is very hard to notice from inside.

Does it matter that the engine may not have searched?

It matters for interpretation, not for validity. A no search reading is a measurement of what the model absorbed, which is still worth having, as long as nobody presents it as a measurement of the web that week. The 27 July 2026 baseline is exactly that case and it is labelled as such.

Reproduce the audit

Everything you need is above: the six groups of five prompts, the blind rule, three named engines, three repeats, the five label states, the recorded conditions, the second labeller, the six discoverability checks, and the rule that every rate is printed as a fraction with its engines and dates. None of it requires our product, and we would rather you ran it by hand first.

Our own capture files are being prepared for publication so that the rows behind the figures above can be checked line by line. Sources: the clawlaw.in programme recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, and the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. Version: 29 September 2026, first publication. Disclosure: we sell an AI visibility product.

What to do first

Run the six discoverability checks before you run a single prompt. They take an hour, they cost nothing, and in our own runs they produced the findings that actually changed outcomes: a pricing page a crawler could not read, two contradicting price lists in public, and a headline number that could not be true. Then write the 30 prompts and book the afternoon.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.