AI Visibility Monitoring Tools Tested: A Reproducible Test Design of Coverage, Citations, and Price

Mentioned once and then gone is usually a repeat count problem, so the number that matters is runs per exact prompt per service.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

If you were mentioned once and then disappeared, the most likely explanation is that you were never reliably mentioned in the first place, and one run told you otherwise. The number that makes any visibility claim interpretable is the count of repeated, documented answer runs per exact prompt per AI service. One run per prompt cannot tell the difference between a real position and a coin landing your way. There is no honest target figure to promise in advance, so here is ours, disclosed: three runs per exact prompt per service per round, in fresh sessions, monthly, which is a schedule we can actually complete. And here is the uncomfortable part, because it is our own data being described: in most of our recorded work to date the repeat count was one. On 27 July 2026 ChatGPT answered 78 blind questions once each. On 18 August 2026 it answered eighteen questions once each. The aiknowsus.com audit of September 2026 covered 24 batches and 72 conversations, which is three conversations per batch rather than three repeats of one prompt. We are naming that limitation rather than presenting those runs as a benchmark.

The only place we have a genuine same day repeat is the one where it changed the conclusion completely. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. Two runs, same questions, same day, opposite conclusions. That is what a single run buys you.

The answer, first

Stated in full, so it can be quoted without the rest of the page. A visibility claim is interpretable when it publishes five things: the exact prompt list with a version and a date, the AI services named individually, the number of fresh runs per prompt per service, the scoring definitions, and the denominator arithmetic written out as a multiplication. A claim missing any of the five cannot be rerun by anybody, including the company that published it.

On the specific question of disappearing, three explanations fit, and they are distinguishable only by repeats.

  • You were never stable. Your mention rate was low, and the run that found you was the lucky one. Three runs per prompt exposes this immediately.
  • The condition changed. The first run retrieved live and the second did not, or the other way round. Log whether any source was cited in each answer and this becomes visible.
  • Something about your site or a third party source changed. The rarest of the three in our experience, and the only one worth a content project.

An assistant has put this better than we can. In batch after batch of the aiknowsus.com audit in September 2026, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is the clearest statement we hold that a single run does not establish a ranking, and it came from the engine rather than from us.

How it was measured

This is the test design in full. It is written so that a reader can run it against any set of monitoring tools, including ours, and get numbers comparable with somebody else's.

One. Build and freeze the prompt set. Between thirty and one hundred prompts in the words buyers type, saved to a file with a version number and a date. The file does not change for the length of the study. If you add prompts, the file gets a new version and every figure you publish names the version it came from.

Two. Apply the blind rule to every prompt used for a visibility rate. No brand names, yours or a competitor's. The 27 July 2026 pair of runs above is the reason, and it is not a preference.

Three. Name the services individually. Write ChatGPT, Claude, Perplexity, Gemini, Copilot, whichever you are testing, each as its own column and its own denominator. Never write all major assistants, because that phrase cannot be rerun.

Four. Set the repeat count and print the arithmetic on the page. Three fresh sessions per prompt per service. With fifty prompts and four services the round is fifty times four times three, which is 600 recorded answers. Every rate you publish from that round carries the multiplication next to it.

Five. Use fresh sessions with no memory. A session that already knows your brand will answer the next prompt differently. Record session state on every row.

Six. Record whether the answer cited any source at all. This field builds the honest denominator for a citation rate, because an answer that cited nothing had no opportunity to cite you. On 27 July 2026 the citable opportunity count across a whole run of 78 questions was zero, because ChatGPT confirmed it had run no live web search for any of them.

Seven. Record search behaviour as a claim, never as a fact. Where the interface shows what it consulted, save it. Where you have to ask the service, mark the answer unverified and keep it in a claimed column beside a verified column. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction. This is not an argument against the engine. It is the reason the field exists.

Eight. Score with fixed definitions, by a person, twice. Mention, citation, clickable link, mentioned through a third party, used without attribution. Have a second person score at least twenty rows blind and publish how often the two agreed.

Nine. Log these fields per answer: date and time, service, mode, session state, location, language, device, prompt version, prompt text, run number, brand mentioned, own domain linked, top source, competitor names present, whether any source was cited, claimed sources, verified sources, and the screenshot file name.

Ten. On the tool side, record capability rather than marketing. For each tool, record which services it queries, whether it exposes the exact prompt text, whether it lets you set the repeat count, whether it reports the denominator, whether it separates mention from citation, whether it exports row level data, and the plan on which each of those is available. Record the date you checked and the URL of the page you read it on.

Eleven. On price, record and do not restate. We do not reproduce any tool's prices here, including our own competitors'. Open each vendor's own pricing page and record the amount, the currency, the plan name, the billing period, the date you read it and the URL, in one row together. The reason is measured: on 17 September 2026 Claude examined two clawlaw.in comparison pages that stated competitors' prices with no link, no date and no source, found the figures correct when it checked them independently, and still treated the pages as advocacy and used official sources instead. A price on a page without a date beside it is discounted even when it is right.

Twelve. Do not restate any vendor's documentation either. Name the document and let the reader open it. A paraphrase of a feature page goes stale and then becomes a false claim about somebody else's product, which is a worse outcome than a missing sentence.

What the numbers were

We publish no tool benchmark on this page, because we have not run one. As of 29 September 2026 we have not tested any monitoring tool, ours included, against the design above, so there is no coverage table, no citation comparison and no price table. That cell is empty and stays empty until the runs exist and the captures are published.

What we hold, with numerators, denominators, services and dates, are these five figures from our own programmes. They are about visibility rather than about tools, and they are in two sectors that must not be mixed.

  • 21 of 78 answers mentioned us, at a repeat count of one. ChatGPT, 27 July 2026, clawlaw.in, legal technology in India. The breadth order in that memory based reading was ProVakil on 28 questions, CLAW on 21 and Legistify on 18.
  • 0 of 78 questions searched live. Same run, same date, so all of the above describes what the model had absorbed rather than what was true that week.
  • 1 of 18 as top source, at a repeat count of one. ChatGPT, 18 August 2026, clawlaw.in.
  • 0 of 6 cited or recommended. Perplexity, September 2026, aiknowsus.com, the AI visibility category. Asked six questions about its own category with no brand named, the assistant audited itself afterwards and reported it had not cited or recommended us in any of the six.
  • 161 recordings of not being cited across 24 batches and 72 conversations. Perplexity, September 2026, aiknowsus.com. That is 72 conversations across 24 batches, which is three conversations per batch, and it is not three repeats of a single prompt. We are printing the construction so nobody can read it as a repeat count.

One more count from the same capture, because it is about tools and it is the only tool level figure we hold. Counted across the whole aiknowsus.com capture of September 2026, the tools named most often in our own category were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a count of how often an assistant named a tool. It is not a quality ranking, it is not a test result, and it should not be read as either.

The tools, and our disclosure

We are the vendor of AI Knows Us, so this section is a disclosure and not a recommendation. We put it first because it is the only tool whose method we can describe from the inside, including its limits, and because everything this page asks of a tool is something we can be held to.

  • AI Knows Us, ours. It runs a frozen blind prompt set across named services on a schedule, records the fields listed above per answer, and reports counts with their denominators. It cannot change an answer, cannot promise a position, and as of 29 September 2026 has not been benchmarked against the tools below by us or by anybody we know of.
  • Semrush, the tool named most often in our capture. Untested by us.
  • Profound. Untested by us.
  • Peec. Untested by us.
  • Otterly. Untested by us.
  • Scrunch. Untested by us.

For each of the five, open the vendor's own current documentation and pricing page, and record the answers to the six capability questions above along with the plan and the date. We are not describing their features here, because a paraphrase written on 29 September 2026 would be a claim about somebody else's product that we cannot keep current.

What this cannot tell you

Six limits.

  • No tool has been tested by us, including our own, against the design on this page. Nothing here is a benchmark result.
  • Our own historical repeat count was one. The figures above are honest counts from single runs, and a single run does not establish a position. The engine said so itself in September 2026.
  • Three repeats is a stability check, not a distribution. It tells you whether a result holds up in one sitting. It does not give you a confidence interval and we will not print one.
  • Sectors do not transfer. The clawlaw.in figures are legal technology in India. The aiknowsus.com figures are the AI visibility category. Neither is evidence about a third sector.
  • Tool coverage changes without notice. Any capability table, ours included, is true on the date it was recorded and no longer, which is why the date belongs in the row.
  • No tool can promise you a position in an AI answer. We do not sell that, and a vendor who offers it is describing something they do not control.

Sources and change log

Every figure above comes from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The tool mention counts and the not cited counts were produced by a script over those capture files rather than from memory, and they cannot yet be checked by a reader from outside, because the captures are not published. The 17 September 2026 discounting of undated competitor prices can be checked from outside today.

Version: 29 September 2026, first publication. The page gets updated when a tool test is completed against the design above and a capability and price table with dates can be published, when any figure is corrected, and when the captures are published.

Common questions

How many runs per prompt is enough?

Three per prompt per service per round is the schedule we disclose and can complete, and it is a stability check rather than a statistical claim. What matters more than the number is that it is stated, held constant between rounds, and used to build the denominator you print. A tool that will not tell you its repeat count is reporting a rate whose meaning you cannot know.

Why did we appear once and then vanish?

Most often because the mention rate was low and one run caught the good outcome. Run the same prompt three times in fresh sessions and record whether each answer cited any source at all. If two of three do not mention you, nothing vanished, and the first result was the outlier.

Can I trust a tool's own mention rate?

Only with five things attached: the prompt list, the services named, the repeat count, the dates, and whether scoring was done by a person or a model. We are a vendor and we would rather you demanded those five from us than not. Without them the number is not auditable, whoever published it.

Is a higher mention count in your capture a sign a tool is better?

No, and it would be wrong to read the Semrush, Profound, Peec, Otterly and Scrunch ordering that way. That count records how often an assistant named a tool in September 2026. It is a measure of what the model had absorbed, in the same way that ProVakil on 28 questions and CLAW on 21 was on 27 July 2026, on a run with zero live searches.

Should we build this in a spreadsheet instead of buying a tool?

For thirty prompts on two services with three repeats, a spreadsheet and two afternoons a month is genuinely enough, and you will understand your own numbers better for having collected them. A tool earns its place when the prompt count, the service count and the repeat count together stop fitting into the time you have.

Why will you not publish a comparison of the tools now?

Because we would have to either test them, which we have not, or describe them from their marketing, which would be a claim about somebody else's product with our name on it. The honest version is the design above plus the instruction to record each vendor's own page with a date. We would rather publish a usable method than an unearned table.

What to do first

Take the prompt that produced your one good mention, run it three times today in three fresh sessions on the same service, and record for each answer whether you were mentioned and whether any source was cited at all. Those three rows will tell you within an hour whether you had a position or a coincidence, and that is the question the rest of this work depends on.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.