How to Report AI Search Visibility: A Three-Run Agency Reporting Format

A visibility report is only readable when the response count is shown as a multiplication the client can check.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

A visibility report is readable only when it states the number of responses in the dataset and shows the arithmetic that produced it: prompts times services times runs, with the dates. Our own recorded datasets, printed the same way, are these. 78 responses on 27 July 2026, from 78 prompts, one service, one run, on ChatGPT for clawlaw.in, plus a second run of the same 78 prompts on the same day that was discarded because its wrapper named the brand. 18 responses on 18 August 2026, from 18 prompts, one service, one run, on ChatGPT. And the aiknowsus.com audit of September 2026, which is 24 batches and 72 conversations on Perplexity, being three conversations per batch rather than three runs of one prompt. Our historical repeat count was one, and that is the weakness of our own reporting, stated here rather than hidden behind a larger looking number.

A response count is the first thing a client should look for and the first thing most reports omit. It is also the number that exposes a report built from a single pass, because prompts times services times runs cannot reconcile to the total unless the runs are real.

The answer, first

Every visibility report should open with one paragraph containing the following seven items, in this order, before any chart or any finding.

  • The response count, as a total.
  • The arithmetic behind it, written out as prompts times services times runs.
  • The prompt file version and date.
  • The services, named individually.
  • The runs per prompt per service.
  • The capture date range, the location setting and the session state.
  • The count of responses excluded, with the reason, because a dataset with silent exclusions does not reconcile and a client who tries the multiplication will notice.

A worked reconciliation, in round numbers rather than as an observation: 100 prompts, four services, three runs is 100 times four times three, which is 1,200 responses. If 37 responses failed to capture and were excluded, the report says 1,163 analysed out of 1,200 attempted, with the 37 named as capture failures. A report that says 1,200 and analysed 1,163 has lost the client's trust for a reason that was avoidable.

How it was measured

The construction rules that make a response count mean something.

One. Freeze the prompt file with a version number and a date, and name that version in every report drawn from it. A prompt set edited between rounds makes the rounds incomparable and nothing can repair it afterwards.

Two. Apply the blind rule to every prompt used for a rate. No brand names. The reason is measured: on 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded. Both runs are counted in our dataset description above, and the discarded one is named as discarded rather than deleted, because a reader is entitled to know it happened.

Three. Name the services and keep them constant between rounds. A service added mid study is a new study, and the response count for each service is reported separately as well as in the total.

Four. Set the runs per prompt per service and hold it constant. Three is a defensible schedule that can be completed. Whatever you choose, it appears in the opening paragraph and it does not change between rounds.

Five. Use fresh sessions with no memory carried over, and record session state per row, because a session that already knows the brand answers the next prompt differently.

Six. Record whether any source was cited at all in each response. This field gives you two honest denominators: all responses, and responses that could have cited anybody. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of the 78 questions, so the citable opportunity count for that dataset was 0 of 78, and a citation rate reported over those 78 responses would have been meaningless.

Seven. Record sourcing as a claim in one column and as verified in another. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction. A report that presents a self reported search count as a fact is reporting a claim as data.

Eight. Score with five fixed definitions and publish the agreement rate between two scorers on a sample of at least twenty responses. Mention, citation, clickable link, mentioned through a third party, used without attribution. The fifth exists because on 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places.

Nine. Keep the row level log and offer it with the report. One row per response, with date and time, service, mode, session state, location, prompt version, prompt text, run number, the five scores, whether any source was cited, claimed sources, verified sources, and the screenshot file name. A client who can recount your numerator will trust the ones they do not recount.

What the numbers were

Our recorded datasets, with their construction shown, and the rates drawn from them. Two sectors, kept apart.

  • 78 responses, 27 July 2026, ChatGPT, clawlaw.in, legal technology in India. 78 prompts, one service, one run. Brand mention rate 21 of 78. Breadth order in the same reading: ProVakil 28 of 78, CLAW 21 of 78, Legistify 18 of 78. Live searches evidenced: 0 of 78.
  • A second 78 response run, same date, discarded. Same prompts, one service, one run, with a wrapper that named the brand. Not used for any rate, and named here so the total attempted reconciles to 156.
  • 18 responses, 18 August 2026, ChatGPT, clawlaw.in. 18 prompts, one service, one run. Top source rate 1 of 18.
  • 72 conversations across 24 batches, September 2026, Perplexity, aiknowsus.com, the AI visibility category. Three conversations per batch. The phrase recording that we were not cited appears 161 times in the engines' own self audits across that capture. In one batch of six questions the assistant reported it had cited or recommended us in 0 of 6 answers.

What we have not run, stated plainly. As of 29 September 2026 we have not completed a dataset with three runs per prompt per service, we have not run a study with a comparison set, and we have not published the capture files. So we publish no rate with a repeat count above one and no attributed change. The only dataset above that a reader can partly check from outside is the clawlaw.in citation of 6 August 2026, when ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline.

The report format, page by page

Six pages. Keep them in this order and the report is hard to misread.

  • Page one: the dataset paragraph. The seven items from the answer section, and nothing else.
  • Page two: the headline rates, each as a numerator over a denominator, per service, never pooled into one figure without the per service figures beside it.
  • Page three: the citation rates, twice. Over all responses, and over responses that cited any source. State both and say which is which.
  • Page four: what moved and what did not, with the comparison set's change printed beside the treated set's change, and a plain sentence if there was no comparison set.
  • Page five: named examples, two or three actual responses quoted with their prompt, date and service, including one where the client did badly.
  • Page six: the limits, copied from the section below, unedited between rounds.

Two rules that hold across every round. Publish the rounds that went backwards, because a series with a missing round reads as a stable round, which is a lie told by omission. And never report a percentage point change smaller than one response in your own denominator, because that change cannot exist in your data. On a denominator of 18, one response is about 5.6 percentage points.

What this cannot tell you

Six limits, and they belong on page six of every report you send.

  • A response count is not a sample of the web. It is a record of what named services returned for named prompts on named dates.
  • Our own historical repeat count was one, and a single run does not establish a position. In September 2026 Perplexity itself declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim.
  • No rate here has a comparison set behind it, so no change in our records is attributed to any action.
  • We print no confidence intervals, because we have not run the repeats that would justify one.
  • Sectors do not transfer. Legal technology in India and the AI visibility category are the only sectors these datasets describe.
  • No report format can promise a position in an AI answer, and we do not sell one.

Sources and change log

The datasets above come from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, and from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. Counts were produced by a script over those files rather than from memory. The 6 August 2026 citation of a named URL can be checked from outside today. The capture wide counts cannot, until the captures are published.

Version: 29 September 2026, first publication. The page gets updated when a dataset with three runs per prompt per service is completed, when a comparison set is run, when any figure is corrected, and when the captures are published.

Common questions

What is the minimum a client should demand in a visibility report?

The response count with its multiplication, the prompt file version, the services named, the runs per prompt, the dates, and the count of exclusions. Six items. Any report missing them cannot be rerun, and a report that cannot be rerun cannot be compared with next month's.

Should the report show one pooled rate or one per service?

One per service, always, with the pooled figure only as an addition. The services select sources differently, so a pooled rate averages two different behaviours and hides which one changed. Each service is its own denominator and should be printed as one.

How do I report a round where the numbers got worse?

Exactly as you report a good one, in the same format, with the numerator visible so the client can see how many responses moved. On a denominator of 18 a loss of one response is about 5.6 percentage points and is not necessarily a change at all. The format protects you here, which is part of why it is worth keeping.

Can I count a response twice if it mentioned the client twice?

No. One response is one row and one unit in the denominator. Count mentions per response in a separate column if the depth of mention matters to you, and label it clearly, because a reader will otherwise assume your numerator is responses.

Is a hundred prompt dataset better than a forty prompt one?

Only if you will rerun it. A hundred prompts across four services with three runs is 1,200 responses per round, which is a large commitment. Forty prompts run every month for six months tells a client far more than a hundred prompts run once, because direction needs repeats and a photograph does not become a series.

What do I do about responses that failed to capture?

Report them. Attempted, analysed, and excluded with the reason, on page one. Silent exclusions are the commonest way a response count stops reconciling, and the client who spots it will reasonably doubt everything else in the document.

What to do first

Open your last client report and try to reconcile its response count by multiplying prompts, services and runs. If you cannot, rewrite page one before you send the next one: the total, the multiplication, the prompt file version, the services, the runs, the dates, and the exclusions. That single paragraph is the difference between a report a client can check and a report they have to believe.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.