Can You Promise AI Search Visibility? A Repeated-Run Outcome Study Design

No, and the reason is a study design nobody has run yet, including us, stated with the two runs we do hold.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

No. Nobody can promise AI search visibility, and an agency or vendor that does is selling something they do not control. What can honestly be promised is a measurement: a citation rate at baseline and at follow up, each with its numerator and denominator, on a frozen blind prompt set, with the same services, the same runs per prompt and a comparison set you did not act on. We have not run that study. As of 29 September 2026 we hold no baseline and follow up citation rate pair on an identical prompt set under identical conditions, so no such change appears on this page. The nearest pair we hold is worth stating precisely, because the reason it is not a rate is the whole lesson. At baseline, on 27 July 2026, ChatGPT answered 78 blind questions for clawlaw.in and named the company nowhere in the litigation due diligence questions it most wanted to win, on a run where it confirmed it had run no live web search for any of the 78. Fourteen days later, on 6 August 2026, ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question. Nothing to something, on one assistant, checkable from outside. It is not a citation rate change, because the follow up was not a rerun of the same 78 prompts under the same conditions, and one run had zero retrieval while the other clearly involved some.

We believe the page was the likely reason for that citation. We cannot show it, because there was no comparison set and fourteen days of assistant behaviour sit between the two observations. The difference between those two sentences is what this page is about, and it is also the line an honest promise sits behind.

The answer, first

Four things can be promised in writing, and four cannot. Put both lists in the contract.

Can be promised. A dated baseline with a numerator and a denominator on a frozen blind prompt set. A stated repeat schedule of runs per prompt per service. Named deliverables: these pages, these facts published as text, these third party records corrected. And every round reported, including the rounds that go backwards.

Cannot be promised. A position in any AI answer. A share of answers that will name the client. A date by which any assistant will change what it says. And an attribution of a change to your work, unless the study carried a comparison set.

The engines themselves take this line on single runs. In batch after batch of the aiknowsus.com audit in September 2026, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. If an assistant will not rank on the strength of its own single runs, a vendor should not promise a position on the strength of a proposal.

How it was measured

This is the outcome study design that would support a defensible before and after. It is published in full so that a client can hold any vendor, including us, to it.

One. Freeze two prompt files before any work begins. A treated set your work targets and a comparison set you will measure and never act on. Allocate them at the start, write the allocation down, and make the two similar in subject and in starting visibility. Allocating after you see results turns a study into a story.

Two. Apply the blind rule to both files. No brand names anywhere. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it measured was our own prompt.

Three. Name the services and hold them constant across both measurements. A service added between baseline and follow up makes the pair incomparable.

Four. Set the runs per prompt and print the arithmetic at both ends. Three fresh sessions per prompt per service. With 50 treated prompts, 50 comparison prompts, two services and three runs, one measurement is 100 times two times three, which is 600 responses, being 300 treated and 300 comparison. Baseline and follow up together are 1,200. Print both multiplications in the report.

Five. Record whether any source was cited in each response. This builds the honest denominator for a citation rate, because a response with no retrieval had no opportunity to cite anyone. At the clawlaw.in baseline the citable opportunity count was 0 of 78, and a citation rate computed over those 78 responses would have described nothing.

Six. Keep claimed sourcing separate from verified sourcing. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction.

Seven. Score with fixed definitions at both ends, by a person, twice, and publish the agreement rate each time. Scoring drift between baseline and follow up produces a change that is entirely your own.

Eight. Hold the window fixed and state it. Ninety days between measurements is a reasonable window because it is long enough for published pages to be found and short enough that the client can act on the result. Say the window in the report and do not move it to catch a better month.

Nine. Report four rates, then the difference. Treated baseline, treated follow up, comparison baseline, comparison follow up, each as a numerator over a denominator. Then the treated change minus the comparison change, in percentage points. If there was no comparison set, print the two treated rates and state in one sentence that the change is not attributable.

Ten. Publish uncertainty only if you have earned it. If your repeat count does not let you describe the spread, say so. We have not, so we do not.

What the numbers were

We publish no baseline and follow up citation rate pair, because no such study exists in our records as of 29 September 2026. The cell is empty and will stay empty until a study with a comparison set and identical conditions at both ends has been completed and its captures published.

The five dated figures we do hold, each with its engine, its date and its sector, printed separately because they cannot be combined into a change.

  • Baseline: 21 of 78 answers mentioned the company, and 0 of 78 involved a live search. ChatGPT, 27 July 2026, clawlaw.in, legal technology in India. Breadth order in that reading: ProVakil 28 of 78, CLAW 21 of 78, Legistify 18 of 78. The company was named nowhere in the litigation due diligence questions it most wanted to win.
  • Follow up observation: one named URL cited, fourteen days after publication. ChatGPT, 6 August 2026. The cited page was clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india and the question was about vendor due diligence. This is checkable by a reader from outside.
  • 1 of 18 as top source. ChatGPT, 18 August 2026, clawlaw.in. A different prompt set from the baseline, so it is not the follow up to it.
  • 0 of 6 cited or recommended. Perplexity, September 2026, aiknowsus.com, the AI visibility category, which is our own domain.
  • 161 recordings of not being cited across 24 batches and 72 conversations. Perplexity, September 2026, aiknowsus.com, being three conversations per batch rather than three repeats of one prompt.

Why we do not subtract any of these from each other. The 78 question set and the 18 question set are different prompt files. The 18 question run and the six question run are different engines, different domains and different sectors. And the baseline run had no retrieval at all while the 6 August observation clearly involved some, which means the two describe different conditions rather than two points in a series. The arithmetic would run. The result would mean nothing, and pages that publish exactly that subtraction are the reason a careful buyer discounts vendor numbers.

Why small denominators make a promise easy to fake

This section is arithmetic in round numbers and is not an observation. With a denominator of 18, one response is about 5.6 percentage points, so a vendor who gains a single mention can report that visibility rose by more than five points. With a denominator of 6, one response is about 16.7 points. With 300 responses, one is about 0.33 points. A promise expressed in percentage points against a small denominator is a promise about one or two answers, and a different afternoon would produce the same move.

Two rules follow. Always print the numerator, so a reader can see how many responses moved. And never report a percentage point change smaller than one response in your own denominator, because that change cannot exist in your data.

What this cannot tell you

Six limits.

  • We hold no repeated run outcome study, and no comparison set has ever been run in our programmes.
  • The 6 August 2026 citation is not attributed to the page that preceded it. We think it is the likely reason. There was no holdout and fourteen days of assistant behaviour sit in between.
  • No rate here has a repeat count above one, which means none of them establishes a position, only what one run returned.
  • We print no confidence intervals. An interval without the repeats behind it is decoration.
  • Sectors do not transfer. Legal technology in India and the AI visibility category are the only two these figures describe.
  • Nothing in this design gives anybody control over an answer. A study tells you whether something moved. It does not give a vendor the ability to promise it will.

Sources and change log

Every figure above comes from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The 6 August 2026 citation of a named URL can be checked from outside today. Counts over our own capture folders cannot, until we publish the captures.

Version: 29 September 2026, first publication. The page gets updated when a ninety day study with a comparison set is completed, at which point it will carry four rates with their denominators and a difference in percentage points, and when any figure here is corrected.

Common questions

What should I ask a vendor who promises results?

Six questions. What is the frozen prompt file and can I have it. Which services, named. How many runs per prompt. What is the comparison set. What is the denominator arithmetic. And will every round be reported including the bad ones. A vendor who cannot answer all six is not measuring, and a vendor who promises a position after answering them is contradicting their own method.

Is a money back guarantee on AI visibility reasonable?

A guarantee tied to delivered work is reasonable. A guarantee tied to an assistant naming you is not, because the vendor is underwriting a system neither party controls, and the argument will end up being about measurement rather than about outcome. Ask instead for the measurement to be contractual.

You cited one page in fourteen days. Why not promise fourteen days?

Because one observation is not a timetable. It was one page, one assistant, one date, in one sector, with no comparison set and no repeats. Turning it into a promise would be the exact behaviour the rest of this page argues against, and a client who checked would be right to discount everything else we said.

How long should the window between baseline and follow up be?

Ninety days is defensible because it is long enough for new pages to be found and short enough to act on, and what matters more is that the window is fixed in advance and stated. A window chosen after the fact, to include a good month, is not a study.

Can a citation rate go down through no fault of ours?

Yes, and that is the single strongest argument for a comparison set. If the untouched set fell by the same amount, the month changed and your work did not fail. Without the comparison set, neither you nor the client can tell those two situations apart.

What is a fair thing to promise a client, in one sentence?

A dated baseline and follow up on a frozen blind prompt set, with numerators, denominators and a comparison set, plus named deliverables, and every round published. That is a promise about work and measurement, both of which you control, and it is the only kind that survives being checked.

What to do first

Split your prompt list in two before you change anything on the site, write down which half is treated and which is the comparison, and measure both three times each this week. Then set the follow up date ninety days out and put it in the calendar. That single decision is what makes the next conversation about evidence instead of about belief.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.