AI Visibility Improvement or Noise?
The only honest answer is your change minus an untouched comparison set's change, and we have never run a holdout.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
You cannot tell whether an improvement in AI visibility is real from your own numbers alone, because an assistant changes its behaviour between your two measurements and that change is invisible to you. The only figure that separates your work from the month is the difference of differences: the change in your treated prompt set, minus the change in a comparison set you deliberately did not touch, in percentage points, with both denominators printed. We have never run a holdout. As of 29 September 2026 no measurement in our records has a comparison group, so we cannot give you a single difference of differences figure for anything, and we will not construct one out of two unrelated runs. What we hold is the material for showing you why that refusal matters: 1 of 18 on ChatGPT on 18 August 2026 for clawlaw.in in legal technology, and 0 of 6 on Perplexity in September 2026 for aiknowsus.com in the AI visibility category. Different engines, different question sets, different sectors, different denominators. Subtracting one from the other would produce a number with no meaning, and pages that do exactly that are the reason this one exists.
The rest of this page is the design that would produce an honest answer, the arithmetic of how small denominators behave, and a plain statement of what has not been run.
The answer, first
Four conditions have to hold before a change in AI visibility can be called real rather than noise. All four, not three.
- The same prompts, word for word, at both measurements. A prompt set edited between rounds makes the two rounds incomparable, and nothing can repair that afterwards.
- The same services, modes, locations and session states. Recorded per row, not assumed.
- More than one run per prompt at each measurement. Without repeats you cannot tell a real position from a single lucky answer.
- A comparison set you did not act on. Prompts, or pages, deliberately left alone for the whole study. This is the one almost everybody skips, and it is the one that makes the difference interpretable.
The engines themselves take this position. In batch after batch of the aiknowsus.com audit in September 2026, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. An assistant refusing to rank on the strength of its own single runs is the clearest statement we hold that one pass does not establish anything.
We also have a same day demonstration of how large the noise can be when a condition changes. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. Same questions, same day, opposite conclusions, with only the wording of the wrapper changed. The first run was discarded.
How it was measured
This is the design that would answer the question, published in full so a reader can run it without us.
One. Split your prompts into two sets at the start, before any work. A treated set, which your content work targets, and a comparison set, which you leave completely alone. Split them so the two are similar in subject and in starting visibility, and write the allocation down before you begin. Allocating after you see the results is how a study becomes a story.
Two. Freeze both files with a version number and a date, and do not edit either one for the length of the study, including the comparison set when it becomes tempting.
Three. Apply the blind rule to every prompt in both sets. No brand names. The 27 July 2026 pair above is the reason.
Four. Name the services individually and keep them the same at both measurements. A service added between rounds is a new study.
Five. Set the repeat count and print the arithmetic. Three fresh sessions per prompt per service at each measurement. With forty treated prompts, forty comparison prompts, three services and three repeats, one measurement is eighty times three times three, which is 720 recorded answers, made of 360 treated and 360 comparison. Baseline and follow up together are 1,440. Print the multiplication next to every rate, both times.
Six. Record whether any source was cited in each answer, so your citation rate has an honest denominator. An answer produced with no retrieval had no opportunity to cite anyone. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of 78 questions, which makes the citable opportunity count for that entire run 0 of 78.
Seven. Treat sourcing statements as claims. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction. Keep claimed and verified sources in separate columns at both measurements.
Eight. Score with fixed definitions, by a person, twice, at both measurements, and publish the agreement rate each time. A scoring rule that drifts between rounds produces a change that is entirely your own.
Nine. Compute four rates, not two. Treated baseline, treated follow up, comparison baseline, comparison follow up, each with its own numerator and denominator printed. Then the difference of differences, in percentage points, as the treated change minus the comparison change.
Ten. Report the uncertainty honestly or not at all. If you have not run enough repeats to describe the spread, say so instead of printing an interval. We have not, so we do not.
What the numbers were
We publish no difference of differences figure, because we have never run a comparison set. As of 29 September 2026 that is true of every measurement in our records, on clawlaw.in and on aiknowsus.com. The cell is empty and stays empty until a study with a holdout has been completed and the captures published.
What we hold are four dated single measurements, each with its numerator, denominator, engine, date and sector. They are printed separately because they cannot be combined.
- 21 of 78 answers mentioned us. ChatGPT, 27 July 2026, clawlaw.in, legal technology in India. Breadth order in that reading: ProVakil 28, CLAW 21, Legistify 18, out of 78.
- 0 of 78 questions searched live. Same run, same date, so the figure above describes what the model had absorbed rather than the web that week.
- 1 of 18 as top source. ChatGPT, 18 August 2026, clawlaw.in. Across eighteen blind commercial questions the company was the top source on exactly one.
- 0 of 6 cited or recommended. Perplexity, September 2026, aiknowsus.com, the AI visibility category.
There is also one dated before and after in our records, and it is worth stating exactly what it is and is not. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. That is a real change from nothing to something, on one URL, on one assistant, checkable by a reader from outside. It is not evidence that publishing the page caused the citation, because there was no comparison set, and something else could have changed on that assistant in those fourteen days. We think the page was the likely reason. We cannot show it, and the difference between those two sentences is the subject of this page.
Why small denominators look like movement
This section is arithmetic in round numbers, not an observation, and it is here because it explains most false improvements.
With a denominator of 18, one answer is one eighteenth, which is about 5.6 percentage points. So a set of eighteen questions where you gain a single mention produces a headline that says visibility rose by more than five points, and nothing has happened that a different afternoon would not also have produced. With a denominator of 6, one answer is about 16.7 percentage points, so a single answer swings the rate by a sixth. With a denominator of 78 one answer is about 1.3 points, and with 720 it is about 0.14.
Two rules follow from that arithmetic. Always print the numerator, because the reader needs to know that the move was one answer. And never report a percentage point change that is smaller than one answer in your own denominator, because that change cannot exist in your data.
What this cannot tell you
Six limits.
- We hold no holdout study, so no page of ours can attribute a change to our work, and this page does not.
- Our measurements cannot be subtracted from each other. 1 of 18 on ChatGPT in August and 0 of 6 on Perplexity in September are different engines, different question sets and different sectors. The arithmetic would run and the result would mean nothing.
- We print no confidence interval anywhere, because we have not run enough repeats to describe the spread, and an interval invented to look rigorous is worse than none.
- A difference of differences still does not prove a mechanism. It tells you the treated set moved more than the untouched one. It does not tell you which change did it, if you made several.
- Sectors do not transfer. Legal technology in India and the AI visibility category are the only two sectors we hold anything in.
- No design here can promise a position in an AI answer, and we do not sell that.
Sources and change log
Every figure above comes from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The 6 August 2026 citation of a named URL can be checked from outside today. The counts over our capture folders cannot, until we publish the captures.
Version: 29 September 2026, first publication. The page gets updated when a study with a comparison set has been completed, at which point it will carry four rates with their denominators and a difference of differences in percentage points, and when any figure here is corrected.
Common questions
Can I use last quarter's numbers as the comparison?
No, and this is the commonest version of the mistake. Last quarter is a different month of assistant behaviour, so a before and after comparison of the same prompts includes every change the assistant made in between. The comparison has to run at the same time as the treated set, not before it.
What should go in the comparison set?
Prompts on subjects you are not working on this quarter, chosen at the start, similar in kind and in starting visibility to the treated ones. Alternatively, pages you deliberately leave untouched while you improve others. Either works. Deciding after you see the results does not.
Is a holdout worth the cost for a small business?
If you are spending money on this work and want to know whether it did anything, yes, and the cost is lower than it sounds because the comparison prompts need no content work, only measurement. If you are not going to run one, the honest report says the change is not attributable, which is still a useful sentence to be able to say out loud.
Our mention rate went from one answer to three. Is that real?
On a denominator of eighteen, that is a move of about eleven percentage points produced by two answers, which is well within what a different afternoon can produce. Run the same prompts three times in fresh sessions at both measurements and look at whether all three runs agree. If they do not, you have noise, and the size of the percentage is irrelevant.
Why will you not give a confidence interval?
Because we have not run the repeats that would justify one, and an interval printed without them is decoration. When we have a study with a holdout and three runs per prompt at both ends, the spread will be published with the raw rows behind it so a reader can recompute it.
Did publishing that clawlaw.in page cause the citation on 6 August 2026?
We think it was the likely reason and we cannot show it. There was no comparison set, the same question set had named the company nowhere at baseline, and fourteen days passed. A reader is entitled to treat it as one candidate explanation among several, including the assistant changing its own behaviour in that period.
What to do first
Before you change anything on your site, split your prompt list in two, write down which half you are going to work on, and measure both halves three times each on the same day. That single hour of extra measurement is what turns next quarter's report from a story into a finding, and there is no way to add it retrospectively.