What makes a page citable in AI answers? A preregistered 12 week experiment, published before it is run
The hypothesis, the randomisation, the treatment definition and the analysis plan, written down in advance.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Nobody can honestly tell you what makes a page citable from a set of pages they wrote and then admired. It needs a treatment group and a control group, decided by chance, scored on a fixed prompt set over a fixed period. We have not run that experiment, so we are publishing the preregistration instead: the hypothesis, the page pairs, how the assignment is randomised, what the treatment is, the 12 week measurement schedule, the primary outcome and the analysis plan, all fixed before the first page goes up. What we hold today is not an experiment: one named page cited by ChatGPT on 6 August 2026 fourteen days after publication, a top source count of 1 of 18 on 18 August 2026, and an engine reporting it had cited or recommended our own domain in none of six answers in September 2026.
Publishing a design before the result is the only way this number is worth anything when it arrives, because it removes the option of choosing the comparison after seeing the data.
The answer, first
What we can say now, and what only an experiment can settle.
What we can say now, because an engine wrote the reason down. On 17 September 2026, during the clawlaw.in programme, Claude ranked two enterprise vendors above the client on a coverage question specifically because they publish explicit court and tribunal coverage lists, and noted that its ordering reflected price transparency and source authority rather than product quality. On the same date it discounted two of the client's comparison pages because they stated competitors' prices with no link, no date and no source; the figures turned out to be correct when it checked them independently, but it had no way to know that at read time, so it used official sources instead. Those are reasons, dated, on one programme.
What only an experiment can settle. Whether adding those features to a page changes how often it gets cited, by how much, and with what uncertainty. Reasons tell you the mechanism an engine described. A controlled comparison tells you the size of the effect, and nothing else does.
What we will not do. Publish a treatment and control rate we did not measure. That number would be quoted back at us for years.
How it was measured, and the preregistration in full
This is the design, fixed in advance. Anything decided after the data is seen is a different study and has to be labelled as one.
The hypothesis. A page that carries dated primary evidence, sourced third party figures, a named enumeration of scope and a stated limit is cited more often by AI assistants than a matched page on the same site answering the same question without those four things.
The unit. One page answering one buyer question. Not a section, not a site.
The sample. Forty buyer questions on one domain, each of which needs a page that does not yet exist. Forty is the smallest number that gives twenty pages per arm, and the analysis plan below states plainly what twenty per arm can and cannot detect.
The assignment. Each of the forty questions is assigned to treatment or control by a coin, recorded before any page is written, with the assignment list published at the start. No swapping after the fact, including when a question looks too important to put in the control arm. That temptation is exactly what the published list exists to remove.
The treatment, defined so a stranger could apply it. A treatment page carries all four of these, each checkable yes or no.
- At least one number, price or figure with the date it was true written next to it.
- Every figure about another company carrying a link and a date.
- Scope stated as a named, counted list rather than as a word like comprehensive.
- An explicit section saying what the page does not cover, or who should not buy.
The control, defined just as tightly. A control page answers the same question, at a similar length, in the same voice, on the same template, published in the same week, and carries none of the four. It is not a bad page and it is not a short page. Making the control deliberately weak would guarantee the result and prove nothing.
What is held constant. One domain, one template, one author, publication dates spread evenly across both arms, no internal link promotion of one arm over the other, no external promotion of either arm during the twelve weeks, and no editing of any page once published. Every one of those is a way a study like this quietly rigs itself.
The primary outcome, chosen before the run. The number of pages in each arm cited at least once by at least one assistant during the twelve weeks, reported as a count over twenty. One primary outcome, so there is no choosing the flattering one later.
The secondary outcomes. Total citations per arm. Days from publication to first citation per page. Pages fetched by a verified search crawler per arm, taken from server logs. Mentions of the company in answers to that page's question. Each reported as a count with its denominator.
The measurement schedule. Three blind prompts per question, written before the pages are, never naming the company or the product. Run weekly for twelve weeks, one prompt per fresh conversation, on at least two assistants, with the location fixed and the searched field recorded for every answer. Raw answers saved as files, one per prompt per week.
The analysis plan. Report the two counts over twenty side by side, with an interval around each computed the same way for both arms and the method named in the report. State the rule before the run: if the two intervals overlap across most of their range, the result is reported as not readable at this sample size, not as a small effect. That sentence is the whole discipline of the thing, because a study of twenty pages per arm can only detect a large difference, and the honest version of a small difference is silence.
The stopping rules. The run ends at twelve weeks whatever the data shows. If the searched share across the weekly probes is low, the primary outcome is reported alongside that share and treated as provisional, because a period where the engines are not searching cannot test a page feature at all. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of 78 blind questions, which is precisely the condition under which this experiment would be measuring nothing.
What would falsify the hypothesis. Control pages cited as often as or more often than treatment pages, with intervals that do not overlap. We commit to publishing that outcome in the same place and with the same emphasis as a positive one.
What the numbers were
The experiment: not run. There is no treatment count, no control count, no interval and no per arm citation total. We are not going to write one of the form so many of twenty against so many of twenty, because it does not exist.
What we hold instead, each with its engine, date and denominator.
- One page cited, ChatGPT, 6 August 2026. clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india was cited as a source for a vendor due diligence question fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. One page, one question, one engine, and the only result in that programme with a URL a reader outside the company can check.
- Top source on 1 of 18, ChatGPT, 18 August 2026. Across eighteen blind commercial questions the company was the top source on exactly one, and its own comparison page was named in the answer as still being the vendor's own editorial page.
- Not cited in any of 6, Perplexity, September 2026. Asked six questions about its own category with no brand named, the engine audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers.
- 161 not cited statements across 24 batches and 72 conversations, September 2026. Counted by script across all 24 second message answers on our own domain. A count of statements in a capture, not of answers.
- Mentions 21 of 78 with a search rate of zero, ChatGPT, 27 July 2026. Breadth order ProVakil 28, CLAW 21, Legistify 18, and the engine confirmed it ran no live search for any of the 78.
Why none of that is evidence about page features. There was no control arm. The pages that got cited were not paired with matched pages lacking a feature, the assignment was not random, and the programme was publishing many pages at once. A before and after with no control cannot separate a page effect from a change in the engines' behaviour, and our own files show the engines' behaviour changing inside six weeks.
What this cannot tell you
- It cannot tell you what makes a page citable. It tells you what one engine said it valued, on two dates, and what our citation counts were on three.
- A completed experiment would still be one domain. Forty pages on one site in one category in one country, run once. It would be evidence, not a law.
- Twenty pages per arm cannot detect a small effect. That is stated in the analysis plan on purpose, so a null result is not later sold as proof of no effect.
- The four treatment features travel together in the wild. Even a clean result would not tell you which of the four did the work, which is what a later study with single feature arms would have to test.
- Engine behaviour is not held constant and cannot be. In three separate batches in September 2026, Perplexity withdrew its own claim to have run a live search, saying it could not honestly substantiate it. The instrument moves during the study.
- No promise of a position. Applying all four treatment features does not guarantee a citation, and nobody, including us, can sell that guarantee.
Common questions
Why not just publish better pages and watch the numbers go up?
Because you will see them go up and you will not know why. Retrieval behaviour changed inside our own programme between July and September 2026, and a before and after with no control credits that change to whatever you happened to publish.
Is a coin toss really necessary for something this practical?
It is the cheapest part of the design and it is the part that makes the rest mean anything. Without it you will put your four best questions in the treatment arm, because everybody does, and the result will measure your question selection.
Twelve weeks feels long. Can I read it at four?
Look at it at four and do not conclude at four. Our one dated interval from publication to first observed citation was fourteen days, by ChatGPT on 6 August 2026, and one interval does not establish a distribution. Stopping when the numbers look good is the most common way an honest study becomes a dishonest one.
Can I run this with ten pages instead of forty?
You can run it, and report it as a pilot with counts and no claim of an effect. Five per arm cannot distinguish anything short of a total difference. A pilot is still useful: it tests whether your prompts, your logging and your crawler verification actually work before you commit forty pages.
What if the control pages start ranking well and I feel bad about them?
That is a sign the control was built properly. The control is a good page missing four specific features, not a weak page. If you cannot bear to publish it, the honest alternative is to reduce the number of questions rather than to weaken the control.
Would you accept an outside team running this on our site?
That is the version worth most. A design published in advance, run by somebody with no stake in the outcome, with the raw answers kept, is the only form of this study that should change anybody's spending. Disclosure: we sell AI Knows Us, so we have an interest in the answer being yes, and that is exactly why the preregistration is published before the run.
Sources and change log
- Tier_1/claude_response_17_09_audit.md. The 17 September 2026 reasons: unsourced competitor prices discounted as advocacy despite being correct, explicit coverage lists outranking a competitor, official portals above commercial products, being named without being linked, and the dropped attribution.
- Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 27 July 2026 blind run of 78 questions, the zero search confirmation, the breadth order, the discarded branded twin, and the 6 August 2026 citation of a named URL fourteen days after publication.
- Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 blind commercial questions of 18 August 2026 and the top source count.
- geo-audits/aiknowsus-com/. September 2026, 24 batches and 72 conversations: the batch of six with no citation in perplexity__B05-b-perplexity.md, the 161 not cited statements counted by script, and the three batches in which the engine withdrew its search claim.
Change log. 27 July 2026 baseline. 6 August 2026 first cited URL. 18 August 2026 commercial run. 17 September 2026 source selection audit. September 2026 own domain audit. Preregistration status: published, not yet started. When it starts, this section will carry the start date, the published assignment list of forty questions, and the week by week probe log. When it ends, it will carry both arm counts and their intervals, including if the control arm wins.
What to do first
If you have forty questions without pages, do the preregistered version: write the assignment list, toss the coin, and publish the list before you write a word. If you have fewer, do the pilot: four questions, two arms, three blind prompts each, weekly for four weeks, and use it to test your logging rather than to claim a result. Either way, the first thing to write down is the primary outcome, because a study that chooses its outcome after the data is not a study.