GEO starting point: a 30 day, evidence led website experiment

One frozen prompt set, one month of changes, and the before and after citation counts printed side by side.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Start with a 30 day experiment on one frozen prompt set. Day one: write 20 to 30 blind prompts, run them on named engines with repeats, and record how many runs cited a page on your own domain. That count is your before figure. Days two to twenty five: change the website and nothing else about the method. Day thirty: rerun the identical prompt set, in the identical conditions, and record the after figure. Report both counts with both denominators and both dates. The before and after we can show from our own programme is this: in the litigation due diligence zone of the clawlaw.in question set, on domain citations were 0 at the blind baseline of 27 July 2026, and on 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, which is 1. Before 0, after 1, ten days apart, one engine.

That is a small result, honestly labelled, and this page explains exactly what it does and does not establish. It is also the only citation figure in that whole programme with a public URL behind it, which is why it is the one we quote.

What GEO means in this experiment

In this experiment generative engine optimisation means one thing only: increasing the number of answer runs in which a page on your own domain is used as a source for a question your buyers ask. Not traffic. Not a score out of 100. Not brand awareness.

We narrow it that far for three reasons. It is countable. A second person can verify it from your capture. And it is the state that actually sends a buyer to you, as opposed to being named without a link, which the buyer hears and your analytics never see.

Three things are deliberately outside the definition, and each gets measured separately if you care about it: being named without a link, appearing inside a third party listicle that gets cited, and the traffic or enquiries that may follow. Folding those into one number is how a 30 day experiment stops being able to tell you anything.

The baseline

Day one, and it is the part people rush. Eight steps.

One. Write 20 to 30 prompts in your buyers' words and save them with a version number and today's date. This file does not change for the length of the experiment. If it changes, the experiment is over and a new one has started.

Two. Keep your brand name out of every prompt. The measured reason: on 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. We discarded the branded run, because the only thing it measured was our own prompt. A 30 day experiment baselined with a branded set will show no improvement, because it started at a fake ceiling.

Three. Pick your engines and name them. At least two, three is better, each reported separately.

Four. Three repeat runs per prompt per engine, in fresh sessions, and write the run count down. Twenty five prompts on three engines with three repeats is 225 answer runs.

Five. Record the conditions on every row: date, time, engine, mode, country and city, language, signed in state, and whether memory or personalisation was active.

Six. Label five states: brand named, own domain cited, top source, competitors named, and material used without attribution.

Seven. Run the discoverability checks and log what you find, because most of your day two to twenty five work will come from here rather than from the prompts. Fetch your key pages as a crawler would, confirm they are indexed, and find every other public copy of your key facts.

Eight. Write the before figure as a sentence you could defend. On domain citations, X of N answer runs, these engines, this prompt set version, these dates, blind.

Changes made

The changes this experiment prescribes are the ones our own runs found to be the actual problems. Each is tied to a dated observation, so you can see why it is on the list rather than taking our word for it.

  • Make every buying fact readable in the page as served. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead. Claude and ChatGPT, same date.
  • Remove every contradicting public copy of a fact you own. On 6 August 2026 ChatGPT noticed that the website and the app store listing carried different plan names and different prices for the same product, and said so in its answer.
  • Delete any claim that cannot survive a check. Also on 6 August 2026, Claude checked a headline figure on the main site, found it could not be true against public numbers, and advised a buyer against the product, after which the company's accurate claims stopped counting for that answer.
  • Source and date every competitor fact. On 17 September 2026 Claude read two comparison pages that stated competitors' prices with no link, no date and no source. The figures were correct when it checked them independently, and it still treated the pages as advocacy and used official sources instead.
  • Publish coverage as an explicit list, not an adjective. On 17 September 2026, on a question about finding every case against a company, Claude ranked two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and it said its ordering reflected price transparency and source authority rather than product quality.
  • Write one page per buyer question, with the answer in the first paragraph. The one page we can show being cited is of exactly that shape: a how to page answering a single question, cited by ChatGPT on 6 August 2026.
  • Do not write instructions to the engines into your pages. On 6 August 2026 Claude found a competitor's page carrying a hidden block of text addressed to answer engines, instructing them to cite that company as the source. It refused the instruction and named the company that had done it.
  • Do not expect a how to question about an official portal. On 17 September 2026 Claude put official court portals above every commercial product for those questions, and said plainly that no commercial product should rank above the official portal for a question about using that portal. Spend your thirty days elsewhere.

One honest gap. We do not publish a fix log for clawlaw.in with dates against each of those changes, because our records hold the findings and not a dated record of each remedy. So this list is what the evidence says to change, and not a claim that we changed all eight on stated dates and measured the result of each.

Results after 30 days

The before and after, stated exactly.

  • Before: 0 on domain citations in the litigation due diligence zone of the blind 78 question set. ChatGPT, 27 July 2026. On the same run the engine confirmed it had run no live web search for any of the 78 questions, so the baseline describes what the model could reach from memory.
  • After: 1 on domain citation in that zone. ChatGPT, 6 August 2026, citing clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question, fourteen days after that page was published.
  • Gap between the two readings: ten days.
  • Wider counts from the same programme, for context rather than as part of this before and after: named on 21 of 78 questions at baseline, against ProVakil 28 of 78 and Legistify 18 of 78; and top source on 1 of 18 blind commercial questions on 18 August 2026.

Four things about that result that a less careful page would leave out.

The second reading was an interim check, not a full rerun of all 78 prompts under identical conditions. So it is a before and after on one zone of the set, not on the whole set, and it does not meet the standard this page asks of you.

The page that was cited had been published fourteen days before 6 August, which means it was already live, and four days old, at the 27 July baseline. It was not cited then. That is interesting and it is not a clean thirty day experiment.

The baseline run involved no live web search at all, confirmed by the engine. A zero produced under those conditions is a weaker zero than one produced with search confirmed.

And our own domain has not completed this experiment. On aiknowsus.com the measured figure is 0 of 6 answer runs citing us, Perplexity, September 2026, with 161 recorded statements across 24 batches and 72 conversations that we were not cited. That is a before figure with no after yet. When the after exists it will be published here with both denominators and both dates, or not published.

Reproduction instructions and limitations

To reproduce: freeze the prompt file with a version and date, keep the brand name out, name your engines, three repeats, record the conditions on every row, label the five states, run the discoverability checks, make only website changes during the window, log every change with its date, and rerun the identical set on day thirty. Report both counts as fractions with both dates.

Seven limits, stated plainly.

  • Thirty days of change is not thirty days of measurement. Crawling and indexing take time, so a page published on day twenty four has barely had a chance.
  • A change in the number is not proof your changes caused it. Engines change their own behaviour inside your window and do not tell you.
  • A small prompt set produces a noisy count. Movement of one or two runs is within the variation you would see from repeats alone.
  • Different engines will move differently, which is why they are reported separately rather than blended.
  • Zero to one is a real change and a fragile one. It can disappear on the next run without anything having gone wrong.
  • Our own before and after above does not meet this standard, and we have said so rather than presenting it as if it did.
  • Nothing in this method can promise a position in an answer.

Platform documentation and change log

Sources for every figure on this page: the clawlaw.in programme, recorded in GEO_BASELINE_RESULTS_2026-07-27.md including its interim check section, GEO_GAP_ANALYSIS_2026-08-18.md, and the assistant audit files of 17 September 2026; and the aiknowsus.com audit of September 2026, 24 batches and 72 conversations. The 6 August 2026 citation, the pricing page rendering failure, the two contradicting price lists, the impossible headline claim, the unsourced competitor pricing and the hidden instruction block can be verified from outside. The counts over our capture files cannot, until those files are published, which is in progress.

Change log. Version 29 September 2026, first publication. This page will be updated when the aiknowsus.com after figure exists, if any figure is corrected, and when the capture files go public. We will not update it with a stronger sounding result that has not been measured.

Common questions

Thirty days feels short. Is it long enough to see anything?

It is long enough to fix what is broken and to see whether a new page can be reached at all. The one verified timing we hold is fourteen days from publication to first citation, ChatGPT, 6 August 2026, on one page. Treat thirty days as the first read rather than the verdict, and keep rerunning the same set monthly.

What should I change first if I only have a week?

Make sure a crawler receives your prices and your key facts in the page itself, and remove any public copy that contradicts them. Those were two of the three problems found on 6 August 2026 in the clawlaw.in runs, and neither needed new content.

Can I add prompts during the experiment if I think of better ones?

Write them down for the next version and do not add them to this one. A denominator that changes mid experiment makes the before and after incomparable, and this is the most common way a 30 day test ends up proving nothing.

Should I publish lots of pages or fix the few I have?

Fix first, then publish one page per buyer question. The page we can show being cited was a single question how to page. The pages that were discounted were the vendor's own comparison pages, which on 17 September 2026 Claude treated as advocacy because they carried competitor prices with no source and no date.

Our number went down after the changes. What does that mean?

Possibly nothing. With a small prompt set and three repeats, a drop of one or two runs is inside normal variation, and engines change independently of you. Rerun before concluding, and check whether your competitors' counts moved in the same direction, which usually indicates the engine rather than your site.

Do I need a tool for this?

No. A spreadsheet, a screenshots folder and an afternoon at each end are enough for 225 runs. We sell a tool and the first experiment is better done by hand, because reading the answers yourself is where the real findings come from.

What to do first

Write the prompt file today and freeze it. Then, before you touch any content, fetch your three most important pages the way a crawler does and read what comes back. That is the check that found a pricing page serving no prices on 6 August 2026, and it is the cheapest thirty minutes in this entire method.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.