How to measure GEO without overclaiming: a before and after measurement protocol

Four measurement layers, both periods reported in full, and the line where observation stops and causation would begin.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

You prove GEO work is working by rerunning the identical prompt set and reporting both periods in full: numerator, denominator and dates for each, per engine, with the prompt set version named. The number that carries the argument is the change in on domain citations. Ours, from the clawlaw.in programme, is this: 0 on domain citations of 78 blind prompt runs in the litigation due diligence zone at baseline, ChatGPT, 27 July 2026, and 1 on domain citation on 6 August 2026, when ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question. Then report business impact separately, as qualified conversions rather than traffic, and if you have not measured conversions, say so. We have not, and we say so below.

The reason to build the protocol this way is that the overclaim is what gets your whole report thrown out. In our own runs we have watched an engine discount pages whose figures were actually correct, purely because they carried no source and no date. A measurement report is a page like any other, and it gets read with the same suspicion.

Define what working means for your business

Write one sentence, before any measurement, naming the outcome you will accept as success. Four common versions, and they need different instruments.

  • Presence. Buyers who ask about our category hear our name. Measured in the answer layer.
  • Reachability. Buyers can get to us from the answer. Measured as on domain citations, then as referrals.
  • Enquiries. More of the right people contact us. Measured in the demand layer, with a source question at the point of enquiry.
  • Revenue. Qualified conversions that closed. Measured in the revenue layer, and the slowest to show anything.

Pick one as the headline and measure the rest as supporting layers. A programme that swaps its definition of success halfway through cannot be evaluated at all, and swapping is easy to do accidentally when the first number does not move.

Four measurement layers

Layer one, the answer layer. What the engines say. Instruments: a frozen blind prompt set, named engines, three repeats, human labelling into five states. Metrics: mention rate, on domain citation rate, top source rate, each as a fraction over answer runs. This is the only layer where you can see a competitor's position as well as your own.

Layer two, the referral layer. Visits arriving from AI surfaces. Instruments: your analytics, with referral sources recorded and kept as raw rows, plus landing page. Metrics: sessions and their landing pages by referrer. The limit is severe and needs stating in every report: an answer that names you without a link produces no referral at all, so this layer under counts by an unknown amount. It is a floor, not a measure.

Layer three, the demand layer. Enquiries. Instruments: one question on your enquiry form asking where the person heard of you, with an option naming AI assistants, plus a note field. Metrics: enquiries by stated source, per month. Self reported and still the best link you have between the answer layer and money.

Layer four, the revenue layer. Qualified conversions. Instruments: your sales records, with a source field filled at qualification rather than at first contact. Metrics: qualified opportunities and closed deals by source, with values. Report this layer as counts of qualified conversions, never as traffic, because traffic is the number most easily moved by things that have nothing to do with buyers.

Our own position across the four layers, stated honestly. We have layer one figures and nothing published in layers two, three and four. On aiknowsus.com the layer one figure is 0 of 6 answer runs citing us, Perplexity, September 2026, with 161 recorded statements across 24 batches and 72 conversations that we were not cited. We publish no conversion figure because we have not measured one. A vendor page that reports layer four results with no method is the exact thing this page tells you to distrust.

Establish the baseline

Nine steps, and the discipline in steps one and two is what makes everything after them mean something.

One. Freeze the prompt set with a version number and a date. It does not change again inside the study.

Two. Keep your brand name out of every prompt. Measured reason: on 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The branded wrapper run came back ranking the company first on almost every question. The blind run put it second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. We discarded the branded run.

Three. Name the engines and report them separately.

Four. Three repeats per prompt per engine, and write down the resulting run count, which is your denominator.

Five. Record conditions on every row: date, time, engine, mode, country and city, language, signed in state, device, personalisation state.

Six. Record the engine's own account of whether it searched, marked unverified. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of 78 questions, which changes how the whole baseline should be read. And in three separate batches of our September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question.

Seven. Label five states with a written rule per state, and have a second person label a sample.

Eight. Snapshot layers two, three and four on the same day, so all four baselines share a date.

Nine. Write the baseline as a sentence per metric, each containing numerator, denominator, engine, prompt set version and date.

Repeat the same test

The rerun is the measurement. Six rules govern it.

  • Identical prompt file, identical version. No additions, no rewording, no reordering.
  • Identical conditions. Same engines, same repeat count, same location and language, same signed in state.
  • A fixed interval, the same day each month, with the date range of each pass recorded.
  • The same labelling rules and, ideally, the same labeller, with the agreement check repeated.
  • A change log covering the whole period: every page published or edited, every technical fix, every off site listing, each with its date. Without this column you will never be able to line a movement up against anything.
  • Report both periods side by side as fractions. Not as a growth percentage, which hides both denominators and is the single most common way these reports mislead.

Separate observed change from causal proof

This is the section that keeps your report credible. An observed change is what you measured. A cause is what you would have to do considerably more work to claim.

Four things can move your number without your work having done anything.

  • The engine changed. Models, modes and retrieval behaviour change without notice or announcement.
  • Whether it searched at all changed. A reading taken with no live search and a reading taken with search are different instruments. We have the first documented: zero searches across 78 questions on 27 July 2026.
  • The web around you changed. A competitor's page appeared or disappeared, a directory reordered, a news article was published.
  • Normal variation. With three repeats and a small prompt set, a movement of one or two runs is inside the noise.

Four things strengthen a causal claim, and none of them make it a proof.

  • A holdout. Leave a matched group of questions with no new pages written for them, and compare the two groups' movement.
  • URL level attribution. The specific page you published being the URL cited, rather than the domain appearing somehow. The 6 August 2026 result is of this kind: a named URL, cited for a named question type, fourteen days after publication.
  • A dated mechanism. The engine itself stating why. On 17 September 2026 Claude said two enterprise vendors outranked clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and that its ordering reflected price transparency and source authority rather than product quality. That is a mechanism you can act on and test.
  • Repetition. The same movement in three consecutive monthly reruns.

The sentence to use in your report: the number moved from A of N to B of N between these dates, on these engines, while we made these changes, and we have not established that our changes caused it. That sentence survives scrutiny. The version without the last clause does not.

Example results, formulas, and limitations

Here is our own before and after, written out the way the protocol demands, with every weakness in it named.

Before. On domain citations 0, in the litigation due diligence zone of a blind 78 question set, ChatGPT, 27 July 2026. Named 21 of 78 overall on the same run, against ProVakil 28 of 78 and Legistify 18 of 78. The engine confirmed zero live web searches across all 78 questions.

After. On domain citations 1, in that zone, ChatGPT, 6 August 2026, citing clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question, fourteen days after that page was published.

Also measured in the same programme. Top source on 1 of 18 blind commercial questions, ChatGPT, 18 August 2026.

The weaknesses, all of them. The second reading was an interim check rather than a full rerun of all 78 prompts under identical conditions. The cited page was already live and four days old at the baseline, so the window is not clean. The baseline involved no live search, which makes it a weaker zero. Layers two, three and four were not snapshotted at either date. And one engine, one zone, two readings is not a trend.

The formulas, so nobody types a rate by hand. On domain citation rate equals runs with an on domain citation divided by total answer runs. Mention rate equals runs naming the brand divided by total answer runs. Prompt coverage equals prompts with at least one qualifying run divided by total prompts. Change equals the two fractions printed side by side, not a single percentage. Qualified conversion count equals opportunities marked qualified with an AI assistant source, counted, with values listed separately and never averaged into a rate over traffic.

Limitations of the whole protocol. It cannot establish causation. It cannot see sources used and not displayed, which on 17 September 2026 Claude confirmed happens, both by ranking pages in its raw results without linking them and by dropping attribution for two arguments it had taken from them. It under counts referrals whenever you are named without a link. It depends on self reported source data in layer three. It is valid only for its dates and its engines. And it cannot promise you a position in an answer, which is why no sentence on this page does.

Sources and change log. Every figure above comes from the clawlaw.in programme, recorded in GEO_BASELINE_RESULTS_2026-07-27.md including its interim check section, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. Only the 6 August 2026 citation can be verified from outside today, because it has a public URL behind it. Version: 29 September 2026, first publication, to be updated when layers two, three and four have dated figures behind them.

Common questions

My management wants one number. What do I give them?

On domain citations, as two fractions side by side with both dates, per engine. One number with no denominator is the thing that will be challenged first, and you will not be able to defend it.

Traffic went up. Can I report that as GEO working?

Report it as traffic, in layer two, with the caveat that answers naming you without a link produce no traffic at all. Then report qualified conversions separately. Traffic is the easiest number to move for reasons unrelated to buyers, which is why the engines' own specs for these pages asked for conversions instead.

How long before the before and after is worth showing anyone?

Two monthly reruns after the baseline, so you have three readings and can see the variation. The single timing figure we hold is fourteen days from publication to first citation of one page, ChatGPT, 6 August 2026, which is one observation and not a schedule.

What is a holdout in this context?

A group of prompts in your frozen set that you deliberately do not write pages for during the period. If the holdout group moves as much as the treated group, your pages are not what moved the number. It costs nothing except the discipline to leave some questions alone.

Should I report the engine's explanation of why it ranked us where it did?

Yes, quoted and dated, as a qualitative row rather than as evidence. The 17 September 2026 coverage list explanation was the single most actionable thing in that audit. It is still the engine's account of itself, and in three separate batches of our September 2026 audit an engine withdrew its own earlier statement about its behaviour when pressed.

Is it worth measuring at all if I cannot prove causation?

Yes. You are measuring whether the outcome you want is happening, which is what a business needs to decide whether to continue. Causation is a much higher standard and almost nobody in this field meets it. Saying so is what makes the rest of your report believable.

What to do first

Write your one sentence definition of working, freeze your prompt file with today's date, and snapshot all four layers this week even if three of them are empty. An empty layer with a date is a baseline. An empty layer with no date is an argument you will lose in six months.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.