Generative engine optimisation: definition, scope, and a reproducible citation test
What GEO is, what it is not, and the counted result of our own blind runs with the dates attached.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Generative engine optimisation is the work of getting a business named, recommended and cited inside the answer an AI assistant writes, and then measuring how often that happens with a repeatable blind test. It is a measurement discipline before it is a content discipline. In our own blind runs, the counted results were these: on 27 July 2026 ChatGPT named clawlaw.in in 21 of 78 questions where no brand was named, and confirmed it had run no live web search for any of the 78. On 18 August 2026, across 18 blind commercial questions, the same domain was the top source on exactly 1 of 18. In September 2026, asked six questions about its own category with no brand named, Perplexity audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six.
Those are small numbers and they are the real ones. The rest of this page is the definition, the boundary of the work, the protocol that produced those counts, and what the counts can and cannot be used to claim.
GEO, in one sentence
GEO is the practice of making a company findable, quotable and citable by an AI assistant that answers a buyer's question in prose, and of measuring the result as a rate over a stated set of questions on stated dates.
Two halves of that sentence carry all the weight. The first half is production work: pages that state a number, a date, a scope or a limit that an assistant can repeat. The second half is measurement work: a fixed question set, a rule that keeps the brand name out of the question, and a count. Without the second half the first half is a belief.
The reason the second half is so easy to get wrong is in our own record. On 27 July 2026 two runs of the same 78 questions happened on the same day. The first used a wrapper that named the brand in the prompt and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. That is the single most useful thing in this page.
What GEO includes, and what it does not
GEO includes the following six kinds of work. Each is named because a vague scope is how this discipline gets sold as magic.
- Crawler eligibility. Whether the assistants' fetchers are allowed to read the pages, and whether the page returns its content without a script having to run first.
- Answerable pages. One buyer question per page, the answer in the first paragraph, and a specific fact under it that can be quoted in one sentence.
- Checkable facts. Prices with a date, coverage stated as a named list rather than as the word comprehensive, and a method described well enough for a reader to repeat it.
- Presence on the sources the assistants already read. Official registers, category directories, review sites and the comparison pages that exist in the sector.
- Consistency across the public record. The same plan names and the same prices on the website, the app listing and the directory entries.
- Measurement. A blind question set, run on more than one assistant, scored the same way each time, with the dates and the raw answers kept.
GEO does not include the following four things, and a page that says otherwise is selling something that does not exist.
- A guaranteed position in an answer. No vendor, ours included, can promise that an assistant will name a company. The assistants do not sell placement in the answer text and they change their source selection without notice.
- Instructing the model from inside the page. We found a competitor's page carrying a hidden block of text addressed to answer engines, instructing them to cite that company as the source. Claude found it, refused it, and named the company that had done it, on 6 August 2026 during the clawlaw.in programme. That is the outcome to expect from this tactic.
- Anything that depends on a ranking formula nobody has published. The assistants document what their crawlers are and what their products do. They do not document a weighted list of ranking factors, and a page that recites one has invented it.
- A single reading treated as a trend. One run on one day on one assistant is an anecdote. In batch after batch during the September 2026 audit, Perplexity itself declined to name a competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim.
How we tested whether a page appeared in AI answers
The protocol below is the one that produced the counts in the first paragraph. It is written out in full so somebody outside the company can run it against their own domain and get a number that means the same thing.
Step one, build the question set. Write the questions a buyer would type before they have decided who to buy from. Split them into four groups and keep the count of each group: the plain category question, the comparison question, the price question, and the how to do it yourself question. Freeze the list. A question set that changes between runs cannot show movement.
Step two, apply the blind rule. No question may contain the company name, the product name, the domain, or a phrase that appears only on the company's own website. This is the rule that the 27 July 2026 pair of runs exists to justify. A named brand in the prompt turns the test into a test of the prompt.
Step three, fix the scoring definitions before the run, not after. We score four outcomes per answer.
- Mention. The company name appears anywhere in the answer text.
- Recommendation. The company appears in the part of the answer that tells the reader what to use or who to contact, not merely in a list of things that exist.
- Citation. The answer carries a link or a named source that resolves to the company's own domain.
- Top source. The company's own domain is the source the answer leans on most for its substance.
Step four, record whether the engine searched. This is the step most tests skip and it decides what the whole run means. Record it two ways: from the engine's own visible search indicator, and from the server logs of the domain being measured for the hour around the run. Then ask the engine directly, in a second message, how many of the questions it searched for. Keep that answer as part of the record, because it can contradict the first message. In September 2026, asked afterwards how many of the questions it had actually searched for, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. That happened in three separate batches.
Step five, count the denominator honestly. A valid answer is one where the engine answered the question asked, in the same session, without an error and without a refusal. Invalid answers come out of the denominator and are reported separately as a count, never silently dropped.
Step six, keep the raw answers. Every rate on this page comes from a file that still exists. A rate without the answers behind it cannot be audited later, including by us.
Results
Four runs, with the scope and date on each. These are the measured numbers we have. Where we have not run something, later sections say so rather than filling the gap.
Run one. clawlaw.in, 27 July 2026, ChatGPT, 78 blind buyer questions. Mentions by breadth came out as ProVakil on 28 questions, CLAW on 21 and Legistify on 18. Live searches: none. ChatGPT confirmed it had run no live web search for any of the 78 questions, so that whole reading describes what the model remembered rather than what it could find. The correct way to read 21 of 78 is as a memory score, not a visibility score.
Run one, discarded twin. clawlaw.in, 27 July 2026, ChatGPT, the same 78 questions with the brand named in the wrapper. It came back ranking the company first on almost every question. It was discarded. It is reported here because a discarded run is evidence about method, and because the gap between the two runs on the same day with the same questions is the strongest argument for the blind rule that we own.
Run two. clawlaw.in, 18 August 2026, ChatGPT, 18 blind commercial questions. The company was the top source on exactly 1 of 18. In the answer, its own comparison page was named as still being the vendor's own editorial page, which is the reason it did not carry more weight.
Run three. aiknowsus.com, September 2026, Perplexity, one batch of six questions about our own category, no brand named. The engine audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers. The count is 0 of 6. There was no position for us to hold.
Run four. aiknowsus.com, September 2026, 24 batches and 72 conversations on our own domain. The phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers. That is a count of statements in a capture, not a count of answers, and it is written that way on purpose.
The one result with a URL behind it. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. That is one citation, on one question, on one assistant. It is the only externally checkable citation in this set and it is not a rate.
What the results can and cannot show
They can show four things. They show that the blind rule changes the result on the same day with the same questions, which means any published mention rate without a blind rule is uninterpretable. They show that an engine can answer 78 buyer questions without searching once, which means a low score can be a memory problem rather than a page problem. They show that an engine's own account of whether it searched is not reliable, because Perplexity withdrew that claim in three separate batches in September 2026. And they show that a single named page can become a cited source within fourteen days of publication, on one question, once.
They cannot show six things, and the list matters more than the numbers above it.
- They are not an industry benchmark. Two domains in two categories in India is not a population. Nothing here supports a sentence beginning with the words most companies.
- They do not establish cause. The page cited on 6 August 2026 was published fourteen days earlier, and other things also changed in those fourteen days. Sequence is not proof.
- They do not transfer across engines. Run two was ChatGPT, run three was Perplexity. Scores from different engines are different instruments and should not be averaged.
- They do not transfer across dates. The assistants change their retrieval behaviour without announcing it. A count from August 2026 describes August 2026.
- Five of the seven counted results cannot be checked from outside yet, because the capture files that hold them are not published. Only the citation with a URL can be verified by a reader today.
- They say nothing about revenue. Being named in an answer is a visibility outcome. Nobody in this record measured what it sold.
Practical checklist for measuring GEO
Ten items, in the order we run them.
- Write between 40 and 100 blind buyer questions and freeze the list.
- Read every question once for brand leakage, including phrases that only exist on your own site.
- Choose at least two assistants and record which product and which mode was used.
- Write the four scoring definitions down before the first run.
- Run the set, one question per fresh conversation, and save every answer as a file.
- Record whether each answer searched, from the indicator and from your own server logs.
- Ask the engine afterwards how many questions it searched, and keep the reply even when it contradicts the run.
- Count invalid answers separately and state the valid denominator.
- Report every rate as a count over a denominator with the date and the engine attached, never as a bare percentage.
- Repeat on a fixed interval with the same frozen set, and publish the change log.
Common questions
How many questions do I need before the number means anything?
Enough that one answer changing does not change your headline. At 18 questions, one answer is worth about six percentage points of your rate, which is why we report 1 of 18 rather than a percentage. Below about 30 questions, treat the result as a list of observations rather than a rate. Above that, the bigger risk is that your question set is unrepresentative rather than too small.
Is GEO just SEO with a new name?
The production work overlaps a great deal: crawlable pages, clear structure, real facts, third party presence. The measurement does not overlap at all. SEO measures the position of a link on a results page. GEO measures whether a name appears inside a written answer, and whether the answer cites you. Those are different instruments and they can disagree. On 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote, and that the company still landed as a name inside a list rather than as a linked recommendation.
Why would an assistant use my content and not link to me?
It happens, and it is worth separating from a ranking judgement. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and had dropped the attribution in both places, and described that as a citation lapse rather than a ranking decision. Scoring mention, recommendation and citation as three separate outcomes is what makes this visible at all.
Should I measure on one assistant to keep it simple?
No, because one assistant can be wrong about itself. The clearest single thing in our September 2026 capture is Perplexity withdrawing its own claim to have searched, in three separate batches. Two assistants also stop you from redesigning a website around one product's behaviour in one month.
Does a low score mean my pages are bad?
Not on its own. Check the search count first. If the engine did not search, as ChatGPT confirmed it had not for any of the 78 questions on 27 July 2026, then your pages were never in the running and the score is a measure of what the model already remembered. Those are two different problems with two different fixes.
Can a tool do this instead?
A tool can run the set on a schedule, keep the captures and do the counting, which is the part people abandon by week three. We sell one, AI Knows Us, so treat that as a disclosure rather than a recommendation. What no tool can do is promise a position in an answer, and any tool that implies it can should be read carefully.
Sources and change log
Everything counted on this page comes from two bodies of work, and both are named so the claims can be traced.
- The clawlaw.in programme, from July 2026. Recorded in Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md for the 78 question baseline, the discarded branded run, the breadth order and the 6 August 2026 citation; Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md for the 18 question run; and Tier_1/claude_response_17_09_audit.md for the 17 September 2026 findings.
- The aiknowsus.com audit of September 2026. 24 batches and 72 conversations, captured to geo-audits/aiknowsus-com/. The 0 of 6 batch is in perplexity__B05-b-perplexity.md. The count of 161 not cited statements was counted by script across all 24 second message answers, not from memory.
What changed and when. 27 July 2026: baseline run of 78 questions, and the branded twin discarded the same day. 6 August 2026: first citation of a named URL, fourteen days after publication, along with the crawler and pricing findings. 18 August 2026: the 18 question commercial run. 17 September 2026: the source selection audit. September 2026: the 24 batch audit on our own domain. This page is dated to the last of those and will carry a new line when a run is added.
What is still missing. We do not have a published dataset a reader can download, so five of the counts above cannot be verified from outside today. We do not have a matched pair of runs designed to test one page feature against a control. Both are named here rather than papered over.
What to do first
Write your 40 blind questions this week and run them on two assistants without naming yourself, one question per fresh conversation, saving every answer. Then write down four numbers: how many answers mentioned you, how many recommended you, how many cited your domain, and how many searched at all. If the last number is low, your first job is eligibility and evidence rather than more pages. If it is high and the first three are low, you have a source selection problem, and the pages that fix it are the ones carrying a dated number, a named coverage list and a stated limit.