GEO vs SEO: what is different, what overlaps, and what we can actually measure
Two instruments, two denominators, and the matched query method that lets you compare them fairly.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
GEO and SEO share most of the production work and almost none of the measurement. SEO counts the position of a link on a results page for a keyword. GEO counts whether your name appears inside a written answer and whether that answer cites your domain. The honest state of our own comparison is this: we have run the GEO side on a frozen set of 78 blind questions on 27 July 2026 and on 18 blind commercial questions on 18 August 2026, and we have not yet run a matched keyword and prompt comparison with both metrics scored on the same list on the same day. This page publishes the full matched query method so anybody can run it, and reports only the numbers we actually have.
Saying that plainly costs us a headline number. It is also the only version of this page worth citing, because a comparison table with invented rates on both sides is exactly what an assistant should discount.
Short answer
They are different in what they count, similar in what they reward, and they can disagree on the same day about the same page. A page can rank in the raw results an assistant reads and still not be named in what it writes. On 17 September 2026 Claude recorded exactly that during the clawlaw.in programme: the pages ranked in its raw results and shaped what it wrote, and the company still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
So the practical answer is that you do not choose between them. You keep both instruments and you never average them.
What SEO measures
SEO measurement has four settled units, and they are settled because the industry agreed on them over twenty years.
- Position. Where a specific URL sits in the list of results for a specific keyword, in a specific location, on a specific date.
- Impressions and clicks. How often that URL was shown and how often somebody clicked it, reported by the search engine itself for your own property.
- Coverage. How many of your pages the search engine has indexed and is willing to show.
- Referral sessions. Visits that arrived from the search engine, attributed in your own analytics.
The strength of these is that the denominator is not in dispute. One keyword, one result page, one position. The weakness is that all four assume the reader will see and click a list.
What GEO measures
GEO measurement has to invent its own units because no assistant reports them to you. We use five, and every one is defined before a run rather than after.
- Mention rate. Answers where the company name appears anywhere, over valid answers.
- Recommendation rate. Answers where the company appears in the part that tells the reader what to use, over valid answers.
- Citation rate. Answers carrying a link or named source that resolves to your own domain, over valid answers.
- Top source rate. Answers where your domain is the source the answer leans on most, over valid answers.
- Search rate. Answers where the engine ran a live web search at all, over valid answers. Without this one the other four cannot be interpreted.
The search rate is the unit with no equivalent in SEO, and it is the one that decides what a run means. On 27 July 2026, asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them. The mention counts from that run therefore describe what the model remembered, not what it could find. In September 2026, asked afterwards how many of the questions it had actually searched for, Perplexity withdrew its own earlier claim in three separate batches, saying it could not honestly substantiate that it had run a live search for each question. Record the search rate two ways, from the engine and from your own server logs, because the engine's account of itself is evidence and not proof.
Where the practices overlap
Six things help both, and this is where most of a budget should go because the work is not duplicated.
- Crawlability. Both need the content in the response, not assembled by a script afterwards. On 6 August 2026 Claude recorded that the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants did quote had come from an app store listing instead of the company's own site.
- One question per page. Good for a results listing and necessary for a quotable answer.
- Clear structure. Headings that match the question, and the answer near the top.
- Specific facts. A dated price, a named coverage list, a counted scope. These earn positions and they earn quotes.
- Third party presence. Directories, registers, review sites and comparison pages are read by both.
- Consistency. On 6 August 2026 the website and the app store listing carried different plan names and different prices for the same product, and ChatGPT noticed the contradiction and said so in its answer. A conflict in public is a problem for both disciplines, but only one of them repeats it back to your buyer in a paragraph.
Where they differ
Five differences that change what you do, not just what you count.
The number of winners. A results page shows ten links. An answer usually names three options. There is no position eleven in a paragraph.
Who judges advocacy. A search engine will rank your comparison page. An assistant will read it and decide whether to treat it as a source. On 17 September 2026 Claude found that two clawlaw.in comparison pages stated competitors' prices with no link, no date and no source. The figures turned out to be correct when it checked them independently, but it had no way to know that at read time, so it treated the pages as advocacy and used official sources instead.
What wins a category question. On 17 September 2026, on a question about finding every case against a company, two enterprise vendors were ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and Claude noted that its ordering reflected price transparency and source authority rather than product quality. A better product lost to a better published scope.
Where an official source is unbeatable. On the same date, on questions about how to look a case up, official court portals took every position above any commercial product, and Claude said plainly that no commercial product should rank above the official portal for a question about using that portal. In SEO you would still chase that keyword. In GEO you should write for the adjacent question instead.
Attribution. A click is logged. A quote may not be. On 17 September 2026 Claude admitted it had used two specific arguments from clawlaw.in pages and dropped the attribution in both places. An analytics report built on sessions records that as nothing happening.
A fair comparison method
Here is the matched query protocol in full. We have not completed a run of it, and the section after this says exactly which parts we have.
Build the matched list. Take between 50 and 100 buyer intents. For each intent write two things: the keyword a buyer would type into a search box, and the question a buyer would type into an assistant. They must be the same intent, not the same words. Freeze both lists together and number them so every row is a pair.
Set the blind rule on the GEO side. No prompt may contain your company name, product name, domain or a phrase that appears only on your site. This is not a preference. On 27 July 2026 two runs of the same 78 questions happened on the same day, and the one with the brand in the wrapper ranked the company first on almost every question, while the blind one put it second by breadth and absent from the litigation due diligence questions it most wanted to win. The branded run was discarded because the only thing it had measured was our own prompt.
Fix the location and the date. Both instruments move with location. Record the country and city setting used, and run both sides within the same 48 hours.
Score the SEO side per row. Position of your best ranking URL in the top ten, yes or no for whether any URL of yours appears at all, and the identity of the first three domains shown.
Score the GEO side per row. Mention, recommendation, citation, top source, search, and the identity of every domain cited.
Report as pairs, not as averages. Four counts, each as a number over the same denominator of rows: rows where you appear in both, rows where you appear in search only, rows where you appear in the answer only, and rows where you appear in neither. The interesting row is the third, because that is content being used without a link, and the second is the one that tells you your ranking is not converting into citation.
State what each result would mean. If you appear in search on many rows and in answers on few, the likely problem is evidence quality and advocacy framing, not eligibility. If you appear in neither and the search rate is near zero, the engines are answering from memory and no amount of new content moves this month's number. If you appear in answers but not in search, check that your citation scoring is not counting a mention of your name as a citation of your domain.
Results and limitations
What we have run. The GEO side, twice, on frozen blind sets. On 27 July 2026, 78 questions on ChatGPT: breadth order ProVakil on 28 questions, CLAW on 21, Legistify on 18, with a search rate of zero for all 78 by the engine's own confirmation. On 18 August 2026, 18 blind commercial questions on ChatGPT: top source on exactly 1 of 18. In September 2026 on our own domain, 24 batches and 72 conversations, in which the phrase recording that we were not cited appears 161 times in the engines' own self audits, and one batch of six questions where Perplexity reported afterwards that it had not cited or recommended aiknowsus.com in any of the six.
What we have not run. The paired SEO side. We do not have position data collected on the same frozen list on the same dates, so we cannot publish a row by row comparison, an agreement count, or any statement of the form appears in search but not in answers as a rate. We are not going to estimate it.
What that means for this page. Treat the method as the deliverable and the numbers as partial. The five limits are these. Two domains in two Indian categories are not a sample of anything. A run in August 2026 describes August 2026, because retrieval behaviour changes without notice. A ChatGPT number and a Perplexity number are different instruments and must not be averaged. Most of our counts cannot be checked from outside yet, because the capture files are not published; the exception is the citation of clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india by ChatGPT on 6 August 2026, fourteen days after that page was published. And none of this measures revenue.
Common questions
If I only have budget for one, which do I measure?
Measure the one that matches how your buyers actually arrive, and be honest that you may not know. The cheap first step is to run 30 blind prompts once and look at the search rate. If the engines are not searching for your category, your buyers are getting answers from model memory and your near term lever is third party presence and consistency, not new pages.
Can a good search ranking be read as evidence of AI visibility?
No, and our own record is the counterexample. On 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote, and that the company still landed as a name inside a list rather than as a linked recommendation. Ranking is a necessary condition in many cases and it is not the outcome.
Why not just report one blended visibility score?
Because the two instruments have different denominators and different failure modes, and a blended number hides the only diagnostic information you have. The whole value of scoring mention, recommendation, citation, top source and search separately is that a drop in one of them tells you what to fix.
Do keyword volumes help me choose prompts?
They help you choose intents and they mislead you on wording. People type longer, more specific questions to an assistant, often with a place or a constraint attached. Use volume to pick which intents matter, then write the prompt the way a person would say it out loud.
How often should the comparison be repeated?
Monthly, with the list frozen, or quarterly if the set is large. The value is in the change on an unchanged list. Any change to the question set resets your history, so if you must add questions, add them as a second numbered block and keep reporting the original block separately.
Does this mean SEO work is wasted?
No. Five of the six overlapping items in the section above are ordinary technical and editorial quality, and they feed both. What is wasted is assuming the ranking is the result. On 6 August 2026 the clawlaw.in pricing page ranked and carried no readable prices for a crawler, and the assistants quoted prices from an app store listing instead.
Sources and update history
- Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 78 question baseline of 27 July 2026, the discarded branded twin, the breadth order, the zero search confirmation, and the 6 August 2026 findings on crawlable prices, the two public price lists and the first cited URL.
- Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 question commercial run of 18 August 2026.
- Tier_1/claude_response_17_09_audit.md. The 17 September 2026 source selection findings on advocacy, coverage lists, official portals, and the dropped attribution.
- geo-audits/aiknowsus-com/. The September 2026 audit of our own domain, 24 batches and 72 conversations, including the batch of six with no citation and the three batches where the engine withdrew its search claim.
Update history. 27 July 2026, GEO baseline. 6 August 2026, crawler and consistency findings and the first citation. 18 August 2026, commercial run. 17 September 2026, source selection audit. September 2026, own domain audit. The matched SEO side is not yet run, and this section will carry the date it is, with the row counts, whatever they say.
What to do first
Build the paired list this week: 50 intents, each with a keyword and a prompt, numbered so the rows line up. Run the prompt side blind on two assistants and record the search rate before anything else. Then pull the position data for the same rows within the same two days. The first number worth acting on is the count of rows where you rank and are not named, because those rows are usually a page problem you can fix in a fortnight: a missing date, a missing source, a coverage claim written as a word rather than a list.