Google Rankings vs AI Mentions: A Matched-Query Study With Search Console and Answer-Level Data
Why a good Google position does not carry into an AI answer, and the matched query study that would measure it properly.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Ranking well in Google and being named by an AI assistant are separate outcomes produced by separate systems, so one does not imply the other. The number that would settle the relationship for your site is the AI mention rate among matched queries where you ranked in Google's top ten, with the query count and the collection dates printed next to it. We have not run that matched study. As of 29 September 2026 we hold no Search Console data joined to answer level data, for any site, so this page publishes no such rate and does not estimate one. What we hold is the finding that explains why the relationship can be zero: on 27 July 2026 ChatGPT answered 78 blind buyer questions for clawlaw.in and afterwards confirmed it had run no live web search for any of the 78. In that condition a Google position cannot influence the answer at all, because no result list was consulted.
That is the first and most important thing a ranking owner needs to hear. A page can be first in Google and irrelevant to an answer, because the answer was written from what the model remembered. The study design below is how you find out which condition your queries are in, and the diagnostics at the end are what to do in each case.
Research question and limits
State the question narrowly or the answer will be worthless. The question is not whether SEO helps AI visibility. It is this: among a named list of queries, on named dates, for one site, what share of AI answers named the site, split by whether the site held a Google top ten position for the matched query on the same dates.
Four limits are built into that question and cannot be removed by any amount of data.
- A correlation here is not a mechanism. Even a strong relationship would not show that the ranking caused the mention. Both can follow from the same third thing, which is usually a page that states specific checkable facts.
- Query matching is a judgement. A Search Console row is the words a person typed into Google. A prompt is a sentence a person typed into an assistant. They are rarely identical and the matching rule decides the result.
- Positions move daily. A ranking recorded a week after the answer was captured is a different fact from a ranking recorded that day.
- The study describes one site. A relationship measured on one domain in one sector is not a finding about search in general, and publishing it as one is the mistake this page is written against.
Sample and query matching
The construction, step by step, with the arithmetic shown.
One. Pull the Google side first. Export Search Console performance data for a fixed date range, by query and by page, with impressions, clicks and average position. Freeze the export as a file with the export date in the name. Never refresh it mid study.
Two. Filter to commercially meaningful queries. Non brand queries only, because a brand query tells you nothing about whether a stranger finds you. Set a minimum impression threshold so you are not matching noise, and write the threshold on the report.
Three. Split by position band. Three bands: positions one to three, four to ten, and eleven or worse. The third band is the comparison group and it is what makes the study worth running. A study of only your good rankings has nothing to compare against.
Four. Sample evenly from the bands. Take the same number of queries from each band, for example forty from each, giving 120 matched queries. Sampling more from the band you hope to prove something about is the commonest way this study becomes useless.
Five. Write the matching rule and publish it. For each Google query, write one assistant prompt that a real person would type to express the same need. Our rule has four parts: keep the same product or service noun, keep the same qualifier such as a city or a sector, convert it into a natural question, and include no brand name. Then have a second person write prompts for twenty of the queries independently and compare, so the reader can see how much of the study is judgement.
Six. Apply the blind rule without exception. No prompt contains your brand. On 27 July 2026 two runs of the same 78 questions happened on one day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. We discarded the first run because the only thing it measured was our own prompt. A matched study run with a named brand would produce a false relationship in both bands at once.
Seven. Set the repeat count and print the denominator arithmetic. Three fresh sessions per prompt per assistant. With 120 matched queries, three assistants and three repeats, the denominator is 120 times three times three, which is 1,080 recorded answers. The per band denominator is 40 times three times three, which is 360. Print both multiplications on the report. A rate whose denominator arithmetic is not shown cannot be checked by anybody.
Eight. Record the position again on the capture day. For every matched query, note the position on the day the answers were captured, not the average over the export window. Store both.
Nine. Log these fields per row: Google query, position band, position on capture day, impressions, prompt text, prompt version, assistant, mode, session state, location, timestamp, whether the brand was mentioned, whether a page on the domain was linked, whether the answer was a top source, claimed sources, verified sources, whether a live search was evidenced, and the screenshot file name.
Dataset and results
No matched query rate is published here, because the study has not been run. The cell is empty as of 29 September 2026 and it stays empty until the design above has been completed and the data published. When it is run, the page will carry the mention rate for each of the three position bands, each with its own numerator and denominator, the query count, the assistants and the collection dates.
What we hold are four dated observations that bear on the question without answering it. Each is named with its engine, its date and its sector, and none of them is a matched study.
- 0 of 78 questions searched live. ChatGPT, 27 July 2026, clawlaw.in, legal technology in India. The assistant confirmed it had run no live web search for any of the 78 blind questions. Where that is the condition, no Google position can affect the answer.
- 21 of 78 mentions, in a memory based reading. Same run, same date. The breadth order was ProVakil on 28 questions, CLAW on 21 and Legistify on 18, which ranks what the model had absorbed rather than what was true that week or what ranked that week.
- Ranking inside the assistant's own results without being linked. Claude, 17 September 2026, clawlaw.in. The assistant confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy. That is retrieval ranking inside an assistant, not a Google position, and it is the closest thing we hold to the mechanism this study is about.
- 1 of 18 as top source. ChatGPT, 18 August 2026, clawlaw.in. Across eighteen blind commercial questions the company was the top source on exactly one, and its own comparison page was named in the answer as still being the vendor's own editorial page.
The third observation is the one worth sitting with. A page that is retrieved, read and used can still fail at the last step, and no ranking metric on either side of the study would have caught that. It was only visible because the assistant said so.
Confounders and limitations
Seven things that will distort this study if you do not handle them, and one of them cannot be handled at all.
- Whether a live search happened. This is the confounder that swamps all the others. Answers written without retrieval cannot respond to rankings. Record the condition per row and report the two groups separately, or your whole rate is an average of two different studies.
- Self reported search claims. Do not take the assistant's word for it. In the aiknowsus.com audit of September 2026, Perplexity withdrew its own earlier statement about having searched for each question, saying it could not honestly substantiate it, and in another batch said its claim to have searched all five was not adequately supported. Three separate batches. Log search behaviour in a claimed column and a verified column.
- Query intent mismatch. A Google query with informational intent and a prompt with buying intent are not matched, however similar the words. Score intent explicitly and drop the mismatches, reporting how many you dropped.
- Location and language. Both systems change their output by country. Fix the location, record it, and do not mix.
- Page level versus site level. Search Console gives you the ranking page. An answer may cite a different page on your domain. Decide before you start whether a citation of any page counts, and say which rule you used.
- Time drift. Capture the answers and the positions on the same days. A month of drift between the two sides makes the join meaningless.
- Selection by your own content history. Your top ten queries are probably the subjects you have written about most, so they differ from your weaker queries in more ways than position. This one cannot be removed by design, only disclosed.
How to reproduce the study
Everything needed is above, and this section is the checklist a second team would work from. Publish the following seven artefacts alongside any rate, or the study is not reproducible.
- The frozen query file, with its version number and export date.
- The prompt file, one prompt per query, with the matching rule stated at the top.
- The position band definitions and the impression threshold.
- The assistants, modes and repeat count, with the denominator multiplication written out.
- The scoring definitions for mention, citation, clickable link, mentioned through a third party, and used without attribution.
- The agreement check, being how often two scorers agreed on a sample of at least twenty rows.
- The full row level log, or a redacted version of it, so a reader can recount the numerator themselves.
Two further rules make reruns comparable. Change nothing between rounds except the date. And publish the rounds that went backwards, because a series with a missing round is read as a stable round.
Practical diagnostics
If you rank well and are never mentioned, work through these six in order. They are the checks that have actually found faults in our own runs.
- Check whether retrieval happened for those prompts at all. If not, the gap is not about your site. It is about whether the assistant is looking, and your work shifts to being present on the sources it already holds.
- Fetch the ranking page as a crawler and read the response. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead. Claude and ChatGPT, same date. A page that ranks can still be empty of the fact the answer needed.
- Test crawler permission and delivery separately for each named crawler, because a page allowed in robots.txt can still be refused by a firewall.
- Read your page as an assistant would judge it. A comparison page that concludes with you every time gets treated as advocacy, which Claude said in as many words on 17 September 2026.
- Check every number on the page for a source and a date. On 17 September 2026 Claude found two clawlaw.in pages with correct competitor prices and no link, no date and no source, and used official sources instead. Correct is not enough.
- Check who owns your subject. On the clawlaw.in how to questions of 17 September 2026, official court portals took every position above any commercial product, and Claude said plainly that no commercial product should rank above the official portal for a question about using that portal. Some subjects are not winnable and the right move is to answer a different question.
One more diagnostic worth running, because it separates the two systems cleanly. Publish one page that answers a single buyer question with specific dated facts, and watch only that URL. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. One page, one assistant, fourteen days. It is not a rule and it is not a promise, and it is a checkable example of the answer side moving on its own timetable.
Sources and change log
Every figure on this page comes from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. No Search Console data has been joined to answer level data in either programme, which is why no matched rate appears. The 6 August 2026 citation and the pricing page reading can be checked from outside today.
Version: 29 September 2026, first publication. The page gets updated when a matched query study is completed and the per band rates can be printed with their denominators, when any figure is corrected, and when the capture files are published.
Common questions
Does good SEO help AI visibility at all?
Some of it clearly does, because a page that cannot be fetched or read cannot be used by either system, and that is the overlap. What we will not tell you is how much, because we have not measured it and any number offered for it is invented. Run the matched study on your own site and you will have an answer about your site, which is the only place it applies.
Why would an assistant ignore a page that is first in Google?
Because it may not be consulting Google results at all. On 27 July 2026 ChatGPT confirmed zero live searches across 78 questions. And even when your page is retrieved it can be used without being credited, which Claude admitted on 17 September 2026 for two specific arguments drawn from clawlaw.in pages.
Can I use average position from Search Console instead of the daily position?
Use both and store them separately. The average tells you what the query looks like over the window. The capture day position is what was true when the answer was generated. If you only keep the average, a reader cannot tell whether your join was valid.
How many matched queries do I need?
Forty per position band is enough to see whether a difference exists worth investigating. It is not enough to publish a confident effect size, and we will not print a confidence interval for a design we have not run. Whatever size you use, print the band denominators separately, never a single pooled number.
Should the comparison band be my low ranking queries or a set of pages I did not touch?
Both, for different questions. Low ranking queries tell you whether position relates to mention. A set of pages you deliberately did not change tells you whether your own work moved anything. If you only have the first, you can describe a relationship and you cannot attribute a change to your actions.
Our Search Console data shows AI traffic separately. Is that enough?
It answers a different question. Referral data tells you about clicks you received. It cannot tell you about the answers where you were named and not linked, or used and not attributed, and both of those are documented above with dates. Referral data is a floor on your visibility, not a measure of it.
What to do first
Export your non brand Search Console queries for the last three months and pick ten from the top three positions and ten from beyond position ten. Write one blind prompt for each, run all twenty on one assistant three times today, and log whether you were mentioned, linked, or neither, along with whether the assistant searched at all. That single afternoon will tell you whether you have a retrieval problem or a page problem, and those need completely different work.