Which sources appear in Google AI Overviews? A dated, clickable source audit
How to count distinct source URLs properly, and why a link count with no query count is meaningless.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
AI Overviews assemble an answer and show links to pages that support it, and the only way to know which sources appear for your subject is to capture the Overviews for a written list of queries, save every supporting URL exactly as given, and count the distinct URLs against that query count and those dates. A count of links with no query count beside it tells you nothing: 400 URLs from 40 queries and 400 URLs from 400 queries are completely different findings. As of 29 September 2026 we have not run an Overviews capture, so we publish no distinct URL count of our own, and this page gives the counting rules plus the source patterns we did measure in assistant answers.
Almost every published claim about where AI answers get their information fails on the denominator. The claim travels as a headline, the query set is never named, and the capture dates are missing, so the number cannot be rerun by anybody. What follows is written so it can be rerun.
What source means in this audit
Source is one of the most abused words in this subject, so it gets four separate definitions here and each is counted separately.
- A displayed supporting link. A URL the Overview shows next to or under the answer. This is the only kind you can observe directly.
- A distinct URL. One unique address after the normalisation rules below are applied. Two rows pointing at the same page are one distinct URL.
- A distinct domain. The registered domain behind those URLs. One domain can supply many URLs, and reporting domains as if they were URLs inflates nothing but confuses everything.
- An unobserved contributor. Any page that shaped the answer and was not shown. You cannot count these, and pretending the displayed links are the full source set is the largest error in this whole area.
The fourth category is real and we have it documented. On 17 September 2026, in the clawlaw.in programme, Claude confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation. A page can be used and not shown. Your source audit will never see it.
The normalisation rules, before any counting
Publish these with your count or the count is not reproducible. We use the following six rules.
- Strip tracking parameters and keep parameters that change the content.
- Treat http and https as the same URL, and the www and non www forms as the same URL.
- Treat a trailing slash as the same URL.
- Follow one redirect and record the destination, noting both the displayed and final address.
- Count a fragment as the same URL as its page, unless the fragment loads different content.
- Count subdomains as separate domains in the domain tally, and say so, because a documentation subdomain and a marketing site behave differently.
Two people who apply different rules to the same capture will report different distinct URL counts from identical data. That is why the rules go on the page.
What Google says about its sources
Google's public help material describes AI Overviews as showing links to pages that support the answer so a reader can go deeper, and it provides a way to send feedback on an Overview. We do not restate its wording as a dated quotation here, because help pages are edited and an undated paraphrase is exactly the kind of source this page asks you to distrust. Open the current help page yourself on the day you read this.
What Google does not publish is the part an audit would need most: how the displayed links are chosen from the pages considered, whether the full set of contributing pages is ever shown, and how the selection varies by country, language and device. Treat the displayed links as a sample of the source set whose selection rule is unknown.
Audit methods
The protocol, in the order you run it.
One. Freeze a query set and version it. Between 20 and 100 queries in buyers' words, saved to a file with a version number and a date. Every count you publish names the version it came from.
Two. Keep your own brand out of every query. The reason is measured. On 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent from the litigation due diligence questions it most wanted to win. We discarded the first run. A source audit run with your brand in the prompt will over collect your own domain in exactly the same way.
Three. Record the conditions on every row. Timestamp, country and city setting, language, signed out or signed in, device, browser version. Source selection is known to vary by location, and an audit that mixes locations cannot be compared with itself later.
Four. Save the raw evidence. Full page screenshot, the Overview text as text, and every supporting URL copied exactly, before normalisation. Keep the raw column and add a normalised column beside it. Never overwrite the raw one.
Five. Classify every source by type, using a fixed list. We use eight types: official or government, platform or vendor documentation, the vendor's own marketing site, independent review or directory, news and trade press, forum or community, educational, and other. A row that needs a ninth type means the list gets extended and the whole capture is relabelled.
Six. Report four numbers together, never one. Distinct URLs, distinct domains, number of queries, and the date range. Then, separately, the count of rows where no link was shown at all, because an Overview with no visible source is a finding and not a blank.
Seven. Rerun the same set on a fixed interval. The interesting result in a source audit is not which domains appeared once. It is which domains appear in every rerun, because those are the pages that decide your subject.
Results
The distinct URL count for AI Overviews is not published here, because we have not captured it. No estimate, no borrowed figure. The cell stays empty until the protocol above has been run and the captures published, and the page will then carry the count, the query count and the dates together in one sentence.
What we did measure, in assistant answers rather than Overviews, is the pattern of which source types reach an answer at all. These four findings are dated and attributed.
- The most cited domains were official pages and vendors' own pages. Across the whole aiknowsus.com capture of September 2026, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Perplexity, 24 batches, 72 conversations.
- Across six questions, no third party review or directory source reached any answer. Claude, 17 September 2026, clawlaw.in. One well known review site did appear in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
- On how to questions, official portals took every position above any commercial product. Claude, 17 September 2026. It said plainly that no commercial product should rank above the official portal for a question about using that portal.
- A source can substitute for your own site when your site is unreadable. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead. Claude and ChatGPT, same date.
The fourth one is the most actionable finding in this batch. Your page being indexed is not the same as your page being readable. If the fact a buyer asks about is injected by a script, the answer will be built from whoever else published that fact about you.
What the audit cannot establish
Six limits, stated plainly because a source audit is unusually easy to over read.
- It cannot list the sources that were used and not shown. Documented above, with a date.
- It cannot establish why a source was chosen. Appearance is not a ranking and the selection rule is not published.
- It cannot generalise beyond your query set. A source audit of 40 buying queries in one sector says nothing about the web.
- It cannot be compared with somebody else's audit unless their normalisation rules, query set and location match yours. They almost never publish those.
- It cannot separate a change in Google from a change in the web. A domain that disappears between reruns may have lost selection, changed its page, or been deindexed.
- Our own counts above are not Overviews counts. They come from ChatGPT, Claude and Perplexity, and they are the nearest real evidence we hold rather than a substitute.
Sources and change log
The findings above come from two recorded programmes: clawlaw.in from July 2026, in GEO_BASELINE_RESULTS_2026-07-27.md and the assistant audit files of 17 September 2026, and the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The domain and source type tallies were produced by a script over those capture files rather than from memory. Of the four findings printed above, the pricing page reading and the app store substitution can be checked from outside today. The capture wide tallies cannot, until the captures are published.
Version: 29 September 2026, first publication. This page will be updated when the Overviews capture is run, when any figure is corrected, and when the capture files are published.
Common questions
How many source URLs does one Overview usually show?
We will not give you a number, because we have not counted it and every number we have seen published for it comes without a query set or a date. Count it yourself on your own 20 queries this week and you will have a figure that is true for your subject, which is the only place it matters.
If my domain never appears as a source, what is the first thing to check?
Whether the fact being asked about is actually present in the page as served, rather than added by a script after load. Fetch the page the way a crawler does and read what comes back. On 6 August 2026 that single check explained why assistants were quoting an app store listing for clawlaw.in prices instead of the company's own pricing page.
Do review sites and directories matter for being cited?
They matter less than most marketing advice assumes, at least for India specific questions. Across six clawlaw.in questions on 17 September 2026 no third party review or directory source reached any answer, and the one well known review site in the raw results was dropped because its list was of American products. A directory listing is still worth having. It is not the thing that decides these answers.
Should I count a source that appears in many queries once or many times?
Both, in two separate columns. Distinct URLs tells you the size of the source pool. Appearances tells you which pages are load bearing for your subject. Reporting only the first hides the fact that three pages might be answering half your queries.
Can I use a tool's source report instead of capturing it myself?
Yes, if the tool gives you the prompt list, the dates, the location, and its normalisation rules. Without those four you have a number you cannot defend in a meeting. We sell a tool and we would rather you asked us those four questions than not.
Download, reproduce, or report a correction
The reproducible part of this page is the protocol and the six normalisation rules, which you can copy and run without us. The capture files behind our assistant findings are being prepared for publication, and when they go up the rows currently marked as uncheckable become checkable line by line.
If you rerun this and get a different pattern, or if you find a sentence here that is wrong, write to us through the contact page on this site with the query, the date and the screenshot. Corrections are made on the page and noted in the change log.
What to do first
Pick your ten most commercially important queries, capture the Overviews for all ten today with your location recorded, and list every supporting URL by type using the eight types above. Then fetch your own most important page as a crawler would and check that the facts a buyer needs are in what comes back. Those two hours will tell you more than any published source study.