Why Your Most-Visited Page Isn’t Cited by AI: A URL-Level Crawl, Indexing, and Citation Test
Traffic, indexing and citation are three different measures, and the denominator that matters is answers that could have cited anything at all.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Your most visited page is usually not cited because it was written to be read by a person who already found you, not to answer one question with facts an assistant can quote, and because traffic, indexing and citation are three separate measures with no necessary relationship between them. The figure that describes it properly is the citation count for that URL divided by the number of answers in which a source could have been cited at all, for a stated period and named assistants. That denominator is the part almost everybody gets wrong: an answer produced with no retrieval had no opportunity to cite anything, so it does not belong in the denominator. We hold one real URL level citation: on 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. That is legal technology in India, one URL, one assistant, one date, and it is the only URL level citation we can point an outsider at. We have not run a full high traffic versus cited comparison, so no rate appears on this page.
The page that got cited was not the site's most visited page. It was a page that answered one question a buyer actually asks, with a title stating that question, and it was two weeks old. That is the whole lesson, and everything below is how to test it on your own site.
Traffic, indexing, and AI citation are different measures
Four measures get conflated into one idea of an important page. They are produced by four different systems and a page can pass any one and fail the others.
- Traffic. Sessions from your analytics. It records people who arrived, which usually means people who searched for your brand, clicked an ad, or came from a newsletter. None of those tell you the page answers a stranger's question.
- Indexing. A search engine holds the page and can return it. Necessary for some retrieval paths and not sufficient for any citation.
- Retrievability by an assistant. A named crawler can request the URL, receives a 200, and gets a body containing the facts. Separate from indexing and separate from permission.
- Citation. The answer credits that URL. This is the only one the reader of an AI answer ever sees.
The gap between the third and the fourth is where most high traffic pages die. A home page or a services page can be perfectly retrievable and still give an assistant nothing to quote, because it contains positioning rather than facts. On 17 September 2026 Claude confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy. Retrieved, read, used, not credited.
Select the pages and comparison group
A test of your top page alone proves nothing, because you have nothing to compare it with. Build three groups and name every URL in your sheet.
Group one: your five highest traffic pages, taken from analytics for a fixed period you write down.
Group two: your five most question shaped pages. Pages whose title is a question a buyer asks and whose first paragraph answers it. If you have none, that is the finding, and it is the most useful result this test can give you.
Group three: five pages you will not touch, chosen at the start and left alone for the whole study. This is the comparison group. Without it, any change you see later cannot be separated from the assistant changing its own behaviour that month.
For each of the fifteen URLs record the following seven fields before you start: the URL, the group, the exact question it answers in one sentence, the publication date, the last substantive update date, whether it carries a specific checkable fact with a date, and its sessions for the period. If the third field is hard to write, you have found why the page is not cited.
Test each URL
Four layers, tested separately, because a page must pass all four and most owners only check the first.
One. Permission. Does your robots.txt allow each named crawler for that path? This is the layer least likely to be the problem and the one everybody checks first.
Two. Delivery. Request the URL with each named crawler's user agent string and record the exact status code. A 403 from a firewall, a 429, a challenge page, a redirect chain or a country redirect all mean no. Record the code, not a yes or no.
Three. Content. Read the body that came back and check whether the facts a buyer needs are in it as plain text. This single test explains more failures than the other three combined. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead of the company's own site. Claude and ChatGPT, same date.
Four. Correctness. Is the body the page you meant? Not a consent wall, not a login screen, not an app download prompt, not a cached older version.
Record one row per URL and per crawler, with the status code, whether the key fact appeared, the byte size of the body, and the date and time of the fetch. With fifteen URLs and four named crawlers that is fifteen times four, which is 60 fetch rows. Print that arithmetic on the report.
One more check belongs here, because it is the one that quietly kills a good page. Look for any claim on the page that cannot be verified. On 6 August 2026 a headline figure on the clawlaw.in main site could not be true, and when Claude checked it against public numbers it advised a buyer against the product, after which the company's accurate claims stopped counting for that answer. An unverifiable number on your best page does not sit there harmlessly. It discounts the page it is on.
Measure citations by page and assistant
Now the answer side. This is where the denominator has to be built carefully.
One. Freeze a prompt set of twenty to forty prompts, written in the words buyers use, saved with a version number and a date. Include at least two prompts aimed at the exact question each group two page answers, so a citation has a chance to occur.
Two. Apply the blind rule. No prompt names your brand. On 27 July 2026 two runs of the same 78 questions happened on one day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win, and the first run was discarded because the only thing it measured was our own prompt.
Three. Name the assistants and set the repeats. Three assistants, three fresh sessions per prompt. With thirty prompts that is thirty times three times three, which is 270 recorded answers per round.
Four. Record, for every answer, whether any source at all was cited. This field builds the real denominator. An answer that cited nothing had no opportunity to cite your URL, and including it inflates your denominator and deflates your rate. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of 78 questions, which means the citable opportunity count for that entire run was 0 of 78. A URL level citation rate calculated over those 78 answers would have been a division by zero dressed up as a small percentage.
Five. Treat every statement about sourcing as a claim. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction. Keep claimed sources and verified sources in separate columns.
Six. Attribute every citation to a specific URL, normalised, with the raw URL kept in its own column. A citation of your home page is not a citation of the page you were testing.
Seven. Compute two rates and print both. Citations of the URL over all recorded answers, and citations of the URL over answers that cited any source. The second is the honest one and it is always higher. Publishing only the first understates your pages, and publishing only the second hides how often retrieval did not happen.
Results: high-traffic versus cited pages
We publish no high traffic versus cited comparison, because we have not run this test on a full page set. As of 29 September 2026 the cell is empty and stays empty until the design above has been completed and the captures published. What we hold are three dated URL level and answer level measurements, in legal technology in India and in the AI visibility category, each with its engine and its date.
- One URL, cited, fourteen days after publication. ChatGPT, 6 August 2026, clawlaw.in. The cited page was clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india, and the question was about vendor due diligence. At baseline the same question set had named the company nowhere. The page was two weeks old and was not the site's most visited page.
- 1 of 18 as top source. ChatGPT, 18 August 2026, clawlaw.in. Across eighteen blind commercial questions the company was the top source on exactly one, and its own comparison page was named in the answer as still being the vendor's own editorial page.
- 0 of 78 citable opportunities. ChatGPT, 27 July 2026, clawlaw.in. No live web search was run for any of the 78 questions, so no answer in that run could have cited any URL from anybody.
Read the first two together. The page that earned a citation was a question shaped article, and the page that was named and discounted was a comparison page the assistant identified as the vendor's own editorial. Same site, same weeks, two different outcomes decided by what kind of page it was.
What to fix and what cannot be guaranteed
Six fixes, in the order we would do them, then the part nobody can promise.
- Put the fact in the served bytes. Every number, date, coverage list and price that a buyer asks about, as plain text in the response to a plain request.
- Give the page one question and answer it in the first paragraph. A page answering six questions is retrieved for none of them cleanly.
- Attach a source and a date to every claim. On 17 September 2026 Claude found two clawlaw.in pages stating competitors' prices with no link, no date and no source. The figures were correct when it checked them independently, and it treated the pages as advocacy and used official sources instead.
- Remove or fix anything unverifiable, because one impossible figure discounts the page it sits on, as the 6 August 2026 finding above shows.
- State your limits on the page. Who should buy something else, what you do not cover. A page that only says good things about the vendor reads as advocacy, which is exactly what Claude said about the clawlaw.in comparison pages on 17 September 2026.
- Write a new page rather than rewriting the top one, where the top page has a different job. A home page has to sell. A citable page has to answer. Asking one page to do both usually produces a page that does neither.
What cannot be guaranteed, by us or anybody: that a specific URL will be cited, that a citation will persist once it appears, that a fix will take any particular number of days, or that an assistant will keep selecting sources the way it did last month. We do not sell a position in an AI answer. The fourteen day observation above is one page on one assistant on one date and it is not a timetable.
Repeat the test
Rerun the frozen prompt set monthly, with nothing changed but the date, and refetch all fifteen URLs with all four crawlers in the same round. Publish every round, including the ones that went backwards, because a missing round is read as a stable round. Keep the untouched comparison group untouched for the whole study, however tempting it becomes, because it is the only thing that lets you say your work rather than the month explains a change.
Three rounds is the minimum before you describe a direction. One round is a photograph.
Sources and change log
The figures above come from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, and from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The 6 August 2026 citation of a named URL and the pricing page reading can be checked from outside today. Counts taken over our own capture folders cannot, until we publish the captures.
Version: 29 September 2026, first publication. The page gets updated when the fifteen URL comparison is completed and per page citation rates can be printed with their denominators, when any figure is corrected, and when the captures are published.
Common questions
Should I rewrite my home page to get it cited?
Usually not. The home page has a commercial job and a page written to be quoted has a different one. The clawlaw.in page that ChatGPT cited on 6 August 2026 was an article answering one specific question about checking a company's court cases, published fourteen days earlier. Write that page instead and leave the home page to sell.
My page is indexed. Why is it not retrievable by an assistant?
Because indexing, permission, delivery and content are four separate layers. Fetch the URL with each named crawler's user agent and read the status code and the body. A page can be indexed by a search engine and refused to an assistant's crawler by your own firewall, and nothing in your search console will tell you that.
Does a page need to be new to be cited?
Nothing we hold says newness is required. What the 6 August 2026 observation does show is that two weeks was enough for one new page to reach one answer on one assistant, in a zone where the site had previously appeared nowhere. Age is not the variable we would work on. Whether the page answers one question with checkable facts is.
How do I stop counting answers where nothing was cited?
Add a field recording whether the answer cited any source at all, and report the two rates side by side. On 27 July 2026 the citable opportunity count for a whole run of 78 questions was zero, because no live search happened. Any rate computed over that run without the field would have been meaningless.
What if a competitor's page is cited instead of mine for the same question?
Open their page and read what it contains that yours does not. In our own runs the answer has usually been a specific list. On 17 September 2026 Claude ranked two enterprise vendors above clawlaw.in on a coverage question specifically because they publish explicit court and tribunal coverage lists, and it said its ordering reflected price transparency and source authority rather than product quality.
Is it worth paying for a tool to track this per URL?
A tool collects at a scale you cannot match by hand, and we sell one. What you must still get from any tool, ours included, is the prompt list, the dates, the location, the repeat count, whether an answer cited any source at all, and whether scoring was done by a person or a model. Without those six its per URL rate cannot be defended in a meeting.
What to do first
Take your single highest traffic page and try to write, in one sentence, the exact question it answers. If you cannot, that is your answer and no crawler test is needed. Then pick the one question your buyers ask most, publish a page whose title is that question and whose first paragraph answers it with a dated checkable fact, and watch that URL specifically rather than watching your site.