Which website pages earn AI citations? A page type benchmark with URL level results
Eight page types, the URL level counts we can actually account for, and the denominators beside each.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
The page shape we can show being cited is the narrowest one: a single question how to page, answering one buyer question, with the answer near the top. In our records the URL level counts are these, each with the number of URLs of that type we can account for: single question how to pages, 1 URL cited of 1 tested, by ChatGPT on 6 August 2026. Vendor comparison pages, 0 cited of 2 examined, by Claude on 17 September 2026, both discounted as advocacy. Pricing pages, 0 cited of 1 examined, on 6 August 2026, because a crawler received the page with no prices in it. Those are small denominators and we print them rather than hiding them, because a page type benchmark with no tested URL count per type is just a list of opinions about content.
We have not run a designed page type benchmark with matched URL counts per type. What follows is the full protocol for one, plus every URL level result our two programmes can actually account for, with its type, its engine and its date.
What earning a citation means here
One definition, applied to a URL rather than to a domain. A URL earned a citation when that exact address appeared as a source in an answer to a blind prompt. Four narrower states are recorded separately, because a page type benchmark that mixes them tells you nothing about what to write.
- Cited. The URL appeared as a source.
- Read and not cited. The page was in the engine's raw results and shaped the answer, and the URL did not appear. This state is real and almost invisible: on 17 September 2026 Claude confirmed clawlaw.in pages had ranked in its raw results and shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
- Named in the answer, page not linked. The business appears, the URL does not.
- Used with attribution dropped. Material from the page appears with no credit. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement.
If your benchmark only records the first state, three quarters of what your pages are doing is invisible to it, and you will draw the wrong conclusion about which types are working.
Pages and page types included
Eight types, defined tightly enough that two people would classify the same URL the same way. Ambiguous pages get the type that matches their main job, and the rule goes in the sheet.
- Single question how to page. One buyer question, answered in the first paragraph, steps underneath.
- Pricing page. Numbers, inclusions, exclusions, and a date.
- Coverage or scope page. An explicit list of what is covered: regions, courts, integrations, certifications, tribunals, categories.
- Vendor comparison page. You against named competitors, published by you.
- How to choose page. A buying guide that does not conclude with you.
- Limits page. What you do not do and who should buy something else.
- Method or process page. How the work is actually done, including what goes wrong.
- Service or product page. The conventional marketing page.
To run this properly you need at least three URLs of each type, live at the same time, on the same domain, and a prompt set that could plausibly pull any of them. Fewer than three per type and you are comparing single pages rather than types.
How URLs were tested
The protocol, in nine steps.
One. List the URLs and classify each by type, before running anything, so the classification cannot be influenced by the result.
Two. Record each URL's publication date and its confirmed indexing date. An unindexed page has not failed to be cited, it has failed to be crawled, and those need separating.
Three. Write prompts that match each type's job, at least two per type, and keep your brand out of all of them. The branded version of a run we did on 27 July 2026 ranked the company first on almost every one of 78 questions, while the blind version placed it second by breadth and absent from the questions it most wanted to win. We discarded the branded run.
Four. Run three repeats per prompt per engine, and name the engines. Report per engine, never blended.
Five. Capture every citation URL exactly, raw, before any normalisation. Then normalise in a separate column using rules you publish: strip tracking parameters, treat http and https as one, treat www and non www as one, treat a trailing slash as the same URL.
Six. Attribute each citation to a URL and therefore to a type, and count both citations and distinct URLs cited, because one page cited nine times is a different finding from nine pages cited once.
Seven. Record the four states above per URL, not just cited or not.
Eight. Report per type as two numbers: URLs cited of URLs tested, and citations of answer runs. Never one without the other.
Nine. Rerun monthly with the same URL list and the same prompts, and log any page edits with dates, because an edited page is a new observation.
Results by page type
This is what we can account for at URL level, and nothing more. Three types have data, five are empty, and the empty ones stay empty.
- Single question how to page: 1 URL cited of 1 accounted for. ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question on 6 August 2026, fourteen days after the page was published, in a zone where the same blind question set had named the company nowhere at baseline on 27 July 2026. This is the only URL level citation in the whole programme with a public address behind it.
- Vendor comparison page: 0 URLs cited of 2 accounted for. Claude, 17 September 2026. Both pages stated competitors' prices with no link, no date and no source. The figures were correct when the engine checked them independently, and it treated the pages as advocacy and used official sources instead. Separately, on 18 August 2026, ChatGPT did name the company's own comparison page in an answer, describing it as still being the vendor's own editorial page, which is the state named in the answer rather than cited.
- Pricing page: 0 URLs cited of 1 accounted for. 6 August 2026. The page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead. Claude and ChatGPT, same date. This is a crawlability result rather than a content quality result, and it is recorded as such.
- Coverage or scope page: no URLs accounted for on our side. The relevant observation is about competitors' pages: on 17 September 2026, on a question about finding every case against a company, Claude ranked two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and noted that its ordering reflected price transparency and source authority rather than product quality. That is evidence about the type and not a count of ours.
- How to choose page, limits page, method page, service page: no URL level data. Empty cells.
One more type level finding worth recording, because it tells you where not to spend. On 17 September 2026, on the questions about how to look a case up, official court portals took every position above any commercial product, and Claude said plainly that no commercial product should rank above the official portal for a question about using that portal. A how to page about somebody else's official process is a page you will lose with.
And one about source types generally: across the whole aiknowsus.com capture of September 2026, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. Perplexity, 24 batches and 72 conversations. Your own pages are in the class of sources that decides these answers, which is why page type is worth benchmarking at all.
What the results can and cannot say
- One URL of one tested is not a rate for the type. It is a single existence proof that this shape of page can be cited.
- 0 of 2 for comparison pages is about those two pages, which had an identifiable defect the engine named: competitor prices with no source and no date. A sourced, dated comparison page has not been tested by us.
- The pricing page result is a technical failure, not a verdict on pricing pages. A pricing page a crawler can read has not been tested by us.
- Page type is confounded with page age, link position and topic. A designed benchmark needs matched URL counts and similar publication dates, which ours does not have.
- Read and not cited is under counted because it depends on the engine volunteering that it happened, which on 17 September 2026 it did.
- Results are per engine and per date, and none of this is evidence about Google AI Overviews, which we have not captured.
- No page type buys you a position in an answer.
Common questions
If I can only write one page this month, which type?
A single question how to page answering a question your buyers actually ask, with the answer in the first paragraph. It is the one shape we can show being cited, on 6 August 2026 by ChatGPT, and it is the cheapest page on the list to write well.
Are comparison pages a waste of time?
Not inherently, and unsourced ones are worse than nothing. The two that failed on 17 September 2026 failed for a stated reason: competitor prices with no link, no date and no source, which made a page with correct figures read as advocacy. If you publish one, source and date every competitor fact and link the competitor's own page for it.
Why did the pricing page not get cited?
Because a crawler received a page with no prices in it. The prices were added by scripts after load. The assistants quoted an app store listing instead, on 6 August 2026. That is a five minute check and it explains more absences than any content theory.
Should I write a coverage page if my coverage is not the best in the market?
Yes, as an explicit list with its limits stated. On 17 September 2026 two enterprise vendors were ranked above clawlaw.in specifically for publishing explicit court and tribunal coverage lists, and the engine said its ordering reflected transparency and source authority rather than product quality. Being explicit and second is better than being vague and first.
How many URLs per type do I need for a real benchmark?
Three minimum, five is better, published in the same period, on the same domain, with prompts that could plausibly reach any of them. With fewer than three you are measuring individual pages, which is still useful and should not be called a page type result.
Does a page get cited more if it is longer?
We have not measured length and we are not going to guess. What the one cited page had was a single question, an answer near the top, and a subject where no official portal owned the answer. Those three are what we would replicate.
Page planning checklist and dataset
Ten checks per page, before you publish it.
- One buyer question per page, in the buyer's words.
- The answer in the first paragraph, in a form that can be quoted in one sentence.
- Every fact in the page as served, not injected by a script. Fetch it and check.
- A date on every number.
- A source link for every competitor fact.
- No claim that a reader could disprove in five minutes.
- Explicit lists where you would otherwise write comprehensive.
- A stated limit: what this does not cover, who should buy something else.
- No hidden instructions to engines. On 6 August 2026 Claude found exactly that on a competitor's page, refused it, and named the company in its answer.
- Publication date and confirmed indexing date logged in your dataset.
The dataset is one row per URL: address, type, publication date, indexing date, prompts it should match, and then per engine per month the four states and the citation count. Sources for the results above: the clawlaw.in programme recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, and the aiknowsus.com audit of September 2026. Version: 29 September 2026, first publication, to be updated when a designed benchmark with matched URL counts has been run.
What to do first
List your existing URLs and classify them by the eight types above. Most businesses find they have five service pages and nothing else, which is the finding. Then write one single question how to page, make sure a crawler can read it, and log its publication and indexing dates so that whatever happens next is measurable.