Why AI cites some websites and ignores others
It cites the page it can read easily and quote confidently.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
An assistant cites the page it can fetch without trouble and quote without hedging. Two barriers stop most pages: the crawler could not read them, or the text gave the assistant nothing specific to repeat. Everything else is detail on top of those two.
What is actually happening underneath
Take the reading problem first, because it is the one that is invisible from the outside. These are the six ways a page becomes unreadable to a crawler while looking perfectly normal in your own browser.
- Text that only appears after JavaScript runs. Many crawlers see an empty page. This is the single most common technical cause.
- robots.txt blocking the crawler. The bots behind AI search are named separately from Googlebot, so a site can be fully indexed by Google and still be closed to them.
- Content behind a login, a paywall, or a form.
- Content inside a PDF or an image when it could have been in HTML.
- Bot protection or rate limiting that returns an error to anything that is not a browser.
- A page that is never linked from anywhere, so nothing leads a crawler to it.
Then the quoting problem. An assistant is writing a short answer and needs a sentence it can stand behind. Pages that give it one tend to share the same habits: the answer near the top, one question per page, facts stated with units and dates, headings that match how the question is asked, and plain admissions of what the thing does not do. Pages that give it nothing tend to be long, general, and written to cover a topic rather than to answer a question.
Site level context matters too, though less than people think. A site with twenty real pages on one subject is read differently from a site with one page on twenty subjects. Being mentioned on other sites also helps, because it makes your page look like a source rather than a stray.
The ten minute readability check
Do this before commissioning anything. It needs no tools beyond a browser and your hosting panel.
One, view source and search for your first sentence. Use view page source and the browser find on a distinctive phrase from the body text. Not present means a crawler probably does not see it either.
Two, open your robots.txt. Type your domain followed by /robots.txt. Read every disallow line and every named user agent. Decide each one on purpose.
Three, search for a quoted sentence from the page. If the page does not come back in ordinary search results, it is not indexed, and citation is not available to it yet.
Four, check the server logs for bot visits. This is direct evidence rather than inference, and most teams have never looked.
Five, read the first paragraph and ask what it says. If it is background rather than the answer, you have found the quoting problem.
Six, count the questions the page answers. More than one, and the page is the best answer to none of them.
A worked example
On 6 August 2026, in a Claude run over the site of clawlaw.in, our sister company, the pricing page turned out to be unreadable in exactly this way.
Its prices rendered only after scripts ran. In a browser the page looked complete. What a crawler received contained no prices at all, so the one fact the page existed to state was missing from the version a machine could read.
The prices the assistants did quote in that reading had come from an app store listing instead of from the company's own site. That is the real cost of failing the first test. You do not go quiet. You hand the answer to whichever page about you does load.
The check that would have caught it is the first one on the list above and it takes a minute. View source on the page, then use the browser find on the price itself. If the number is not in the raw HTML, no crawler is getting it either, however good the page looks to you.
What it means for a business
Most pages that never get cited fail the first test, not the second. That is good news, because the first test is cheap to check and cheap to fix. It is worth checking before you commission a single new article.
It also means that length is not the lever. A short page that answers one question with a dated fact gets quoted more often than a long page that covers everything vaguely. Length helps only when it is more real content: another named item, another worked example, another honest limit.
What this does not explain
Readability and specificity get you into the running. They do not decide the outcome, and three things sit outside your control.
Competition for the slot. A perfectly built page still loses to a better known page answering the same question. Being eligible is not being chosen.
The retrieval logic. It is not published. Everything here is inferred from which pages get cited and what changes when a page changes.
Whether a search happened at all. For an answer built from training data, none of this applies, because no page was fetched.
And one thing worth saying plainly: nobody can guarantee a citation. A page that is readable and specific is in the running. A page that fails either test is not, and that is the part you control.
Common questions
Should I block AI crawlers?
It is a real choice with a real cost. Blocking the training crawler keeps your text out of future models and does not remove you from live search. Blocking the search crawler removes you from the path you have the most influence over. Decide each bot separately and write down the reason, because the person who inherits the site will not guess it.
Does schema markup make me more likely to be cited?
Treat it as helping a machine confirm what your page already says, not as a way to make a vague page quotable. Add it where it genuinely applies and keep it consistent with the visible text. Markup that disagrees with the page is worse than none.
Is my single page site enough?
Usually not, because one page cannot be the best answer to several questions. A small number of pages, each answering one question, does more than one long page covering everything.
My page is indexed but never cited. What now?
You are past the reading problem and stuck on the quoting one. Move the answer into the first paragraph, add one checkable fact with a date, cut the sentences that could appear on a competitor's site unchanged, and add a line on who the product is not for.
Can bot protection settings block this without anyone noticing?
Yes. Protection that challenges anything without a browser fingerprint will also turn away crawlers you wanted. If your logs show no bot visits at all, check this before rewriting anything.
How long should I wait before deciding a page has failed?
Give it a couple of monthly readings after you can confirm it is indexed. If it has been indexed for two months and never appears in a citation list for the question it was written for, treat it as a content problem and rewrite the top of it.
What to do first
- View source on your three most important pages. If you cannot find the body text in the raw HTML, fix that first.
- Read your robots.txt. Decide on purpose which crawlers you allow, rather than discovering the default later.
- Check the server logs for bot visits, so you are working from evidence rather than a guess.
- On each of those three pages, move the answer into the first paragraph.
- Add one checkable fact with a date, and a line saying who the product is not for.
Then wait, and check with a question that page is the natural answer to. Nobody can guarantee a citation, but a page that is readable and specific is in the running, and a page that fails either test is not.