Can AI search crawlers read your website? A dated test of robots.txt, HTTP access and named crawlers
The named crawler list, the fetch test, the fields to record, and a plain statement of what we have run and what we have not.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
For most sites the answer is yes for some crawlers and no for others, and the only honest way to know is to fetch your own pages as each named crawler and record what came back. A crawler can be allowed by your robots.txt and still be blocked at your CDN. It can receive a 200 and still get a page with none of your prices in it. The number that settles it is the count of crawler and URL combinations that returned the page you meant, out of the number you tested, on a stated date. This page publishes that test in full: the named crawler inventory, the URL set, the ten fields to record, and the count to report. It also says plainly that we have not yet published a run of it on our own domain, so no combination count appears here yet.
What "read" actually means
"Can it read my site" hides four separate questions, and a crawler has to pass all four. Most site owners check the first one and assume the rest.
- Permission. Your robots.txt allows that user agent for that path. This is the only layer most people check, and it is the layer least likely to be the problem.
- Delivery. Your server, CDN, web application firewall or bot manager actually returns a 200 to a request carrying that user agent string, rather than a 403, a 429, a 503, a challenge page or a redirect loop.
- Content. The bytes that came back contain the words you care about. A page whose prices are written in by scripts after the first response has no prices in the first response.
- Correctness. The text is the page you meant. Not a cookie wall, not a consent interstitial, not a login screen, not a version redirected by country, not an app download prompt.
A crawler that fails any one of the four has not read the page. Three of the four are invisible in your normal analytics, because analytics runs in a browser and these requests are not browsers.
Named crawler inventory
Test by name. "AI crawlers" is not a thing you can allow or block. These are the identities to put in your test list, grouped by the company that publishes them. Two of them are not crawlers at all, which is the most common mistake in this area.
- OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User. Three identities with three jobs, and blocking one is not blocking the others.
- Anthropic: ClaudeBot, Claude-SearchBot, Claude-User. Same split, with a separate identity for a fetch caused by one person's question.
- Perplexity: PerplexityBot, Perplexity-User. An indexing crawler and a user triggered fetcher.
- Google: Googlebot, and the Google-Extended token. Googlebot is the crawl that feeds Google Search, including the AI answers shown there. Google-Extended is a robots.txt token that controls a use of the content, not a crawler that makes requests, so it will never appear in your logs.
- Microsoft: Bingbot. The Bing index sits behind Microsoft's assistant answers, so this old name still matters for a new channel.
- Apple: Applebot, and the Applebot-Extended token. Again, one fetches and one is a control token.
- Amazon: Amazonbot.
- Meta: Meta-ExternalAgent and Meta-ExternalFetcher. A bulk agent and a fetcher.
- Others worth a row each: CCBot, Bytespider, DuckAssistBot, MistralAI-User, cohere-ai. CCBot belongs to Common Crawl, whose archive is read by many people other than its owner, which makes it the row that matters most in this list.
That is nineteen identities that make requests and two tokens that do not. Sort them by job rather than by company, because the job is what decides whether a block costs you anything:
- Bulk crawlers that gather pages at scale. Blocking these is a decision about how your text is used, and it does not by itself remove you from today's answers.
- Index builders that feed the assistant's own search. Blocking these does remove you from answers that use that search.
- User triggered fetchers that request one page because one person asked one question. These are the ones that matter on the day a buyer asks about you, and they are the ones most often caught by a bot manager, because they arrive singly and from data centre addresses.
Reproducible access test
This is the whole procedure. It needs a terminal and about two hours, and anybody can repeat it on any site.
Step one: fix the URL set at six pages. Use the same six shapes on every site so results can be compared. Your home page. Your pricing page. One product or service page. One long answer page such as a guide or a blog post that you would want quoted. Your sitemap.xml. Your robots.txt itself. Nineteen fetching identities across six URLs is 114 combinations. That number is the design of the test, not a result of it.
Step two: choose one target string per URL before you start. Something that must be present for the page to have been read: a price with its currency on the pricing page, a city name on a location page, a coverage count on a coverage page. Deciding this before the test stops you grading your own homework afterwards.
Step three: request each combination with the real user agent string, unchanged, with no browser headers added. Do not spoof a browser to get a nicer result. The point is to see what that crawler sees.
Step four: record ten fields per request. Anything less and the row cannot be rechecked later.
- Timestamp in UTC.
- The exact user agent string sent.
- The URL requested.
- The final URL after redirects.
- The HTTP status of every hop, not only the last one.
- Response size in bytes.
- Content type returned.
- Whether the target string was present in the raw response body, yes or no.
- Whether robots.txt allowed that path for that agent, decided by a parser and not by reading it yourself.
- The network the request came from, including the country and whether it was a home connection or a data centre.
Step five: run the whole set twice, from two different networks. One home or office connection and one cloud server. Bot managers treat those differently, and a test run only from your own office can pass while every real crawler is being refused.
Step six: compare raw against rendered for the same URL. Fetch the HTML with no scripts running, then fetch it again through a headless browser that runs them, and record whether the target string appeared in the first, the second or both. A string that appears only in the second is a string that most crawlers will not have.
Step seven: read your own edge logs for the last thirty days and count real visits by each of the nineteen identities. Your test says what happens when you knock. The logs say who has actually been knocking, which is a different and equally useful fact.
Results by crawler and page
Report one row per combination with all ten fields, then one headline count in this exact form: the number of the 114 combinations that returned the page we meant, with the date of the run and the two networks used. Beside it, report three failure counts, because they need different fixes: combinations refused before delivery, combinations delivered but missing the target string, and combinations delivered to the wrong page.
What we have run, and what we have not, as of 29 September 2026. We have not published a full 114 combination run of this test on our own domain, so this page carries no combination count. Publishing the protocol without the result is deliberate. The count is only worth anything with a date attached, and a date we have not reached is not a date.
What one site's own fetch check did show. On 6 August 2026, in our work on clawlaw.in, the pricing page rendered its prices only after scripts had run, so what a crawler received contained no prices at all. The prices the assistants quoted for that company had come from an app store listing instead of from the company's own site. Claude's run recorded it. The same day, ChatGPT noticed that the website and the app store listing carried different plan names and different prices for the same product, and said so inside its answer. That is layer three and layer four of the four above, failing together, on a site whose robots.txt was never the problem.
Current primary source documentation and last checked dates
Do not trust anybody's summary of a crawler list, including this one. Each of these companies documents its own identities, changes them without notice, and adds new ones. So this section names the document to read rather than restating what it says, and the discipline that goes with it is to write the date you read it beside the row.
- OpenAI publishes its crawler names and their purposes in its own platform documentation on bots.
- Anthropic documents its crawlers and how to control them in its support documentation.
- Perplexity documents its bots in its own help documentation.
- Google documents Googlebot and its other crawlers in Search Central, and documents the Google-Extended control separately.
- Microsoft documents Bingbot in Bing Webmaster help.
- Apple documents Applebot and Applebot-Extended in its own support pages.
- Amazon documents Amazonbot in its developer documentation.
- Common Crawl documents CCBot on its own site.
Two rules make this section useful instead of decorative. Record a last read date per row, not one date for the section, because you will check some of them and not others. And verify the vendor's identity claim by IP where the vendor publishes verifiable ranges, because a user agent string is a line of text that anyone can copy.
Safe change checklist and limitations
Eight changes, in the order that gets the most back for the least risk.
- Decide training and answering separately. They are different identities and different business decisions. Blocking a bulk crawler is a choice about your text. Blocking an index builder or a user fetcher removes you from answers.
- Check the CDN before the robots file. In practice most refusals come from a bot rule at the edge that nobody remembers switching on.
- Allow by verified IP range, never by user agent string alone, if you whitelist at all.
- Put the important text in the first response. Prices, coverage, cities, dates. If scripts have to run first, most of the nineteen will not see it.
- Keep one price list in public. Two public price lists that disagree is not a small untidiness. On 6 August 2026 ChatGPT pointed the contradiction out inside its own answer about clawlaw.in.
- Date every page that carries a fact, so a crawler that reads it in six months knows what it is holding.
- Never put hidden text addressed to the models. On 6 August 2026 a competitor's page in the clawlaw.in category carried a hidden block of text instructing answer engines to cite that company as the source. Claude found it, refused it, and named the company that had done it.
- Re-run the test after any change to your CDN, your framework or your robots file, and keep the dated rows.
What this test cannot show. It cannot tell you whether any assistant used your page in an answer. It cannot tell you whether a model retained it. It cannot tell you what is inside anybody's training set. It cannot tell you how a real session behaved, because real sessions are personalised and yours is not. And it is a photograph of one moment from one or two networks, so a pass today is not a pass next month. Passing this test is necessary and not sufficient: it only means the door was open.
Common questions
Our robots.txt allows GPTBot. Why would a fetch still fail?
Because robots.txt is a request and delivery is a decision made by other software. A bot rule at your CDN, a rate limit, a country block, a challenge page or a consent interstitial can all return something other than your page to a request that robots.txt permitted. That is why the test records the status of every hop and the final URL.
Should we block AI crawlers?
Treat it as two decisions. If your worry is your text being used to train models, that points at the bulk crawlers and at the two control tokens. If you want to be named when a buyer asks a question in your category, blocking the index builders and the user triggered fetchers works directly against that. Write down which of the two you are choosing, because most sites have blocked the second one by accident while meaning to do the first.
Does publishing an llms.txt file help?
It is a proposed convention rather than something the major vendors commit to honouring in their own crawler documentation, so treat it as cheap and unproven. It costs almost nothing to publish one. It is not a substitute for any of the four layers above, and no result on this page depends on it.
Our site is rendered by JavaScript. Is that fatal?
Not fatal, and it is the most common cause of a silent failure at layer three. The fix is not to rebuild the site. It is to make sure the specific facts you want quoted, the price, the coverage list, the city, the date, are present in the HTML of the first response, even if the rest of the page fills in later.
How often should we re-run the 114 combinations?
Once a quarter, and immediately after any change to the CDN, the bot rules, the framework or the robots file. The rows are worth keeping because a change log of dated access results is the only way to catch a block that appeared on a day nobody was looking.
Can we see the crawlers in Google Analytics?
No, and expecting to is a common trap. These requests do not run your analytics script. Server logs and CDN logs are where they appear, and that is why step seven of the test reads the logs rather than the dashboard.
What to do first
Block out one afternoon. Fix the six URLs, pick the six target strings, run the 114 combinations from two networks, and publish the rows with the date. Then fix the delivery failures first, because a refused request is a total loss, while a delivered page missing its prices is a page you can improve. Re-run and publish the second count beside the first, so the change is on the record rather than in somebody's memory.