Which AI crawlers should you allow? A dated robots.txt and access control test
Sixteen named crawler tokens, four configurations to test, and how to count real fetches from your own server logs.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Decide this by crawler and by purpose, not by allowing or blocking AI as a single thing. The measurement that answers it is a count, taken from your own server logs, of which named crawlers actually fetched your test pages under each configuration you published, with the dates and the test conditions recorded. We have not run that four configuration test as of 29 September 2026, so this page publishes no fetch counts per configuration and no list of which crawlers obey a rule. It publishes the full protocol, the sixteen named tokens to test, and the four configurations to test them under. The nearest measured fact we hold is a warning about what a fetch count would not tell you: on 6 August 2026, for clawlaw.in, the pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants did quote had come from an app store listing instead of from the company's own site. Access was not the problem. Readability was.
The general advice you will read is to allow the crawlers that can send you a citation and block the ones that only take content for training. That is a reasonable starting position and it is not a measurement. Which bucket each crawler belongs in is stated by its operator, changes over time, and is not something we are going to restate here as a fact of ours.
How it was measured
Step one. Write down the sixteen tokens you will test, grouped by operator, and verify each one against the operator's own published crawler documentation on the day you run the test. We name the tokens so you know what to look for in your logs. We do not restate what each one is used for, because that is the operator's documentation and a restated document goes stale into a false claim. Open each operator's crawler page, record the URL and the date, and keep a screenshot beside your log data.
- OpenAI: GPTBot, OAI-SearchBot, ChatGPT-User.
- Anthropic: ClaudeBot, Claude-User, Claude-SearchBot.
- Google: Googlebot, Google-Extended.
- Perplexity: PerplexityBot, Perplexity-User.
- Microsoft: Bingbot.
- Apple: Applebot, Applebot-Extended.
- Others worth counting: CCBot, Amazonbot, Bytespider.
Three things about that list matter more than the list. Some of these tokens control training use rather than fetching, so a token you have blocked may still legitimately appear nowhere in your logs and mean nothing about access. Some crawlers fetch because a person asked a question at that moment, which is a different event from a scheduled crawl. And the list changes, so a token list without a date is already wrong.
Step two. Build the test pages. Six pages, each with a unique string on it that exists nowhere else, so you can later prove a fetch produced content and not just a status code.
- A plain page with the fact in the server rendered HTML.
- A page where the same fact only appears after scripts run. This is the arm that catches the 6 August 2026 failure above.
- A page behind a soft consent or interstitial layer.
- A page in a section disallowed in robots.txt.
- A page allowed in robots.txt but blocked at the network edge by your CDN or firewall rules.
- A page with a noindex directive but no crawl block, so you can separate crawling from inclusion.
Step three. Publish four configurations, one at a time, each for a fixed number of days chosen in advance. Two weeks per configuration is the minimum that lets a scheduled crawler come round. Record the exact robots.txt text for each, with the timestamp it went live and the timestamp it came down.
- Configuration A, allow all. The baseline. Without it you cannot tell a crawler that obeys a block from a crawler that was never coming anyway.
- Configuration B, allow the crawlers your operators document as serving answers with links, disallow the rest. Name every token in the file rather than using a wildcard, so the log analysis is unambiguous.
- Configuration C, disallow every token on the list, keep ordinary search crawlers allowed.
- Configuration D, disallow everything for the AI tokens and enforce it at the edge for half the pages. This is the arm that separates a request from an enforcement, which is the single most misunderstood part of this subject.
Step four. Count fetches from your own server logs, not from a tool. Log lines are the evidence. For every request, record twelve fields.
- Timestamp with time zone.
- Full user agent string, not a shortened label.
- Source IP, and whether it passed the operator's published verification method for its crawler.
- The path requested.
- HTTP status returned.
- Bytes returned. A 200 with a tiny body is not a successful content fetch.
- Whether the unique string was in the response body.
- Which configuration was live at that moment.
- Whether robots.txt was fetched by the same agent first, and when.
- Whether the request came through the CDN or reached origin directly.
- Referrer, where present.
- Any rate limit or challenge response you served.
Step five. Verify identity before you count. A user agent string is self declared and anybody can send one. Each operator publishes a verification method, usually a reverse lookup or a published address range. Count a fetch as belonging to a named crawler only when verification passes, and keep the unverified ones in a separate column. Reporting unverified user agent counts as crawler behaviour is the most common error in published crawler studies.
Step six. Report per configuration, per crawler, as a fraction. Numerator is the number of test pages that crawler fetched with the unique string present. Denominator is the six test pages times the days the configuration was live. Print the arithmetic. A crawler that fetched two of six pages on eight of fourteen days is a different fact from a crawler that fetched everything once.
Step seven. Run a separate, parallel answer test, because a fetch is not a citation. Ask each engine the question your test pages answer, blind, with your brand named in none of the prompts, three times each, and record whether the unique string appears in the answer. This is the arm that tells you whether allowing a crawler bought you anything. Record any claim the engine makes about having searched, and mark it unverified: in three separate batches of our September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question.
What the numbers were
What has not been run, as of 29 September 2026. We have not published the four configurations, we have not collected the logs, and we have not counted verified fetches per crawler per configuration. So this page contains no fetch counts, no obedience rates, and no statement about which crawler respects which directive. Those are exactly the numbers the question needs, and inventing them would make this page worthless for the purpose it was written for.
Here is what we do hold, and every line names its sector so nothing is carried across.
- Allowed and still unreadable. Legal technology, clawlaw.in, 6 August 2026, Claude. The pricing page rendered prices only after scripts ran, so what a crawler received contained no prices at all. The assistants quoted prices taken from an app store listing instead of from the company's own site. Nothing was blocked. The content simply was not in what a crawler got.
- Two public price lists, and the engine noticed. Legal technology, clawlaw.in, 6 August 2026, ChatGPT. The website and the app store listing carried different plan names and different prices for the same product, and the engine said so in its answer. Access control decisions have this side effect: if you block your own pages, the copy of your information that engines can reach is whichever third party listing you do not control.
- Allowed and not searched at all. Legal technology, clawlaw.in, 27 July 2026, ChatGPT. Asked 78 blind buyer questions, the engine then confirmed it had run no live web search for any of them, so 0 of 78 answers depended on fetching anything that day. A crawler being allowed does not mean a fetch happens for your question.
- The sources that actually decide answers. AI visibility tooling, aiknowsus.com, September 2026, Perplexity, counted across 24 captures. The domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. Blocking your own site removes you from the category that wins most often.
What this cannot tell you
- Robots.txt is a request, not a lock. Only a server side or edge rule enforces anything, which is why configuration D exists in the protocol above.
- A log count cannot prove intent. You can see that a verified crawler fetched a page. You cannot see what the fetch was used for.
- A blocked token can mean two different things and the logs cannot separate them: the crawler obeyed, or it was never going to come. That is why configuration A is run first.
- User agent strings are self declared. Counts without operator verification are counts of claims.
- Token lists and operator behaviour change. A result is valid for the dates it covers, and this page's list needs rechecking against each operator's documentation before you use it.
- Allowing everything does not get you cited. On 27 July 2026 the engine searched for none of 78 questions, and across our own September 2026 capture the engines recorded 161 times that they had not cited us while nothing on our site was blocked.
- Nothing in this test promises a position in an AI answer.
Sources and change log
Documents to read yourself. Each operator publishes its own crawler documentation listing its tokens, what each is used for, and how to verify a request really came from it. Read OpenAI's, Anthropic's, Google's, Perplexity's, Microsoft's and Apple's crawler pages directly, record the URL and the date for each, and keep screenshots. We name the documents rather than restating them, because a restated crawler document is a false claim waiting for an update.
Our own figures. The clawlaw.in programme of July and August 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md including the Claude run section, and the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The 6 August 2026 rendering and pricing observations are externally checkable, since the pages involved are public.
Version. 29 September 2026, first publication. Updated when the four configuration test completes, when any operator's token list changes, and whenever a reader shows us an error.
Common questions
Should I just block all AI crawlers to protect my content?
Decide it per crawler and know what you are giving up. The domains cited most often across our September 2026 capture were official documentation and vendors' own websites. If your site is unreachable, the version of your business that engines can describe is whatever third party listing they can reach, and we have a dated case of that going wrong: on 6 August 2026 the prices ChatGPT quoted for clawlaw.in came from an app store listing that disagreed with the company's own site.
If I allow a crawler, will my pages be cited?
No, and access and citation should be measured separately for this reason. Allowing a crawler removes a blocker. It does not create a reason to quote you, and on 27 July 2026 the engine did not search at all for any of 78 questions.
Does blocking training crawlers hurt my visibility in answers?
We have not measured it and we are not going to guess. Read each operator's documentation to see which token governs training and which governs answering, record what it says with the date, and if you want the answer for your own site, run configurations A and B for two weeks each and compare the answer test results. That is a real measurement you can own.
Some crawler traffic in my logs looks fake. What do I do?
Verify before you count. Each operator publishes a verification method, and requests that fail it go into a separate column rather than into your crawler totals. Unverified user agent counts are the reason two published crawler studies can disagree completely.
My robots.txt blocks a crawler and it still appears in my logs. Is it ignoring me?
Possibly, or the request is a different kind of event, such as a fetch triggered by a person asking a question, which several operators document as a separate token. Check the token, check the verification, and check whether robots.txt was fetched by that agent before the request. If it is a verified crawler ignoring a published rule, you need an edge rule rather than a robots rule.
What is the cheapest useful version of this test?
Two pages, two weeks, two configurations, and your existing access logs. One page with the fact in the HTML and one with the fact only after scripts run, allowed for two weeks and disallowed for two weeks. That alone tells you whether your important content is reachable at all, which is the failure we actually have on record.
What to do first
Before touching robots.txt, fetch your three most commercial pages with scripts switched off and read what comes back. If your prices, your coverage list or your key numbers are missing from that text, your access policy is not your problem yet. Fix the rendering, then run configuration A for two weeks and look at your logs to see which verified crawlers come at all. You cannot make a sensible blocking decision about crawlers that were never visiting.