How to check whether ChatGPT search can access a new page: a timestamped crawl test log

Eligibility is documented, a timetable is not. Here is the log to keep, and what our own record does and does not contain.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Nobody publishes a timetable for this, and you should not accept one. What you can do is prove eligibility and record timestamps yourself: confirm the search crawler is allowed, confirm your page returns its content without scripts, then log the first request from that crawler in your own server logs and the first time the page appears as a citation. Our own record holds the second half of that and not the first. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. We do not hold a verified crawler request timestamp for that page, so we publish the protocol for capturing one and say plainly that our own run of that part is not done.

Fourteen days is one interval, for one page, on one engine, for one question. It is not a schedule and we will not present it as one.

Short answer: eligibility is documented, a timetable is not

Three separate facts get merged into the question how long does it take.

  • Whether the crawler may fetch your page. Documented, and controlled by you.
  • When the crawler did fetch it. Observable, in your own access logs, with a timestamp. This is first party evidence and most teams never look for it.
  • When an answer will cite it. Not documented, not promised, and not under your control. A page can be fetched and never cited.

So the honest measurement is two intervals, kept separately: publication to first crawler request, and publication to first observed citation. Conflating them produces the sentence people want and cannot support, that a page takes so many days to appear in ChatGPT.

Check crawler access

Six checks. Do them in this order, because the later ones are pointless if an earlier one fails.

  • Read your own robots file as text. Not a plugin's summary of it. Look for a wildcard disallow added at some point to stop scrapers, which is the most common accidental block.
  • Check the rules apply to the search crawler specifically. The assistants run more than one fetcher, and the one that serves search answers is not the one that gathers training data. Blocking one is a different decision from blocking the other, and you should make each one on purpose.
  • Check for blocks above the application. A content delivery network, a firewall rule or a bot protection service can refuse a fetch that your robots file allows. The application never sees it, so only the edge logs show it.
  • Check the response code. A soft failure that returns a success code with an empty body looks fine in a browser and is empty to a fetcher.
  • Check your sitemap actually lists the new page, with a last modified date that is true.
  • Search your access logs for the crawler by name, for the last thirty days, and note whether it appears at all. If it never appears, the question is not how long, it is whether.

Check the page itself

A page can be reachable and still carry nothing a fetcher can use. Five checks.

  • Fetch it without running scripts and read what comes back. This is the check that found the clearest problem in our own programme. On 6 August 2026 Claude recorded that the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants did quote had come from an app store listing instead of the company's own site.
  • Confirm the key facts are text. A price, a registration number or a coverage list inside an image does not exist for a reader that only has the text.
  • Confirm the answer is near the top. A page whose first useful sentence is in the ninth paragraph is harder to quote.
  • Confirm the page says when it was last updated, in text, in a form a person can read.
  • Confirm the facts on it agree with your other public pages. On 6 August 2026 the website and the app store listing carried different plan names and different prices for the same product, and ChatGPT noticed the contradiction and said so in its answer.

Run a documented page publication test

This is the protocol we are publishing because it is the thing that would produce the measured interval the question deserves. Ten steps.

One. Choose three to five named test pages. Name them in your own log, with their full URLs. One page is an anecdote; five give you a range. Pick pages that answer a real buyer question, because a page nobody would ask about will never be cited and will teach you nothing beyond the crawl interval.

Two. Record the publication timestamp to the minute, from your own deployment record, not from the page's displayed date.

Three. Confirm eligibility before publishing, using the two checklists above, so a null result is not just a blocked fetch.

Four. Turn on log retention. Most default configurations discard access logs faster than this test runs. Keep at least sixty days.

Five. Watch for the search crawler by name in the logs, and record the first request per test page: the timestamp, the URL, the response code and the bytes returned. A request that got a redirect or an error is not a successful fetch and should be logged as what it was.

Six. Verify the requester rather than trusting the user agent string. A user agent can be set by anybody. Confirm the request came from the published address ranges of the company that operates the crawler, and record that you did. An unverified line in a log is not evidence.

Seven. Run a fixed citation probe on a schedule. Write three blind prompts per test page, the questions that page answers, with no brand name in them. Run them daily or every second day, one per fresh conversation, on at least two assistants. Record for each: date, engine, whether it searched, whether the page appeared as a citation.

Eight. Record two intervals per page. Publication to first verified crawler request. Publication to first observed citation in the probe. Keep them in separate columns, because they answer different questions and one can happen without the other.

Nine. Record the non events too. A page fetched on day two and never cited by day sixty is a result, and it is the most common one. A test that only reports successes is an advertisement.

Ten. Say what each outcome would mean before you start. No crawler request at all points at eligibility, and you go back to the first checklist. A request within days but no citation for weeks points at the page's content rather than its access, and the usual gap is a missing dated fact or a missing named list. A citation with no matching crawler request in the window means the answer used something else, possibly a cached copy or a third party page quoting you, and your citation interval is measuring the wrong thing. A citation that appears and then disappears from later probes is normal, and it is the reason the probe is repeated rather than run once.

Results and limitations

What we have. One dated citation interval, on one named page, on one engine. ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question on 6 August 2026, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. This is the one result in the whole clawlaw.in programme with a URL behind it that a reader outside the company can check.

What we do not have, stated plainly. We do not hold a verified search crawler request timestamp for that page or for any other named test page. The interval above is publication to first observed citation, not publication to first fetch. We did not have log retention and crawler verification running as a designed test in July and August 2026, so there is no first party fetch record to report, and we are not going to reconstruct one after the fact. Until that test is run, this page's answer to the crawl timing question is a method, not a number.

Why one interval cannot be a timetable. Six reasons. It is one page. It is one engine. It is one question in one zone. The publication timestamp in our record is held as an elapsed interval rather than as a minute, so the fourteen days is as precise as the record is. Whether the engine searched for that question on other days was not probed daily, so 6 August 2026 is the first day we observed a citation and not necessarily the first day one existed. And retrieval behaviour changed across our own programme: on 27 July 2026 ChatGPT confirmed it had run no live web search for any of 78 blind questions, which means a page published that week could not have been cited by that run at all, whatever its crawl status.

One more caution about self reports. In September 2026, asked afterwards how many of the questions it had actually searched for, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. That happened in three separate batches. If you cannot see the fetch in your own logs, you do not know that it happened.

Troubleshooting checklist

  • Crawler never appears in the logs: check robots, then the edge, then the firewall, then whether logs are being kept at all.
  • Crawler appears but only on the home page: check internal links to the new page and whether the sitemap lists it.
  • Crawler appears and gets a redirect: fix the link you published so it points at the final URL.
  • Crawler gets a success code and a near empty body: the page needs scripts to render its content, which is the failure recorded on 6 August 2026.
  • Page is fetched and never cited: check whether it answers a question anyone asks, and whether it states a fact worth quoting with a date on it.
  • Page is cited once and not again: normal. Keep probing on a schedule instead of concluding from one reading.
  • The engine says it read your page and your logs disagree: trust the logs, and keep the engine's claim in the file as a note.

Common questions

How long does it actually take for a new page to be usable in an AI answer?

We will not give you a number, because we have one interval from one page. The honest answer is that the two things you can control are eligibility and quotability, and the interval you should measure on your own site is publication to first verified crawler request. Our single observed publication to citation interval was fourteen days, by ChatGPT, on 6 August 2026, and one interval is not a timetable.

Can I make a crawler come sooner?

You can remove the reasons it cannot come, and you can make the page easy to find: list it in your sitemap with a true last modified date, link to it from a page that is already fetched often, and make sure it returns its content on a plain request. Beyond that, nobody outside the company that runs the crawler controls its schedule.

Is a crawler visit good news on its own?

It is necessary and it is not the outcome. A page can be fetched on day two and never cited, and in our experience that is the ordinary case rather than a failure state. The gap between fetched and cited is usually about what the page states: a dated number, a named list, or a limit somebody would quote.

Should I block the training crawler and allow the search one?

That is a business decision and this page will not make it for you. What matters for measurement is that you know which one you have allowed, and that you do not read a low citation count as a content problem when it is a blocking decision somebody made two years ago.

Why verify the crawler rather than trusting the user agent?

Because anybody can send any user agent string. If you report an interval from an unverified log line, you may be reporting a scraper imitating a crawler. Verification against the operator's published address ranges is the difference between a log and evidence.

Does an older page get re fetched when I update it?

Often, and you should measure it the same way rather than assuming it. Record the update timestamp, watch for the next verified request, and run the citation probe again. An update with no re fetch and no change in citations is a useful null result, and it is worth knowing before you rewrite fifty pages.

Primary sources and update history

  • Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The interim check section holds the 6 August 2026 citation of clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india, fourteen days after publication. The same file holds the 27 July 2026 confirmation that no live web search was run for any of the 78 blind questions, the unreadable pricing page finding, and the two conflicting public price lists.
  • geo-audits/aiknowsus-com/. The September 2026 audit of our own domain, 24 batches and 72 conversations, including the three batches in which the engine withdrew its own claim to have searched.

Update history. 27 July 2026: baseline, search rate zero. 6 August 2026: first observed citation of a named page, fourteen days after publication, plus the crawler readability findings. September 2026: own domain audit and the withdrawn search claims. Not yet run: the timestamped crawl test described above, with named test pages, verified crawler requests and a daily citation probe. When it is run, this section will carry the per page intervals, including the pages that were fetched and never cited.

What to do first

Before you publish anything else, do the two minute version: read your robots file as text, fetch your most important page without scripts and read what comes back, and search your last thirty days of access logs for the search crawler by name. If it is not there, the timing question does not apply to you yet. If it is, extend log retention to sixty days today, pick three named pages you are about to publish, and record the publication minute for each. That is the whole start of a test that will give you a real number in a month.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.