Where ChatGPT gets its sources
Training data when it does not search, fetched pages when it does, and you can tell which.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
ChatGPT draws on two things: the text the model was trained on, and the pages it fetches during the answer when it searches. When you see citations, those are the fetched pages. When you see none, the answer came from training and there are no sources to inspect.
What ChatGPT does differently
Searching is not automatic. Some assistants search for almost everything and cite by default. ChatGPT decides, so the same question can produce a cited answer one day and an uncited one the next. This matters because the two kinds of answer reward completely different work.
The crawling is also split across separate bots with separate jobs. One fetches pages for the search index behind answers. One fetches a page when a user or the assistant follows a link during a conversation. One collects text used in training. They are named separately in robots.txt, which has a consequence most sites have not thought through: you can be open to one and closed to another without intending it.
When it does search, the fetched set is small. A handful of pages, usually the kind that are already a list of options: comparison articles, directories and review platforms, forum threads, and vendor pages that carry something specific.
We saw how much the distinction matters in our own work. In a baseline run on 27 July 2026, ChatGPT confirmed it had run no live search for the questions in that reading, so the whole baseline described what the model remembered rather than what it could find.
What that means for whether you are named
Your robots.txt is now a business decision rather than a technical default. Blocking the training crawler is a legitimate choice, and it does not block the search crawler. Blocking the search crawler removes you from the path you have the most influence over. Whatever you choose, choose it on purpose and write down why.
It also means an uncited answer and a cited answer are two different problems. If you are absent from cited answers, the fix is indexing, third party presence and quotable pages. If you are absent from uncited ones, you are missing from what the model remembers, and that changes slowly however much you publish this month.
The four kinds of page that get fetched
Named, because a vague total is no use when you are deciding where to spend an afternoon.
One, comparison and best of articles. Already in the shape the answer needs, so they are read first for buying questions.
Two, directories and review platform category pages. Lists of options with structured fields. In India the sector specific ones are cited more often than people expect and are easier to join.
Three, forum and community threads. They carry independent opinion, which no vendor page can supply.
Four, vendor pages, in a narrow role. Yours gets fetched to confirm facts: coverage, price, who you serve. A vendor page of adjectives fails even that job.
A worked example
On 6 August 2026, in a Claude run over the site of clawlaw.in, our sister company, we found out where the facts in the answers had come from, and it was not from the company's own site.
Its pricing page rendered the prices only after scripts ran, so what a crawler received contained no prices at all. In that reading, the prices the assistants did quote had come from an app store listing instead.
So the source set for those answers was decided by what loaded. The company had published its prices. The page a machine could read had not, and a third party listing became the authority on what it charged.
That is the four kinds of page above in practice, with the fourth kind missing. Your own pages get fetched to confirm facts, and when they cannot supply a fact, something else supplies it. Check which of your important facts survive in the raw HTML before you spend time on anything further up this page.
What is not public
Four things nobody outside these companies can tell you, so treat any confident claim about them with suspicion.
What decides whether it searches. There is no published rule, and asking the assistant is not evidence, because the reply is generated text.
What is in the training data. You cannot check whether your pages are in it, and you cannot request inclusion.
How the fetched pages are ranked. The retrieval and ordering logic is not documented.
Whether the shown citations are all the pages used. The list is what the app chose to display.
What is public and checkable is the bot names in robots.txt, the citations shown in an answer, and your own server logs. Build your plan on those three and you will not have to unwind it when a guess turns out wrong.
Common questions
How do I know whether it searched?
Look for citations and any line in the answer saying it searched. That is evidence. Asking it directly is not reliable, because it can produce a plausible account of itself that is wrong in either direction.
Should I allow the training crawler?
It is a genuine choice. Allowing it may help you become part of what the model holds later, with no way to verify or measure that. Blocking it protects your text and costs you nothing in the live path. Decide it deliberately and record the reason.
Can I stop my competitor's page being used as a source?
No. You have no control over what pages are fetched. What you can do is be present and accurate on the same pages, and correct anything wrong about you on them.
Is a sitemap enough to get fetched?
It helps with indexing and it is not enough on its own. Internal links from pages that already get visits, plain HTML, and a page that is the clear answer to a real question all matter more.
Do citations mean a visitor will arrive?
Sometimes. Assistant referrals tend to be small in number and unusually well informed. Check your referral report for the assistant domains rather than assuming either a flood or nothing.
Why is my page cited for one question and not a similar one?
Because the search step runs per question and small wording changes pull different pages. This is normal, and it is why a set of questions tells you something a single question cannot.
What to do first
- Read your robots.txt today and list which AI related crawlers it allows. Decide each one deliberately and write the reason down.
- Serve your important pages as plain HTML. A page whose text needs JavaScript is often fetched as an empty page.
- Be present on the four kinds of source above, not only on your own site.
- Keep one page per buyer question, with the answer near the top, so a fetched page is worth quoting when it is fetched.
- Ask one buying question with search on, copy out every citation, and count how many of the cited pages are yours and how many are third party.
Then ask the same question with search off. If you appear only in the searched version, your presence is entirely in the live path. That is a useful and quite common finding for a business that has been working on this for less than a year.