Where Claude gets its information

From training for general knowledge, and from live web search when the question needs it.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Claude answers from two places: what it learned during training, and web pages it fetches when a question needs current information. Which one it used changes what the answer means for you, so it is worth asking directly, and it will tell you.

The two modes, and why you should separate them

Training knowledge is months old at best. If you published pages recently, they are not in it. Live search is today's web, and your new pages can appear there quickly.

This is not a small distinction. On 27 July 2026 we put 78 buyer questions to ChatGPT for clawlaw.in, our sister company, and it confirmed afterwards that it had run no live web search for any of them, so the entire reading described a market as the model remembered it. Fourteen days after one new page went up, on 6 August 2026, ChatGPT cited that page for a vendor due diligence question in a part of the market where the same question set had named the company nowhere at baseline. Anybody measuring without checking which mode was used is likely reading the wrong thing.

Anthropic publishes the names of its crawlers, which is the other half of this story. ClaudeBot is the crawler that collects content, Claude-User covers fetches made when a user's request causes a page to be read, and Claude-SearchBot relates to search indexing. If you block these in robots.txt or at your firewall, you are removing yourself from the half of this engine that can react to anything you publish. Decide that deliberately rather than inheriting it from a copied file.

What it cites when it does search

In our own audits the pattern is consistent and, for a small company, sobering. The following five kinds of source do the work.

  • Official and government pages, which dominate wherever the subject touches public records, rules or definitions with a legal meaning.
  • Vendors' own pages, appearing frequently, usually those that publish real pricing or documentation rather than positioning.
  • Technical documentation, which is quoted with noticeably more confidence than marketing material from the same company.
  • Independent review sites, which appear far less than most people assume, and in some markets barely at all.
  • Community and forum threads, for questions about what a thing is like in practice and what goes wrong.

That fourth point matters for anyone told to go and get third party reviews. In several categories those reviews effectively do not exist, and the realistic route is to be the vendor whose own pages are the most verifiable.

How to find out which mode you got

Ask, in the same conversation: for how many of these questions did you run a live search, and which sources did you actually retrieve. The answer is specific and it reframes everything above it. A confident answer built on no searches is a statement about the past.

Write the count in your sheet next to the result. Two readings without that number cannot be compared, because you do not know whether you are looking at a change in the world or a change in the instrument.

A worked example

Here is the reading that makes the point, and the number in it is real. On 27 July 2026 we put 78 buyer questions to ChatGPT for clawlaw.in, our sister company, with no brand named in any of them. Asked afterwards, it confirmed that it had run no live web search for any of the 78. Not most of them. None.

So all 78 answers were a memory reading. Nothing published in the weeks before could have appeared in any of them, because nothing published had been read. A team looking at that set without knowing the search count would have concluded that recent work had failed, when recent work had simply never been in the room.

The same caution applies to what an engine says about its own searching. In September 2026, on our own site aiknowsus.com, we asked Perplexity how many of the questions it had really searched for. It withdrew its earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question. That happened in three separate batches.

So ask Claude for the count, write it down, and split your sheet on it. Searched questions are a live reading: they reflect today's web, your recent pages could have appeared in them, and a change you make can show up there. Unsearched questions are a memory reading: they reflect how widely your category was written about before the model was built, and nothing you publish this quarter will move them. Being named on one searched question and on none of the unsearched ones is not one small number. It says your recent work is starting to be findable and your long term presence in the record is thin, and those are two different jobs on two different timescales.

If you average everything into one figure, you get a number that tells you nothing about what to do on Monday.

How to use this

If you want to influence the training side, you need to be written about in many places over a long time, and that is slow. If you want to influence the search side, publish pages that answer real questions and can be fetched, and you can see movement in weeks. Do the second one first.

The one piece of work that helps both is the same: get your name, stated accurately, onto pages you do not own. Those pages are retrievable today and they are part of the record that the next model reads.

What this page cannot tell you

Anthropic does not publish what is in the training data, how the decision to search is made, or how sources are weighed once retrieved. The patterns above come from reading many answers and from asking the model what it did, which is better evidence than most articles have and is still not documentation.

The citation list also shows what was used, not everything read, so being absent from it is weak evidence. And behaviour differs between surfaces and product versions, so a reading in one place does not stand for all of them.

Common questions

Can I just ask it whether it searched?

Yes, and you should, every time. It answers specifically, including how many searches and which pages. This is the single most useful habit available on this engine and it takes one extra line in your prompt.

How often should I check?

Monthly for the full set. If you check more often, expect the memory based questions to sit still, which is not a sign that your work failed. For fast feedback on new pages, use an engine that always retrieves.

Why do two runs disagree?

Usually one searched and one did not. Then, among searching runs, the retrieved pages can differ. Record the search count and most of the mystery disappears.

It cited a competitor for a fact about my company. What now?

Ask which page, open it, and check the fact. Wrong facts get a polite correction request and a dated copy kept for your records. Correct facts mean the gap is on your side: publish the fact on a page built for that question, in text, with a date, and confirm the crawlers are not blocked.

Is being cited the same as being recommended?

No. Your page can supply the facts of an answer that recommends somebody else, which is the most common outcome for vendor pages here. Count them separately.

Should I block the crawlers to protect my content?

That is a real decision with a real cost, and it belongs to you rather than to an article. Blocking protects content from being used and also removes you from the answers your buyers read. What you should not do is block by accident, which is what happens when a robots.txt is copied from a blog post. Read your own file today.

What to do first

Add one line to your test prompt: ask it to tell you, after answering, how many live searches it ran and which pages it retrieved. Then run your three most important questions and write that number in your sheet. Everything else on this page is easier once you can tell a live reading from a memory one.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.