Claude's information sources: Anthropic's published statements, web search, and what users can verify

Two paths in, one of them inspectable, and a claim-level audit for counting citations that actually support the answer.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Claude's answers come from two places: what the model absorbed up to its knowledge cutoff, and pages fetched during the conversation when a search or browsing tool is available. Only the second path leaves evidence a user can check, which is a displayed citation that opens and contains the claim it is attached to. The measurement worth running is therefore claim-level rather than answer-level: how many tested claims had a displayed citation that directly supported them, out of all claims tested. As of 29 September 2026 we have not completed that audit with a published count, so this page carries the full protocol and the dated observations we do hold, one of which is an assistant confirming it used our material and dropped the credit.

The answer, first

Four sources, ranked by how much you can verify about them.

  • Training knowledge, fixed at the cutoff. Not inspectable from a chat, and the reason a fluent description of your company can be two years out of date.
  • Web search or browsing, when the tool is available. This is the inspectable path. Anthropic documents a web search tool for the API, and the consumer product has its own search behaviour. When it runs, citations appear.
  • What you put in the conversation. An uploaded file, a pasted URL, an earlier turn. A real source, and in a test a contaminant, which is why every test session starts clean.
  • Connected tools and data, where a user or a business has connected them. Whatever those tools return is a source, and it is visible only to the person who connected them.

The practical rule for a business: only the second path can be influenced by publishing this month. Work aimed at the first path is slower and less certain, and work aimed at the third is not visibility work at all.

How it was measured

Why the unit is a claim and not an answer

An answer-level count tells you a source list existed. A claim-level count tells you whether the specific sentence about your price, your coverage or your suitability had anything behind it. Those come apart constantly, and the gap between them is where wrong facts about businesses live.

The protocol, in eight steps

  • Write twenty questions whose answers must contain checkable facts about your category and your company. Ten with your name in them, ten blind.
  • Run each in a clean session, with no uploads, no pasted links and no earlier turns.
  • Record the conditions: date, time, time zone, surface, model name or model string, and whether a search step was visible.
  • Break every answer into claims. One factual statement per row. A sentence with three facts becomes three rows. This is the slow part and it is the part that produces the number.
  • For each claim, record whether a citation is attached in the displayed sources.
  • Open that citation and decide one thing: does the page contain the claim, in substance, or not. Write the deciding line from the page into the row.
  • Score each claim into one of four states: supported by a citation, cited but not supported, correct with no citation, or wrong.
  • Report two figures: claims supported by a citation over claims tested, and claims correct but uncited over claims tested. Both with the dates.

The four states, with why each matters

  • Supported by a citation. The strong state. A reader can check it and so can the next assistant.
  • Cited but not supported. A link is attached and the page does not say it. This is the state that misleads a reader who trusts the presence of a link.
  • Correct with no citation. Right today, unverifiable, and liable to drift. This is where most confident statements about mid-sized companies sit.
  • Wrong. Record the exact wording, then go and find the public page that says it.

What the numbers were

What we have not run: a claim-level audit with a published count. There is no "X of Y claims were citation-supported" figure on this page because we have not produced one, and we are not going to estimate one.

What we hold, dated, and each of it bears on the same mechanism:

  • Our material used, the credit dropped. The assistant admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, which it described as a citation lapse rather than a ranking judgement. Claude, 17 September 2026, on clawlaw.in. Two claims, both sourced from our pages, neither citation-supported on screen.
  • Unsourced figures were set aside even when correct. Two comparison pages stated competitors' prices with no link, no date and no source. The figures turned out to be correct when the assistant checked them independently, but it had no way to know that at read time, so it treated the pages as advocacy and used official sources instead. Claude, 17 September 2026, on clawlaw.in.
  • Facts taken from the wrong public source. A pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants did quote had come from an app store listing instead of from the company's own site. Claude, 6 August 2026, on clawlaw.in.
  • What the source mix looks like across a whole capture. The domains cited most often across the whole capture were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Perplexity, September 2026, on aiknowsus.com.
  • How often we were not cited at all on our own domain. Across 24 batches and 72 conversations, the phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers. Perplexity, September 2026, on aiknowsus.com.
  • An assistant's account of its own searching, withdrawn. Asked afterwards how many of the questions it had actually searched for, the assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Perplexity, September 2026, on aiknowsus.com.

The last one is the reason the protocol above never asks the model anything about its own sources. It asks the interface, and then it opens the links.

What this cannot tell you

  • It cannot see training data. Nothing a user can run reveals what the model absorbed about your business before the test.
  • A citation is not proof of use. The model may have known the claim and cited a page that agrees with it. The audit records support, not causation.
  • An absent citation is not proof of nothing fetched. Interfaces differ in what they display, and display behaviour changes.
  • Claim splitting is a judgement. Two people will break the same paragraph into a slightly different number of claims, which is why the denominator must be published with the log rather than quoted on its own.
  • It is one account, one location, one set of days. Your figures are yours, and somebody else repeating the prompt set may get different ones.
  • It says nothing about other assistants. Gemini, ChatGPT, Copilot and AI Overviews each need their own run.

Common questions

Can I ask Claude where it got something?

You can ask, and treat the reply as a statement rather than a record. We hold a dated case of an assistant withdrawing its own account of whether it had searched. The citations displayed with the original answer are the evidence, and opening them is the check.

Why is our price wrong when our website shows it?

Check what a machine receives from your page rather than what your browser shows. In our own programme the prices existed on screen and not in what a crawler received, and the figures the assistants used came from an app store listing instead. Claude, 6 August 2026, on clawlaw.in.

Does the knowledge cutoff mean old information about us is permanent?

No, but it does mean two separate jobs. Retrieval can be reached this month by publishing pages worth fetching. Model memory changes when models change, on a timetable nobody outside controls.

If Claude uses our page without linking it, is that worth anything commercially?

It is worth something and less than a link. Your argument reaches the buyer, your name does not. In our audit that was described by the assistant as a citation lapse rather than a judgement about our quality, which suggests the page was doing its job and the credibility signals around it were not.

How many claims do I need to audit?

Twenty questions usually yields well over a hundred claims, which is enough to write both figures with denominators. Do not start with a hundred questions; nobody finishes the claim splitting, and an unfinished audit produces no number at all.

Are review sites a route into these answers?

On what we hold, they are weaker than official pages and vendors' own pages. Across six questions in one audit no third party review or directory source made it into any answer, and one well known review site was discarded for offering American products against an India question. Claude, 17 September 2026, on clawlaw.in.

Sources and change log

Read Anthropic's own pages for the documented behaviour: the web search tool documentation for the API, the model overview pages including knowledge cutoffs, the usage policies, and the pricing page when you need a cost figure, read on the day and dated in your notes. We deliberately restate no prices here.

Our observations come from the clawlaw.in programme from July 2026, the Claude response audit of 17 September 2026, the interim checks of 6 August 2026, and the aiknowsus.com audit of September 2026 covering 24 batches and 72 conversations. Where an observation concerns a public page it can be checked from outside. Counts taken over our capture files cannot be, until we publish the captures.

Change log. First published 29 September 2026 with the protocol complete and no claim-level count. When the audit is run, the numerator, the denominator, the question set and the dates will be added here, and this note will stay so a reader can see what the page said before.

Disclosure. AI Knows Us sells AI visibility measurement.

What to do first

Ask one question whose answer must contain your price, in a clean session, and split the answer into claims on paper. For each claim, look for a citation and open it. Most businesses find within twenty minutes that the statements about them are correct-with-no-citation at best, and that tells you exactly which page to publish next.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.