Does Claude search the web for this answer? A surface-by-surface verification guide
Five surfaces, five different answers, and how to verify retrieval on the one in front of you.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
It depends on the surface and on the request, and you should verify rather than assume. Web search is a tool that has to be available and then used, so the same question can be answered from memory on one surface and from fetched pages on another. On the API it is switched on by the developer and billed per thousand searches, and the current price is published on Anthropic's own pricing page, which is where you should read it on the day you need it. We do not restate that figure here, for the reason this whole site keeps repeating: a price copied into an article with no date is exactly the kind of claim an assistant discounts, and we have a dated case of that happening to unsourced figures.
The behaviour question has a better number behind it anyway: the count of answers in your own test run that showed a search step and produced citations you could open. That is measurable this afternoon, and the method is below.
The answer, first
Five surfaces, treated separately, because a finding on one does not transfer to another.
- The claude.ai web interface. Search behaviour is a product decision and can change. Evidence of retrieval appears in the interface and in the citations attached to the answer.
- The Claude mobile apps. Same product family, and worth testing separately because what an interface displays about its own sources is not identical everywhere.
- The Claude desktop app, which may also have connected tools available, and anything a connected tool returns is a source too.
- Claude Code in a terminal, where fetching and searching depend on which tools are enabled in that environment.
- The API. Nothing is searched unless the developer supplies the web search tool. This is the only surface where you fully control the condition, which makes it the right instrument for a controlled test.
Two rules that hold across all five. A confident answer is not evidence of retrieval. And the model's own account of whether it searched is a statement, not a log.
How it was measured
Where to get the price, and what to record
For the API, open Anthropic's pricing page, find the web search tool line, and write down three things: the price per thousand searches as stated, the date you read it, and the URL. Put those three in your own cost note. If you publish the figure anywhere, publish all three together. A price without a date ages into a wrong claim, and we have a dated case of correct figures being set aside precisely because they carried no source and no date. Claude, 17 September 2026, on clawlaw.in.
The behaviour test, in seven steps
- Write twenty questions: ten that cannot be answered correctly without current information, ten that can be answered from general knowledge. The contrast is what makes the result readable.
- Pick your surfaces and test each one separately, never pooling the results.
- Run each question three times on three separate days, in a clean session every time, with nothing uploaded and nothing pasted.
- Record per answer: date, time, time zone, surface, model name or model string, whether a search step was visible, the number of citations, every URL, and the full answer text.
- Open the citations. Decide for each whether the page contains the claim it is attached to, and write the deciding line into your log.
- Score one strict outcome: the answer counts as retrieval-verified only if at least one citation opens and supports the claim it sits next to.
- Report counts per surface, split by the two question types, with the dates. Never one blended figure.
Why the strict definition matters
A loose test counts any link. A link to a home page, or to a page that does not contain the claim, tells a reader that the answer is grounded when it is not. The strict definition produces a smaller number and it is the only number worth publishing, and it is also the definition to apply when you check whether your own pages were used properly.
What the numbers were
We have not run the five-surface test and publish no per-surface fraction. There is no Claude search rate on this page, because we have not measured one.
What we hold, dated, on the underlying question of whether assistants retrieve when they appear to:
- ChatGPT, 27 July 2026, on clawlaw.in. Asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. Breadth in that memory reading was ProVakil on 28 questions, CLAW on 21, Legistify on 18.
- Perplexity, September 2026, on aiknowsus.com. Asked afterwards how many of the questions it had actually searched for, the assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches.
- Claude, 17 September 2026, on clawlaw.in. Retrieval happened and the credit did not follow: our pages ranked in the assistant's raw results and shaped what it wrote, and the site landed as a name inside a list rather than as a linked recommendation. In the same run the assistant admitted using two arguments from our pages and dropping the attribution in both places.
- ChatGPT, 6 August 2026, on clawlaw.in. A verified live-source answer with a URL in it: clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india cited as a source for a vendor due diligence question, fourteen days after publication.
Read together, those four say something a per-surface rate would not. Retrieval is not guaranteed, self reports are unreliable, and retrieval happening is not the same as your domain being credited. A test that measures only the first of those three will mislead you about the other two.
What this cannot tell you
- It cannot establish what Claude does in general. Twenty questions on one surface on three days describe sixty answers.
- It cannot see inside the product. You are inferring from what the interface displays, and interfaces change what they display without notice.
- It cannot use the model as a witness. The dated retraction above is the reason.
- It cannot price your own usage. The API price per thousand searches is on Anthropic's page, it changes, and only your own account statement tells you what you spent.
- It cannot be pooled across surfaces. The claude.ai interface and an API call with a tool enabled are different conditions, and adding them together produces a number describing neither.
- It cannot promise a citation. Even a verified retrieval answer may not cite you, as our own 17 September 2026 run shows.
Common questions
Why does Claude sometimes say it cannot browse?
Because on that surface, in that session, the tool was not available. That is a real condition and it is exactly why a test records the surface for every answer rather than assuming one behaviour across the product.
Is the API cheaper than the consumer product for running a test?
They are different things. The consumer product is bundled and the API is billed per use, including per thousand searches. Read the current figures off Anthropic's pricing page on the day and record the date, and remember that the API half of a test is cleaner while the consumer half is closer to what your buyers experience.
If I see a search step, is my page definitely readable?
No. A search step means queries ran. Whether your page was reachable is a separate question, and in our own programme a pricing page gave a crawler no prices at all because they only rendered after scripts ran. Claude, 6 August 2026, on clawlaw.in.
How do I prove an answer used live sources, for a report?
Save the answer text, the full citation list, and a note for each citation saying which line on that page supports the claim. That is defensible. A screenshot of a link icon is not.
Does asking for sources change the result?
Often yes, which is why a measurement must not do it. Ask the question your buyer would ask. If you add "with sources", you are measuring your own instruction rather than the default behaviour.
Sources and change log
Read directly: Anthropic's documentation for the web search tool, including how results and citations are returned; Anthropic's model overview pages for knowledge cutoffs; and Anthropic's pricing page for the web search price per thousand searches and for token pricing, on the day you need it, with the date recorded next to whatever you copy.
Our observations come from the clawlaw.in programme from July 2026, including the baseline of 27 July 2026, the interim check of 6 August 2026 and the Claude response audit of 17 September 2026, and from the aiknowsus.com audit of September 2026, 24 batches and 72 conversations. The 6 August citation is checkable from outside because it names a URL; counts taken across our capture files are not, until we publish them.
Change log. First published 29 September 2026. It carries no API price figure on purpose, and no per-surface search rate, because we have measured neither. Both will be added with dates and denominators if and when we have them.
Disclosure. AI Knows Us sells AI visibility measurement.
What to do first
Ask the same question twice today, once on claude.ai and once anywhere else you use Claude, and for each answer write down whether a search step appeared and whether any citation opens to a page containing the claim. Two answers is not a rate, and it will settle in five minutes whether the surface your buyers use retrieves anything at all about your category.