Claude web search and SaaS recommendations: a dated citation and repeatability test

The fraction that matters is runs where a product was named and supported by a citation, printed per prompt with its sample size.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

When Claude answers a SaaS buying question it may or may not run a web search, and a product can be named in the answer with a supporting link, named with no link, or used without being named at all. So the only fraction worth publishing is narrow: for one exact prompt, the number of repeated runs in which a named product appeared and was supported by a citation, over the number of runs made, with the date. As of 29 September 2026 we have not run a repeated run citation test on Claude for SaaS prompts, so no such fraction appears on this page. What appears instead is the protocol, the four states you must label, and the Claude measurements we do hold from 6 August and 17 September 2026 in Indian legal research software.

The reason this fraction is so narrow is that the wider versions are all misleading. A share of answers mentioning a product counts unlinked mentions, which send nobody anywhere. A share of citations counts links without saying whether the product was recommended. Only named and cited together tells you that a buyer heard the name and could act on it.

Scope and definitions

Four states, and they are not interchangeable. Most published visibility numbers quietly mix them, which is why none of them can be compared with each other.

  • Named and cited. The product appears in the answer text and a link on the product's own domain supports it. The only state that both tells the reader who you are and sends them to you.
  • Named, not cited. The name appears with no link, or with a link to somebody else's page. On 17 September 2026, Claude confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
  • Cited, not named. A link to your page supports a point that does not mention you. Useful, and invisible to a brand monitoring tool watching for your name.
  • Used, not attributed. The answer clearly uses material from your page and neither your name nor your URL appears. In that same 17 September 2026 run, Claude admitted it had used two specific arguments drawn from clawlaw.in pages and had dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement.

Two more labelling rules keep this stable across hundreds of rows. A product named inside a third party listing that the answer links to is the listing being cited, not the product, so it gets its own state. And a product named in a follow up answer, after the reader asked a second question, is a separate observation with its own prompt text, never merged into the first.

How the test was run

Nine steps. This is the protocol we will run, written so you can run it first if you want to.

One. Freeze the prompt set. Between 10 and 40 prompts in buyers' words, in a file with a version number and a date, not edited for the length of the study.

Two. Keep your brand out of every prompt. Measured reason: on 27 July 2026, two runs of the same 78 questions happened on one day in the clawlaw.in programme. The wrapper that named the brand ranked it first on almost every question. The blind run put the company second by breadth and absent from the litigation due diligence questions it most wanted to win. The first run was discarded because it had measured our own prompt. ChatGPT, 27 July 2026.

Three. Fix the repetition count at five runs per prompt as a floor, and decide it before you see any results. Five is the smallest count that distinguishes a prompt that always names a product from one that names it occasionally.

Four. Make each run independent. A new conversation every time. No projects, no custom instructions, no memory features, and each of those recorded as off. Runs spread across the day rather than fired in a minute.

Five. Record whether a search ran, and treat it as a claim. Claude's web search is a tool that may or may not be invoked for a given question. Where the interface shows you that it searched, save that. Where you have to ask, mark it unverified, because an assistant's account of its own work has failed before: in the aiknowsus.com audit of September 2026 Perplexity withdrew its own statement about how many questions it had searched for, saying it could not honestly substantiate the claim, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches.

Six. Log these thirteen fields per run. Prompt set version, prompt id, exact prompt text, run number, timestamp, surface used, model version as the interface reports it, whether search was shown as having run, country, language, device, the answer text file name, and the screenshot file name. Then one row per product named, carrying the product name, the state from the four above, and the domain of any supporting link.

Seven. Save the answer text and every link exactly as given, before any cleaning. Your conclusion is not evidence. The answer text is.

Eight. Label twice. A second person labels at least 20 rows blind, and you publish how often you agreed. The four states above are easy to apply inconsistently, especially the difference between named and cited.

Nine. Rerun the frozen set on a fixed interval, same conditions, and report each pass as its own fraction rather than as a growth percentage, because a growth percentage hides both denominators.

How to write the fraction

Per prompt, always. Prompt 4 of version 1, five runs, product named and cited in 2 of 5, named without a citation in 1 of 5, absent in 2 of 5. That is one line and it is complete. The study wide line comes after it: across 12 prompts at 5 runs, that is 60 runs, and the product was named and cited in some number of those 60. If runs were dropped for a refusal or an error, the dropped count gets its own line so the denominator shrinks where the reader can see it. This is arithmetic, not an observation.

Named SaaS categories and products

We do not print product samples we have not tested. The two samples with real runs behind them, each with its sector named so nobody reads one as the other:

  • Indian legal research software. ProVakil, CLAW and Legistify, appearing on 28, 21 and 18 of 78 blind questions in the ChatGPT baseline of 27 July 2026. On that same run ChatGPT confirmed it had run no live search for any of the 78, so those counts describe what the model had absorbed rather than the web that week.
  • AI visibility and generative engine optimisation tooling. Semrush, then Profound, then Peec, then Otterly, then Scrunch, by how often they were named across the whole aiknowsus.com capture of September 2026. Perplexity, 24 batches and 72 conversations.

Six categories that deserve separate frozen prompt files, because buyers ask about them in different language and the source sets barely overlap: accounting and GST filing, payroll and HR, helpdesk and customer support, contract lifecycle management, legal research, and logistics or field service.

Findings

Repeated run citation fraction for Claude on SaaS prompts, as of 29 September 2026: not run. Zero runs exist. We publish no fraction, no estimate and no range in its place.

The Claude measurements we do hold come from the clawlaw.in programme in Indian legal research software, and they are printed here because they describe the same mechanism this page is about.

  • 0 of 6. Across six questions on 17 September 2026, the number in which any third party review or directory source reached the answer was zero. One well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
  • Used and not attributed, twice in one run. 17 September 2026. Claude admitted it had taken two specific arguments from clawlaw.in pages and dropped the attribution both times, calling it a citation lapse rather than a ranking judgement.
  • Named without a link. Same run. The site's pages ranked in the raw results and shaped the answer, and it still appeared as a name in a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
  • Two enterprise vendors ranked above a better product for a stated reason. Same run, on a question about finding every case against a company. Claude placed two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and noted that its ordering reflected price transparency and source authority rather than product quality.
  • Official portals above every commercial product on how to questions. Same run. Claude said plainly that no commercial product should rank above the official portal for a question about using that portal.
  • An unreadable page hands the fact to somebody else. On 6 August 2026 the clawlaw.in pricing page rendered prices only after scripts ran, so a crawler received no prices, and the prices the assistants quoted came from an app store listing. Claude and ChatGPT, same date.
  • A claim that could not be true cost the recommendation. Also 6 August 2026. A headline figure on the main site could not be true, and when Claude checked it against public numbers it advised a buyer against the product, after which the company's accurate claims stopped counting for that answer.

The fourth finding is the one a SaaS company can act on this week. Claude ranked competitors above a product it appeared to rate, and said the reason out loud: they published explicit coverage lists and clearer prices. Publishing a named, counted, dated scope list is a smaller job than building a feature, and on that run it was the difference.

What Anthropic documents about web search

We do not restate Anthropic's documentation here. It is edited, and an undated paraphrase turns into a false claim silently. The behaviour we are describing, whether a search runs for a given question, is exactly the kind of thing that changes between product versions, so a restatement on this page would be the least trustworthy sentence on it.

Read these three yourself, and note the URL and the date beside anything you take from them.

  • Anthropic's help centre pages for the Claude apps, for what the consumer product does about web search and citations.
  • Anthropic's developer documentation for the web search tool, if you are testing through the API, because the API and the app are different surfaces and a result from one is not a result about the other.
  • Anthropic's published notes on model versions, so the version string you log means something to a reader a year from now.

If any of those contradicts a sentence here, the sentence here is wrong. Tell us and we will change it.

Raw data and limitations

Six limits.

  • There is no Claude SaaS repeat run data behind this page. The Claude findings above are from six questions in Indian legal research software on two dates, and that is a small, specific sample.
  • A fraction from repeated runs is not a probability. 2 of 5 on one prompt on one day is five draws, not a forty per cent chance for your buyer next week.
  • One observation is not a median and not an average. The nearest thing we hold to a timing result is a single citation of a named clawlaw.in URL by ChatGPT on 6 August 2026, fourteen days after that page was published. Fourteen days is one observation. It is not a median time to first citation and we will not present it as one.
  • Some rows cannot be checked by you yet. The 6 August 2026 citation of a named URL can be verified from outside. The counts over our capture folders cannot, until the captures are published.
  • A change after a change is not a cause. An engine can change its own behaviour in the same month you publish pages.
  • Nothing here can promise a position in an answer. We do not sell that and would not believe anybody who did.

Raw data status: the protocol, the four states and the thirteen log fields are fully reproducible today without us. Our own answer texts and screenshots for the dated findings are being prepared for publication.

Version: 29 September 2026, first publication. This page updates when the Claude repeat run set is completed, when a figure is corrected, and when the capture files go up.

Common questions

How do I tell whether Claude searched for my question?

Look at what the interface shows for that answer and save it. Do not rely on asking the assistant afterwards. In September 2026 Perplexity withdrew its own statement about how many questions it had searched for, in three separate batches, and any engine's account of its own work has to be logged as a claim rather than as a measurement.

Why separate named from cited at all?

Because they have different consequences and different fixes. An unlinked mention means the model knows of you and had no page worth linking, which is a content problem. A citation with no mention means your page was useful and your name was not the point, which is a positioning problem. Merged into one number, neither is visible.

Our comparison page is accurate. Why would it be treated as advocacy?

Because accuracy is not observable at read time. On 17 September 2026 Claude checked two clawlaw.in comparison pages that stated competitors' prices with no link, no date and no source, found the figures correct on independent checking, and still used official sources instead. The fix is not better prose. It is a date, a link and a source on every competitor fact you state, and pointing the reader at the competitor's own page.

Should I test through the app or the API?

Whichever your buyers use, and say which one you used. They are separate surfaces with separate behaviour around search and citations, and a fraction from one does not transfer to the other. If you test both, report them as two studies.

Five runs feels small. Why not fifty?

Fifty is better if you can afford the time. Five is the floor because it is where variance becomes visible at all. The honest move is not a bigger number, it is printing whatever number you used next to every fraction, so the reader can judge it instead of guessing.

What to do first

Pick your five highest value buying prompts. Run each five times in five separate conversations today, logging the thirteen fields, and label every product mention with one of the four states. Then take the single most important fact a buyer asks you about, check that it appears in your page as served rather than after scripts run, and add a date and a source to every competitor fact on your comparison pages. Those three jobs address, in order, the measurement and the two failures Claude named out loud on 6 August and 17 September 2026.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.