Claude web search for marketing agency recommendations: a dated prompt and citation study

Print the repeat run count next to every prompt, and report cited mentions as a fraction of those runs.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Claude may or may not run a web search for a given agency question, and the agency names it produces can change between two identical runs, so the only defensible output is a line per prompt: the exact prompt text, the number of repeated runs behind it, and the number of those runs in which an agency was named with a citation that actually supports it. As of 29 September 2026 we have not run a marketing agency prompt set on Claude, so the repeat run count on this page is zero and no cited mention rate is published. What is published is the protocol, the run log fields, the arithmetic, and the Claude measurements we hold from 6 August and 17 September 2026 in Indian legal research software.

A cited mention rate without its run count beside it cannot be read. Three of five and thirty of fifty can be written as the same percentage and they carry completely different weight, and the run count is the only thing that tells them apart.

Study scope and definitions

Five labels. Each is counted separately, and a study that merges any two of them produces a number nobody can act on.

  • Named and cited to the agency's own site. The name is in the answer text and a supporting link sits on the agency's domain.
  • Named and cited to a third party. A directory, a listicle, a client page or a trade publication supports the name.
  • Named with no support. The name appears and no link supports it, or the link offered does not mention the agency at all when you open it.
  • Cited without being named. A page on the agency's domain supports a point in the answer and the agency's name never appears. Invisible to a brand monitoring tool watching for the name.
  • Used without attribution. Material from the agency's pages shapes the answer and neither the name nor the URL appears anywhere. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement.

The headline fraction this page is about uses the first two labels together, because both send a reader somewhere and both let them check the claim. The third label is tracked separately and never added in.

Test protocol

Nine steps, Claude specific where it matters.

One. Freeze three prompt files, one for local intent with the city named, one for industry intent with the sector named, one for task intent describing the work. Version number and date on each, and no editing during the study.

Two. No agency name inside any prompt. Measured reason, from a different sector: on 27 July 2026 two runs of the same 78 questions happened on one day in the clawlaw.in programme, and the run whose wrapper named the brand ranked it first on almost every question while the blind run placed the company second by breadth and absent from the questions it most wanted to win. The branded run was discarded. ChatGPT, 27 July 2026.

Three. Fix the repeat count before the first run. Five per prompt is our floor. Deciding the count after seeing results is how a study becomes an argument.

Four. One new conversation per run. No projects, no custom instructions, no memory features, each recorded as off. Spread runs across the day.

Five. Record whether a search ran, as a claim rather than a fact. Claude's web search is a tool that may not be invoked for a given question, so an answer with no search behind it is a different observation and must be labelled. Where the interface shows what it consulted, save that. Where you have to ask the assistant, mark it unverified: in the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement about how many questions it had searched for, saying it could not honestly substantiate the claim, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches.

Six. Log these thirteen fields per run. Prompt file version, prompt id, exact prompt text, run number, timestamp, surface, model version as the interface reports it, whether search was shown as having run, country, city setting, language, device, and the answer text filename. Then one row per agency named, with the agency name, the label from the five above, the supporting domain, and whether the destination page mentions the agency.

Seven. Open every citation. A supporting link that does not mention the agency is recorded as named with no support. Studies that skip this step overstate their cited rate.

Eight. Label twice, with a second person doing at least 20 rows blind, and publish the agreement rate. The line between named and cited to the agency's own site and named with no support is where labellers disagree most.

Nine. Rerun each frozen file on a fixed interval, reporting each pass as its own fraction with its own run count, never as a growth percentage.

How the fraction is written

One line per prompt, in this shape: prompt 3 of local file version 1, Pune, five runs, agency named and cited in 2 of 5, named with no support in 1 of 5, absent in 2 of 5, dates stated. Then the file wide line: six local prompts at five runs is 30 runs, and the cited mention count over 30. Dropped runs, whether from a refusal or an error, get their own line so the denominator visibly shrinks. This is arithmetic, not an observation, and none of these numbers is a result of ours.

Named sample

We print no agency names, because we have run no agency prompts. A list assembled from directories and placed next to a protocol reads like a finding and is not one.

The sample generation rules, which are the transferable part: derive the list from a first pass rather than from your own view of the market; record each name's first appearance with prompt id, run number and date; resolve every name to a working website and keep unresolvable names as their own rows; write the merging rule for groups and local offices before labelling; and publish the sample with the prompt file version that produced it.

The two samples we do hold, each generated from real runs, with sectors named so they are not misread: ProVakil, CLAW and Legistify in Indian legal research software, appearing on 28, 21 and 18 of 78 blind questions in the ChatGPT baseline of 27 July 2026; and Semrush, then Profound, then Peec, then Otterly, then Scrunch in AI visibility and generative engine optimisation tooling, by frequency of mention across the whole aiknowsus.com capture of September 2026, counted by Perplexity across 24 batches and 72 conversations.

Findings

Repeat runs completed on Claude for marketing agency prompts as of 29 September 2026: zero. Cited mention rate: not measured. No estimate, no range, nothing borrowed.

What we hold from Claude, in Indian legal research software, on two dates. These are printed because they describe the same mechanism an agency study would measure, not because they are agency results.

  • 0 of 6. Across six questions on 17 September 2026, the number of answers in which any third party review or directory source appeared was zero. One well known review site reached the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
  • Coverage lists beat the product. Same run. On a question about finding every case against a company, Claude ranked two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and noted that its ordering reflected price transparency and source authority rather than product quality.
  • Named, not linked. Same run. clawlaw.in pages ranked in the raw results and shaped what Claude wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
  • Used, attribution dropped, twice. Same run. Claude admitted taking two specific arguments from clawlaw.in pages and dropping the attribution in both places.
  • Official portals above every commercial product on how to questions. Same run. Claude said plainly that no commercial product should rank above the official portal for a question about using that portal.
  • Unsourced but correct figures were discounted. Same run. Two comparison pages stated competitors' prices with no link, no date and no source. The figures were correct when Claude checked them independently, and it still treated the pages as advocacy and used official sources instead.
  • An unreadable page hands the fact to somebody else. On 6 August 2026 the clawlaw.in pricing page rendered prices only after scripts ran, so a crawler received no prices, and the prices the assistants quoted came from an app store listing. Claude and ChatGPT, same date.

For an agency the second and sixth findings are the ones to act on. The vendors that won that answer had published an explicit, named, dated scope list and clear prices; the one that lost had a comparison page with correct but unsourced figures. The agency equivalent of a coverage list is a named list of services, named sectors worked in, named platforms with certification dates, and a price or a range with a date on it. It is a smaller job than most of what agencies publish instead.

Anthropic documentation

We do not restate Anthropic's documentation. It is edited, and a paraphrase with no date next to it becomes a false statement that our page cannot notice. Whether a search runs for a given question is exactly the sort of behaviour that changes between product versions, so any restatement here would be the least reliable sentence on the page.

Read these three and record the URL and the date beside whatever you take from them: Anthropic's help centre pages for the Claude apps, on web search and citations; Anthropic's developer documentation for the web search tool, if you are testing through the API rather than the app, because they are separate surfaces; and Anthropic's published model version notes, so the version you log still means something to a reader next year.

If any of those contradicts a sentence on this page, the sentence here is wrong and we want to be told.

Raw data and limitations

Seven limits.

  • No Claude agency data exists behind this page. Every Claude finding above is from six questions in Indian legal research software on two dates.
  • A cited mention rate belongs to its prompt, its city and its dates. It is not a probability and not a platform rate.
  • One observation is not a median or an average. The nearest timing result we hold is one citation of a named clawlaw.in URL by ChatGPT on 6 August 2026, fourteen days after publication. Fourteen days is a single observation, not a median time to first citation, and we will not present it as one.
  • The app and the API are different surfaces, and a fraction from one does not transfer to the other. Say which you ran.
  • Used without attribution cannot be measured from outside. Documented above with a date, and it means every cited rate is a floor rather than a full account of your influence.
  • A change after a change is not a cause. The engine changes on its own schedule.
  • No study here promises a position in an answer. We do not sell that and would not believe anybody who did.

Reproducible today without us: the five labels, the nine protocol steps, the thirteen log fields, the sample generation rules and the arithmetic. Not yet available: our own Claude agency runs, which do not exist, and the capture files behind the dated findings, which are being prepared for publication. Of the observations above, the 6 August 2026 citation of a named URL can be checked from outside today.

Version: 29 September 2026, first publication. Updated when the Claude agency set is run, when a figure is corrected, and when the capture files are published.

Common questions

How many runs per prompt is enough for an agency study?

Five to see the variance, more before you publish a rate and defend it. No count converts a small sample into a general truth. Print the count beside every fraction and let the reader judge it.

Claude gave a different set of agencies each time. Is the test broken?

No, that is the result. Instability is the thing the repeat count exists to expose, and an agency that appears in 1 of 5 runs is in a different position from one that appears in 5 of 5. A single run would have shown you one of those five answers and told you nothing about the other four.

Should I count a mention where the link goes to a listicle?

Count it, in its own column, as cited to a third party. It is real and it is not yours. Merging it with a citation to your own domain hides the difference between an asset you own and a page that can drop you tomorrow.

What is the fastest thing an agency can publish to become more citable?

A page that states what you do in named, countable terms with a date on it: services named individually, sectors you have actually delivered in, platforms and certifications with dates, and a price or a range. On 17 September 2026 Claude placed two enterprise vendors above clawlaw.in specifically because they published explicit coverage lists, and said the ordering reflected price transparency and source authority rather than product quality.

Do I need to tell readers that our comparison page is ours?

Yes, and it helps rather than hurts. On 18 August 2026 ChatGPT named clawlaw.in's comparison page in an answer while describing it as still being the vendor's own editorial page. The engine will make that observation whether or not you do, and the pages that survive it are the ones that state the relationship, date every competitor fact and link to the competitor's own page.

Can I quote another study's cited rate for agencies?

Only if it publishes the prompt list, the run count per prompt, the cities, the dates and whether the links were opened and checked. Without those, the rate is not comparable with yours and cannot be defended in a meeting.

What to do first

Choose your five most valuable buying prompts, one per city you actually serve, with your agency name in none of them. Run each five times in five separate conversations today, log the thirteen fields, label every agency mention with one of the five labels, and open every supporting link to check it mentions the agency. Then write your own scope page: services named, sectors named, platforms and certifications with dates, and a price or a range with a date. The first job gives you a baseline. The second addresses the failure Claude named out loud on 17 September 2026.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.