How to Run AI Visibility Across Multiple Clients: A Workspace and QA Method

We do not yet track labour hours per client, so this page publishes the task list, the timing method and a plain statement of what is unmeasured.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Running AI visibility for several clients at once works when the prompt sets are frozen per client, the measurement is identical across all of them, and every answer lands in one row level log with the client in a column. The number an agency needs in order to price and staff this is the measured monthly labour hours per client for a declared workflow, with the client count, the prompt volume and the task definitions published beside it. We do not track it. As of 29 September 2026 no time log exists in our records, for any client, so this page publishes no hours figure and will not estimate one. What it publishes instead is the declared workflow in named tasks, the arithmetic that sets the volume, the timing method that would produce the hours, and the quality checks that stop a multi client programme from producing numbers that cannot be compared with each other.

We can state the volume side honestly, because volume is arithmetic rather than a measurement. With 20 clients, 40 prompts each, three services and three runs per prompt, one round is 20 times 40 times three times three, which is 7,200 recorded answers. Monthly rounds make that 7,200 answers a month. That number is the reason this work fails without a method, and it is also the reason an hours figure matters: without it, nobody can tell whether 7,200 answers a month is a two person job or a ten person one.

The answer, first

Six rules decide whether a multi client programme produces comparable numbers or a pile of unrelated spreadsheets.

  • One frozen prompt file per client, versioned and dated, never edited mid round, and handed to the client as their property.
  • One identical measurement protocol for every client. Same services, same runs per prompt, same scoring definitions, same fields. A protocol that varies by client makes cross client learning impossible and makes each client's series fragile.
  • One row level log with a client column, not one spreadsheet per client. Cross client questions become possible and per client reports remain a filter.
  • The blind rule enforced centrally, with a check that no prompt file contains its own client's brand name.
  • A comparison set per client, chosen at the start and never worked on, so a change can be attributed to something.
  • A timing log from day one, because hours per client is the number you will need to price the next ten clients and it cannot be reconstructed afterwards.

How it was measured

Volume is arithmetic. Hours are a measurement we have not taken. Here is both, separated so nobody can mistake one for the other.

The declared workflow, in eleven named tasks. This is the task list an hours figure would have to be measured against, because hours per client means nothing without it.

  • Prompt set construction, once per client, from sales calls, enquiry emails and support tickets.
  • Blind rule check, once per prompt file version.
  • Comparison set allocation, once per client, written down before any work starts.
  • Answer capture, per round, at prompts times services times runs.
  • Scoring, per answer, against the five fixed definitions.
  • Second scorer agreement check, per round, on at least twenty rows per client.
  • Source verification, per answer that names a source, opening the URL rather than trusting the claim.
  • Crawler and content checks on the client's key URLs, per round.
  • Deliverable production, per named deliverable rather than per hour of content.
  • Third party record corrections, per record raised and per record followed up.
  • Reporting, per client per round, including the rounds that went backwards.

The timing method that would produce the hours. Log start and stop times against those eleven task names, per client, in the same sheet as the work. Do it for two full rounds before you publish any figure. Then report median minutes per task per client alongside the prompt volume and the client count, and state the round in which the timing was collected. Two rounds is the minimum because a first round always includes setup that never recurs, and reporting it as a monthly figure would overstate the ongoing cost.

The measurement protocol, identical for every client. Three fresh sessions per prompt per service, named services, blind prompts, fresh sessions with no memory, and per row logging of date and time, client, service, mode, session state, location, prompt version, prompt text, run number, the five scores, whether any source was cited, claimed sources, verified sources, and the screenshot file name.

Two protocol rules that exist because of measured failures. First, the blind rule. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win, and the first run was discarded because the only thing it measured was our own prompt. Across twenty clients, one contaminated prompt file produces one flattering client report that cannot be defended. Second, the claimed versus verified split. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement when asked how many questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction.

What the numbers were

Measured monthly labour hours per client: not measured, as of 29 September 2026. No time log exists in our records. We publish no hours figure, no range and no per task estimate, and nothing on this page should be read as one.

The volume figures we can state, because they are arithmetic and we show the multiplication.

  • 7,200 answers per round at 20 clients, 40 prompts, three services, three runs. Monthly rounds make that 7,200 a month.
  • 400 second scorer rows per round at twenty rows per client across twenty clients.
  • 60 prompt files to keep frozen if each client has a treated file, a comparison file and an archive of the previous version.

And the largest real capture we hold, stated with its construction so it cannot be mistaken for a multi client benchmark: the aiknowsus.com audit of September 2026 covered 24 batches and 72 conversations on Perplexity, being three conversations per batch, on one domain in the AI visibility category. Across that capture the phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers, and in one batch of six questions the assistant reported it had cited or recommended us in 0 of 6 answers. That was one client, ourselves, and it was not timed.

Quality checks that catch the real failures

Seven checks, each of which has a named failure behind it. Run them per round, per client.

  • Brand name check on every prompt file. Any client brand appearing in its own prompt file invalidates that client's round.
  • Prompt drift check. Compare the prompt file hash or contents with the previous round. A silently edited file makes two rounds incomparable.
  • Session state check. Any row without fresh session state recorded is dropped from rates and reported as an exclusion.
  • Denominator reconciliation. Answers analysed plus answers excluded must equal prompts times services times runs. If it does not, the report is not sent.
  • Scorer agreement. Publish the agreement rate per client per round. A falling agreement rate means the definitions are drifting, not that the web changed.
  • Source verification sample. Open a sample of claimed source URLs per client per round and record how many resolved to what the answer said.
  • Advocacy check on the client's own pages. On 17 September 2026 Claude confirmed that clawlaw.in pages had ranked in its raw results and shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy. A page that only says good things about the client is a known failure mode and is worth flagging before the client pays for ten more of them.

What this cannot tell you

Six limits.

  • We publish no hours per client, because we do not track time, as of 29 September 2026.
  • Volume arithmetic is not effort. 7,200 answers a month tells you the size of the job, not how long it takes, and the difference is exactly the number we have not measured.
  • A first round is not a monthly cost. Setup does not recur, so any timing collected in round one overstates the ongoing figure.
  • Our largest capture is one domain. 24 batches and 72 conversations on aiknowsus.com is not a multi client benchmark and cannot be divided by a client count.
  • Sectors do not transfer. The observations here are legal technology in India and the AI visibility category.
  • No workflow can promise a client a position in an AI answer, and running twenty clients does not change that.

Sources and change log

The observations above come from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md and the assistant audit files of 17 September 2026, and from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. Counts were produced by a script over those files rather than from memory, and cannot yet be checked from outside, because the captures are not published. No time or labour data has been collected in either programme.

Version: 29 September 2026, first publication. The page gets updated when two full rounds of task level timing have been logged, at which point it will carry median minutes per task with the client count, the prompt volume and the round in which the timing was collected.

Common questions

How many clients can one person run?

We cannot tell you, because we have not measured hours per client, and a number invented here would be used to staff a team. What we can tell you is how to find out in two rounds: log start and stop times against the eleven named tasks, then report median minutes per task. You will have your own answer faster than you would get a trustworthy one from anybody else.

Should every client use the same prompt count?

The protocol should be identical and the prompt count can differ, as long as each client's denominator is printed with their own multiplication. What must not differ is the services, the runs per prompt, the scoring definitions and the fields logged, because those are what make one client's numbers readable next to another's.

Can we reuse a prompt set across clients in the same sector?

Partly, and carefully. Shared category prompts are useful and make cross client comparison possible. Each client still needs prompts drawn from their own sales language, and no client's prompt file may contain any brand name, theirs or a competitor's, if the results are going into a visibility rate.

What is the first thing that breaks at scale?

The denominator. Somebody adds a prompt for one client mid round, or a capture fails silently, and the report no longer reconciles to prompts times services times runs. Make the reconciliation a gate that blocks the report rather than a check somebody remembers to run.

Do we need a second scorer for every client?

For at least twenty rows per client per round, yes, and publish the agreement rate. Across twenty clients that is 400 rows, which is real work and is the cheapest defence you have against a client asking how the scoring was done. A single scorer's rate is one person's reading.

Is a tool necessary at this volume?

At 7,200 answers a month, capture by hand stops being possible, and we sell a tool so treat that as a disclosure. What you must still hold yourself, whatever tool you use, is the frozen prompt files, the row level log with the client column, the scoring definitions, and the timing log. Those four are the programme. The tool is the collection.

What to do first

Start the timing log today, even if you have three clients rather than twenty, with the eleven task names as the only categories. Two rounds from now you will have a median minutes per task figure that is yours, and you will be able to price the next ten clients from evidence rather than from a number somebody published without a method.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.