B2B software in AI search: a 100 query citation and recommendation benchmark, with raw results
The query set, the row schema for raw results, the four outcome levels, and a count of how many of our observations a reader can check from outside.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
B2B software does get cited and recommended by AI assistants, and the only version of that claim worth publishing is one with the raw rows attached. The number that matters is not a score, it is the count of fully documented, repeatable query and platform observations behind it. Ours is 21 dated observations across two programmes, of which 7 can be checked from outside our company today, and 1 of those 7 is a named URL cited by a named assistant on a named date. We have not run the 100 query, three platform, three repetition benchmark set out below, which would produce 900 observations. This page publishes the query set, the row schema, the scoring definitions and the exact arithmetic, so that anyone can run it and so that our own run, when it happens, can be checked line by line against this description.
The answer, first
Three things are true at once, and a benchmark that hides any of them is not worth reading.
Citation and recommendation are different outcomes and need separate columns. An assistant can recommend you without citing you, and cite you without recommending you. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, and described that as a citation lapse rather than a ranking judgement. In the same run clawlaw.in pages ranked in the raw results and shaped what it wrote, and it still landed as a name inside a list rather than as a linked recommendation. One study, two different failures, and a single visibility score would have shown neither.
Most observations are of absence, and absence is the honest baseline. Across 24 batches and 72 conversations in our own category in September 2026, the phrase recording that we were not cited appears 161 times in the assistant's own audits of its answers. In one batch of six questions, with no brand named, the assistant reported afterwards that it had not cited or recommended aiknowsus.com in any of the six.
Raw rows are the deliverable. A benchmark without published rows is a claim. With rows, a reader can recompute your numbers, find your coding mistakes, and rerun your queries. That is the difference between a page an engine will quote and a page it will discount.
How it was measured
The query set: 20 topics times 5 shapes, which is 100 queries. Choose the twenty topics from your own category and freeze them. The five shapes are fixed, so every topic gets the same five treatments and the results can be compared across topics.
- Shape one, category best. "Best [topic] software for [buyer type] in India."
- Shape two, alternatives. "Alternatives to [a named competitor]." Your own name never appears.
- Shape three, constraint. "[Topic] tool that must support [requirement], for a team of [size]."
- Shape four, task. "How do we [job to be done], and what do teams use for it."
- Shape five, evidence. "What should we ask a [topic] vendor before signing, and who publishes their pricing."
The denominator arithmetic, published on the page. 100 queries times 3 platforms is 300 query and platform pairs. Times 3 repetitions is 900 observations in a wave. Every rate on the page divides by a number a reader can rebuild from those three figures. If you run two platforms instead of three, say so, and the denominator becomes 600.
The blind rule. No brand of yours in the prompt, the wrapper, the system instruction or an earlier turn. We hold the reading that shows what happens otherwise: on 27 July 2026 the same 78 questions were run twice in one day for clawlaw.in, and the run whose wrapper named the brand had ChatGPT ranking it first on almost every question, while the blind run on the same day put it second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded.
Session hygiene, five rules. A fresh session per observation. Memory and personalisation off. No follow up questions before scoring. Region and language recorded. The same operator instructions for every run, saved in a file with a date.
Four outcome levels, scored in order. Each observation gets all four, because they are not a ladder you can infer upwards from.
- Mentioned. The product name appears in the answer.
- Recommended. The answer puts it forward as a suitable option for the buyer described, rather than listing it as an also exists.
- Cited. A URL on the product's own domain appears in the sources.
- Top source. That domain is the first source listed.
Plus one more column that is not an outcome and matters more than any of them. Accuracy: whether every factual claim the answer made about the product is correct, with the wrong claims copied out. An answer that recommends you while getting your price wrong is a problem, not a win.
The raw results row schema: sixteen fields per observation. Publish these as the raw file.
- Observation id, and the wave number.
- Timestamp in UTC.
- Platform, and the model or mode shown in the interface.
- Topic id and shape id.
- The exact prompt text, character for character.
- Repetition number, one to three.
- Region and language of the session.
- Whether a live search was visible in the interface.
- The number of sources listed.
- The ordered list of source domains.
- Products named, in the order named.
- The four outcome scores for the subject product.
- Accuracy verdict, and any wrong claim quoted.
- The full answer text, or a link to the capture file.
- Coder initials.
- Notes, including anything the assistant said about its own reasoning.
Recording the search claim as a claim. Log what the interface showed, then ask the assistant afterwards how many queries it searched for, and store that as a separate field. Do not reconcile them. In three separate batches of our September 2026 audit the assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Two disagreeing records are more informative than one tidy one.
What the numbers were
The headline count: 21 dated observations, of which 7 are checkable from outside. That is the whole of our documented evidence base across two programmes as of 29 September 2026. Every one of the 21 has a date, a named assistant, a named subject and a file it was recorded in. Seven of them a reader outside this company can verify today, because they concern pages, listings and answers that are or were public. The other fourteen are counts over our own capture files, which nobody can check until we publish the captures, and this page says so rather than presenting all 21 as equal.
The breakdown by programme, platform and shape.
- clawlaw.in, legal technology in India. A blind reading of 78 questions on ChatGPT on 27 July 2026, with the assistant confirming no live web search for any of them. Breadth order in that reading: ProVakil on 28 questions, CLAW on 21, Legistify on 18. Eighteen blind commercial questions on ChatGPT on 18 August 2026, with the company the top source on exactly one. Six questions on Claude on 17 September 2026. One named citation on ChatGPT on 6 August 2026, fourteen days after the cited page was published.
- aiknowsus.com, our own category. 24 batches and 72 conversations on Perplexity in September 2026. 161 recordings of not being cited. One batch of six questions with 0 of 6 cited or recommended.
Repetition, stated because it is the weakness. Repetition per query in everything above is one. Not three. That is exactly the limitation the protocol above is designed to remove, and until it is removed none of these counts should be read as a rate.
What has not been run, as of 29 September 2026. The 100 query set has not been run. No three platform wave has been run. No 900 observation denominator exists, and no rate is computed on this page. There is no raw results file to download yet. When there is, its row count will be 900 per wave and it will carry all sixteen fields.
What the source lists already tell us, without a benchmark. Across the whole September 2026 capture the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. And across the six questions of 17 September 2026 no third party review or directory source made it into any answer at all: one well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question. For a B2B software vendor selling in India that is the most actionable thing on this page, and it came out of source lists rather than out of scores.
What this cannot tell you
- It cannot be compared with a vendor's visibility score. Without the prompt set, the blind rule, the repetition count and the four outcome definitions, two numbers that look alike mean different things.
- It cannot establish why. On 17 September 2026 Claude ranked two enterprise vendors above clawlaw.in on a question about finding every case against a company, and said its ordering reflected price transparency and source authority rather than product quality. That is the model's account of itself, and it is a lead to test, not a cause that has been proved.
- It cannot tell you about the buyers who never asked. A benchmark measures answers to questions you chose.
- It cannot survive a changed prompt. Edit one word mid wave and the wave is no longer one wave.
- It cannot promise anything. We do not guarantee a place in any assistant's answer and neither should anybody selling you this. Disclosure: we are the vendor of AI Knows Us, and this page is the standard we want to be held to.
Sources and change log
Sources. Two bodies of work, both dated. The clawlaw.in programme from July 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the audit response files of 17 September 2026. The aiknowsus.com audit of September 2026, 24 batches and 72 conversations, captured to geo-audits/aiknowsus-com/. The 21 observation count and the 7 externally checkable count are counts over that record, not estimates. The single result any reader can check today is ChatGPT citing clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india on 6 August 2026.
Change log. 29 September 2026: first published with the query set, the sixteen field row schema, the four outcome levels, the 900 observation arithmetic, and the statement that the benchmark has not been run. On the first wave this page will carry the raw file, the denominator and the four outcome counts, with this entry left in place.
Common questions
Why 100 queries rather than 20 or 1,000?
Twenty topics is about the number a B2B category actually has once you strip the synonyms, and five shapes per topic is what it takes to see the difference between being named in a list and being cited as a source. A thousand queries mostly buys you duplicates and a coding backlog you will abandon in week three.
Do we need three platforms?
Two is the minimum that tells you anything, because the most common finding in this work is two assistants answering the same question from different sources on the same day. If you run one, your result is about that one assistant and the page should say so in the heading.
Should we publish the raw rows if they look bad?
Yes, and the bad ones are the citable part. Our own most quoted number is 161 recordings of not being cited on our own domain. Publishing that is what makes the rest of what we say about measurement believable.
How do we score "recommended" without arguing about it?
Write the rule before the run and apply it mechanically: the answer must put the product forward as suitable for the buyer described in the prompt. Being listed under "other tools in this space" is mentioned, not recommended. Then have two people score the same fifty observations and publish how often they disagreed.
What if the assistant refuses to name products?
Record the refusal as an outcome rather than as a failed run, with the wording it used. A category where assistants hedge and send buyers to official sources is a real finding about the category, and on 17 September 2026 we recorded exactly that for the questions about how to look a case up, where official court portals took every position above any commercial product.
How long does one wave take by hand?
Nine hundred observations by hand is a week of work for one person, which is why most people skip the repetitions. If a week is not available, cut topics rather than repetitions, because a stable result on ten topics is worth more than an unstable one on twenty.
What to do first
Pick your twenty topics and write the hundred prompts into a file with today's date, then run shape one and shape two on two platforms with three repetitions. That is 240 observations and about a day. Publish the rows with your denominator, including the empty ones, and you will already have a more checkable benchmark than most of what is published in this field.