How to test whether Gemini recommends a business: a reproducible 30-day benchmark
The full benchmark design, and an honest statement of which parts we have run and which we have not.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
You test whether Gemini recommends a business by running a fixed set of blind buyer questions on a fixed schedule for 30 days, and counting recommendations as a number over the completed responses in that window. The count is the measurement. We publish the full design below so anybody can run it, and we state plainly that as of 29 September 2026 we have not completed this benchmark on Gemini, so this page carries no Gemini recommendation rate.
We could have written a number here. Every page that does, without a dated log behind it, is the reason the assistants discount vendor pages in the first place. What follows is the design, the scoring rules, and the runs we have actually finished on other assistants with their dates.
What this benchmark measures, and does not measure
It measures one thing: how often a named business appears as a recommendation in the answers to one fixed prompt set, on one surface, in one 30-day window. Everything else is outside it.
It does not measure Gemini. A benchmark of 60 questions describes those 60 questions. It does not measure quality, and a business ranked below another in an answer is not worse. It does not measure causation, so a rate that rises after you publish pages does not by itself show the pages did it. And it does not predict next month, because the model, the retrieval and the pages it reads all change without notice.
Test design
The five fixed elements
- A frozen prompt set. Written once, before the first run, and not edited during the 30 days. A prompt changed mid-window destroys the comparison.
- A fixed schedule. The same set, on the same weekday, at roughly the same hour, in a stated time zone. Weekly is the usual choice, which gives five runs in a 30-day window.
- A fixed surface, recorded. The consumer app, the web interface, and the API are different products and must be counted separately. Write down the model name the interface shows for every run.
- At least three assistants. Gemini alone tells you about Gemini. A business needs to know whether it is missing from one answer engine or from all of them, and that is a different diagnosis with different work behind it.
- A clean session per question. New chat, no memory, no custom instructions, no earlier question in the same thread. Previous turns steer the next answer, which turns a benchmark into a conversation.
The blind rule
Your business name appears nowhere in the prompt. This is not a preference. On 27 July 2026, on clawlaw.in, two runs of the same 78 questions happened on the same day. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. That was ChatGPT, and the mechanism is not specific to ChatGPT.
The four scoring definitions
- Completed response. The assistant produced a full answer to the question asked. Refusals, errors and truncated answers are logged and excluded from the denominator, and the exclusions are reported.
- Named. The business name appears in the answer text.
- Recommended. The answer presents the business as a suggested option for the asker, rather than mentioning it in passing or listing it as an example of something to avoid.
- Cited. A link to the business's own domain appears in the displayed sources.
The headline figure is recommendations over completed responses, written as a count over its denominator with the window's dates attached. Named and cited are reported alongside it and never merged into it.
What gets logged per answer
Date, time, time zone, surface, model name on screen, the question, the full answer text, every displayed source link, whether a search step appeared, and the scorer's decision on each of the four definitions with a one-line reason. Two people scoring the same log should reach the same numbers, and if they do not, the definitions are not tight enough yet.
The prompt set
Sixty questions, twelve in each of the following five groups. The proportions matter, because a set weighted towards "best" questions will make any business look either brilliant or invisible.
- Category selection. "Best contract management software for a mid-sized Indian company."
- Constrained selection. The same question with a real constraint: a budget, a city, a compliance requirement, a team size.
- Task and how-to. "How do I keep track of contract renewal dates across 200 contracts." These questions often belong to official sources rather than vendors, and it is important to know which of yours do.
- Comparison. Two named competitors compared, with your own name absent.
- Risk and verification. "What goes wrong with these tools", "how do I check a supplier is real". These decide whether you are described as trustworthy when you are described at all.
Write every question the way a buyer types it, including the awkward phrasing. A set written in marketing language measures how the assistants answer marketing language.
Results
What we have run on Gemini: nothing, as of 29 September 2026. We hold no dated Gemini capture file, so we publish no Gemini recommendation rate, no Gemini citation count and no Gemini trend. When we run this benchmark the numerator, the denominator, the window dates and the prompt set will be published in this section.
What we have completed, on other assistants, with dates:
- ChatGPT, 27 July 2026, on clawlaw.in. Asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. The order by breadth in that reading was ProVakil on 28 questions, CLAW on 21 and Legistify on 18.
- ChatGPT, 6 August 2026, on clawlaw.in. ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline.
- ChatGPT, 18 August 2026, on clawlaw.in. Across eighteen blind commercial questions the company was the top source on exactly one of them, and its own comparison page was named in the answer as still being the vendor's own editorial page.
- Perplexity, September 2026, on aiknowsus.com. Across 24 batches and 72 conversations on our own domain, the phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers. In one batch of six questions the assistant reported that it had not cited or recommended aiknowsus.com in any of the six answers.
Read those as what they are: results from three assistants that are not Gemini, on two domains, in a specific market. They are the reason we think the design above is the right one, and they are not a Gemini result.
Limitations and interpretation
Seven things this benchmark cannot do, whoever runs it.
- It cannot generalise beyond its prompt set. Change the sixty questions and you have a different benchmark with a different rate.
- It cannot prove that anything you did caused a change. Publishing, a competitor's change, a model update and a retrieval change all land in the same window. Report the sequence and resist the word "because".
- It cannot rely on the model's own account of what it did. Asked afterwards how many questions it had actually searched for, one assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Perplexity, September 2026, on aiknowsus.com.
- It cannot produce a reliable percentage from a small set. Five weekly runs of sixty questions is 300 observations at most, and a single unstable question can move the headline by a full point.
- It cannot see personalisation. Your result comes from your account, your location and your history. Somebody else's Gemini may answer differently and the benchmark has no way to know.
- It cannot be compared with somebody else's benchmark unless they publish their prompt set, their scoring definitions and their exclusions. Two rates built on different definitions are two different quantities.
- It cannot promise a position. No supplier, including us, can guarantee that an assistant will recommend a business. Anyone who does is selling something they do not control.
Reproduce or download the test
You do not need us to run this. The whole method is above, and the four things you need to assemble are the following.
- A prompt sheet with sixty rows: question, group, and the date each run was made.
- A log sheet with one row per answer, carrying the fields listed under test design.
- A scoring sheet with the four definitions written out at the top, so the person scoring in week four uses the same rules as the person scoring in week one.
- A folder of saved answers, one file per answer, with the source list intact. This is the part everybody skips and the only part that makes the numbers checkable later.
If you want the same set run across several assistants on a schedule with the sources kept, that is what our product does, and we are the vendor saying so.
Common questions
Why does this page have no Gemini number on it?
Because we have not run the benchmark on Gemini, and a rate we did not measure would be the one thing on this site that cannot be corrected later. The design, the scoring rules and the log format are the reusable parts, and they are all here. The numbers we do hold are from ChatGPT, Claude and Perplexity, and each one is dated and labelled with the assistant it came from.
Thirty days or ninety?
Thirty days gives you four or five runs, which is enough to see whether an answer is stable and not enough to see a trend. Ninety days gives you a trend and brings in more product and model changes that are not yours, which is why the change log of what happened during the window matters more the longer you run.
Can I run this on the Gemini app alone?
You can, and record it as an app result. App answers can be shaped by your account, your location and your saved instructions, so they describe your experience rather than a reproducible condition. If you want a number somebody else can repeat, run the API half as well and report the two separately.
What counts as a recommendation if the answer hedges?
Use the definition above and stick to it: the answer has to present the business as a suggested option for the asker. A business mentioned as an example of the category, or as something to be careful about, is named and not recommended. Write your decision and a one-line reason in the log, so a second person can check your judgement rather than guess at it.
Our name never appears at all. Is the benchmark still worth running?
Yes, because the rest of the log is the diagnosis. Which businesses are named, which sources are cited, and which questions produce no company names at all. In our own capture the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. Perplexity, September 2026, on aiknowsus.com. That tells you where to go and publish.
Should the prompt set change as we learn more?
Not during a window. Freeze it, finish the 30 days, then start a second benchmark with the revised set and report the two separately. A set edited mid-window destroys the only comparison the benchmark exists to make.
Google's published guidance and sources
Read the primary pages rather than commentary about them. The four that matter for this question are Google's own Gemini app help pages, Google's Search documentation on how content is used in AI experiences, Google's guidance for site owners on content and structured data, and the model documentation for the Gemini API where the grounding and search tool behaviour is described. Note the date you read each one, because these pages change and an undated quotation from them ages badly.
We are not aware of any Google document that publishes a formula for which businesses Gemini names. A separate page on this site covers that question and the search we did for such a document.
Change log. First published 29 September 2026, with the design complete and the Gemini results section empty and marked as empty. Any Gemini figures added later will carry their window dates and will be added to this page rather than replacing what is here.
What to do first
Freeze twelve questions today, one from each group, with your name in none of them. Run them on Gemini and on one other assistant, save every answer and every source list, and write the date on the folder. That is week one of the benchmark, it costs you an hour, and it is the only version of this number that belongs to you.