Which SaaS companies does Gemini recommend? A dated, repeatable citation study

The repeat run count is the method. A share of runs with no run count beside it is not a measurement.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Gemini does not have a fixed list of SaaS companies it recommends. Ask the same buying question twice and you can get two different sets of names, so the only honest output of a study like this is one line per prompt: the exact prompt text, the number of independent repeated runs behind it, and for each named product the number of those runs it appeared in, written as a fraction with the date. As of 29 September 2026 we have not run a repeated run study on Gemini, so we publish no Gemini share at all. What follows is the protocol in full, the denominator arithmetic printed so you can check ours, and the nearest measurements we do hold, which come from ChatGPT, Claude and Perplexity in two named sectors.

A share of runs is not a probability. If a product appears in 3 of 5 runs of one prompt on one day, that is a fact about five draws from one account in one place on one date. It is not a sixty per cent chance that Gemini will recommend that product to your buyer next week. Every page we have seen on this subject makes that jump, and the jump is the error.

What this study measures, and what it does not

It measures four things and nothing else.

  • Appearance. Whether a named product appears in the answer text of a given run.
  • Support. Whether that appearance came with a link the reader could open, and whose page it was: the vendor's own site, a third party listing, an official source, or nothing at all.
  • Stability. How often the same product appears across repeated runs of the identical prompt, which is the only measure of whether an answer is a habit or a coin toss.
  • Drift. Whether the same prompt set gives the same answers when you rerun it weeks later, with nothing else changed.

It does not measure the following five things, and no study run from outside Google can.

  • Why a product was named. Selection reasons are not published and an answer's own explanation of itself is not evidence.
  • What Gemini will do for a different person. Account history, language, country and the model version in use all change the answer.
  • Whether a live search ran. You can record what the interface shows and what the model claims. Neither is proof.
  • Any platform wide rate. Your prompt set is the subject of your study. The platform is not.
  • A position you can be promised. Nobody can sell you a place in an answer, including us.

Test protocol

Ten steps. None of them need a paid tool, and two people running them separately should be able to compare results.

One. Write the prompt set and freeze it. Between 10 and 40 prompts in the words a buyer actually types, taken from sales calls, from the questions your support inbox repeats, and from enquiry emails. Save the list to a file with a version number and a date. Do not edit that file for the length of the study. If you add a prompt later, the file gets a new version and every fraction you publish states which version it came from.

Two. Apply the blind rule. No prompt may contain your own brand name. This is measured, not stylistic. On 27 July 2026, in the clawlaw.in programme, two runs of the same 78 questions happened on the same day. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. ChatGPT, 27 July 2026, on clawlaw.in.

Three. Set the repetition count before you start, and never after. We use five independent runs per prompt as the floor. Five is not a magic number and we do not claim it is enough for a published rate. It is the smallest count that shows you whether a prompt is stable at all, because a product that appears in 5 of 5 and a product that appears in 1 of 5 are telling you different things and a single run cannot tell them apart.

Four. Make the runs independent. A new conversation each time, not a follow up in the same thread. Any memory or personalisation feature switched off and recorded as off. Spread the runs across the day rather than firing five in one minute, because five answers generated in one minute share more than five answers generated across eight hours.

Five. Fix the conditions and write them on every row. The exact product surface you used, the model version name as the interface reports it, signed in or signed out, the country and city setting, the language, the device, and the timestamp of each run. A study that mixes a signed in phone in Pune with a signed out desktop in Delhi cannot be compared with itself later.

Six. Save the evidence, not your conclusion. For every run save the full answer text as text, every link exactly as given, and a screenshot. A spreadsheet cell reading yes is not evidence. If a number is disputed six months later you will need the answer text.

Seven. Log these fourteen fields per run. Prompt set version, prompt id, exact prompt text, run number, timestamp, surface, model version as reported, signed in state, country, city, language, device, whether a search was shown as having run, and the screenshot filename. Then one row per product named, with the product name, whether a link supported it, and the link's domain.

Eight. Record any search count as a claim, never as a fact. Where the interface tells you what it consulted, save that. Where you have to ask the assistant, mark the answer unverified. In the aiknowsus.com audit of September 2026, Perplexity withdrew its own earlier statement when asked how many of the questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction.

Nine. Have a second person label a sample blind. At least 20 rows, without seeing the first labels, and publish how often the two of you agreed. Twenty minutes of work, and it is the difference between a number and one person's reading.

Ten. Rerun the frozen set on a fixed interval. Same day of the month, nothing else changed. One pass is a photograph. Only the rerun tells you whether anything moved.

The denominator arithmetic, printed

This is arithmetic rather than an observation, and it is worked through here so that you can check the shape of any figure we ever publish. Twelve prompts at five runs each is 60 runs. If a product appears in 21 of those 60 runs, the study wide line is 21 of 60 on the stated dates, and that line is close to useless on its own. The useful part is the per prompt breakdown: the product might be 5 of 5 on two prompts, 3 of 5 on one, and 0 of 5 on the remaining nine. Those three numbers describe a product that owns two questions and is invisible on nine, which is a completely different situation from 21 appearances spread thinly everywhere.

So the reporting rule is: per prompt first, study wide second, and never a percentage without the fraction next to it. If you also drop runs because of an error or a refusal, the dropped count gets its own line and the denominator shrinks visibly rather than quietly.

SaaS categories and named test cases

We will not print a list of SaaS products we have not tested. A named sample we had not run would be an invented finding wearing a table, and this page exists because an engine asked us for the opposite.

What we can name is the two sectors where we hold real measurements, with the products that actually appeared in those runs.

  • Indian legal research software. In the blind baseline of 27 July 2026, ChatGPT's breadth order across 78 questions was ProVakil on 28 questions, CLAW on 21 and Legistify on 18. Those three names came out of the runs, not out of a marketing list.
  • AI visibility and generative engine optimisation tooling. Counted across the whole aiknowsus.com capture of September 2026, the tools named most often in our own category were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. Perplexity, 24 batches and 72 conversations.

For your own category, build the sample the same way rather than from a directory. Run your frozen prompt set once, write down every product that appears anywhere in the answers, and that list becomes your tested sample. Adding a product because you think it is a competitor tells you nothing, because the question is what the assistant names, not what you would name.

Six categories worth testing separately, because buyers ask about them in different words: accounting and GST filing software, payroll and HR software, customer support and helpdesk software, contract lifecycle management, legal research, and field service or logistics software. Keep each category's prompt set in its own frozen file. A mixed file gives you a number that belongs to no category.

Results

The number of independent repeated runs we have completed on Gemini as of 29 September 2026 is zero. So there is no Gemini share of runs on this page, no estimate, no range, and no figure borrowed from anybody else's study. The cell is empty on purpose and it stays empty until the protocol above has been run and the run logs published.

Here is what we do hold, with denominators, engines and dates. Read each one as belonging to the sector named beside it.

  • 28, 21 and 18 out of 78. Breadth order in the blind baseline of 27 July 2026, ChatGPT, Indian legal research software: ProVakil on 28 of 78 questions, CLAW on 21, Legistify on 18.
  • 0 of 78 searched. Same run, same date. ChatGPT then confirmed it had run no live web search for any of the 78 questions, so the whole reading described what the model remembered rather than what it could find.
  • 1 of 18. Across eighteen blind commercial questions on 18 August 2026, ChatGPT made clawlaw.in the top source on exactly one, and named the company's own comparison page in the answer as still being the vendor's own editorial page.
  • 0 of 6. Asked six questions about its own category with no brand named, Perplexity audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers. September 2026, AI visibility tooling.
  • 161 across 24 batches and 72 conversations. In the aiknowsus.com capture of September 2026, the phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers. Perplexity.

One distinction matters more than any of those figures. The 78 question numbers are breadth, not repetition. They count how many different questions a product appeared on, once each. They do not count how often the same question gives the same answer. So they cannot stand in for the repeat run share this page is about, and we are not offering them as a substitute. We hold no repeat run share for any engine.

There is a related finding worth quoting, because it came from an engine rather than from us. In batch after batch of the September 2026 capture, Perplexity declined to name a competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is the clearest statement we have that a single run does not establish a ranking, and it came from the machine the ranking would have been about.

What the official documentation says

We deliberately do not restate Google's documentation here. Help pages are edited, an undated paraphrase becomes a false claim without the page noticing, and we have watched an assistant discount correct information for exactly that reason: on 17 September 2026 Claude checked two clawlaw.in comparison pages that stated competitors' prices with no link, no date and no source, found the figures were correct when it checked them independently, and still treated the pages as advocacy and used official sources instead.

So read these three yourself, on the day you read this, and write the date and the URL into your notes beside whatever you take from them.

  • Google's Gemini Apps help pages for what the consumer product does, including anything it says about links and about how answers are produced.
  • Google's developer documentation for the Gemini models if you are testing through the API rather than the app, because the surfaces behave differently and a result from one is not a result about the other.
  • Google Search Central documentation for anything about how your pages are crawled, indexed and used, which is a separate system from the assistant.

If what you find there contradicts a sentence on this page, that sentence is wrong and we want to know. Tell us and we will change it and note the change in the log.

Limitations, raw data, and how to reproduce the study

Six limits, each one of which has cost somebody a wrong conclusion.

  • No Gemini data exists behind this page. The findings above are ChatGPT, Claude and Perplexity. The systems select sources differently and a result from one is not evidence about another.
  • A share of runs is a sample of that prompt, on that day, from that account. It is not a probability and it is not a platform rate.
  • Five runs is a floor, not a sufficient sample. It shows you variance. It does not let you publish a defended rate.
  • Counts over our own capture files cannot be checked by you yet. Of the observations behind this page, the citation of a named URL by ChatGPT on 6 August 2026 can be verified from outside. The counts over our capture folders cannot, until we publish the captures.
  • A change after a change is not a cause. If a product starts appearing after you publish pages, the pages are one candidate explanation among several, including the engine changing its own behaviour that month.
  • Nothing here can promise you a position in an answer. We do not sell that and we would not believe anybody who did.

The reproducible part is the whole protocol above plus the fourteen log fields, which you can copy and run this week without us. What is not yet downloadable is our own run data, because the Gemini runs do not exist and the capture files behind the ChatGPT, Claude and Perplexity figures are still being prepared for publication.

Version of this page: 29 September 2026, first publication. It will be updated when the first Gemini repeat run set is completed, in which case the empty cell gets a numerator, a denominator, a prompt set version and dates; when any figure above is corrected; and when the capture files are published, which will move several rows from unverifiable to checkable.

Common questions

How many runs per prompt do I actually need?

Five to see whether the prompt is stable, and more than five before you publish a rate and defend it in a meeting. There is no count that turns a small sample into a general truth. What grows with the count is how little the fraction moves when you rerun it, which is the only thing you can actually observe. Whatever count you use, print it next to every share.

Why not publish an estimate with a caveat underneath?

Because the number gets quoted and the caveat does not. A figure lifted out of a page travels without its footnote. And an unsourced number does not gain trust from being hedged: the two clawlaw.in comparison pages Claude checked on 17 September 2026 had correct figures and were discounted anyway, because at read time there was no way to tell.

Does it matter if Gemini did not search the web for my question?

It changes what your result means. An answer produced without a live search describes what the model absorbed during training, which is a real and useful thing to measure, as long as nobody reads it as a measurement of the web this week. On 27 July 2026 ChatGPT confirmed it had run no live search for any of 78 questions, and every breadth count from that run has to be read that way.

Can I use one engine's result to predict another's?

No, and this is the most common mistake in published studies. The four systems people test most select their sources differently and change on their own schedules. Run each one separately or say which one you ran.

Our product name is also a common word. How do we count appearances?

Write the matching rule into your labelling sheet before the first run, with an example of a match and an example of a non match, and have the second labeller apply the same rule. Then publish the rule. Most arguments about a visibility number are really arguments about string matching.

Is a mention without a link worth anything?

It has value and it is hard to bank. The reader hears the name and can search for it later. It sends no visit and will not appear in your analytics, so a business measuring only referrals concludes nothing happened. Count it as its own state and report it separately from a supported mention.

What to do first

Write ten prompts in your buyers' words with your brand name in none of them, and save the file with today's date. Run each one five times in five separate conversations, spread across the day, logging the fourteen fields above. Then write the per prompt fractions out by hand. By this evening you will have a dated baseline for your own category, which is worth more than any published percentage, because it is about your product and you know exactly how it was made.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.