Does AI search recommend SaaS products? A reproducible 90 day benchmark across ChatGPT, Gemini, Perplexity and Google AI Overviews

The protocol in full, the observation counts we actually hold with their dates, and a plain statement of the parts we have not run.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Yes. Assistants name SaaS products by name, and they do it from a small set of sources, so whether they name yours is a countable fact rather than an opinion. The number that answers it is not a visibility percentage. It is a count of answer observations, reported with the platform, the query and the repetition beside it. The counts we hold are these: in our own audit of aiknowsus.com in September 2026 we recorded 24 batches and 72 conversations, in which the phrase recording that we were not cited appears 161 times in the assistant's own audits of its answers, and in one batch of six questions about our own category the assistant reported afterwards that it had not cited or recommended us in any of the six. On clawlaw.in we hold three dated readings of 78, 18 and 6 questions. We have not yet completed the full 90 day, four platform, three repetition benchmark set out below, and this page publishes the protocol so that you can run it and so that anybody can check ours when we do.

The answer, first

Three findings, stated as plainly as we can put them.

Assistants do recommend named SaaS products. In our own category the tools named most often across the whole September 2026 capture were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a count over our own capture files, and it is a count of mentions, not of quality.

Not being named is the normal state, and it has a number. 161 recordings of not being cited, across 24 batches and 72 conversations, on our own domain, in our own category, in September 2026. Nobody enjoys publishing that. It is the single most useful number on this page, because it is the baseline almost every SaaS company is actually starting from.

A single run is not a measurement. In batch after batch of that same audit the assistant declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. When the machine you are measuring tells you your sample is too small, that is worth listening to, and it is why the protocol below repeats everything.

How it was measured

This is the design. Anything we have not run is marked in the next section, not hidden here.

The four platforms. ChatGPT, Gemini, Perplexity and Google AI Overviews. They are named because they behave differently: three answer in a chat and one answers above a page of results, and Google AI Overviews cannot be prompted in a conversation at all, so it needs a query typed into search rather than a prompt.

The prompt set: 40 questions in five shapes, eight of each. Fixed wording, written once and never edited mid study, because editing a prompt resets the study.

  • Category best. "Best [category] software for [buyer type] in India."
  • Named alternative. "Alternatives to [a competitor, never us]."
  • Constraint led. "[Category] tool that supports [specific requirement] under [budget]."
  • Task led. "How do I [job the software does], and what do people use for it."
  • Comparison. "[Competitor A] compared with [Competitor B], which suits [buyer type]."

The blind rule, which is the rule the whole study stands on. Your own brand name never appears in any prompt, in any wrapper, in any system instruction, or in any earlier turn of the same session. We have the two readings side by side that prove why. On 27 July 2026 the same 78 questions were run twice in one day for clawlaw.in. The first run used a wrapper that named the brand, and ChatGPT came back ranking it first on almost every question. The second named nothing, and it put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt.

Repetition: three runs per question per platform per wave, and seven waves across 90 days. Fresh session each time, memory and personalisation off, region recorded, no earlier turn in the session. Three runs is not statistics. It is the minimum that tells you whether an answer is stable, which is a different question from whether it is favourable.

Denominator arithmetic. 40 questions times 4 platforms times 3 runs is 480 observations in one wave, and 3,360 across seven waves. Publish that arithmetic on your own page. A reader who cannot see how the denominator was built cannot check any rate you compute from it.

Six scoring definitions, decided before the first run. Each one is scored per observation, yes or no, by a person reading the answer.

  • Named. Your product appears anywhere in the answer text.
  • In the first five. Named within the first five products the answer lists. This is the one that behaves like a ranking.
  • Cited. A URL on your own domain appears in the answer's sources.
  • Top source. Your domain is the first source listed.
  • Described accurately. Every factual claim the answer makes about you is correct. Score this separately, because being named wrongly is worse than not being named.
  • Search mode. Whether the answer shows evidence of a live fetch, and how many sources it lists.

Recording the search count, and why it needs a second question. Ask the assistant afterwards, in the same session, how many of the questions it searched for, and log the answer as a claim rather than as a fact. On 27 July 2026, asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. And in the September 2026 aiknowsus.com audit the assistant withdrew its own earlier statement in three separate batches, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. So record the claim, and record the visible source list beside it, and treat disagreement between the two as data.

What the numbers were

Here is every count we hold, with the platform, the date and the repetition beside it. Read the repetition column first, because it is where every one of these is weak.

  • 72 conversations across 24 batches, on aiknowsus.com, our own domain and our own category, September 2026, on Perplexity. Repetition per question: one. This is the observation count for that audit.
  • 161 recordings of not being cited, counted across all 24 of those batches in the assistant's own audits of its answers.
  • 0 of 6. One batch of six questions about our own category, with no brand named. The assistant audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers, so there was no position for it to hold. Perplexity, September 2026.
  • 78 questions in one reading, clawlaw.in, ChatGPT, 27 July 2026, with no live search behind any of them by the assistant's own confirmation. Breadth order in that reading: ProVakil on 28 questions, CLAW on 21, Legistify on 18. That is a ranking of what the model had absorbed, not of what was true that week.
  • 1 of 18. Eighteen blind commercial questions, ChatGPT, 18 August 2026, in which clawlaw.in was the top source on exactly one. Its own comparison page was named in the answer as still being the vendor's own editorial page.
  • 6 questions, Claude, 17 September 2026, in which no third party review or directory source made it into any answer at all. One well known review site did appear in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
  • One dated citation with a URL behind it. On 6 August 2026, fourteen days after the page was published, ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, in a zone where the same question set had named the company nowhere at baseline.

What is not run, as of 29 September 2026. We have not run the 40 question set across all four named platforms. We have not run three repetitions per question. We have not run seven waves across 90 days. So there is no 3,360 observation denominator on this page and no rate computed from one. Every count above comes from single runs on one or two platforms, which is why the protocol exists in this much detail: the next version of this page should carry a real denominator, and it should be checkable against this one.

Which sources decided these answers. Across the whole September 2026 capture the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. For a SaaS company that is the practical finding: the pages that decide these answers are mostly official pages and vendors' own pages, so your own site is not a weak source in this channel, provided what it says can be checked.

What this cannot tell you

Six limits, and the third is the one people most often miss.

  • It cannot produce a market share. A count of mentions in 72 conversations is a count of mentions in 72 conversations.
  • It cannot separate cause from coincidence. A page published on Monday and cited on Friday is a sequence, not a proof. The fourteen day citation above is the strongest single result we hold and it is still one page and one question.
  • It cannot be compared with anybody else's percentage. Different prompt sets, blind rules and scoring definitions produce numbers that look alike and mean different things. A visibility score with no prompt set behind it is not a measurement.
  • It cannot see personalisation. Your customers' sessions carry history, location and account context. Yours does not, by design, and that is a limit rather than a fix.
  • It cannot tell you the model's reason. Assistants explain themselves plausibly and inconsistently. On 17 September 2026 Claude said its ordering on one question reflected price transparency and source authority rather than product quality, which is useful and is still the model's account of itself.
  • It cannot promise a position. Nobody can guarantee a place in any assistant's answer, and we do not. Disclosure: we are the vendor of AI Knows Us, which sells this measurement, so treat our numbers exactly as strictly as we are asking you to treat everybody else's.

Sources and change log

Sources. The clawlaw.in programme, run from July 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the two audit response files of 17 September 2026. The aiknowsus.com audit of September 2026, 24 batches and 72 conversations, captured to geo-audits/aiknowsus-com/. Every count on this page was produced by a script over those files rather than from memory. One result on this page can be checked from outside today: the citation of clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india by ChatGPT on 6 August 2026.

Change log. 29 September 2026: first published, carrying the full protocol, the seven counts we hold with their dates and repetitions, and the statement that the 90 day four platform benchmark has not been run. When wave one runs, this page will carry the 480 observation denominator, the six scores against it, and the date, and this entry will stay where it is so the two can be compared.

Common questions

Why not just report a visibility percentage?

Because a percentage hides the three things that decide whether it means anything: how many questions, on how many platforms, repeated how many times. 0 of 6 on one platform in September 2026 tells you what we measured and how weak it is. "Visibility improved 40 percent" tells you nothing you can check, and the engines discount it for the same reason you should.

Ninety days is a long time. Can we get a reading in a week?

You can get wave one in a week, and that is worth having as a baseline. What you cannot get in a week is drift, which is the point of the seven waves. Answers change without your site changing, and a single wave cannot tell the two apart.

Can we include our brand name in the prompts to save effort?

No, and this is the one rule with no flexibility. On 27 July 2026 the branded wrapper made ChatGPT rank clawlaw.in first on almost every question, and the blind run on the same day put it second by breadth and absent from the questions that mattered most. A branded prompt measures your prompt.

Which platform should we start with if we can only afford one?

Start with the one your buyers actually use, which you can find by asking ten of them rather than by guessing. If you have no idea, start with two rather than one, because the most common surprise in this work is two assistants answering the same question from different sources on the same day.

How do we know the assistant really searched?

You do not, and you should record that uncertainty rather than resolve it. Log the visible source list, then ask it afterwards how many questions it searched for, and keep both. In three separate batches of our September 2026 audit the assistant withdrew its own claim to have searched, which is the clearest reason to treat its account as a claim.

Is 40 questions enough?

It is enough to be useful and not enough to be precise. Forty questions times four platforms times three runs gives a wave you can act on. If your category has several distinct buyer types, add eight questions per buyer type rather than stretching the same forty, and keep the sets separate in the scoring.

What to do first

Write the 40 prompts today and freeze the wording in a file with a date on it. Then run wave one on two platforms with three repetitions, score the six definitions by hand, and publish your denominator alongside your counts. Freezing the wording before you see any result is the cheapest honesty available in this work, and it is the part almost everybody skips.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.