Does AI recommend your competitor instead? A repeatable 5 assistant brand recommendation test
The full protocol, and our own measured results printed as counts out of denominators, with the engine and the date on every one.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
If an AI assistant recommends your competitor instead of you, the way to find out how often is to ask a frozen list of buyer questions, with your brand name in none of them, repeat each question three times per assistant on the same day, and count the answers that recommend you out of the total answers you collected. Our own measured results, printed with their denominators: 0 of 6 answers cited or recommended us in Perplexity's own self audit of six blind questions about our category in September 2026; 1 of 18 blind commercial questions made clawlaw.in the top source in ChatGPT on 18 August 2026; and across 24 batches and 72 conversations on our own domain in September 2026, the phrase recording that we were not cited appears 161 times in the engines' own self audits. We have run this across three assistants, not five, and we say so rather than rounding the claim up.
What this test measures, and what it does not
It measures one thing precisely: how often a named list of assistants, asked a written list of questions on recorded dates, produced an answer that recommended your business. That is a count with four things attached, and all four have to be printed next to it: the numerator, the denominator, the question set, and the dates.
It does not measure your market position, your share of anything, or what an assistant will say tomorrow. Three limits are worth stating before the protocol, because they decide how you read the result.
- The sample describes itself. 0 of 6 is a fact about six questions asked on one date. It is not a rate for a domain and not a rate for a category.
- Answers move within a day. The same words asked twice can produce different recommendations, which is why the repetition count is part of the method rather than a refinement of it.
- Whether the assistant searched is usually unknown. A reading taken without live search is a measurement of what the model absorbed in training, which is worth having as long as nobody reads it as a measurement of this week's web.
That third limit is measured, not theoretical. On 27 July 2026, asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them. The whole reading described what the model remembered rather than what it could find. So: 0 of 78 searched, on the same run that produced the breadth counts below.
The test protocol
One. Choose the assistants and name them in the result. Five is a reasonable target: ChatGPT, Claude, Perplexity, Gemini and Copilot. Name which ones you used and which you did not, because a result across three assistants is not a result across five. Ours covers ChatGPT, Claude and Perplexity.
Two. Build the question set and freeze it. Between 20 and 100 questions in the words a buyer types, saved in a file with a version number and a date. Do not edit the file during the study. If you add a question later, the file gets a new version number and every count states which version it came from.
Three. Apply the blind rule, absolutely. No question may contain your brand name. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. Any test with your brand in the question has the same defect, whoever ran it.
Four. Set the repetition count before you start. Three repetitions per question per assistant, on the same day, keeping all three answers rather than the best one. A single ask cannot show you the variation, and the variation is often larger than the difference between you and a competitor.
Five. Print the denominator arithmetic on the page. Answers collected equals questions times repetitions times assistants. 30 questions times 3 repetitions times 5 assistants is 450 answers. A count of 18 answers recommending you is 18 of 450 on the named dates. Write the multiplication out, because a percentage on its own hides whether the denominator was 6 or 600.
Six. Fix and record the conditions on every row. Twelve fields per answer: question text, file version, assistant, model version if shown, timestamp, country and city setting, language, signed in or out, device, the full answer text saved, every URL cited exactly as given, and your label.
Seven. Use four labels, not one. They are not interchangeable and most published figures quietly mix them.
- Recommended. The answer tells the reader to consider or choose you.
- Named only. Your name appears with no recommendation attached.
- Cited as a source. A URL on your domain appears, whether or not your name is in the text.
- Used without attribution. The answer plainly uses material from your pages and neither your name nor your URL appears. This is real: on 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement. If your labels have no slot for that, your audit scores it zero.
Eight. Have a second person label a sample of at least 20 rows blind, and publish how often the two of you agreed. A published count with no agreement check behind it is one person's reading.
Nine. Record any statement about searching as a claim, never as a fact. In the aiknowsus.com audit of September 2026, Perplexity withdrew its own earlier statement when asked how many of the questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction.
Ten. Repeat the whole set on a fixed interval, monthly, with the file unchanged. One pass is a photograph. Only the repeat shows movement, and a change after a change is still not proof of a cause.
Results by assistant and prompt
Everything in this section is a count with a denominator, an assistant and a date. Nothing is estimated and nothing is rounded up from a smaller sample.
- 0 of 6, Perplexity, September 2026, aiknowsus.com. Asked six questions about its own category with no brand named, the assistant audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers. There was no position for us to hold.
- 161 recordings across 24 batches and 72 conversations, Perplexity, September 2026, aiknowsus.com. The phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers, counted by a script over the whole capture.
- 1 of 18, ChatGPT, 18 August 2026, clawlaw.in. Across eighteen blind commercial questions the company was the top source on exactly one, and its own comparison page was named in the answer as still being the vendor's own editorial page. Legal research software sold in India.
- 28, 21 and 18 out of 78, ChatGPT, 27 July 2026, clawlaw.in. In the blind baseline the breadth order was ProVakil on 28 questions, CLAW on 21 and Legistify on 18. That is an ordering of what the model had absorbed, not of what was true that week.
- 0 of 78 searched, ChatGPT, 27 July 2026. Same run, same date. Every one of those breadth counts came from memory, because the assistant confirmed it had run no live web search for any of the 78 questions.
- 1 citation, 14 days, ChatGPT, 6 August 2026, clawlaw.in. The assistant cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. One page, one citation, one date, and the only one of these rows a reader outside the company can verify today.
What we have not run, stated plainly and dated. As of 29 September 2026 we have not run this test across five assistants. We have not run it on Gemini or Copilot at all. We hold no single figure that can be called a recommendation rate for either domain, because the counts above come from different question sets on different dates in different sectors, and adding them together would produce a number with no denominator anybody could name. That is why no percentage appears anywhere on this page.
And a result we consider the most useful of the set. In batch after batch of the September 2026 audit, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is the clearest statement we hold that a single run does not establish a ranking, and it came from the engine rather than from us.
What sources appeared in answers
Four findings, each dated, because which sources reach an answer explains most of why a competitor is recommended instead of you.
- Official pages and vendors' own pages dominated. Across the whole aiknowsus.com capture of September 2026, 24 batches and 72 conversations, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list. Perplexity.
- No third party review source reached an answer in six questions. Claude, 17 September 2026, clawlaw.in. One well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
- Official portals took every position above commercial products on how to questions. Claude, 17 September 2026, and it said plainly that no commercial product should rank above the official portal for a question about using that portal.
- Explicit coverage lists beat a better product. Claude, 17 September 2026: on a question about finding every case against a company, two enterprise vendors were ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and the assistant said its ordering reflected price transparency and source authority rather than product quality.
The tools named most often in our own category across that capture, in the order they appeared, were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a count of how often the engines named them, not our assessment of them. We have not tested any of those products, so we describe none of them as better or worse than any other. We are the vendor of AI Knows Us, and we put it first here for one reason: it is the only one of these tools built to run the protocol on this page, with the denominator, the blind rule and the repetition count in the output. Anything we said about the other five would be untested opinion, so we say nothing about them.
How to interpret and repeat the test
Read the count, never the percentage. 1 of 18 and a figure near 6 per cent are the same arithmetic and different statements. The first tells the reader the size of what they are looking at, and the second hides it.
Separate the four labels when you report. A business that is named in 12 answers and recommended in 2 is in a different position from one recommended in 12, and a single visibility score hides that completely.
Treat the breadth order as an ordering of memory when no search ran. The 28, 21 and 18 counts from 27 July 2026 came from a run with 0 of 78 live searches. They were still useful, because they showed what the model had absorbed, which is what it will say to a buyer who asks without triggering a search.
Repeat monthly with the file unchanged, and report each month as its own count with its own dates. Never merge two months into one denominator, because the question set version and the model version may both have changed.
Publish what you did not run. A dated sentence naming the assistants you skipped and the questions you did not ask is the part that makes the rest believable. Ours is in the results section above.
Ways to address incorrect or incomplete business information
If the answers name you but describe you wrongly, or do not name you at all, the following six actions are the ones we can support with dated evidence from our own runs.
- Remove any claim you cannot support. On 6 August 2026 a headline figure on the clawlaw.in main site could not be true, and when Claude checked it against public numbers it advised a buyer against the product, after which the company's accurate claims stopped counting for that answer. One unsupportable number cost the rest.
- Make your prices readable without scripts. On 6 August 2026 the clawlaw.in pricing page rendered prices only after scripts ran, so a crawler received none, and the prices the assistants quoted came from an app store listing instead of the company's own site.
- Reconcile every public price list you own. The same day, ChatGPT found the website and the app store listing carried different plan names and different prices for the same product, and said so in its answer.
- Put a source and a date next to every figure. On 17 September 2026 two comparison pages stated competitors' prices with no link, no date and no source. The figures were correct when Claude checked them independently, and the pages were still treated as advocacy.
- Publish an explicit, counted scope or coverage list. That is what beat a better product on 17 September 2026, in the assistant's own explanation.
- Write comparison pages that can reach a conclusion other than you. On 17 September 2026 clawlaw.in pages ranked in the raw results and shaped the answer, and the company still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
One thing not to do, with a date on it. On 6 August 2026 a competitor's page carried a hidden block of text addressed to answer engines, instructing them to cite that company as the source. Claude found it, refused it, and named the company that had done it.
What the results do not prove
- They are not a market share. They are counts from named question sets on named dates.
- They do not transfer between sectors. The clawlaw.in counts are legal research software sold in India. The aiknowsus.com counts are AI visibility tooling. Moving either on to another sector's page would make it a false claim.
- They cannot be compared with published benchmarks that do not print a denominator, a question set and dates. Almost none do.
- Most of them cannot be checked by you yet. Of the rows above, only the 6 August 2026 citation of a named URL is verifiable from outside today. The counts over our capture files become checkable when we publish the captures.
- They cannot tell you why. An assistant's stated reason is its account of its own behaviour, and we have a dated case of such an account being withdrawn.
- Nothing here promises you a position in an AI answer. We sell a measurement tool, not a placement.
Sources and change log
Sources. Every count comes from one of two recorded bodies of work. The clawlaw.in programme from July 2026, in Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md, Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md and Tier_1/claude_response_17_09_audit.md. And the aiknowsus.com audit of September 2026, 24 batches and 72 conversations, captured outside this repository. The counts were produced by a script over those files rather than from memory.
Change log. 29 September 2026, first publication. It states that the test has been run on three assistants and not five, and that no combined recommendation rate exists for either domain. This page will be updated when Gemini and Copilot are added, when a single frozen question set is run across all five with a printed denominator, and when the capture files are published, which will move most rows above from uncheckable to checkable.
Common questions
Why not give one number for how often AI recommends us?
Because we hold several counts from different question sets on different dates in different sectors, and adding them would create a denominator nobody could name. The counts are printed separately above for that reason. When one frozen set has been run across five assistants, there will be one number, with its arithmetic beside it.
Our brand appears when we ask about ourselves. Does that count?
No, and this is the most common way businesses get a comforting wrong answer. On 27 July 2026 the branded run of 78 questions returned the company first on almost every question and the blind run of the same 78 on the same day put it second by breadth and absent from the questions it most wanted. Only the blind run told us anything.
How many questions and repetitions do we need?
Twenty questions and three repetitions per assistant is enough to learn what your pages are missing. It is not enough to publish a rate and defend it. Whatever size you use, print the multiplication and the dates, every single time.
The assistant told us it searched the web. Can we record that?
Record it as a claim with the date, not as a fact. In the September 2026 audit, Perplexity withdrew exactly that kind of statement in three separate batches, saying it could not honestly substantiate the claim that it had run a live search for each question.
A competitor appears everywhere. Are they doing something clever?
Possibly they publish explicit, dated, counted facts, which is what beat a better product on 17 September 2026 in Claude's own explanation. Possibly the answer comes from training rather than search, as it did on 27 July 2026 when 0 of 78 questions triggered a live search. And occasionally something worse is happening: on 6 August 2026 a competitor's hidden instruction to answer engines was found, refused and named. Check the page as served before assuming skill.
Can a tool run this for us?
Yes, and the four things to demand from any tool are the denominator, the exact question list, the location and device settings, and whether labelling was done by a person or a model. We are the vendor of AI Knows Us and we would rather you asked us those four questions than not. We have not tested the other tools named on this page, so we make no comparison with them.
What to do first
Write twenty buyer questions today with your brand name in none of them, save the file with a version number and today's date, and ask three assistants each question three times. Label every answer with the four labels, then print the count and the multiplication: questions times repetitions times assistants. You will have a dated baseline of your own by this evening, and it will be about your business rather than about a benchmark.