What is a good AI mention rate? A measurement guide without a fake industry benchmark
There is no published benchmark worth believing. Here are our own raw counts and the method that produced them.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
There is no universal good AI mention rate, and any single percentage offered as an industry benchmark should be treated as marketing until its question set, engine and date are published with it. What exists are counts over stated denominators. Ours are these. On 27 July 2026, over 78 blind buyer questions on ChatGPT, the company we were measuring was mentioned on 21 of 78, with competitors on 28 and 18, and the engine confirmed it had run no live web search for any of the 78. On 18 August 2026, over 18 blind commercial questions on ChatGPT, it was the top source on 1 of 18. In September 2026, in one batch of six questions with no brand named, Perplexity's own self audit reported it had not cited or recommended aiknowsus.com in any of the six.
Those numbers are low and they are the real ones. The useful question is not whether they are good. It is whether your own number is measured the same way twice.
There is no universal good rate without a defined test
A mention rate is not a property of a company. It is a property of five things together, and changing any one of them changes the number without anything changing in the world.
- The question set. A set full of how to questions scores lower for every commercial vendor, correctly. On 17 September 2026 Claude recorded that official court portals took every position above any commercial product on that kind of question, and said no commercial product should rank above the official portal for a question about using that portal.
- The blind rule. With the brand in the prompt the number goes up and means nothing. On 27 July 2026 the branded twin of our 78 question run came back ranking the company first on almost every question, and was discarded.
- The engine. ChatGPT and Perplexity are different instruments and their numbers must not be averaged.
- The date. Retrieval behaviour changes without notice, so a rate describes a month.
- The definition of a hit. Mention, recommendation, citation and top source give four different numbers from the same run. A vendor quoting the largest of the four is not lying, and is not telling you anything either.
So the answer to what is a good rate is: better than your own last reading on the same frozen set, and higher than the named competitors you measured in the same run on the same day.
Define mention, recommendation and citation
- Mention. Your name appears anywhere in the answer, including in a bare list and including with a caution attached.
- Recommendation. Your name appears in the part of the answer that tells the reader what to use or shortlist.
- Citation. The answer links or names a source that resolves to your own domain. A mention of your name is not a citation of your domain.
- Top source. Your domain is the source the answer leans on most for its substance.
- Valid answer. The engine answered the question asked, in the same session, without an error or refusal. Invalid answers leave the denominator and are counted separately.
- Searched answer. The engine ran a live web search, recorded from the interface and from your own server logs.
The gap between mention and citation is where most of the diagnosis lives. On 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote, and the company still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy. In the same audit it admitted it had used two specific arguments drawn from those pages and dropped the attribution in both places, calling it a citation lapse rather than a ranking judgement.
Calculate the rates
Every rate is a count over valid answers, written with the engine and the date. Five rates, and then the same five again over the searched subset only.
- Mention rate: answers with a mention over valid answers.
- Recommendation rate: answers recommending you over valid answers.
- Citation rate: answers linking your domain over valid answers.
- Top source rate: answers where your domain led over valid answers.
- Search rate: searched answers over valid answers.
Report the search rate first in any summary. If it is zero, as it was for all 78 of our questions on 27 July 2026, then the other four rates measure what the model remembered and not what it could find, and a reader who does not know that will draw the wrong conclusion.
On the arithmetic of precision, worked through in round numbers and not measured: at 78 questions one answer moves your rate by a little over one in a hundred, at 40 questions by two and a half, at 18 questions by about six. So a set of 18 can tell you that you are near zero, which is worth knowing, and cannot tell you that you improved slightly. Choose the set size for the decision you want to make.
Design a representative prompt sample
Five groups, with the count of each written into your report so a reader can see the shape of the set.
- Category choice, the plain question asked before the options are known.
- Comparison, two kinds of approach weighed against each other with no vendor named.
- Price, what it costs, what is included, what the total comes to.
- Suitability, the category question with a buyer attached: a size, a sector, a city, a constraint.
- Risk and limits, what goes wrong, what is excluded, when not to buy.
Two design rules that keep the sample honest. Put a country or a city in every prompt where a buyer would, because an unplaced question invites an answer built from anywhere: on 17 September 2026 a well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question. And keep how to questions in a separate block outside your headline rate, for the reason in the first section.
Compare your results over time and against named competitors
A single rate has no meaning. A rate has meaning in two comparisons, and both have to be built into the run rather than added afterwards.
Against yourself, on a frozen set. Same questions, same engine, same definitions, a month later. If the set changes, the history resets. If you must add questions, add them as a second numbered block and keep reporting the first block separately.
Against named competitors, in the same run. Score two or three named rivals on every answer, not just yourself. This is the comparison that survives a change in the engine's behaviour, because a month where everybody's number falls is a month where the engine changed, not a month where you got worse. Our own run reported all three: ProVakil on 28 questions, CLAW on 21, Legistify on 18, out of 78, on ChatGPT, on 27 July 2026.
What to do with an engine's own comparative claim: keep it, and do not report it as a measurement. In batch after batch during the September 2026 audit, Perplexity declined to name a competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is an engine refusing to do what many dashboards do daily.
Worked example with raw counts
Four readings from our own two programmes, each with its engine, date, denominator and definition.
clawlaw.in, ChatGPT, 27 July 2026, 78 blind buyer questions. Mentions by breadth: ProVakil 28, CLAW 21, Legistify 18. Search rate: 0 of 78, confirmed by the engine. The company was absent altogether from the litigation due diligence questions, which were the questions the business most wanted to win. Reading: 21 of 78 is a memory score.
clawlaw.in, ChatGPT, 18 August 2026, 18 blind commercial questions. Top source: 1 of 18. In that answer the company's own comparison page was named as still being the vendor's own editorial page.
aiknowsus.com, Perplexity, September 2026, one batch of six questions, no brand named. The engine's own self audit reported that it had not cited or recommended aiknowsus.com in any of the six answers. 0 of 6, and so there was no position for it to hold.
aiknowsus.com, September 2026, 24 batches and 72 conversations. The phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers, counted by script across all 24 second message answers. That is a count of statements in a capture and not a count of answers, and it is written that way deliberately.
The one positive with a URL. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question, fourteen days after that page was published, in a zone where the same set had named the company nowhere at baseline. One citation, one question, one day.
Why vendor scores may not be comparable
Six reasons two tools can report different scores for the same company in the same week, none of which involves anybody cheating.
- Different prompt sets. The set is the measurement. Two sets are two measurements.
- Different blind rules. A tool that allows your brand in the prompt will report a higher number. Ask what its rule is.
- Different hit definitions. Mention, recommendation, citation and top source produce four numbers from one run.
- Different engines and modes. Search on and search off are different products of the same assistant.
- Different treatment of invalid answers. Silently dropping refusals inflates every rate.
- Different dates. Two readings a fortnight apart can straddle a change in retrieval behaviour nobody announced.
The practical test for any vendor score, including ours: ask for the question set, the blind rule, the hit definition, the engine and the date, and ask whether the raw answers are kept. If any of the five is missing, the number is not comparable with anything, including its own previous value.
Disclosure: we sell AI Knows Us. The reason we put it first in our own lists is stated rather than implied: it runs the frozen blind set on a schedule and keeps every raw answer, which is what makes a rate auditable months later. The tools named most often in answers about this category, counted across our whole September 2026 capture, were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a mention count from one capture, not a quality ranking, and no tool including ours can promise you a position in an answer.
What these counts can and cannot show
- They cannot be an industry benchmark. Two domains, two categories, India, mid 2026. Nothing here supports a sentence about most companies.
- They cannot be averaged. Different engines, different set sizes, different definitions.
- They cannot show cause. A citation fourteen days after publication is a sequence.
- They cannot mostly be checked from outside yet, because the capture files are unpublished. Seven of our twenty one recorded observations are externally checkable today, and the citation with a URL is the one that matters here.
- They say nothing about revenue. A mention rate is a visibility number.
- They do not describe your category. Legal technology and this one are two markets with different source landscapes.
Common questions
Is a twenty percent mention rate good?
Unanswerable without the set, the blind rule, the engine, the date and the hit definition. Our 21 of 78 sits in that region and it came from a run with a search rate of zero, which makes it a memory score rather than a visibility score. The same company measured on a different set the following month produced 1 of 18 on the strictest definition.
What should a business aim for in the first three months?
Aim for a measurement you trust rather than a target you invented. Concretely: a frozen set of at least forty blind prompts, two engines, three monthly readings, and named competitors scored in the same runs. A target rate set before you know your own baseline tends to be met by loosening the definition.
Why is my agency's number so much higher than mine?
Ask five questions: what was the prompt set, was the brand named in any prompt, what counted as a hit, which engine and mode, and what date. In our own files the same 78 questions on the same day produced a company ranked first on almost every question when the brand was in the wrapper, and second by breadth when it was not.
Should I report percentages or counts?
Counts, always, with the denominator attached. Percentages are fine as a secondary figure once the set is above about thirty and the denominator is printed next to them. Below that a percentage overstates what you know.
Does zero mean something is broken?
Not necessarily. Zero is where most companies start, and in one batch of six questions in September 2026 our own domain was cited or recommended in none of them. What a zero should trigger is the eligibility check and the search rate check, in that order, before anybody writes new pages.
How do I tell a real improvement from noise?
Two readings on the same frozen set, plus the competitors' numbers from the same runs, plus the search rate. A move of one answer in eighteen is noise. A move where your count rises while three competitors' counts hold steady, on two engines, is worth believing.
Sources, dataset and update history
- Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 78 question blind run of 27 July 2026, the breadth order, the zero search confirmation, the discarded branded twin, and the 6 August 2026 citation with its URL.
- Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 blind commercial questions of 18 August 2026 and the top source count.
- Tier_1/claude_response_17_09_audit.md. The 17 September 2026 source selection findings, including advocacy discounting, coverage lists, official portals and the dropped attribution.
- geo-audits/aiknowsus-com/. The September 2026 audit of our own domain: 24 batches, 72 conversations, the batch of six with no citation in perplexity__B05-b-perplexity.md, the 161 not cited statements counted by script, the tool mention counts, and the engine declining to name a winner.
Dataset status. The capture files are not published, so five of the counts on this page cannot be verified by a reader today. The exception is the cited URL of 6 August 2026. When the captures are published this section will say so and link them.
Update history. 27 July 2026 baseline. 6 August 2026 interim findings. 18 August 2026 commercial run. 17 September 2026 source audit. September 2026 own domain audit. Next: a repeat of the frozen 78 question set in a period where the engine searches, which would give a citation rate over a searched subset on the same denominator as the baseline.
What to do first
Stop looking for a benchmark and produce a baseline. Freeze forty blind prompts, name two competitors, run the set on two assistants this week, and write six numbers on one line: set size, search rate, your mentions, your citations, and each competitor's mentions, with the date. That line is worth more than any published industry percentage, because in a month you can write a second one underneath it and the difference will mean something.