How to measure whether ChatGPT mentions your company: a repeatable prompt test protocol
Three definitions, a frozen prompt set, a log sheet, and rates reported as counts over a denominator.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
You measure it by writing a fixed set of buyer questions that never name your company, running each one in a fresh conversation, logging every answer, and counting three things separately: mentions, recommendations and citations of your own domain. On our own frozen set of 78 blind buyer questions, run on ChatGPT on 27 July 2026, the company being measured was mentioned in 21 of 78 answers, and the engine confirmed it had run no live web search for any of the 78. On a separate set of 18 blind commercial questions on 18 August 2026, the same domain was the top source in exactly 1 of 18.
Everything else on this page is how to get a number like that for yourself, and how to avoid the one mistake that makes the number meaningless.
Define mention, recommendation and citation
Most measurement arguments are really definition arguments. Fix these four before your first run and write them at the top of your log sheet.
- Mention. Your company name appears anywhere in the answer text, including in a list of things that exist and including with a caution attached.
- Recommendation. Your company appears in the part of the answer that tells the reader what to use, who to contact or what to shortlist. A name in a background paragraph is a mention, not a recommendation.
- Citation. The answer carries a link or a named source that resolves to your own domain. A mention of your name is not a citation of your domain, and conflating the two is the most common scoring error we see.
- Top source. Your own domain is the source the answer leans on most for its substance. This is the strictest outcome and the one worth tracking separately, because it is the difference between being used and being listed.
Two supporting definitions decide your denominator. A valid answer is one where the engine answered the question you asked, in the same session, without an error and without a refusal; invalid answers leave the denominator and get reported as a count. A searched answer is one where the engine ran a live web search, recorded from the interface indicator and, separately, from your own server logs.
Why the last one is not optional: on 27 July 2026 ChatGPT answered 78 buyer questions and then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. A mention rate with no search rate next to it cannot be interpreted.
Build a representative prompt set
Aim for 40 to 100 questions. Below 40 you have observations rather than a rate. Above 100 the run becomes hard to repeat by hand, which matters more than the extra precision.
The blind rule, first. No prompt may contain your company name, your product name, your domain, or a phrase that only appears on your own website. Here is why, from our own record. On 27 July 2026 two runs of the same 78 questions happened on the same day. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt.
Then fill five groups and record how many prompts are in each.
- Category choice. The plain question a buyer asks when they do not yet know the options.
- Comparison. Two kinds of thing weighed against each other, without naming a vendor.
- Price. What it costs, what is included, and what the total comes to.
- Suitability. The same category question with a buyer attached: a size, a sector, a city, a constraint.
- Risk and limits. What goes wrong, what is not covered, when not to buy.
Two things to leave out, on purpose. Keep how to questions in a separate block rather than in your headline set, because official sources usually win them and they will drag your rate down for a reason you cannot fix. On 17 September 2026 Claude recorded that on questions about how to look a case up, official court portals took every position above any commercial product, and said plainly that no commercial product should rank above the official portal for a question about using that portal. And put a country or a city in every prompt where a buyer would. On 17 September 2026 a well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
Run and log the tests
Seven rules that make a run repeatable.
- One prompt per fresh conversation. A follow up inherits the earlier context and is no longer an independent observation.
- Record the product and the mode. Which assistant, which model if it is shown, and whether search was on.
- Record the location setting, because these answers move with it.
- Save the full answer text as a file, one file per prompt, named with the prompt number and the date. A rate with no answers behind it cannot be audited later, including by you.
- Log the citations as domains, not as titles, so you can count them.
- Ask for a self audit in a second message once the run is done: how many questions it searched for, whether it cited your domain, which company it would put first. Keep the reply, including when it contradicts the run. In September 2026, asked exactly this, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. It happened in three separate batches.
- Run on at least two assistants. One assistant can be wrong about itself, and two stop you redesigning a website around one product's behaviour in one month.
The log sheet needs eleven columns: prompt number, prompt group, prompt text, assistant, date, valid, searched, mention, recommendation, citation of your domain, and every domain cited. Add a twelfth free text column for the reason the answer gave, because that is where the work instructions come from.
Calculate mention, recommendation and citation rates
Each rate is a count over valid answers, written with the engine and the date attached. Never publish a bare percentage, because a bare percentage hides the denominator and the denominator is the whole argument.
- Mention rate. Answers with a mention, over valid answers. Ours was 21 of 78 on ChatGPT on 27 July 2026.
- Recommendation rate. Answers where you were in the part that tells the reader what to use, over valid answers.
- Citation rate. Answers linking your domain, over valid answers.
- Top source rate. Ours was 1 of 18 on ChatGPT on 18 August 2026.
- Search rate. Searched answers over valid answers. Ours was 0 of 78 on 27 July 2026, by the engine's own confirmation.
Then compute each rate twice: once over all valid answers, and once over the searched subset only. The second is the one that describes the current web. If your search rate is near zero, say so at the top of your report, because every other number in it is a memory score.
On the arithmetic of small sets, worked through in round numbers rather than measured: at 18 questions one answer is worth about six in every hundred of your rate, so a move from 1 of 18 to 2 of 18 is not a trend. At 78 questions one answer is worth a little over one in a hundred, which is why we treat the 78 question set as our reporting set and the 18 question set as a probe.
Worked example using a dated test
The clawlaw.in programme, ChatGPT, 27 July 2026, 78 blind buyer questions, one question per conversation, no brand named anywhere in the prompts. Counted by breadth of mention, the order was ProVakil on 28 questions, CLAW on 21 and Legistify on 18. The engine then confirmed, when asked, that it had run no live web search for any of the 78.
Read that log sheet properly and it says three things. First, the mention rate is 21 of 78 and it is a memory score, not a visibility score, because the search rate was zero. Second, the competitor was ahead by seven questions on a measure that no page published that month could have influenced. Third, the branded twin of the same run, on the same day, with the same 78 questions, put the company first on almost every question, and was discarded for that reason.
Three weeks later the picture was narrower and sharper. On 18 August 2026, across 18 blind commercial questions, ChatGPT made the company the top source on exactly 1 of 18, and in that answer named its own comparison page as still being the vendor's own editorial page. Two runs, two denominators, two dates, and no averaging between them.
The one result in this programme with a URL behind it: on 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. That is a single citation on a single question, and we report it as one, not as a rate.
Tool options, and why we do not print their prices here
You can run this by hand with a spreadsheet, and for a first read you should, because doing one run by hand teaches you what the definitions above actually cost to apply. After that, the part people abandon by week three is the logging, and that is what a tool is for.
The tools named most often in our own category, counted across the whole September 2026 capture of 24 batches, were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a count of how often an assistant named them in answers about this category, and nothing more. It is not a quality ranking and it is not a shortlist we are endorsing.
Disclosure: we sell one of these, AI Knows Us. We put it first in our own lists for one reason we can state plainly: it runs the blind prompt set on a schedule and keeps every raw answer, which is the part that makes a rate auditable later. Every other tool named above is a real product and should be judged on its own pages.
On pricing: we do not restate other vendors' prices on this page. The reason is the same rule we apply to ourselves. On 17 September 2026 Claude found two clawlaw.in comparison pages stating competitors' prices with no link, no date and no source; the figures turned out to be correct when it checked them independently, but it had no way to know that at read time, so it treated the pages as advocacy and used official sources instead. An undated price copied onto a third party page is the thing that gets a page discounted. Read each vendor's own pricing page on the day you are deciding, and note the date you read it.
Limits and interpretation
- A rate describes a set, a date and an engine. Change any of the three and you have a different number, not a trend.
- Zero search makes every other number a memory score. Our 21 of 78 is the clearest example we own.
- An engine's account of itself is evidence, not proof. Perplexity withdrew its own search claim in three separate batches in September 2026.
- Small sets cannot show small movements. One answer in 18 is about six in every hundred.
- Counts do not prove cause. A page published fourteen days before a citation is a sequence, not a mechanism.
- Nothing here measures revenue. Being named is a visibility outcome and this method says nothing about what it sold.
- No method and no product, ours included, can promise you a place in an answer.
Common questions
Can I just ask ChatGPT whether it mentions my company?
You can, and treat the answer as one more observation rather than as a report. Our September 2026 capture contains an engine withdrawing its own claim about its own behaviour in three separate batches. A self audit is most useful when it disagrees with your log sheet, because then you know which fields to check.
How many prompts is enough?
Forty to a hundred for a reporting set. Under about thirty, publish counts and skip percentages. We publish 1 of 18 and 0 of 6 exactly that way.
My name appears but there is no link. Does that count?
It counts as a mention and not as a citation, and keeping them apart is the point. On 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote while the company still landed as a name inside a list rather than as a linked recommendation, and in the same audit that it had used two of the site's arguments and dropped the attribution in both places.
Should I include questions where a government site is the right answer?
Keep them, in a separate block, and do not count them in your headline rate. On 17 September 2026 official court portals took every position above any commercial product on that kind of question, and Claude said no commercial product should rank above the official portal for a question about using that portal. Those prompts are useful for finding adjacent questions to write for, not for scoring yourself.
How often should I re run the set?
Monthly, with the set frozen. If you must add prompts, add them as a second numbered block and keep reporting the first block separately, or you lose your history.
What is the first fix when the rate is low and the search rate is high?
Open the pages that were cited instead of yours and find the checkable thing each one publishes. In our own programme that thing was twice the same: a named coverage list and a price with a date. On 17 September 2026 two enterprise vendors ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and Claude noted that its ordering reflected price transparency and source authority rather than product quality.
Template and sources
The log sheet template, twelve columns: prompt number, group, prompt text, assistant, date, valid, searched, mention, recommendation, citation of own domain, all domains cited, reason given. One row per answer. One file per answer, named with the prompt number and the date, stored alongside.
The report template, five lines: the set size and the date range, the search rate as a count over valid answers, the mention, recommendation and citation rates each as a count over valid answers, the same three over the searched subset, and a named list of the domains cited most often instead of yours.
Sources. Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md holds the 78 question run of 27 July 2026, the zero search confirmation, the breadth order, the discarded branded twin and the 6 August 2026 citation. Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md holds the 18 question run of 18 August 2026. Tier_1/claude_response_17_09_audit.md holds the 17 September 2026 source selection findings. geo-audits/aiknowsus-com/ holds the September 2026 audit of our own domain, 24 batches and 72 conversations, including the tool mention counts and the three batches where the engine withdrew its search claim.
What is not in this record. We do not have a run where the engine searched for every question, so we cannot yet publish a citation rate over a searched subset of a full set. That run is the next one, and this page will carry its counts and its date whatever they are.
What to do first
Write 40 blind prompts today, put them in a numbered sheet, and run the first ten on two assistants in fresh conversations. Before you score anything, fill in the searched column. If it is mostly empty, your report is a memory score and your next month goes into consistency and third party presence. If it is mostly full, score the citations and read the pages that were cited instead of yours.