How to test whether ChatGPT recommends your business: a 20 prompt benchmark
The 20 prompts written out, how to run them, and the ChatGPT results we have measured with their dates.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Run 20 prompts, none of them containing your brand name, three times each in fresh chats, and count three things separately: how many of the 20 named your business, how many linked a page on your own domain, and how many made your page the top source. Report each as a fraction over 20, with the date and the number of repeat runs beside it. The measured ChatGPT results we hold are these: on 18 August 2026, across eighteen blind commercial questions, clawlaw.in was the top source on exactly 1 of 18, and on 27 July 2026, across 78 blind questions, ChatGPT named the company on 21 of 78 while confirming it had run no live web search for any of them. We have not yet run the 20 prompt version on our own domain, and we are not going to publish a number for a run we have not done.
The reason those two figures are quoted instead of a tidier one is that they are the real ones. Both come from a recorded programme with files behind them. The 1 of 18 is the number you should compare yourself against if you are a business that has published pages and is wondering whether they are working, because that company had published pages.
What this benchmark measures
It measures whether ChatGPT reaches for your business when a buyer asks a question in your category without knowing your name. It does not measure what ChatGPT says about you when asked directly, which is a different test with a different purpose.
Four outcomes get counted separately, because they are worth different amounts to you.
- Named. Your business name appears in the answer. The buyer now knows you exist.
- Linked. A URL on your own domain appears as a source. The buyer can reach you and you can see the visit.
- Top source. Your page is the first or most heavily used source. This is the strongest position available and it is rare.
- Used without credit. The answer clearly contains your material and neither your name nor your URL appears. It happens: on 17 September 2026, Claude admitted it had used two specific arguments drawn from clawlaw.in pages and had dropped the attribution in both places, calling it a citation lapse rather than a ranking judgement.
Keep the fourth column even though it is subjective, and label it only where the wording is close enough that you would recognise it as yours. It stops you concluding that pages did nothing when they were doing quite a lot.
Define your business, market, and customer
Write these five lines down before you write a single prompt. Every prompt is generated from them, and a vague version produces prompts nobody types.
- The category word your buyers use, not your internal one. If they say tally operator and you say accounting services, the prompt uses their word.
- The geography that matters, at the level a buyer would state it. A city, a state, or India.
- The buyer's role and the decision they are making, for example an HR manager choosing between an in house payroll person and a bureau.
- The three competitors a buyer would compare you with, named. You will use their names in the prompts and never your own.
- The commercial constraint, the thing that actually decides the purchase: budget, turnaround, compliance, integration, or location.
The 20 prompt test set
Here are the 20, written as templates with square brackets to fill from the five lines above. They are split into five groups of four, because the groups behave differently and a set made only of the first group flatters everybody.
Group A, category questions. The buyer knows what they want to buy.
- 1. Best [category] in [city] for a [buyer role], and why?
- 2. Who are the leading [category] providers in India in 2026?
- 3. What should I look for when choosing a [category] provider?
- 4. Recommend three [category] options for a [company size] business in [city].
Group B, comparison questions. Name competitors, never yourself.
- 5. [Competitor 1] vs [Competitor 2]: which is better for a [buyer role]?
- 6. What are the alternatives to [Competitor 1] in India?
- 7. Is [Competitor 2] worth the price for [use case]?
- 8. Who competes with [Competitor 1] and [Competitor 3] in [category]?
Group C, problem questions. The buyer does not yet know your category exists. These are the ones most businesses never test and most often win.
- 9. How do I [the job your product does] without hiring somebody full time?
- 10. My [specific problem] keeps happening. What are my options?
- 11. Is it cheaper to do [the job] in house or outsource it in India?
- 12. What goes wrong when businesses try to [the job] themselves?
Group D, commercial questions. Money, terms and total cost.
- 13. What does [category] cost in India in 2026, and what is included?
- 14. What hidden charges should I expect when buying [category]?
- 15. What contract terms are normal for [category] in India?
- 16. How do I compare quotes for [category] fairly?
Group E, trust and risk questions. These decide the shortlist more often than features do.
- 17. How do I check whether a [category] provider is legitimate?
- 18. What are the risks of using a [category] provider for [sensitive task]?
- 19. Which [category] providers publish their [coverage, pricing or compliance detail]?
- 20. What do buyers complain about most with [category] providers in India?
Two rules about the set. Your brand name appears nowhere in any of the 20. And once you have filled the brackets, freeze the file with a version number and a date, because every number you report afterwards belongs to that version.
How to run the test in ChatGPT
Eight steps. It takes one focused afternoon for 60 runs.
One. Start a fresh chat for every single run. One prompt per chat, no follow ups before you have captured the answer. A continued conversation carries context and the context changes the answer.
Two. Turn off or account for memory and personalisation, and write down which state you used. An account that has discussed your company before is not a clean instrument.
Three. Run each prompt three times, in three separate chats, giving 60 runs. Identical prompts produce different answers, so a single response is one draw.
Four. Record the conditions on every row: date, time, model or mode selected, whether browsing or search was used, country and language, and signed in state.
Five. Capture the full answer as text and the citations exactly, plus a screenshot. Text matters because you will want to search it later for your own phrases.
Six. Label each run with the four outcomes above, and also record which competitors were named, because that is the most useful side output of the whole exercise.
Seven. Record what it says about searching, and mark it unverified. You may ask. Do not trust the answer. In three separate batches of our September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question. And on 27 July 2026 ChatGPT confirmed it had run no live search for any of 78 questions, which means that whole reading described what the model remembered.
Eight. Never name your brand to check something mid test. If you must, do it in a separate labelled session that is excluded from the numbers. On 27 July 2026 the branded version of our own run came back ranking the company first on almost every question, while the blind version put it second by breadth and absent from the questions it most wanted to win. We discarded the branded run.
Results: the measured benchmark
We have not run the 20 prompt set on aiknowsus.com, so there is no 20 prompt figure to publish here. The cell is empty on purpose, and it will be filled with a fraction over 20, a repeat count and a date when the run is done.
The ChatGPT measurements we do hold, each with its denominator and date, are these.
- Top source on 1 of 18. Eighteen blind commercial questions, ChatGPT, 18 August 2026, clawlaw.in. The company's own comparison page was named in one answer as still being the vendor's own editorial page, which is the engine telling you why a vendor page underperforms.
- Named on 21 of 78. The blind baseline of 27 July 2026, ChatGPT. The order by breadth was ProVakil on 28 of 78, CLAW on 21 of 78 and Legistify on 18 of 78. On the same run, ChatGPT confirmed it had run no live web search for any of the 78, so those counts measure what the model had absorbed rather than the web that week.
- One verified first citation, 14 days after publishing. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, in a zone where the same question set had named the company nowhere at baseline. One page, one engine, one date, and the only figure in that programme with a public URL behind it.
- On our own domain, on a different engine: 0 of 6. Perplexity, September 2026, six blind questions in our own category, and it reported afterwards that it had cited or recommended us in none of them. Included here so you can see that a zero is a normal starting point, including for the company writing this page.
Read those four together and the useful pattern is this: being named is much easier than being the top source, and the gap between 21 of 78 and 1 of 18 is where most of the work sits.
Limitations
Seven, and the first three are the ones that cause wrong decisions.
- 20 prompts is a diagnostic, not a rate. It is enough to find what is missing on your site. It is not enough to publish a percentage about your category.
- Your account is not a neutral instrument. Memory, history and location all shape answers, which is why those fields are recorded.
- A result is valid for its date only. Models and modes change, and a benchmark is a photograph.
- The engine's account of its own behaviour is unreliable, documented above with three separate retractions.
- Use without credit is under counted because it depends on a person recognising their own material.
- ChatGPT results are not evidence about other engines. Run the same set on the others and report separately.
- No method here can promise you a position in an answer. Anybody who offers that is selling something else.
Sources and change log
All figures come from the clawlaw.in programme, recorded in GEO_BASELINE_RESULTS_2026-07-27.md and GEO_GAP_ANALYSIS_2026-08-18.md, and from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The 6 August 2026 citation is verifiable from outside; the counts over our capture files are not, until those files are published. Version: 29 September 2026, first publication, to be updated when our own 20 prompt run completes.
Common questions
How many of the 20 should a healthy business expect to be named in?
We will not give you a target, because we have no dataset that would support one and neither does anybody else we have read. Use your own first run as the baseline and measure against yourself. For scale, the strongest single figure we hold for a company with published pages is 1 of 18 for top source, on 18 August 2026.
Why three repeat runs rather than one?
Because answers vary between identical prompts. In the aiknowsus.com capture of September 2026 the assistant itself declined, batch after batch, to name any competitor as winning most often, saying its own earlier answers had not produced a comparable live tested result to support such a claim. If the engine will not draw a conclusion from one run, you should not either.
Can I ask ChatGPT to tell me why it did not mention us?
You can ask and the answer is a story, not a measurement. It may also be useful: in the clawlaw.in audit of 17 September 2026 Claude explained that two enterprise vendors were ranked above the company specifically because they publish explicit court and tribunal coverage lists, and it noted its ordering reflected price transparency and source authority rather than product quality. That was actionable. It was still the engine's account of itself.
Should the prompts mention my city?
In some of them, yes, because that is how buyers ask. Keep the local and the national versions as separate rows rather than mixing them, since the answers draw on different sources and a mixed set hides which one you are losing.
Our brand is named but never linked. Is that a win?
Partly. The buyer hears your name and no visit arrives, so it will not appear in your analytics. On 17 September 2026 Claude confirmed clawlaw.in pages had ranked in its raw results and shaped what it wrote, while the site landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy. The fix is usually the page, not the prompt.
Is 60 runs enough to report to my management?
It is enough to report as a dated baseline with its denominator printed, and that is exactly how you should present it: 20 prompts, three repeats, 60 runs, these dates, these three counts. Present it as a rate for your category and the first sceptical question will undo it.
Downloadable worksheet and what to do first
The worksheet is three sheets: the 20 filled prompts with a version and date, one row per run with the condition and label columns listed above, and a change log of every site change with its date. Build it in twenty minutes.
What to do first: fill the five lines about your business, fill the brackets in the 20 prompts, and run Group C, the problem questions, before anything else. There is a measured reason to start there rather than with the how to questions: on 17 September 2026 Claude put official court portals above every commercial product for the questions about how to look a case up, and said plainly that no commercial product should rank above the official portal for a question about using that portal. Group C is the group where a commercial page is the natural answer.