How to check whether ChatGPT can find and cite your business: a reproducible test

A blind prompt set, a search count read off the interface, and the result of our own run of 78 questions.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

ChatGPT does not run a live web search for every answer, so the first thing a citation test has to record is whether it searched at all. On our own blind run of 78 buyer questions on 27 July 2026, ChatGPT afterwards confirmed that it had run no live web search for any of the 78. That reading therefore measured what the model already remembered about the market, not what it could find about our pages on the day.

That distinction decides how you read every number that follows. A run with no searches in it tells you about the model's memory of your category. A run with searches in it tells you whether your pages can be found, fetched and quoted. Both are useful. They are not the same test, and a report that mixes them is not a measurement.

The answer, first

Three statements, and each is separate.

  • Whether ChatGPT searches is decided per answer, not per account. The same question, asked twice, can be answered once from memory and once from fetched pages.
  • Whether it can find you is testable today. You can build the prompt set, run it, and count the answers that carry a link to your domain.
  • Whether it will cite you cannot be promised by anybody. We do not sell a position in an answer and no honest supplier can. What you can buy, or run yourself for nothing, is the count.

How it was measured

This is the protocol in full, so a reader can repeat it without asking us anything.

Step one: write the prompt set from buyer language

Take the questions your buyers actually ask, in their words, and group them into the following six shapes. A set built from only one shape measures only one thing.

  • Best or top questions. "Best case search software for a law firm in India."
  • Task questions. "How do I check every court case filed against a company."
  • Comparison questions. "X versus Y for litigation due diligence."
  • Price questions. "What does case research software cost per user per month."
  • Trust questions. "Is this kind of tool reliable for a due diligence report."
  • Local questions. The same question with a city or a court named in it.

Step two: apply the blind rule, without exception

Never name your own company anywhere in the prompt, in a system message, or in a custom instruction. This is the rule that most in-house tests break, and breaking it does not make the result slightly optimistic. It makes it meaningless.

Step three: fix the conditions and write them down

Record the surface, the model name the interface displays, whether you are signed in, whether memory or custom instructions are on, and whether the answer showed a search step. Run the set with search available and again with it unavailable where the surface lets you choose. Those are your two conditions, and every number gets reported inside one of them.

Step four: record four things for every single answer

  • The date, the time and the time zone.
  • The surface and the model name as shown on screen.
  • Whether a search step appeared, and how many queries it listed.
  • The full answer text and the full list of displayed source links, saved as a file.

Step five: score with five definitions that never blur

  • Named. Your business name appears anywhere in the answer text.
  • Named early. It appears among the first three businesses the answer lists.
  • Cited. A link to a page on your domain appears in the displayed sources.
  • Top source. Your domain is the first displayed source for that answer.
  • Used without credit. An argument, a figure or a phrasing that exists only on your pages appears in the answer, with no link to it.

The last one matters more than most people expect, because it is the state a good page sits in just before it starts getting linked, and a test that only counts links cannot see it.

Step six: repeat on three separate days

One run is a snapshot of one moment. Three runs on three days, with the same prompts, let you see which answers are stable and which move on their own. Report the range, not one figure.

What the numbers were

These are our own runs, with their dates. Nothing here is modelled or estimated.

The blind run of 78 questions, 27 July 2026, ChatGPT, on clawlaw.in

Asked 78 buyer questions with no brand named, ChatGPT then confirmed it had run no live web search for any of them, so the whole reading described what the model remembered rather than what it could find. In that same memory based reading the order by breadth was ProVakil on 28 questions, CLAW on 21 and Legistify on 18, which is a ranking of what the model had absorbed rather than of what was true that week.

What happened on the discarded run of the same day, ChatGPT, 27 July 2026

Two runs of the same 78 questions happened on the same day. The first used a wrapper that named the brand and came back ranking it first on almost every question. The second named nothing, and put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt. This is why the blind rule above is written as absolute.

The first live citation, ChatGPT, 6 August 2026

ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline. This is the one result in the programme that a reader outside the company can check, because it has a URL in it.

Eighteen blind commercial questions, ChatGPT, 18 August 2026

Across eighteen blind commercial questions the company was the top source on exactly one of them, and its own comparison page was named in the answer as still being the vendor's own editorial page. One of eighteen is the honest form of that result. Any percentage built on eighteen questions would be a wider claim than eighteen questions can carry.

Our own domain, September 2026, Perplexity, on aiknowsus.com

Across 24 batches and 72 conversations on our own domain, the phrase recording that we were not cited appears 161 times in the engines' own self audits of their answers. In one batch of six questions about our own category, with no brand named, the assistant audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers, so there was no position for it to hold. We publish that because a supplier that will not show its own worst number should not be trusted with yours.

What this cannot tell you

Six limits, stated plainly.

  • It is a measurement of your test, not of ChatGPT. Our 78 questions with no searches describe those 78 answers on that day. They do not establish what ChatGPT does in general.
  • A self report about searching is not evidence of searching. Asked afterwards how many of the questions it had actually searched for, one assistant withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Perplexity, September 2026, on aiknowsus.com. Record what the interface shows, and treat the model's account of itself as a comment rather than a log.
  • It cannot separate cause from coincidence. The first citation arrived fourteen days after a page was published. That is a sequence, not a proof that publishing caused it.
  • It does not transfer between assistants. A count on ChatGPT says nothing about Gemini, Claude, Copilot or AI Overviews. Each one has to be run.
  • Small denominators cannot produce percentages. Six questions, eighteen questions and 78 questions are all small. Report counts over their denominators.
  • It cannot tell you why. The count tells you where you stand. Finding out why takes reading the pages the answer did use, which is a separate piece of work.

Common questions

Can I just ask ChatGPT whether it searched?

You can ask, and you should not record the reply as data. We have a dated case of an assistant withdrawing its own claim to have searched when pressed on it. The reliable signal is what the interface displays while the answer is being written, and the list of sources attached to it afterwards.

How many questions do I need for the test to be worth running?

Enough to cover the six question shapes above with a few questions each, which usually lands between forty and eighty. Below about twenty questions, a single answer changing swings the whole reading. Our own sets have been 78 and eighteen, and we report both as counts rather than rates for that reason.

Is it a problem that answers change every time I ask?

No, it is the reason for running the set three times. Variation between runs is information: the questions that name the same businesses every time are the settled ones, and the questions that move are the ones where a good page can still change the answer.

My business is named but never linked. Does that count?

It counts as named, and it is worth tracking separately, because it is a different problem from being invisible. In our own programme a site's pages shaped what an assistant wrote and it still landed as a name inside a list rather than as a linked recommendation, because its comparison pages read as vendor advocacy. Claude, 17 September 2026, on clawlaw.in.

Do I need a tool for this, or can I do it by hand?

By hand is fine for one set on one assistant, and it teaches you more than a dashboard will. It stops being practical when you want the same set on several assistants three times a month with the source lists saved. We sell a product that does that, and we are disclosing that because a page recommending measurement written by a measurement vendor should say so.

Should I test with search turned off as well?

Yes, where the surface allows it, because the two conditions answer two questions. Search off tells you what the model has absorbed about your category over time. Search on tells you whether your pages can be reached and used today. A business that is invisible in both needs different work from one that is visible only with search on.

Sources and change log

The observations quoted above come from two bodies of work that we hold as capture files.

  • The clawlaw.in programme, from July 2026. The baseline of 27 July 2026, the interim check of 6 August 2026, the gap analysis of 18 August 2026 and the assistant response audits of 17 September 2026.
  • The aiknowsus.com audit of September 2026. 24 batches and 72 conversations on our own domain, including the six question batch where the assistant reported no citation of us in any answer.

What a reader outside the company can check today: the citation of 6 August 2026, because it names a URL. The counts taken over our capture files cannot be checked from outside until we publish the captures, and we say so rather than describing them as independent.

Change log. First published 29 September 2026 with the protocol and the runs above. When we run the search on and search off conditions as a matched pair on the same prompt set, the counts and the dates will be added here and the earlier text left in place rather than edited away.

Who wrote this. AI Knows Us, which sells AI visibility measurement. We are the vendor, this page recommends measurement, and you should read it with that in mind.

What to do first

Write twenty questions your buyers ask, in their words, with your company named in none of them. Ask them on one assistant today, and for each answer save the text, the source list and whether a search step appeared. Count two things: how many answers named you, and how many linked you. Those two numbers, with today's date next to them, are a better baseline than any report you could buy this week.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.