Should you ask ChatGPT about your own company? A fair test protocol

Yes, but never only by name. A named prompt and a blind prompt measure two different things.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Yes, ask, but never only by name, and never report the named result as your visibility. We have run both versions of the same test on the same day. On 27 July 2026 two runs of the same 78 buyer questions were done on the clawlaw.in programme. The run whose wrapper named the brand came back ranking the company first on almost every one of the 78. The run that named nothing put the company second by breadth, mentioned on 21 of the 78, and absent altogether from the litigation due diligence questions it most wanted to win. The branded run was discarded, because the only thing it had measured was our own prompt.

Both prompts are useful. They answer different questions, and the mistake that ruins a measurement is reporting one as if it were the other.

Short answer: yes, but not only by asking your company's name

A named prompt tests what the engine can say about you when it already knows who you are. A blind prompt tests whether it brings you up when nobody has. Your buyer is doing the second one. You should do both, keep the results in two separate columns, and publish only the blind one as your visibility number.

Keeping the two apart is also the single cheapest way to stop a dashboard from flattering you. Our own branded run produced the best looking result in the whole programme and was thrown away on the day it was produced.

What direct name prompts test

A prompt that names you is a test of your public record, and it is genuinely worth running. Five things it can find that a blind prompt cannot.

  • Whether the engine has anything about you at all. An answer that hedges on every specific is telling you your facts are not published anywhere it can read.
  • Whether your facts are correct in its account. On 6 August 2026 during the clawlaw.in programme, the website and the app store listing carried different plan names and different prices for the same product, and ChatGPT noticed the contradiction and said so in its answer. A named prompt is how you find that.
  • Whether your prices are readable at all. On 6 August 2026 Claude recorded that the pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants did quote had come from an app store listing instead of the company's own site.
  • Whether a claim of yours fails a check. On 6 August 2026 Claude found a headline figure on the main site that could not be true, checked it against public numbers and advised a buyer against the product, after which the company's accurate claims stopped counting for that answer.
  • Which sources it uses to describe you. The domains it cites for your own company are the pages you need to fix or answer.

What a named prompt cannot do is tell you whether you are in the running. Politeness is the default. An engine asked whether a company is any good will usually find something positive to say about it.

What neutral category prompts test

A blind prompt tests the thing you actually care about: whether your name arrives unprompted in the answer a buyer gets. Four things it finds.

  • Whether you are named at all, as a count over a denominator.
  • Whether you are recommended or merely listed. On 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote, and the company still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.
  • Which competitors arrive instead, and why. On 17 September 2026, on a question about finding every case against a company, two enterprise vendors ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and Claude noted that its ordering reflected price transparency and source authority rather than product quality.
  • Whether the engine searched. On 27 July 2026 ChatGPT confirmed it had run no live web search for any of the 78 blind questions, which reframed the whole reading as what the model remembered rather than what it could find.

A fair test in four steps

Step one. Write the blind set and freeze it. Forty to a hundred buyer questions. No company name, no product name, no domain, no phrase that appears only on your own site. Read the list again specifically for leakage, because the leak is usually a phrase you invented and forgot was yours.

Step two. Write the named set as a mirror. For each blind prompt, write the version that names you. Same intent, same wording otherwise. This pairing is what makes the two columns comparable.

Step three. Run both, one prompt per fresh conversation, on the same day and on at least two assistants. Record for every answer: valid, searched, mentioned, recommended, cited, top source, and every domain cited. Save the full answer text as a file.

Step four. Report the two columns separately and never combine them. Publish the blind counts as your visibility. Use the named counts as a factual accuracy report: what the engine believes about you, what it got wrong, and which source it got it from. If you ever see the two numbers averaged in a report, the report is worthless.

State in advance what each pattern means. Named high and blind zero is the ordinary starting position, and it means your facts exist somewhere readable but you are not in the consideration set. Named wrong and blind zero means fix the facts first, because being brought into an answer with a wrong price is worse than being absent. Named low, as in the engine cannot describe you at all, means your basic facts are not published as text anywhere it reads, and that is a one week job. Blind present but never cited means your pages are read as advocacy or your evidence is unsourced.

Example results and limitations

The clawlaw.in programme, 27 July 2026, ChatGPT, the same 78 questions run twice on the same day.

Blind run, sample size 78. The company was mentioned on 21 of the 78 questions. By breadth the order was ProVakil on 28, CLAW on 21, Legistify on 18. It was absent altogether from the litigation due diligence questions, which were the questions the business most wanted to win. Live searches: none, for any of the 78, confirmed by the engine.

Named run, sample size 78. The wrapper named the brand and the run came back ranking the company first on almost every question. Our record describes it that way, as almost every question, and does not retain a per question count, because the run was discarded on the day and was not scored question by question. We report it as the record holds it rather than putting a number on it now.

What the pair shows. Same day, same engine, same 78 questions, two results that would support opposite conclusions. The blind one is the one a buyer would experience. That is the whole case for the blind rule, and it is the most useful single thing in our own files.

What the pair cannot show. It cannot tell you how large the gap would be on another engine, in another category, or on another date. It is one pair of runs on one domain. It also cannot tell you that the blind number is correct in an absolute sense, only that it is the honest one: with a search rate of zero it is a measure of what the model had absorbed, not of what was published that week. And five of the seven counted results in this programme cannot be checked by a reader from outside, because the capture files are not published. The exception is the citation of clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india by ChatGPT on 6 August 2026, fourteen days after that page went up.

One more limit worth stating, because it applies to any self audit. In September 2026, asked afterwards how many of the questions it had actually searched for, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. That happened in three separate batches. An engine asked about itself is a witness, not an instrument.

When a tracking service is worth considering

Do the first run by hand. It is the only way to learn what your own definitions cost to apply, and it takes an afternoon for forty prompts. Consider a service at the point where three things are true: the set is large enough that a manual run takes more than a day, you want the same set run on a schedule so the comparison is real, and you need the raw answers kept so a number can be audited months later.

The tools named most often in answers about this category, counted across our whole September 2026 capture of 24 batches, were Semrush, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a mention count from one capture, not a quality ranking.

Disclosure: we sell AI Knows Us, and we put it first in our own lists for one stated reason, that it runs the blind set on a schedule and keeps every raw answer so a rate stays auditable. The other tools are real products and should be judged on their own pages. We do not restate their prices here, because an undated price on a third party page is exactly the thing that gets a page discounted: on 17 September 2026 Claude discounted two clawlaw.in comparison pages whose competitor prices were in fact correct, because at read time they carried no link, no date and no source.

Common questions

Is asking about my own company harmful in any way?

Not harmful, only misleading if you report it as visibility. Run it deliberately as an accuracy check and label the column that way. The harm is a dashboard that shows you first on everything, which is precisely what our discarded run of 27 July 2026 produced.

What about asking it to compare me with a named competitor?

Useful, and it is a third column, not a blind result. A prompt naming both companies tests how the engine argues between two options it has been handed. It tells you nothing about whether either would have come up.

Does logging in or using my own account skew the result?

Assume it can, and record it. Personalisation and memory features mean an account that has discussed your company before is not a clean instrument. Where you can, run the blind set in a fresh session with memory features off, and note in the log which account and which mode were used.

How do I catch brand leakage in a prompt?

Read every prompt for three things: your name, your product name, and any phrase you coined. The third is the one that slips through. If a phrase gets no results anywhere except your own site, it is your brand in disguise.

Is one named prompt enough for the accuracy check?

No. Ask at least five differently framed named questions: what the company does, what it costs, who it is for, what it does not do, and whether it is credible. The price question and the limits question are where the contradictions turn up, as they did on 6 August 2026 when ChatGPT pointed out that the website and the app listing disagreed.

Should I correct the engine when it gets something wrong?

Correct the page, not the engine. A correction inside a conversation ends with the conversation. A dated, readable fact on your own site, matching your app listing and your directory entries, is the thing the next answer can read.

FAQ notes and sources

  • Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The method note recording the two runs of the same 78 questions on 27 July 2026 and the decision to discard the branded one; the breadth order; the zero search confirmation; and the 6 August 2026 findings on unreadable prices, the two public price lists, the unsupportable headline figure and the first cited URL.
  • Tier_1/claude_response_17_09_audit.md. The 17 September 2026 findings on advocacy discounting, published coverage lists and being named without being linked.
  • geo-audits/aiknowsus-com/. The September 2026 audit of our own domain, 24 batches and 72 conversations, including the tool mention counts and the three batches in which the engine withdrew its own search claim.

Update history. 27 July 2026, the paired blind and branded runs. 6 August 2026, the accuracy findings and the first citation. 17 September 2026, the source selection audit. September 2026, the own domain audit. Still to be run: a paired blind and named set scored question by question on both sides, which would let us publish the branded column as a count instead of as the phrase our record uses.

What to do first

Write your forty blind prompts and five named prompts this week, and run the five named ones first, because they take twenty minutes and they usually surface a factual problem you can fix immediately: a price the engine cannot read, two different price lists in public, or a claim that does not survive a check. Then run the blind forty, and report only those as your visibility, with the search rate printed next to them.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.