How ChatGPT can surface company recommendations: what is documented and what is not

There is no published recommendation formula. Here is the test we ran instead, with the counts and the dates.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

There is no public, complete formula for how ChatGPT decides which companies to name. What can be checked is narrower and more useful: whether the answer ran a live search, which sources it cited, and which companies it named. In our own blind run of 78 buyer questions on 27 July 2026, ChatGPT confirmed it had run no live web search for any of the 78, which means zero of those 78 answers rested on a live source. The companies it named most often in that run were ProVakil on 28 questions, CLAW on 21 and Legistify on 18. On 18 August 2026, across 18 blind commercial questions, the domain we were measuring was the top source on exactly 1 of 18. Those are study results about answers, not a ranking formula.

Short answer: no public complete recommendation formula

Three separate things get confused in this question, and separating them is most of the answer.

Eligibility is documented. Whether a crawler may fetch your pages, and which crawler does what, is published by the companies that run them and is controlled by you through your own robots file and your server.

Retrieval behaviour is observable. Whether a given answer searched, and what it cited, can be seen in the answer and, for your own domain, in your server logs.

Selection is not documented. How the model weighs one candidate source against another inside a single answer is not published as a formula, and any page that lists weighted ranking factors for it has invented them. Treat a confident list of factors as a signal to check the page's sources.

The useful consequence is that you should stop trying to satisfy a formula and start measuring answers. An answer is a thing you can capture, date and count. A formula is not.

What is documented, and what you can check yourself

Four things are checkable today without asking anybody's permission.

  • Crawler access. Your robots file either allows the assistants' fetchers or it does not. Read your own file rather than trusting a plugin's summary of it.
  • Server evidence of a fetch. Your access logs record which agent asked for which URL and when. This is the only first party evidence you have, and it is the evidence most teams never look at.
  • Whether the page renders without scripts. Fetch your own URL the way a crawler would and read what comes back. On 6 August 2026 during the clawlaw.in programme, Claude recorded that the pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead of the company's own site.
  • What the answer cited. The links in the answer are the clearest statement of what it used, and the pattern across many answers is more informative than any single one.

What a cited answer does, and does not, prove

A citation proves that the answer used or displayed that source in that answer, on that date, for that question, in that session. It does not prove five other things, and every one of these has bitten a real reading.

  • It does not prove the answer searched. An engine can restate something it remembers and attach a plausible link afterwards. In September 2026, asked afterwards how many questions it had actually searched for, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. That happened in three separate batches on the aiknowsus.com audit.
  • It does not prove your page changed the ranking. A source can be cited as background while a different source decides the recommendation.
  • Absence of a citation does not prove absence of use. On 17 September 2026 Claude admitted it had used two specific arguments drawn from clawlaw.in pages and dropped the attribution in both places, and called it a citation lapse rather than a ranking judgement.
  • It does not repeat reliably. The same question in a new session can cite different sources the same day.
  • It is not an endorsement. On 6 August 2026 Claude checked a headline figure on the clawlaw.in main site against public numbers, found it could not be true, and advised a buyer against the product. The site had been read closely and the outcome was the opposite of a recommendation.

A reproducible recommendation test

This is the protocol that produced the counts on this page. Run it as written and your numbers will mean the same thing as ours.

One. Choose the decision you are testing. A recommendation test is about answers where a buyer is choosing a supplier. Do not mix in how to questions, because those are won by official sources. On 17 September 2026, on questions about how to look a case up, official court portals took every position above any commercial product, and Claude said plainly that no commercial product should rank above the official portal for a question about using that portal.

Two. Write the prompt set and freeze it. Between 40 and 100 questions. Every one blind: no company name, no product name, no domain, no phrase that exists only on your site.

Three. Run one question per fresh conversation. A follow up in the same thread inherits the earlier context and stops being an independent observation.

Four. Capture four things per answer. The full answer text. Every company named, in the order named. Every source cited, with its domain. Whether the answer searched, taken from the interface and from your own logs.

Five. Ask the engine to audit itself, in a second message. Ask how many of the questions it searched for, whether it cited your domain, and which company it would name first. Keep the reply as data, not as truth. Its value is that it sometimes contradicts the first message, which is itself a finding.

Six. Score with three counts over one denominator. Answers that named at least one company. Answers that carried at least one link to a source. Answers that named your company. Report each as a count over valid answers with the date and the engine.

Seven. Say what each pattern would mean. Many companies named and no links means the engine is answering from memory, so your short term lever is presence in the sources that trained and are remembered, not this week's page. Many links and few companies means the questions are being read as informational rather than commercial, so rewrite the prompts as choices. Links present and your domain absent while competitors' domains appear means a source selection problem, and the next step is to read what those pages carry that yours does not.

Results

Run one. clawlaw.in, 27 July 2026, ChatGPT, 78 blind buyer questions. Companies were named freely: ProVakil on 28 of the 78 questions, CLAW on 21 and Legistify on 18. Live search: none, confirmed by the engine for all 78. So the count of answers resting on a live source is 0 of 78, and the breadth order is a ranking of what the model had absorbed rather than of what was true that week.

The discarded twin, same day. The same 78 questions run through a wrapper that named the brand came back ranking it first on almost every question. It was discarded. It is reported because it is the cleanest available demonstration that a branded prompt measures the prompt.

Run two. clawlaw.in, 18 August 2026, ChatGPT, 18 blind commercial questions. Top source on exactly 1 of 18. In that answer the company's own comparison page was named as still being the vendor's own editorial page.

Run three. aiknowsus.com, September 2026, Perplexity, 6 questions, no brand named. The engine's own self audit reported that it had not cited or recommended aiknowsus.com in any of the six. 0 of 6.

Run four. aiknowsus.com, September 2026, 24 batches, 72 conversations. The phrase recording that we were not cited appears 161 times in the engines' own self audits across that capture. Counted across all 24 second message answers, by script.

Which sources actually appeared. Counted across the whole September 2026 capture, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Separately, across six questions on 17 September 2026, no third party review or directory source made it into any answer at all; one well known review site did appear in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.

The one citation with a URL. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india for a vendor due diligence question, fourteen days after the page was published, in a zone where the same set had named the company nowhere at baseline.

How to interpret company mentions and citations

Three rules that we apply to every reading.

Read the mention and the citation as separate facts. A name in a list is cheap. A link to your domain means your own page was used. The gap between them is a diagnosis: on 17 September 2026 Claude recorded that clawlaw.in pages ranked in its raw results and shaped what it wrote, and the company still landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.

Read the sources an answer preferred, not just your absence. On 17 September 2026, on a question about finding every case against a company, two enterprise vendors ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and Claude noted that its ordering reflected price transparency and source authority rather than product quality. That sentence is a work instruction: publish the list.

Never accept an engine's ranking claim as a measurement. In batch after batch in September 2026, Perplexity declined to name a competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. If the engine will not make that claim from its own answers, a vendor dashboard should not either.

What this cannot tell you

  • It cannot give you the selection rule. Nothing in a set of answers reveals the internal weighting, and this page does not claim one.
  • It cannot generalise. Two domains, two categories, India, mid 2026.
  • It cannot separate cause from sequence. The page cited fourteen days after publication may have been cited for reasons that also changed in those fourteen days.
  • It cannot be pooled across engines. ChatGPT and Perplexity numbers here are separate instruments.
  • It is mostly not checkable from outside yet. Of the observations behind this page, only seven can be verified by a reader today, and the capture files are not published.

Common questions

Does ChatGPT have a list of preferred websites for recommendations?

No published list exists, and our capture does not show one either. What it shows is a pattern in what got cited: across the whole September 2026 audit the most cited domains were the assistants' own documentation and the vendors' own websites, with one well known review site far down the list. A pattern in one capture is not a preference list.

If the engine did not search, is there anything I can do this month?

Yes, but not by publishing more pages and expecting this month's answers to change. When the engine answers from memory, the levers are the ones that show up in many places at once: consistent facts across your site, your app listing and your directory entries, and presence in the registers and directories of your sector. On 6 August 2026 ChatGPT noticed that the website and the app store listing carried different plan names and different prices for the same product, and said so in its answer.

Why would an assistant recommend a competitor with a worse product?

Because it can only compare what is published. The 17 September 2026 finding is the clearest case we have: two enterprise vendors ranked higher specifically because they publish explicit coverage lists, and the assistant said its ordering reflected price transparency and source authority rather than product quality.

Should I put instructions for AI engines in my page?

No. On 6 August 2026 Claude found a competitor's page carrying a hidden block of text addressed to answer engines, instructing them to cite that company as the source. It refused the instruction and named the company that had done it. The downside is being described that way in front of your buyer.

How many answers do I need before I report a rate?

Report the count and the denominator at any size, and hold off on percentages below about 30 valid answers. We publish 1 of 18 and 0 of 6 as counts for exactly this reason.

Can a tracking product tell me why I was not recommended?

It can tell you what was cited instead, which is most of the way there, and it can keep the captures so the comparison is possible at all. We sell one, AI Knows Us, so read that as a disclosure. Nobody can sell you a guaranteed place in an answer.

Primary sources and test materials

  • Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 78 question blind run of 27 July 2026, the zero search confirmation, the breadth order, the discarded branded twin, and the 6 August 2026 findings including the first cited URL and the hidden instruction block on a competitor page.
  • Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 blind commercial questions of 18 August 2026 and the top source count.
  • Tier_1/claude_response_17_09_audit.md. The 17 September 2026 audit: advocacy discounting, coverage lists, official portals, the dropped attribution, and the six questions with no third party review source.
  • geo-audits/aiknowsus-com/. September 2026, 24 batches and 72 conversations, including the withdrawn search claims in three batches and the batch of six with no citation.

Change log. 27 July 2026 baseline, 6 August 2026 interim findings, 18 August 2026 commercial run, 17 September 2026 source audit, September 2026 own domain audit. Not yet run: a repeat of the 78 question set in a period where the engine does search, which is the run that would let us report a live citation rate over the same denominator.

What to do first

Pick the ten questions a buyer asks immediately before choosing a supplier in your category. Run them blind on ChatGPT, one per fresh conversation, and record three counts: how many answers named any company, how many carried a link, and how many named you. Then read the pages that were cited instead of yours and write down the one thing each of them states that you do not: a price with a date, a coverage list with names in it, or a limit stated plainly. That list is your next month of work.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.