ChatGPT and Gemini product recommendations: the 100 prompt comparison protocol, and what we have run so far

How to measure top pick agreement between two assistants, and an honest account of which part of it we have done.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Two assistants asked the same buying question often name different companies, and the only honest way to say how often is to run both on the same frozen prompt list on the same day and count how many times their top pick matches. We have not run that paired comparison between ChatGPT and Gemini, so we are not going to publish an agreement rate. What we have is single engine readings on different sets and dates: ChatGPT on 78 blind buyer questions on 27 July 2026, ChatGPT on 18 blind commercial questions on 18 August 2026, Claude on a source selection audit on 17 September 2026, and Perplexity across 24 batches and 72 conversations in September 2026. None of those pairs up into an agreement count, and this page publishes the full protocol instead, so the number can be produced properly by us or by anybody else.

A fabricated agreement rate would be the easiest number on this site to invent and the hardest to defend. The protocol is the deliverable.

The answer, first

Three things can be said with confidence, and one cannot.

They can differ, and our own files show cross engine divergence even without a paired test. On 27 July 2026 ChatGPT answered 78 blind buyer questions and confirmed it had run no live web search for any of them, so its answers came from memory. On 17 September 2026 Claude, asked about the same category, was reading live sources closely enough to check a competitor's prices independently and to notice that two comparison pages carried no link, no date and no source. Two engines, two months apart, behaving in ways that would produce different top picks on the same question for reasons that have nothing to do with the products.

The reason they differ is mostly source selection, not preference. An engine that searches reads today's pages. An engine that does not reads what it absorbed. An engine that discounts vendor advocacy will drop the same comparison page that another engine happily quotes.

The measurement has to be paired to mean anything. Same prompts, same day, same location, same definitions, both engines. Anything less is two separate readings being compared by eye.

What cannot be said is how often they agree, by us, today. That is a count we do not have.

How it was measured, and how to measure it

The protocol below is written so that two people running it separately would produce comparable numbers. It is the 100 prompt design the question deserves.

One. Build 100 paired prompts. One hundred buyer intents, each written once, used identically on both engines. Five groups of twenty, and the group counts go into the report: category choice, direct comparison of two approaches, price, suitability for a stated buyer, and risk or limits.

Two. Apply the blind rule to every prompt. No company name, no product name, no domain, no phrase that appears only on one vendor's site. Our own record is the argument for this rule. On 27 July 2026 two runs of the same 78 questions happened on the same day; the one whose wrapper named the brand came back ranking it first on almost every question, while the blind one put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The branded run was discarded, because the only thing it had measured was our own prompt. A branded prompt would break a cross engine comparison twice over, because each engine would be flattering the same name.

Three. Fix the conditions. Same location setting on both. Both runs inside the same 24 hours. Search enabled on both, and recorded. Memory and personalisation off where the product allows it, and the setting recorded where it does not. One prompt per fresh conversation, on both engines.

Four. Define the top pick before the run. The top pick is the first company the answer names in the part that tells the reader what to use or shortlist. Not the first company named anywhere in the text, because background mentions come earlier and are not recommendations. Write down four handling rules in advance: if the answer names no company, the top pick is none; if it presents two options as equally suitable, record a tie and list both; if it recommends a category of product rather than a company, record none with a note; if it recommends an official body or a government portal, record that, because it is the true top pick and it happens more than vendors expect. On 17 September 2026 Claude recorded that on questions about how to look a case up, official court portals took every position above any commercial product, and said plainly that no commercial product should rank above the official portal for a question about using that portal.

Five. Score five fields per answer, per engine. Valid, searched, top pick, every company named in order, every domain cited.

Six. Compute agreement three ways. Strict agreement: prompts where both engines gave the same single top pick, over valid paired prompts. Loose agreement: prompts where each engine's top pick appears anywhere in the other's named list. Set overlap: the average number of companies appearing in both lists, reported with the list lengths. Report all three, because strict agreement alone will look low for a reason that is partly definitional.

Seven. Test stability before you believe either engine. Run 20 of the 100 prompts three times on each engine in fresh conversations. Count how often each engine agrees with itself on the top pick. Self agreement is the ceiling on cross engine agreement. If an engine only agrees with itself on 12 of 20, then a cross engine agreement count is mostly measuring variance, and the honest report says so.

Eight. Say what each result would mean. High strict agreement with high self agreement means the category has a settled published consensus, and the way to enter it is to get into the sources both engines read. Low strict agreement with high self agreement means the two engines are reading different source sets, and the actionable finding is which domains each one cites. Low self agreement makes every other number provisional, and the correct response is more repeats, not a bolder headline. And if the searched count is low on either engine, that engine's answers are memory and the comparison is between one live reading and one recollection.

What the numbers were

The paired ChatGPT and Gemini comparison: not run. We have no agreement count to publish, strict, loose or set overlap, and no self agreement count for either engine. We are not estimating one.

What we do hold, engine by engine.

  • ChatGPT, 27 July 2026, 78 blind buyer questions, clawlaw.in programme. Companies named by breadth: ProVakil on 28 questions, CLAW on 21, Legistify on 18. Search rate: 0 of 78, confirmed by the engine itself. Reported as a memory reading, not a live one.
  • ChatGPT, 18 August 2026, 18 blind commercial questions. The measured company was the top source on exactly 1 of 18, and its own comparison page was named in the answer as still being the vendor's own editorial page.
  • Claude, 17 September 2026, source selection audit. On a question about finding every case against a company, two enterprise vendors ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and Claude noted that its ordering reflected price transparency and source authority rather than product quality. Across six questions, no third party review or directory source made it into any answer at all.
  • Perplexity, September 2026, aiknowsus.com, 24 batches and 72 conversations. The phrase recording that we were not cited appears 161 times in the engines' own self audits. In one batch of six questions the engine reported afterwards that it had not cited or recommended aiknowsus.com in any of the six. In three separate batches it withdrew its own claim to have run a live search, saying it could not honestly substantiate it.

Why those cannot be turned into an agreement rate. Different prompt sets, different denominators, different months, and in one case a search rate of zero. Comparing 21 of 78 with 1 of 18 is not a comparison, and neither of those is a top pick count.

The one cross engine finding we can state. Both ChatGPT and Claude, in the same programme, ended up pointing at the same missing documents rather than at each other. On 6 August 2026 Claude found the pricing page rendered its prices only after scripts ran, so a crawler received no prices at all, and the prices the assistants quoted had come from an app store listing instead. The same week ChatGPT noticed that the website and the app store listing carried different plan names and different prices for the same product, and said so in its answer. Two engines, two findings, one cause.

What this cannot tell you

  • It cannot give you an agreement number today. That is the honest state of this page and the reason the protocol is the main content.
  • A completed run could not be generalised either. One hundred prompts in one category in one country on one pair of days describes that.
  • Agreement is not correctness. Two engines can agree because they read the same directory page, and that page can be out of date or written about another country. On 17 September 2026 a well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
  • Disagreement is not evidence of bias. The commonest cause in our own files is whether the engine searched at all.
  • An engine's self report is not an instrument. Perplexity withdrew its own search claim in three separate batches in September 2026. Record self reports and verify them against your own server logs.
  • Nothing here promises a position. No method and no product, ours included, can guarantee that either engine names a company.

Common questions

Which engine should I optimise for if they disagree?

Neither, at the level of engine specific tricks. The work that helps on both is the same: pages that are readable without scripts, facts with dates and sources, coverage stated as a named list, and consistent numbers across your site, your app listing and your directory entries. Where our own engines differed, they still pointed at the same missing documents.

If they disagree, whose answer do buyers see?

Whichever one your buyer opened, which is why the practical answer is to measure both and report them separately rather than blending. Two engines are also a cheap guard against redesigning a website around one product's behaviour in one month.

Why is a top pick better to measure than a mention?

Because a mention is cheap and a top pick is the decision. Keep both: mention, recommendation, citation, top source. In our 18 question run of 18 August 2026 the strictest measure was 1 of 18, and the looser measures would have read very differently from the same answers.

How many repeats do I need to trust an engine's answer?

Three per prompt on a subset of twenty is enough to see whether the engine agrees with itself. If self agreement is low, your headline number needs more repeats rather than a stronger claim. This is the step most published comparisons skip.

Can I compare an engine's answer today with a screenshot from last month?

No. Retrieval behaviour changes without notice, and the mode you used may have changed too. A comparison needs both readings inside the same short window, which is why the protocol fixes a 24 hour rule.

Does a tie count as agreement?

Record ties separately and report them as their own count. If you fold ties into agreement you will inflate it, and if you fold them into disagreement you will understate it. Publishing the tie count is what lets a reader recompute your rate under their own rule.

Sources and change log

  • Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 78 question blind run of 27 July 2026, the zero search confirmation, the breadth order, the discarded branded twin, and the 6 August 2026 findings on unreadable prices and the two conflicting public price lists.
  • Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 blind commercial questions of 18 August 2026.
  • Tier_1/claude_response_17_09_audit.md. The 17 September 2026 source selection audit, including coverage lists, official portals, and the discarded American product list.
  • geo-audits/aiknowsus-com/. September 2026, 24 batches and 72 conversations, including the withdrawn search claims and the batch of six with no citation.

Change log. 27 July 2026 ChatGPT baseline. 6 August 2026 interim findings. 18 August 2026 ChatGPT commercial run. 17 September 2026 Claude source audit. September 2026 Perplexity own domain audit. Not run: the paired 100 prompt comparison, the self agreement subset, and therefore any agreement rate. When it is run this section will carry the date, the valid paired denominator, the strict and loose agreement counts, the tie count and the self agreement counts, whatever they turn out to be.

What to do first

Write twenty paired prompts rather than a hundred, and run them on both engines in one afternoon, one prompt per fresh conversation, with the searched field filled in before you score anything. Then run five of those prompts three times each on both engines and count self agreement. If an engine does not agree with itself, stop building a comparison and fix the sample size first. That one afternoon tells you whether a hundred prompt run is worth the week it costs.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.