Do AI answers favour big brands? A dated test of 100 category prompts

How to define brand size groups before you measure, the rate difference arithmetic, and the only dated counts we hold.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

The question can only be answered with a rate difference between groups you defined in advance, reported as raw counts over their own denominators. We have not run that test as of 29 September 2026, so this page publishes no difference in recommendation rate between brand size groups, and no headline percentage. What we hold is one dated breadth count in one sector, and it is a count by vendor rather than by size: in legal technology on 27 July 2026, across 78 blind questions on ChatGPT with one run each, ProVakil was named in 28 of 78 responses, CLAW in 21 of 78 and Legistify in 18 of 78. As arithmetic those are about 36, 27 and 23 in every 100 responses. That is a ranking of three named vendors, not a comparison of brand size groups, and the engine confirmed it had run no live web search for any of the 78, so it describes what the model had absorbed rather than what was true that week. The more useful dated finding is about the mechanism: on 17 September 2026 Claude ranked two enterprise vendors above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and said its ordering reflected price transparency and source authority rather than product quality.

That last sentence is the answer most people are actually looking for. What looks like a preference for big brands is, at least in the one run we can point to, a preference for published evidence. Big companies tend to publish more of it. That is a different problem from size, and unlike size it is something a smaller business can change.

How it was measured

The design is 100 category prompts, four engines, three runs each. That is 100 times 4 times 3, which is 1,200 responses. The arithmetic is printed because the denominator decides everything: a recommendation rate over 1,200 responses and one over 100 prompts differ by a factor of twelve, and two studies using different denominators cannot be compared even when both are honest.

Step one. Define the brand size groups in writing, before any prompt is run. This is the step that decides whether the finding means anything. Group membership must be decided by a rule a second person could apply from public information, and the rule has to be recorded with the date it was applied, because companies move between groups.

  • Employee count from a filed or published source, with the source and date recorded. Not an estimate from a website.
  • Disclosed external funding or public listing status, as a yes or no with the source.
  • Count of independent third party sources that cover the brand, counted by a written rule. This is the proxy for public evidence, and it usually correlates with size, which is why it must be recorded separately rather than assumed.
  • Domain age and indexed page count, both recorded on the date of grouping.

Assign every brand you will count into a small, medium or large group using those four fields, and publish the assignment table with the grouping date. If you cannot publish the table, you cannot publish the difference.

Step two. Build 100 category prompts and freeze them. All of them are category questions, because mixing in comparison prompts that name large brands guarantees the large group wins and tells you nothing. Save the set with a version number and the date. Vary the phrasing, the city, and whether a price or a size constraint is included, so the set is not 100 rewordings of one question.

Step three. Name no brand you are measuring, in any prompt. The reason is measured. On 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it had measured was our own prompt.

Step four. Count with a closed brand list plus an open column. Before you start, list every brand you will track. During labelling, also record any brand named that is not on your list, because the brands you had not thought of are half the finding about whether the answers concentrate on a few names.

Step five. Score five states separately. Named, own domain linked, top source, recommended and cited together, and used without attribution. Report the difference between groups on each state, because a large brand being named and a large brand being cited are different advantages with different causes.

Step six. Do the rate difference arithmetic on the page, in the open. For each group: numerator is the number of responses naming at least one brand from that group, denominator is 1,200 responses, since every response could have named any brand. Report the two rates as counts first and as a percentage second, then the difference in percentage points, and state the denominator again next to the difference. A difference reported without both denominators is not a result.

Step seven. Record every claim an engine makes about searching, as a claim. Two dated reasons. On 27 July 2026 ChatGPT, asked 78 blind questions, then confirmed it had run no live web search for any of them: 0 of 78. And in three separate batches of the September 2026 aiknowsus.com audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported. Whether an answer came from memory or from the live web changes what a size effect would even mean, so this column is not optional.

Step eight. Log fourteen fields per response and keep the captures: prompt identifier, prompt text, prompt set version, engine, run number, date and time with time zone, country and city, language, signed in state, brands named in order, URLs offered, the five states, the search claim, the labeller's name, and the capture link.

What the numbers were

Everything we hold, with the sector named on each line, because a legal technology count is not evidence about retail, healthcare or manufacturing.

  • Breadth of naming across 78 blind questions. Legal technology, clawlaw.in, 27 July 2026, ChatGPT, one run per question. ProVakil in 28 of 78 responses, CLAW in 21 of 78, Legistify in 18 of 78. As arithmetic, roughly 36, 27 and 23 in every 100. The same run recorded 0 of 78 questions for which the engine would confirm a live search, so this is a memory based reading.
  • Why an ordering came out the way it did. Legal technology, clawlaw.in, 17 September 2026, Claude. On a question about finding every case against a company, two enterprise vendors were ranked above clawlaw.in specifically because they publish explicit court and tribunal coverage lists, and the assistant noted that its ordering reflected price transparency and source authority rather than product quality.
  • Top source on 1 of 18 responses. Legal technology, clawlaw.in, 18 August 2026, ChatGPT, eighteen blind commercial questions, and in that one answer the company's own comparison page was named in the answer as still being the vendor's own editorial page.
  • 0 of 6 responses cited or recommended us. AI visibility tooling, aiknowsus.com, September 2026, Perplexity, six blind questions in our own category. We are a small brand in that category, and that is our own number.
  • 161 recorded statements that we were not cited, across 24 batches and 72 conversations, September 2026, Perplexity, in the engines' own self audits.
  • Mention order in our own category, counted across the whole September 2026 capture, Perplexity: Semrush most often, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a count of how often each tool was named, not a quality ranking, and it is the closest thing we hold to a size effect. It is not a size effect, because we did not group those tools by size before counting.

What has not been run, as of 29 September 2026. We have not defined brand size groups for any sector, and we have not run the 100 prompt, 4 engine, 3 run test. So this page publishes no recommendation rate for a small group, no rate for a large group, no difference in percentage points, and no significance claim. Those numbers do not exist in our files and we are not going to produce them by reasoning.

There is one more dated finding that belongs here, because it is the opposite of a size story. On 6 August 2026 Claude found a competitor's page carrying a hidden block of text addressed to answer engines, instructing them to cite that company as the source. It refused the instruction and named the company that had done it. Whatever advantage exists in these answers, it is not bought by tricking the reader.

What this cannot tell you

  • It cannot tell you whether AI answers favour big brands in general. Nobody can, from outside, across all sectors. A grouped test tells you about your prompts, in your sector, on those dates.
  • Three vendors in one sector are not size groups. The 28, 21 and 18 counts above are named vendors in legal technology. Reading them as evidence about company size would be inventing a finding, and we will not do it.
  • Size and published evidence are tangled together and a simple grouping cannot separate them. That is why the third grouping field, the count of independent third party sources, is recorded separately. Without it you will attribute an evidence effect to size.
  • A memory based reading and a live search reading are different measurements. On 27 July 2026 the engine had searched for none of the 78 questions. A size effect in memory says something about training data, which you cannot change. A size effect in live retrieval says something about published sources, which you can.
  • Percentages from a single week are not stable. Answers vary between identical runs, engines change, and a difference of a few points from one pass is noise until it repeats.
  • Nothing here promises you a position in an AI answer, at any company size.

Sources and change log

Figures come from the clawlaw.in programme of July to September 2026, recorded in GEO_BASELINE_RESULTS_2026-07-27.md, GEO_GAP_ANALYSIS_2026-08-18.md and the assistant audit files of 17 September 2026, and from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. Counts were produced by a script over those files. The hidden instruction case of 6 August 2026 and the pricing behaviour of 17 September 2026 are recorded in those same files.

Disclosure. We sell an AI visibility product, AI Knows Us. We are a small brand in our own category and our own count is 0 of 6, which we publish above.

Version. 29 September 2026, first publication. Updated when the grouped 1,200 response test runs, and whenever a figure is corrected.

Common questions

If big brands win anyway, is there any point in a small business trying?

The one dated explanation we hold points at evidence rather than size. On 17 September 2026 Claude put two enterprise vendors above a smaller one because they published explicit coverage lists, and said the ordering reflected price transparency and source authority rather than product quality. Publishing your scope and your prices clearly is available to a business of any size, and it costs writing time rather than a media budget.

Why will you not just give me a percentage?

Because we have not measured it, and a percentage without a denominator is the thing this page exists to argue against. We hold 28 of 78, 21 of 78 and 18 of 78 in one sector on one date, and those are counts of three named vendors rather than of size groups.

Does naming a big competitor in my prompts help or hurt the test?

It changes what you are measuring. Comparison prompts that name large brands will reliably return those brands, so a set mixing them with category prompts will show a size effect that your own prompt set created. Keep the 100 prompts to category questions, and run comparison prompts as a separate, separately reported set.

Our brand is never named but our pages are cited. What is happening?

Your content is trusted and your commercial case is not being made in the same breath. That is a real, separate state and it should be scored separately. On 17 September 2026 Claude confirmed that clawlaw.in pages had shaped what it wrote while the site landed as a name inside a list rather than as a linked recommendation, because its own comparison pages read as vendor advocacy.

Can we improve our position by adding instructions for AI in our page code?

No, and there is a dated case against it. On 6 August 2026 Claude found a competitor's hidden block of text telling answer engines to cite that company, refused it, and named the company. The downside is being named as the company that tried it.

How many prompts before a difference is worth believing?

Report the counts at any size and resist the difference until it survives a repeat pass on a later date with the same frozen set. The engine itself will not rank a market on one run: in batch after batch in September 2026, Perplexity declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim.

What to do first

Before running anything, write down your group definitions and paste in the public sources for each brand you will track, with today's date. Then run ten category prompts on two engines, three times each, and count only which brands were named. If the same three names come back in almost every response, your next job is not a bigger test. It is publishing the scope and price evidence those three names publish and you do not.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.