Why a handful of questions cannot give you a percentage
Five questions give you five answers. They do not give you a rate.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
You need enough questions to cover every kind of buying question you sell into, and more than a handful before a percentage means anything: with five questions you have five answers, not a rate. Two separate problems make a small set unreliable, and adding more readings of the same five fixes neither.
What is actually happening underneath
Problem one is coverage. Five questions cannot represent four kinds of buying question across several cities, customer sizes and product lines. Whatever number comes out describes the five things you happened to ask. If all five were problem questions, you have measured your position on problem questions and learned nothing about category questions, which is usually where the money is.
Problem two is variation. Ask the same question twice and the answer can differ. So each reading is one draw out of several possible ones, and a small number of draws produces a number that moves on its own.
These are different problems with different fixes. Coverage is fixed by adding questions of kinds you are missing. Variation is fixed by asking more questions and repeating them. Neither is fixed by reading the same five questions every day.
Batch 01 has the arithmetic of the second problem: how a range is calculated on a rate, and why 2 of 5 honestly means somewhere between very low and quite high. Read it before you report a change to anybody, because the most common misleading report in this field is a small sample run twice.
How to size a set without any statistics
You do not need a formula. You need a grid and one rule.
List your zones. A zone is a part of the business where you could plausibly be strong or absent independently of the others: a city, a customer size, a product line, a buyer role. Most businesses have between three and six.
Put five to ten questions in each zone. Spread across the four question kinds so no zone is measured only on the easy ones.
Then apply the rule. Report per zone as counts, and only report a percentage for the whole set once the whole set is large. If the biggest number you can honestly write is 4 of 30, write that.
This gives you a set that is defensible for a reason that has nothing to do with statistics: every number you publish has a named denominator and a named scope, so anybody can check what it describes.
A worked example
We have one count of our own that only a large capture could have produced. Across the whole September 2026 capture on aiknowsus.com, our own site, 24 batches and 72 conversations with Perplexity, we counted which tools were named most often in our category.
In order, they were Semrush, then Profound, then Peec, then Otterly, then Scrunch. The established search tool appeared alongside the specialist ones rather than below them, which is not what we expected before counting.
That ordering needed the whole capture. On a handful of questions the same names come back in a different order almost every time, because each answer names only a few tools and any one of them can be missing from a given draw. The order is a property of the set, not of any single answer in it.
So the rule on this page holds in both directions. A small set cannot give you a percentage, and it cannot give you a ranking either. What it can give you is a list of names worth counting, which is a good enough reason to run a small set now and grow it.
What it means for a business
Until the set is large, report counts and not percentages. Named in 2 of 5 answers is honest. Forty per cent implies a precision you do not have, and it invites a comparison next month that the data cannot support.
Build the set by zone rather than by total. Then you can report per zone, which is far more useful than one blended figure. A company can be strong in one zone and invisible in another, and a single number averages that insight away.
There is a real cost to a bigger set, in time or in credits. The honest trade off is fewer readings of a larger set rather than frequent readings of a tiny one.
What a larger set does not fix
Four things, because the answer to every problem here is not simply more questions.
A biased set stays biased. Thirty easy questions produce a flattering number just as five do. Coverage is about which questions, not only how many.
It does not remove variation. It narrows the range. Individual answers still differ between runs, and a single question is still an anecdote.
It does not make tools comparable. Another vendor's percentage came from another set with another definition of a mention. Size does not make two different measurements the same one.
It does not tell you about demand. A larger sample of questions you invented is still questions you invented. Nobody knows how often buyers ask them.
And there is a practical ceiling. A set so large that the reading stops happening is worse than a moderate set read every month.
Common questions
What is the minimum number of questions?
There is no threshold that turns a set from wrong to right, and any single number quoted as the minimum is invented. The useful floor is coverage: five to ten questions in every zone you sell into. If that produces twenty five questions, you have twenty five.
Is it better to ask five questions ten times or fifty questions once?
Fifty questions once, for coverage, which is usually the bigger problem. Repetition helps with variation, so the ideal is a larger set repeated monthly rather than either extreme.
Can I report a percentage if my board insists?
Print the two numbers next to it and one line saying the set is a sample of your own questions. That line is what stops a normal wobble being read as a collapse, and it is the difference between a number that survives and one that gets discarded.
How do I explain a small set honestly?
Write the scope into the sentence: named in 4 of 20 answers, from 10 questions on 2 assistants, covering two of our four zones, brand questions excluded. Anybody reading that knows exactly what it does and does not describe.
Does asking the same question twice in a reading help?
It gives you a feel for how much the answer moves, which is worth doing once to calibrate expectations. As a routine it costs more than it returns. Spend the same effort on more questions.
Should different zones have the same number of questions?
Roughly, unless one zone is much more of your revenue. Weight by importance if you must, and write down the weighting, because an unexplained weighting is indistinguishable from a mistake.
What to do first
- List your zones. Most businesses have between three and six.
- Write five to ten questions per zone, in buyer wording, across the four question kinds.
- Report per zone, as counts, with the denominator next to each.
- Add a one line scope statement to the top of the report, naming the set, the assistants and the exclusions.
- If you can only afford a small set, keep it and say plainly in the report that it is too small for a percentage.
We print a range next to every rate we publish, which sometimes makes our own results look less impressive than a competitor quoting a bare percentage. It is the only version that is true.