AI visibility reports explained: a metric dictionary with worked examples
Fourteen metrics defined by their denominators, and how to read a score that has none.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Read an AI visibility report from the bottom up, not the top down. Find the sample first: how many prompts, how many repeat runs, which engines, which location, which dates. Then read the score. Every rate in a visibility report is a fraction, and a fraction without its denominator is not a measurement of anything. There is no universal percentage that a good report contains, so the thing to demand is the denominator behind every single number on the page, printed next to it. If a report gives you a visibility score of any size and will not tell you how many answer runs it was calculated over, the correct response is to stop reading it.
This page defines fourteen metrics by their denominators, shows how the same underlying result can be presented as three different numbers, and works through real examples from our own runs with their dates and engines attached.
Start with the sample, not the headline score
Six questions, in this order, before you look at any chart. If the report cannot answer the first four, nothing else in it can be interpreted.
- How many prompts, and can I see them? The exact text, not a description of themes.
- Does any prompt contain our brand name? If yes, the score is inflated. On 27 July 2026 we ran the same 78 questions twice on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent from the litigation due diligence questions it most wanted to win. We discarded the branded run, because it had measured our own prompt.
- How many repeat runs per prompt? One run is a draw from a distribution, not a behaviour.
- Which engines, reported separately or blended? A blend hides which engine you are absent from.
- Which location, language and signed in state? Answers differ by all three.
- Which dates, and was the prompt set the same as last month? A changed prompt set makes a trend line meaningless.
The metric dictionary
Fourteen metrics, each with the denominator that defines it. Where two reports use the same name for different denominators, that is the source of almost every disagreement between tools.
- Mention rate. Runs where your brand name appears in the answer text, divided by total answer runs. Denominator: answer runs, which is prompts times engines times repeats.
- Prompt coverage. Prompts where your brand appeared in at least one run, divided by total prompts. Denominator: prompts. This is always the flattering version of mention rate, because one lucky run out of three counts the whole prompt.
- Citation rate, or on domain citation rate. Runs where a URL on your own domain appeared as a source, divided by total answer runs. The strongest and strictest of the presence metrics.
- Top source rate. Runs where your page was the first or most heavily used source, divided by total answer runs.
- Link share. Your cited URLs divided by all cited URLs across the capture. Denominator: total citations, not runs. Easy to confuse with citation rate and quite different.
- Share of voice. Your mentions divided by all brand mentions across the capture. Denominator: brand mentions. It moves when a competitor gets more mentions even if you do not change at all.
- Position or rank when named. Where in the list you appear, averaged only over the runs where you appeared. Denominator: runs where you were named, which is not the same as total runs, and reporting it against the wrong one is a common error.
- Sentiment. How favourably you are described. Denominator: runs where you were named. Worth having and the least reliable, because the labelling is a judgement.
- Answer presence rate. Runs where any answer of the relevant type was produced, divided by prompts tried. On some surfaces no answer appears for some prompts, and those rows must be counted, not dropped.
- Search rate, claimed. Runs where the engine said it had searched, divided by total runs, and marked unverified. In three separate batches of our September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question.
- Competitor set size. Distinct competitor brands named across the capture, divided by nothing. It is a count, not a rate, and reporting it as a percentage of anything is meaningless.
- Source type mix. Citations by source type, divided by total citations. Across the whole aiknowsus.com capture of September 2026 the most cited domains were the assistants' own documentation and the vendors' own websites, with a single well known review site far down the list.
- Use without attribution. Runs where your material appears and your name and URL do not, divided by total runs. Under counted by definition, and real: on 17 September 2026 Claude admitted using two specific arguments from clawlaw.in pages and dropping the attribution in both places.
- Inter rater agreement. Runs where two labellers agreed, divided by runs double labelled. Not a visibility metric, and the one that tells you whether any of the others can be trusted.
Check the denominator, engine, location, and date
Here is why the denominator is not a technicality. One underlying result can be reported three ways, all arithmetically correct, and they will not sound like the same thing.
Consider a capture of 10 prompts, on 1 engine, with 3 repeats. That is 30 answer runs. Your brand appeared in 6 of those 30 runs, spread across 4 different prompts. Three honest numbers follow: a mention rate of 6 of 30, a prompt coverage of 4 of 10, and, if somebody chooses to report only the prompts where you appeared at all, a claim that you appear for 40 per cent of buying questions. That last version is the one that ends up on a slide. It is the same six runs. This paragraph is arithmetic, not an observation, and it is the arithmetic that most visibility disputes turn out to be about.
Four more checks that matter as much as the denominator.
- Engine. Never accept a blended score. Ask for the same table per engine and you will usually find one engine carrying the whole result.
- Location and language. A report for India run from a default location elsewhere is a report about somewhere else.
- Date range, and whether the prompt set version changed inside it. If it changed, the trend is between two different instruments.
- Who labelled it. A person, a model, or a keyword match. A keyword match will score a mention of a similarly named company as you.
Worked examples
Four real ones from our own programmes, written the way a report should write them, with the engine and date in every line.
Example one, the correct way to report a zero. Citation rate 0 of 6 answer runs, which is 0 per cent. Six blind prompts about our own category, one engine, Perplexity, September 2026, on aiknowsus.com. The engine audited its own answers afterwards and reported it had not cited or recommended us in any of the six, so there was no position for us to hold. Notice what that sentence contains: numerator, denominator, blind rule, engine, date, domain and the source of the judgement.
Example two, a rate that needs its context to be honest. Mention rate 21 of 78 prompts, ChatGPT, 27 July 2026, clawlaw.in, blind. Competitor comparison from the same run: ProVakil 28 of 78, Legistify 18 of 78. And the context that changes its meaning entirely: on the same run, ChatGPT confirmed it had run no live web search for any of the 78 questions. So that reading is a measurement of what the model had absorbed, not of the web that week. A report that printed 21 of 78 without that line would be misleading while being arithmetically correct.
Example three, the gap between two metrics on the same business. Top source rate 1 of 18 prompts, ChatGPT, 18 August 2026, clawlaw.in, eighteen blind commercial questions. Compare it with the mention figure above. Being named is much easier than being the deciding source, and a report that only shows mention rate will make a business feel further along than it is. In that single winning answer the engine also named the company's own comparison page as still being the vendor's own editorial page, which is the qualitative note that explains the quantitative result.
Example four, a result that a rate cannot capture. On 6 August 2026 ChatGPT cited clawlaw.in/blog/how-to-check-a-companys-court-cases-in-india as a source for a vendor due diligence question, fourteen days after that page was published, in a zone where the same question set had named the company nowhere at baseline on 27 July 2026. One URL, one engine, one date. It is the only figure in that programme with a public URL behind it, and it is not a median, a rate or a trend. A good report carries observations like this one as named rows rather than folding them into a percentage.
What the report cannot prove
Seven things, and any report claiming otherwise is overreaching.
- It cannot prove causation. A number that moved after you published pages is consistent with the pages working and with the engine changing.
- It cannot prove market share of attention. Your prompt set is not a sample of what buyers asked.
- It cannot be compared with another company's published score, because you do not have their prompt set, engines, location or repeat count.
- It cannot see sources used and not displayed, documented above with a date.
- It cannot separate an engine change from a web change when a competitor appears or disappears.
- It cannot validate its own labelling without a second labeller and a published agreement rate.
- It cannot promise a position in any answer, and no vendor, ours included, should offer that.
There is one more limit worth quoting to anybody who wants a single winner declared from a single run. In batch after batch of the aiknowsus.com capture of September 2026, the assistant declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is an engine refusing to do what many dashboards do every month.
Common questions
Which single metric should I put in front of my management?
On domain citation rate, as a fraction over answer runs, per engine, with the date. It is the strictest of the presence metrics, it is the hardest to inflate, and it corresponds to something a buyer can actually click.
Two tools give us different visibility scores. Which is right?
Possibly both. Different prompt sets, different repeat counts, different locations and different labelling rules produce different numbers from the same reality. Pick one method, freeze it, and measure your own movement rather than trying to reconcile the two.
Is share of voice worth tracking?
Only alongside your own absolute counts, because it moves when competitors move. A share of voice that falls while your citation count rises means your market got noisier, which is a different problem from your pages failing.
Should sentiment be in the report?
Include it, clearly labelled as a judgement, with its denominator being the runs where you were named. It is useful for spotting a specific negative characterisation early, and it is the metric most likely to be wrong.
Our tool reports a score out of 100. What do I do with it?
Ask what the inputs are and how they are weighted. If the vendor cannot show you the formula and the underlying fractions, treat the score as an internal index for tracking your own movement, and never quote it externally as a fact about your market.
How small a sample is too small to report?
No sample is too small to report, as long as the denominator is printed beside it. 0 of 6 is a perfectly honest sentence. What is not honest is turning 0 of 6 into a percentage, dropping the six, and putting it on a chart with a trend line.
Download the example dataset and calculation sheet
The calculation sheet is the fourteen definitions above, each with its denominator, plus one row per answer run carrying the label columns, plus a formula row that computes each rate from those columns so nobody types a number by hand. That structure is the whole point: a rate that is typed rather than computed is a rate that will eventually disagree with its own data.
Our own capture files behind the four examples are being prepared for publication, at which point the 0 of 6 and the 161 not cited statements across 24 batches and 72 conversations become checkable row by row. Sources: the clawlaw.in programme, GEO_BASELINE_RESULTS_2026-07-27.md and GEO_GAP_ANALYSIS_2026-08-18.md plus the assistant audit files of 17 September 2026, and the aiknowsus.com audit of September 2026. Version: 29 September 2026, first publication.
What to do first
Open your most recent visibility report and try to write one sentence in the form of example one above: metric, numerator, denominator, engine, date, blind or not. If you cannot complete that sentence from what the report contains, you have found the thing to fix, and it is not your website.