Short answer
Measure AI visibility by asking a fixed set of blind buyer questions (never naming your brand) on each AI app, with live search switched on, many times over. For each question and AI app, count how often you are named (mention rate), how often your site is a source (citation rate), and where you appear in the list (position). Report each rate with a 95% range such as a Wilson interval. To judge whether your work caused a change, compare questions you worked on with a control group you did not touch.
On this page
Why this is harder than rank tracking
Ask an AI app the same question twice and you may get two different lists. This is called non-determinism. It comes from how models sample words, from different search results, and from changes the provider makes. A single screenshot therefore proves very little.
The fix is the same one used in any survey: take a sample, count, and show how uncertain the count is.
The core metrics
| Metric | What it counts | Why it matters |
|---|---|---|
| Mention rate | Share of checks where your brand, or a known alias, is named | The closest thing to "did the buyer hear about us" |
| Citation rate | Share of checks where a page on your domain is a source | Shows whether your own content is being used |
| Position | Where your name first appears in the list of names, counting from 1 | Being first in a list of five is not the same as being fifth |
| Share of voice | Your mentions divided by all brand mentions across the question set | Useful for tracking against named competitors |
| Searched flag | Whether the AI app actually ran a web search for that check | Separates fresh answers from memory |
Rule one: blind questions
A tracking question must never name your brand, domain or products. If it does, the AI app will talk about you because you asked, and your "visibility" will look far better than what a real buyer sees.
| Blind (good) | Not blind (bad) |
|---|---|
| What is the best billing app for a small pharmacy in India? | Is Tallybook a good billing app for a pharmacy? |
| Which inventory tools work offline? | Compare Tallybook with other inventory tools. |
Keep a small separate set of brand questions, such as "what does Tallybook do", to check accuracy. Do not mix them into your visibility numbers.
Rule two: force live search and record it
Ask each AI app to search before answering, and record from the response whether it did. OpenAI, Google and Anthropic each document how web search appears in their API responses. Keep answers from memory in a separate bucket. They reflect older training data, not your recent work.
Build a question universe
- Collect questions from sales calls, support tickets, reviews and your own search data.
- Group them by what the buyer is doing, such as looking, comparing, checking the price, or finding someone nearby.
- Tag each with a tier. Tier 1 is buying intent and winnable. Tier 3 is long tail.
- Check Tier 1 often, and the others less often, to control cost.
- Freeze the wording. Changing a question resets its history.
Show every rate with a range
A rate from a small sample is uncertain. The Wilson score interval is a standard way to put a 95% range around a proportion, and it behaves well with small samples and with rates near 0% or 100%.
Decide if a change is real
To compare two periods, use a two-proportion z-test. A common bar is a z score of 1.96 or more either way, which matches the 95% level.
Then add a control group. Pick similar questions you did not work on. If the treated questions rise and the controls stay flat, your work is the likely cause. If both rise, look for an outside cause such as a model update.
Know the limits
- API answers are not identical to what a logged-in person sees in an app, which may use memory or a different setup.
- Location, language and account history change real answers. You can set a country, but not reproduce every user.
- Models change. Log the model name returned with each answer.
- Your questions are a sample. Better questions give better numbers.
How we help
These rules are on by default in AI Knows Us: blind questions, forced search with a searched flag, tiered checks, Wilson ranges on every rate, and a two-proportion test with control groups for experiments. The full method is on our methodology page.
Frequently asked questions
How many checks do I need?
There is no single number. The range narrows as checks add up. Look at the width of the Wilson range and keep checking until it is narrow enough for your decision.
Why not just use a screenshot?
One answer is a sample of one. The next run may differ. Rates with ranges are the only fair way to compare.
What is a good mention rate?
It depends on the category and question. Compare yourself with named competitors on the same questions, and track your own trend, rather than chasing a universal target.
Should I track every AI app together?
Track them separately. Each AI app searches and writes differently, so a blended number can hide a big gap on one AI app.
What if the AI app did not search?
Record it and keep those answers separate. They reflect training data and will not show your recent changes.
Sources
Last verified: 17 September 2026. Facts on this page come from the public documents below. Brands such as Tallybook are fictional and used only as examples.