Collection, one AI app at a time
We check the real ChatGPT and Gemini apps, the same ones your buyers use, through a third-party data provider. Perplexity and Claude are checked through their official APIs, with live web search. Each AI app is asked to search the web before it answers.
| AI app | Source |
|---|---|
| ChatGPT | The real ChatGPT app, through a third-party data provider |
| Gemini | The real Gemini app, through a third-party data provider |
| Perplexity | Official Perplexity API (Agent API) |
| Claude | Official Anthropic API, with live web search |
| AI Overviews | Licensed search-results provider |
Google AI Overviews data comes from a licensed search-results provider, which records the overview shown on a Google results page for the question.
Model versions
Each AI app is called with a fixed model setting where the source allows one. The model name reported back is logged with every answer whenever the source provides it. When a provider changes its model, the change shows in your history, so you can tell a model update apart from a real change in your visibility.
The blind question rule
A tracking question never names your brand, your domain or your products. If we named you, the AI app would talk about you because we asked, and the result would say nothing about what a buyer hears.
Every tracking check sends one message with the same structure:
- It says the asker is a buyer in one country: the first country on your list of places. If you sell in more than one, we check only that first one, and the others are not measured at all. AI apps answer differently in different countries, so a result from one is not a result from another.
- It asks the AI app to run a live web search first, not to answer from memory, and to say so on the first line if it could not search.
- It gives the buyer question, word for word.
- It asks for a normal answer that names the specific products or companies the AI app would recommend, in order, with one line on why, using only public information.
- It asks for two closing lines: one listing every name given, in order, and one listing every source domain used.
Only a small, separate set of brand questions may name you. These are used to check what AI apps say about you, not whether they recommend you.
Forced live search and the “searched” flag
Answers from memory reflect old training data. Answers after a search reflect the web today, which is what your work can change. So each check records a searched flag, read from the response: for example, whether a web search call was made or grounding queries were returned. Checks where the AI app did not search are marked, so you can see them separately.
How often we check
| Priority | Questions | Checked |
|---|---|---|
| Priority 1 | Buying intent and winnable. About 15% of questions. | Weekly (every 7 days) |
| Priority 2 | Important, but less urgent. | Monthly (every 30 days) |
| Priority 3 | Long tail. | Quarterly (every 90 days) |
Why weekly, not daily
Buyers and AI apps do not change their minds overnight. The same question can get a slightly different answer minutes apart, but that is random variation, not a trend. The things that move your visibility, such as a page being indexed, new reviews, a place on a best-of list or a competitor's new page, show up over weeks.
Daily single checks mostly record that day-to-day noise, and they multiply the cost. We check weekly, combine the checks into rates with 95% ranges, and only call a change real when it beats the noise. You can still press Run checks at any time, for example right after you publish a page or fix a wrong fact.
What we record
- Mentioned: the AI app named your brand, or one of its known aliases, in the answer. A vague description does not count.
- Cited: the AI app used a page on your own domain as a source.
- Position: where your name first appears in the list of names the AI app gave, counting from 1.
- We also record the competitors named and every source domain. Checking each answer against your fact sheet for wrong claims is on our roadmap.
95% ranges (Wilson)
A rate such as “named in 30% of checks” is shown with a 95% range, using the Wilson score interval. The Wilson method behaves well with small samples and with rates near 0% or 100%, where simpler methods give ranges that go below zero or above 100%.
Example: named in 12 of 40 checks is 30%, with a 95% range of 18% to 45%. The range narrows as checks add up.
Calling a change real
To compare two periods, or the questions we worked on with the ones we left alone, we use a two-proportion z-test. A change is called real only when the test clears the 95% level (a z score of 1.96 or more either way). Anything less is labelled “Not distinguishable from noise”, and we keep measuring.
Experiments compare questions your work targeted with similar questions it did not. If both groups rise, the cause was probably something outside your work, such as a model update.
Reusing a recent answer
One AI answer can be used for more than one customer, and each customer's business is scored separately against it. If the same question has been asked on the same AI app for the same country recently, we reuse the answer we already collected instead of asking again. The window is half the question's check frequency: about 3.5 days for Priority 1, 15 days for Priority 2 and 45 days for Priority 3. Question text is matched exactly, after trimming and ignoring case.
Known limits
- Logged-out sessions. ChatGPT and Gemini answers come from the real apps, in a logged-out session in the country you choose. They can differ from what a logged-in user with memory and chat history sees. That is true of any tracking tool.
- Personalisation. Real users see answers shaped by location, language, history and account. We set a country, but we cannot reproduce every user.
- Non-determinism. The same question can get a different answer a minute later. That is why we use many checks and show ranges, not single screenshots.
- AI Overviews. This data comes through a licensed search-results provider, not from Google directly. Coverage and timing depend on that provider.
- Timing. New pages typically take 4 to 8 weeks to show on Perplexity and 2 to 4 months on ChatGPT. That is typical, not guaranteed.
- Question choice. Results are only as good as the questions tracked. We write them blind and by buyer job, but they are still a sample of what buyers ask.
Questions about the method? Contact us. See also how the product works.