Why one AI app is not enough to measure with
The same question, the same day, two assistants, two different shortlists.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Yes, ChatGPT and Gemini regularly give different recommendations for the same question on the same day, and the difference is often large. So a number from one app is not your AI visibility. It is your visibility in one app, which is a smaller and less useful thing.
What is actually happening underneath
The assistants differ in five ways that all affect which companies get named.
Where the pages come from. Gemini is grounded in Google's index and reaches for what Google already ranks. ChatGPT uses its own search path. Perplexity searches for almost everything and shows its sources by default. Different indexes, different candidate pages.
Whether it searched at all. Some answers are fetched, some are recalled from training. A recalled answer favours companies that have been written about for years. A fetched answer favours pages that exist now and answer the question.
What each model was trained on, and when. The cut off dates differ, so the same old fact can be current in one and stale in another.
How the answer is shaped. Some apps give a tidy numbered list of three. Some give prose with links. Being named in prose and being third in a list are not the same result, and a log that treats them as identical loses the distinction.
How much the account changes things. Saved memory, history and custom instructions shape answers differently in each app, so the same person can get very different readings from two apps while signed in.
Which assistants to include, and why
You do not need all of them. You need enough variety that the differences show up. Three kinds are worth covering, and a fourth is worth knowing about.
One that reaches through Google. Gemini, and Google's own AI answers in search. This is the column that turns an indexing problem into something you can see.
One with its own search path. ChatGPT. It also sometimes does not search at all, which makes it the app where the difference between remembered and fetched visibility is easiest to observe.
One that searches for nearly everything and cites by default. Perplexity is the usual choice. Because it almost always shows sources, it is the cheapest way to collect a map of which third party pages your category is answered from.
The assistants built into other products. Copilot, and the assistants inside browsers and phones. Worth checking occasionally rather than tracking monthly, because they change quickly and there is less public detail about what they retrieve.
Two is a real measurement. Three is comfortable. Adding a fourth adds less than adding ten more questions to your set.
A worked example
On 6 August 2026 two assistants read the same pricing page, on clawlaw.in, our sister company. They did not report the same problem, and neither report contained the other.
In the Claude run, the prices on that page appeared only after scripts ran, so what a crawler received contained no prices at all. The prices the assistants did quote had come from an app store listing instead of from the company's own site.
In the ChatGPT run the finding was a different one. The website and the app store listing carried different plan names and different prices for the same product, and ChatGPT noticed the contradiction and said so in its answer.
One reading would have produced one of those two facts. With only the Claude run you would have gone to the developers about rendering. With only the ChatGPT run you would have gone to whoever writes the pricing copy. Both were true on the same day, about the same page, and only the pair describes the whole problem.
What it means for a business
Three things follow, and the third is the one people get wrong.
First, one app is one opinion. If you are strong in ChatGPT and absent in Gemini, you have a Google indexing problem, and measuring only ChatGPT would have hidden it.
Second, the fixes are not identical. Gemini rewards being indexed and ranked by Google. ChatGPT rewards being present on the third party pages its search path reaches. The overlap is large but not total.
Third, do not average the apps into one headline figure without saying so. A blended number hides exactly the difference you needed to see. Keep a per app column, then add a total if you want one, and label it as a total.
What this approach does not solve
Adding apps improves coverage. It does not remove three problems, and one of them it makes slightly worse.
Variation stays. Each app still answers the same question differently between runs. More apps means more rows, not a steadier number.
Weighting is guesswork. You do not know which app your buyers use, so any weighted blend of the apps is a guess dressed as a metric. Report them side by side instead.
Comparability across apps is limited. One app returns a ranked list of three, another returns prose. Position two does not mean the same thing in both, so record what you saw rather than forcing a shared scale.
Effort per reading rises. Three apps is three times the work by hand. If that means the reading stops happening, two apps done every month beats four apps done once.
Common questions
Which two should I pick if I can only do two?
One that reaches through Google and one that does not. That pairing gives you the diagnosis you cannot get any other way: whether your problem is indexing and ranking, or presence on third party pages.
Do I need Perplexity if I already have ChatGPT and Gemini?
Not for the rate. It is useful as a source finder, because it cites by default, so one pass through your questions there produces a fuller map of the pages your category is answered from.
Should I use the paid tiers?
Use whatever your buyers are likely to use, and keep it the same every month. Paid tiers can behave differently, especially about searching and about answer length. Mixing tiers between readings is a change you will mistake for a result.
The two apps disagree completely. Which one is right?
Both are right about themselves. There is no true answer they are approximating. Disagreement is the finding, and it usually points to one retrieval path you are present in and one you are not.
Is it worth tracking Google's AI answers in search separately from Gemini?
Yes, if you have the time, because they are different surfaces and they do not always name the same companies. Treat them as two columns and do not merge them.
Can I automate this with the APIs instead?
You can, and it is steadier and cheaper to repeat. The catch is that an API can behave differently from the app your buyer uses, particularly about whether it searches. Label API rows as API rows and never compare them with app rows in the same total.
What to do first
- Pick at least two assistants, three if you have time. Include one that is grounded in Google and one that is not.
- Ask the identical question text in each. Do not reword it to suit the app.
- Log per app. Same sheet, extra column.
- Look at the gap between apps before you look at the total. The gap is the diagnosis.
- On the app where you are weakest, open one answer and read the cited pages. That is where your next month of work is.
We run our own readings across 5 assistants for this reason, and our case study sets are built that way. It is also the part of the work that a person checking by hand can copy exactly: open two tabs instead of one.