Do review sites change AI software recommendations? The 12 week audit method, and what our own captures show
How to count review site citations by named service, and the honest state of our own counts.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
In our own captures, review and directory sites were a small part of what actually reached the answer. On 17 September 2026, across six questions in the clawlaw.in programme, no third party review or directory source made it into any answer at all: one well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question. Counted across our whole September 2026 audit of 24 batches and 72 conversations, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Our records do not name which review service that was, so we cannot report a count by service, and we will not guess. This page publishes the 12 week audit design that would produce per service counts, and reports only what we hold.
The counted fraction we can state is 0 of 6 questions carrying any third party review or directory source into the answer, on Claude, on 17 September 2026, in one category in India. That is a small denominator and it is a real one.
The answer, first
Four statements, separated by how much support each has.
Review site citations were rare in our captures. Zero of six questions on 17 September 2026, and a review site placed far down the ordering across the September 2026 capture of 24 batches.
The reason recorded was relevance, not quality. The list the review site offered was of American products, so it was not an answer to an India question. For an Indian buyer's question that is the mechanism to understand: a large international listing can lose to a smaller local one on fit.
What did reach the answers was official pages and vendors' own pages. That is the ordering from the September 2026 capture, and on 17 September 2026 official court portals took every position above any commercial product on how to questions, with Claude stating that no commercial product should rank above the official portal for a question about using that portal.
What we cannot say. Whether reviews on a named service change a recommendation, in either direction, for any category. That needs the longitudinal design below, and we have not run it. We also cannot report our one counted fraction by service, because the record says a well known review site and does not name it.
How it was measured, and how to measure it properly
The design below is a 12 week longitudinal audit with per service counts. It is written so somebody with no stake in the answer could run it.
One. Name the services in advance. List every review and directory service you will count, by name, before collecting anything. In software that list usually includes the large international review services, the app stores, and the India specific directories for the category. Counting them separately is the whole point, because a combined review site figure hides which one matters in your market.
Two. Build the prompt set and freeze it. Forty to a hundred blind buyer prompts, in five recorded groups: category choice, comparison, price, suitability with a place or a sector attached, and risk or limits. No vendor name, no product name, no domain, no phrase that appears only on one vendor's site. Put a country in the prompts where a buyer would, because our own record shows an answer discarding a list for being about the wrong country.
Three. Fix the run conditions. One prompt per fresh conversation. Two assistants at least, reported separately. Location recorded. Search recorded per answer. Weekly for twelve weeks, same prompts every week.
Four. Define a recommendation before you count one. A recommendation is the part of the answer that tells the reader what to use or shortlist. A company named in a background paragraph is a mention, not a recommendation. Count both, in separate columns.
Five. Record citations at two levels. Per answer: did this answer cite each named service, yes or no. Per citation: every cited domain, so you also have the citation level denominator. Report the fraction both ways, because so many of so many answers and so many of so many citations are different numbers and both get quoted.
Six. Separate raw results from citations. If your method shows you what the engine looked at as well as what it used, keep them in different columns. On 17 September 2026 a review site appeared in the raw results and was discarded from the answer. A method that conflates the two would have reported a review citation that never reached the reader.
Seven. Add the change component, which is what makes it longitudinal. Pick a set of products and record, weekly, their review counts and ratings on each named service, alongside the citation counts. Twelve weeks of paired readings lets you ask whether a change on a service is followed by a change in citations. State in advance that this is observational: you are not assigning reviews, so a relationship found here is a correlation and nothing more.
Eight. Say what each result would mean. A service cited in many answers means a listing there is a source in your category and worth maintaining properly. A service cited in almost none, as in our own 0 of 6, means a listing there may still be useful to buyers and is not what decides these answers. Ratings moving with citations following is suggestive and confounded, because the same period includes product launches, news and engine changes. And if the searched share is low, review citations will be near zero for a reason that has nothing to do with reviews.
What the numbers were
The 12 week per service audit: not run. We have no weekly series, no per service fractions, and no paired review count and citation readings. We will not publish a figure of the form so many of so many recommendations cited a named service.
What we hold.
- 0 of 6 questions, Claude, 17 September 2026, clawlaw.in programme. Across six questions no third party review or directory source made it into any answer at all. One well known review site appeared in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question.
- Ordering across 24 batches and 72 conversations, September 2026, aiknowsus.com. The domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Which means the sources that decided these answers were mostly official pages and vendors' own pages rather than independent reviews.
- A directory standing in for a vendor's own site, 6 August 2026. The clawlaw.in pricing page rendered its prices only after scripts ran, so a crawler received no prices at all, and the prices the assistants did quote had come from an app store listing instead of the company's own site. An app store is a directory, and on that question it was the price source. In the same week ChatGPT noticed that the website and the app store listing carried different plan names and different prices for the same product, and said so in its answer.
- What did decide a coverage question, Claude, 17 September 2026. Two enterprise vendors ranked above the client specifically because they publish explicit court and tribunal coverage lists, with the engine noting that its ordering reflected price transparency and source authority rather than product quality. Not reviews. Published scope.
- Our own citation counts in the same period. Top source on 1 of 18 blind commercial questions, ChatGPT, 18 August 2026. Not cited or recommended in any of 6 questions, Perplexity, September 2026. Across 24 batches and 72 conversations, the phrase recording that we were not cited appears 161 times in the engines' own self audits.
Why we cannot split our one fraction by service. Our record describes a well known review site without naming it. Assigning that to a specific company now, to make the page look more precise, is exactly the kind of small invention that makes everything else on a site unreliable. So the fraction stands as 0 of 6 for third party review and directory sources combined, on one engine, on one date, in one category.
What this cannot tell you
- It cannot tell you whether reviews help. Absence of a citation is not evidence that reviews do not influence a buyer, a search ranking, or a model's memory of a category.
- Six questions is a very small denominator. It supports the sentence review citations were rare in this set and nothing broader.
- One category, one country, one period. Legal technology in India in September 2026. Software categories with a large international review culture may look completely different, and the way to find out is the audit, not an assumption.
- No per service split. Our record does not name the service, so any claim about a named review company would be invented.
- Correlation only, even in the full design. Nobody is assigning reviews at random, so a relationship between rating changes and citations would remain observational.
- Engine self reports are not instruments. In three separate batches in September 2026 Perplexity withdrew its own claim to have run a live search. Record what the engine says and verify against your own logs.
- No promise. Getting reviews does not guarantee a recommendation, and no tool including ours can promise a position in an answer.
Common questions
Should we still collect reviews on the big software review services?
Yes, for the reasons that existed before assistants: buyers read them, they surface in ordinary search, and they are part of how a category is described publicly. What our own captures do not support is treating them as the lever that decides AI recommendations. In our set the answers were decided by official pages and vendors' own pages.
Why would an assistant discard a big review site?
The recorded reason was fit. On 17 September 2026 the list the review site offered was of American products, so it was not an answer to an India question. That is worth knowing, because it means an accurate, India specific listing can outrank a much larger source on an India question.
Is a directory listing worth more than a review profile?
In one dated case a directory listing carried the company's prices into answers when its own site could not. On 6 August 2026 the prices the assistants quoted came from an app store listing because the pricing page rendered only after scripts ran. Treat that as an argument for fixing your own page and keeping listings consistent, not as a ranking of channels.
How many prompts before a review citation fraction means anything?
Forty or more per weekly wave, counted per answer and per citation, with the two denominators printed. At six questions you can say it was rare in that set, which is what we say.
What should we do if reviews are not the lever in our category?
Publish what the answers were actually using. In our own record that meant three things: a price that a crawler can read, a coverage list with names in it, and a source next to every figure about a competitor. On 17 September 2026 two comparison pages with correct competitor prices were discounted purely because there was no link, no date and no source at read time.
Can you run this audit for a category you do not sell into?
The design is public and category neutral, so anybody can. Disclosure: we sell AI Knows Us, which runs frozen blind prompt sets on a schedule and keeps the raw answers, which is the part that makes a twelve week series possible at all. The named review and directory services are real businesses and should be judged on their own pages; we do not restate their prices here, for the same reason we tell readers to distrust an undated price on a third party page.
Sources and change log
- Tier_1/claude_response_17_09_audit.md. The 17 September 2026 findings: no third party review or directory source in any of six answers, the review site discarded for offering American products, official portals above commercial products, the coverage lists that outranked the client, and the two comparison pages discounted as advocacy.
- geo-audits/aiknowsus-com/. September 2026, 24 batches and 72 conversations: the most cited domain ordering, the 161 not cited statements counted by script across all 24 second message answers, the batch of six with no citation, and the three batches where the engine withdrew its search claim.
- Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md. The 6 August 2026 findings on the unreadable pricing page, the prices sourced from an app store listing, and the two conflicting public price lists.
- Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md. The 18 blind commercial questions of 18 August 2026.
Change log. 27 July 2026 baseline. 6 August 2026 interim findings. 18 August 2026 commercial run. 17 September 2026 source selection audit, which is the source of the 0 of 6 fraction. September 2026 own domain audit. Not run: the 12 week per service audit, the weekly citation series, the paired review count readings, and therefore any per service fraction. When it is run this section will carry the service list, the weekly denominators, and the per service counts, including services cited zero times.
What to do first
Before you spend a quarter collecting reviews, spend an afternoon counting. Take twenty of your own buyer questions, run them blind on two assistants, and write down every cited domain. Then count how many answers cited each named review or directory service. If the count is near zero in your category, move the budget to the three things our own record shows deciding these answers: a readable dated price, a named coverage list, and a source next to every competitor figure. If the count is high, your listings are a source and they deserve the same care as your own pages.