How often do AI assistants recommend local accounting firms? A reproducible India benchmark, and what we have and have not run as of September 2026
The city and service grid, the own domain citation definition, the 324 observation arithmetic, and the nearest X of Y we actually hold.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Nobody knows yet, and the shape of the answer is a fraction: the number of independent runs in which an accounting firm's own domain was actually cited, out of the number of runs, with the date range and the test conditions beside it. We have not run that benchmark for accounting firms, so this page publishes no X of Y for the sector. What it publishes is the whole test: six cities crossed with six services, three assistants, three repetitions, which is 324 observations, with the exact scoring definitions and the confounds that make local answers unlike any other kind. The nearest real fractions we hold come from our own domain and from a legal technology company, not from an accounting firm: 0 of 6 on Perplexity in September 2026, and top source on 1 of 18 on ChatGPT on 18 August 2026.
The answer, first
Three things to fix in your head before running anything.
The outcome to count is a citation of the firm's own domain, not a mention. Mentions are easy to inflate and hard to compare. A URL on the firm's own website appearing in the sources of an answer is unambiguous, and it is the thing that sends a person to you.
Expect the official portals and the directories to win most of these answers, and count that rather than resenting it. We have this pattern on record in an adjacent sector. On 17 September 2026, on questions about how to use an official portal, Claude put official court portals above every commercial product and said plainly that no commercial product should rank above the official portal for a question about using that portal. A GST or income tax question has the same shape, and the sensible target is to be the page that explains the procedure and links the portal.
Read your own profession's rules first. In India, advertising and solicitation by chartered accountants are restricted under the Chartered Accountants Act 1949 and the ICAI Code of Ethics, and the ICAI has issued guidelines covering members' websites. Read the current text at the ICAI's own site, write down the date you read it, and take borderline questions to the Institute rather than to a marketing agency.
How it was measured
The query grid: six cities crossed with six services, which is 36 queries. Fix both lists in writing before you start.
- Six cities: Mumbai, Delhi, Bengaluru, Pune, Hyderabad, Ahmedabad. Substitute your own market, but keep six and keep them fixed.
- Six services: GST registration and monthly filing, income tax return for a small company, statutory audit, company incorporation and annual ROC filing, TDS compliance, and outsourced bookkeeping.
The query wording, fixed once. Two shapes per pair, to separate a recommendation question from a procedure question: "Who can help with [service] in [city]" and "How do I do [service], and should I hire someone in [city]". Freeze the wording in a dated file. Editing a prompt mid study ends the study.
The blind rule. No firm name in any prompt, wrapper, system instruction or earlier turn, including the firm running the test. The reason is on our record: on 27 July 2026 the same 78 questions were run twice in one day in the clawlaw.in programme, and the branded wrapper had ChatGPT ranking that brand first on almost every question while the blind run on the same day put it second by breadth and absent from the questions it most wanted to win. The branded run was discarded.
Three assistants, three repetitions, and the arithmetic in public. 36 queries times 3 assistants times 3 repetitions is 324 observations per wave. Report every count against 324, or against whatever your own grid produces, and print the arithmetic so a reader can rebuild it.
Test conditions, six of them, recorded per observation. These decide the result more than anything in the prompt.
- Fresh session, with memory and personalisation switched off.
- The location the session reports, and whether you set it or the network did.
- The interface used, because an app, a browser and an API answer differently.
- Whether a live search was visible, and the number of sources listed.
- The language of the query, and whether you also ran a Hindi or regional language version.
- The timestamp in UTC, because answers drift within a day.
Five scores per observation.
- Own domain cited. A URL on an accounting firm's own website appears in the sources. This is the primary outcome and the numerator of the X of Y.
- Any firm named. Whether the answer names any accounting firm at all, which in local queries is often no.
- Directory cited. Whether a business directory or listing site supplied the answer instead of a firm.
- Official source cited. Whether the answer pointed at the GST portal, the income tax portal, the MCA site or the ICAI. Count this as a correct answer to a procedure question, not as a loss.
- Accuracy. Whether everything the answer said about a named firm is correct, with any wrong claim copied out.
How to record search counts. Log what the interface showed, then ask the assistant, in the same session, how many of the queries it searched for, and store that as a claim in its own field. Do not reconcile the two. In three separate batches of our own September 2026 audit the assistant withdrew its earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported.
What the numbers were
Accounting firm runs completed: none, as of 29 September 2026. The 36 query grid has not been run on any assistant, so the X of Y for this sector does not exist and is not published here. Writing 0 of 324 would also be wrong, because 0 of 324 is a result and we have not taken the measurement.
The nearest fractions we do hold, with the sector stated so nobody transplants them.
- 0 of 6. Six questions about our own category, with no brand named, on Perplexity in September 2026. The assistant audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers, so there was no position for it to hold. That is our own domain, in software, not an accounting firm.
- 1 of 18. Eighteen blind commercial questions on ChatGPT on 18 August 2026, in which clawlaw.in was the top source on exactly one. That is legal technology, not an accounting firm.
- 161 recordings of not being cited, across 24 batches and 72 conversations on our own domain in September 2026, in the assistant's own audits of its answers. Repetition per question in all of these was one, not three, which is the weakness the protocol above is built to remove.
Why we publish the protocol without the result. An X of Y for accounting firms would be the most quotable number on this page, and it is precisely the number we do not have. Producing one would take about a week of careful work and about four minutes of invention, and only one of those two is worth reading.
What this cannot tell you
- It cannot tell you what a client in your city sees. Real sessions carry location, history and account context. The test switches those off deliberately, which makes it comparable and not lifelike.
- It cannot separate a local answer from a maps answer. Several assistants lean on map and listing data for "near me" questions, and a citation of a listing is not a citation of a firm.
- It cannot be compared with somebody else's visibility score unless they publish the grid, the blind rule, the repetitions and the five scoring definitions.
- It cannot tell you whether being cited brought work. That needs enquiry tracking on your own site, and it is a different measurement.
- It cannot promise a position. Nobody can guarantee a place in an assistant's answer. Disclosure: we are the vendor of AI Knows Us, which sells this kind of measurement.
- It is not compliance advice. What an ICAI member may publish is decided by the Institute's current rules, not by this page.
Sources and change log
Sources. The professional conduct material to read is the Chartered Accountants Act 1949, the ICAI Code of Ethics and the ICAI's guidelines on members' websites, at the Institute's own site, with your own date recorded. The dated results quoted come from the aiknowsus.com audit of September 2026, captured to geo-audits/aiknowsus-com/, and the clawlaw.in programme of July to September 2026, recorded in Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md, Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md and Tier_1/claude_response_17_09_audit.md.
Change log. 29 September 2026: first published with the 36 query grid, the 324 observation arithmetic, the five scores, the six recorded conditions, and the statement that no accounting firm run has been completed. When wave one runs, this section will carry the X of Y, the date range, the three assistants used and the raw rows, and this entry will stay in place so the two can be compared.
Common questions
Why not just ask ChatGPT who the best accountants in my city are and see?
Do that once, by all means, and then notice how little it tells you. One run on one assistant on one day is not repeatable, is affected by your own session history, and cannot show drift. Three repetitions across three assistants is the minimum that separates a stable answer from a lucky one.
Our firm cannot advertise. Is this test even useful to us?
Yes, because most of what wins these answers is not advertising. A clear procedural explanation of a filing, with the official portal linked and the dates and thresholds stated, is information rather than solicitation. Confirm the boundary with the ICAI's current rules first, and note that the test also tells you when the honest answer is that no firm is named at all.
What if the answers only ever name directories?
Then you have measured something worth knowing, and your action changes: get your listing details correct and consistent everywhere, because the directory is the source the answer is using. Count directory citations as their own score for exactly this reason.
Should we run this in Hindi as well?
If your clients ask in Hindi or in a regional language, yes, and keep it as a separate grid rather than mixing it into the English one. The sources an assistant reaches for can differ by language, and averaging the two hides that.
How long does one wave take?
Three hundred and twenty four observations scored by hand is roughly three days for one person. If that is too much, cut to three cities and three services, which is nine queries and 81 observations, and say so in your heading rather than implying a bigger study.
What counts as an independent run?
A fresh session, with no earlier turn, no memory, and no prompt reuse within the session. Two questions asked in the same conversation are not two independent runs, and treating them as two is the most common way these fractions get inflated.
What to do first
Write your six cities and six services into a dated file today and run one repetition on two assistants, which is 72 observations and about half a day. Count only own domain citations and whether any firm was named at all. That first count tells you whether your problem is ranking or whether, as in most local categories, the answer currently names no firm at all.