AI recommendations for recruitment agencies: a dated, reproducible India benchmark by role and city
The role and city grid, separate employer and candidate waves, the own domain citation definition, and an honest account of what has not been run.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
The answer has to be a fraction with its conditions attached: X of Y independent runs in which a recruitment agency's own domain was cited, with the role and city mix and the collection dates beside it. We have not run that benchmark, so no fraction for this sector appears on this page. What appears instead is the whole design: eight roles crossed with six cities, which is 48 queries, run separately for employers and for candidates because they are different questions, on three assistants with three repetitions, which is 432 observations per wave and 864 across both. The nearest real fractions we hold are from our own domain and from a legal technology company: 0 of 6 on Perplexity in September 2026, and top source on 1 of 18 on ChatGPT on 18 August 2026.
The answer, first
Four things decide whether this benchmark tells you anything.
Count own domain citations, not mentions. A URL on the agency's own site appearing in the answer's sources is unambiguous and comparable. "We got mentioned" is neither.
Run employers and candidates as two separate waves and never pool them. An employer asks who can fill a vacancy. A candidate asks how to get hired. They produce different answers from different sources, and one combined number describes neither.
Expect the large job platforms to supply most of these answers. Naukri, LinkedIn and Indeed are established platforms in India with far more indexed pages about roles and cities than any agency has, so a benchmark that only counts agencies will look like a flat zero and teach you nothing. Score platform citations as their own column.
Check the licensing position before you publish anything about overseas placement. Recruitment for overseas employment from India requires a recruiting agent licence from the Ministry of External Affairs, and the eMigrate system sits behind that process. What may be charged to a candidate is set by those rules, not by the agency. If you place workers abroad, your licence number and registration are the most checkable facts you own, and the rules are the first thing to read at the Ministry's own site, with your date recorded.
How it was measured
The grid: eight roles crossed with six cities, which is 48 queries. Fix both lists in a dated file before running anything.
- Eight roles, chosen to span the market rather than to flatter one desk: software engineer, sales manager, accountant, HR manager, customer support executive, warehouse supervisor, civil site engineer, and staff nurse.
- Six cities: Mumbai, Delhi, Bengaluru, Pune, Hyderabad, Ahmedabad. Substitute your own markets, keep the count fixed, and say what you changed.
Two waves with different wording, scored separately.
- Employer wave: "We need to hire a [role] in [city]. Which recruitment agencies can help, and what do they charge."
- Candidate wave: "I am looking for a [role] job in [city]. Should I go through a recruitment agency, and which ones."
The blind rule. Your agency's name never appears in any prompt, wrapper, system instruction or earlier turn. On 27 July 2026, in the clawlaw.in programme, the same 78 questions were run twice in one day: the run with a wrapper naming the brand had ChatGPT ranking it first on almost every question, and the blind run on the same day put it second by breadth and absent altogether from the questions it most wanted to win. The branded run was discarded, because the only thing it had measured was our own prompt.
The arithmetic, printed on the page. 48 queries times 3 assistants times 3 repetitions is 432 observations in one wave. Two waves is 864. Every count divides by a number a reader can rebuild.
Conditions recorded per observation, six of them. Fresh session with memory and personalisation off. The location the session reports. The interface used, app, browser or API. Whether a live search was visible and how many sources were listed. The language of the query. The UTC timestamp, because answers drift within a day and hiring queries are seasonal within a quarter.
Six scores per observation.
- Own domain cited. A URL on a recruitment agency's own website appears in the sources. This is the numerator.
- Any agency named. Whether the answer names any agency at all.
- Job platform cited. Whether a large job platform supplied the answer instead.
- Fee information present. Whether the answer says anything about what agencies charge, and whether it is sourced.
- Official source cited, which for overseas placement means the Ministry's own material or eMigrate.
- Accuracy. Whether every claim about a named agency is correct, with wrong claims copied out.
Recording the search claim. Log the visible sources, then ask the assistant afterwards how many of the queries it searched for, and store that separately as a claim. In three separate batches of our own September 2026 audit the assistant withdrew its earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question, and in another batch that its claim to have searched all five was not adequately supported.
What the numbers were
Recruitment agency runs completed: none, as of 29 September 2026. Neither wave has been run, so there is no X of Y for this sector, no role and city mix to report and no collection dates. That sentence is the result this page has.
The nearest fractions we hold, with their sectors named so nobody transplants them.
- 0 of 6, our own domain in our own software category. Six questions with no brand named, on Perplexity in September 2026. The assistant audited itself afterwards and reported that it had not cited or recommended aiknowsus.com in any of the six answers, so there was no position for it to hold.
- 1 of 18, legal technology. Eighteen blind commercial questions on ChatGPT on 18 August 2026, in which clawlaw.in was the top source on exactly one, with its own comparison page named in the answer as still being the vendor's own editorial page.
- 161 recordings of not being cited, across 24 batches and 72 conversations on our own domain in September 2026, counted in the assistant's own audits of its answers.
The weakness in all three: repetition of one. Every count above came from a single run per question. That is exactly why the protocol specifies three, and it is why none of those three numbers should be described as a rate.
The methodological finding worth more than any of them. In batch after batch of that September 2026 audit the assistant declined to name a competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. That is the clearest statement we hold that a single run does not establish a ranking, and it came from the machine being measured.
What this cannot tell you
- It cannot tell you what an employer in your city sees, because real sessions carry location and history that the test deliberately switches off.
- It cannot tell you whether a citation produced a mandate. That needs enquiry tracking on your own site and is a separate measurement.
- It cannot handle role titles cleanly. One market's sales manager is another's business development lead, and the grid freezes eight titles rather than solving that problem.
- It cannot see seasonality in one wave. Hiring questions move with quarters and with campus cycles, so a single wave is a photograph.
- It cannot be compared with a vendor's visibility score that does not publish its grid, blind rule, repetitions and scoring definitions.
- It cannot promise a position. We do not guarantee a place in any assistant's answer. Disclosure: we are the vendor of AI Knows Us, which sells this measurement.
Sources and change log
Sources. For overseas placement rules, the Ministry of External Affairs material on recruiting agent licensing and the eMigrate system, read at the Ministry's own site with your date recorded. For the dated results, the aiknowsus.com audit of September 2026 captured to geo-audits/aiknowsus-com/, and the clawlaw.in programme of July to September 2026 recorded in Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md and Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md.
Change log. 29 September 2026: first published with the eight role by six city grid, the two separate waves, the 432 and 864 observation arithmetic, the six scores, and the statement that neither wave has been run. When wave one runs, this section will carry the X of Y, the role and city mix, the collection dates, the assistants used and the raw rows.
Common questions
Why separate employer and candidate queries at all?
Because they are answered from different sources and they mean different things commercially. An agency that is cited for candidate queries and never for employer queries has a real problem that a pooled number would hide completely.
If job platforms win everything, what is the point?
The point is to find the questions where they do not. Fee structures, notice period practice, replacement guarantees, salary bands for a specific role in a specific city, and how a contract to hire arrangement works are all questions the platforms answer badly and an agency can answer precisely. Score platform citations so you can see which queries are worth contesting.
Should we publish our fees?
Publish the structure with its conditions: the basis, what triggers the invoice, the replacement or guarantee period, and what changes the number. In our own dated work, price transparency has been named by an assistant as part of why it ordered vendors as it did, on 17 September 2026, and unsourced price claims on two pages were discounted as advocacy even though they turned out to be correct.
How do we handle the same role having three different titles?
Freeze eight titles for the benchmark so results stay comparable, then run a separate small set of synonym queries as an experiment. Do not swap titles mid wave, because the wave stops being one wave.
How long does a wave take?
Four hundred and thirty two observations scored by hand is about three days for one person. A reduced grid of four roles and three cities is 12 queries and 108 observations, which is a day, and the honest thing is to say in your heading that it was four roles and three cities.
What counts as an independent run?
A fresh session with no earlier turn, no memory and no reuse of a prompt within the session. Two questions in one conversation are one run, and counting them as two is the most common way these fractions get inflated.
What to do first
Write your eight roles and six cities into a dated file and run the employer wave once on two assistants, which is 96 observations and about half a day. Count own domain citations and whether any agency was named at all. If the answer names no agency anywhere, your first job is not ranking, it is publishing the fee structure and the role and city pages that give an assistant something specific to quote.