AI visibility tools compared: a reproducible benchmark across 6 tools

A 15 prompt test spec, what each tool must reproduce, and the full cost of running it.

Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026

Do not compare these tools on their dashboards. Compare them on whether each one can reproduce the same result from the same 15 prompts. Run the 15 prompts by hand first, on the engines you care about, three times each, and label the answers yourself. That hand audit is your ground truth. Then give every tool the identical 15 prompts and check how closely its numbers match yours and whether it hands you the raw answers. The total cost of executing that disclosed 15 prompt test by hand across three engines is zero in software: 15 prompts times 3 engines times 3 repeats is 135 answer runs, all of which can be run in the engines' own free interfaces, and the only cost is the person doing it. Any tool you then add has to justify its own cost on top of zero.

We make one of the six tools below and we list it first, which you should discount accordingly. We have not benchmarked these products against each other, so this page contains no comparative scores for them, and it says so in the results section rather than filling the gap with feature claims from their own marketing.

Choose by budget and workflow

Five buyer situations, and the honest recommendation for each. None of them start with buying software.

  • One brand, one country, first measurement ever. Do the hand audit. 15 prompts, three engines, one afternoon, no spend. You will learn more about your own gaps from reading 135 answers than from any dashboard, and you will then know what to demand from a tool.
  • One brand, monthly reporting to a board or an investor. Buy a tool, and make raw answer export a hard requirement. The first question a sceptical reader asks is how the number was produced, and a chart cannot answer it.
  • An agency with several clients. Seat and client separation costs will dominate your bill, not prompt volume. Price the tenth client, not the first.
  • Several markets or languages. Per market pricing is where these bills grow. Count the markets you will genuinely report on every month, not the ones you might.
  • You already pay for a large search platform. Check what its AI reporting already covers before buying a second product. Procurement time has a cost too.

Dated public pricing and plan limits

We publish no prices for these vendors on this page, as of 29 September 2026, because we have not verified them on a dated basis. That is a deliberate refusal, and we have watched what happens to pages that do it the other way. On 17 September 2026, Claude read two clawlaw.in comparison pages that stated competitors' prices with no link, no date and no source. It checked the figures independently, found them correct, and still treated the pages as vendor advocacy and used official sources instead. Unsourced pricing loses the page even when the pricing is right.

Collect the prices yourself, into these eleven fields, one row per tool, each with the date you read it and a screenshot.

  • Plan name, list price, currency, and the date you checked.
  • Billing term, and whether the headline price assumes an annual commitment.
  • Included prompt volume per month, stated in the vendor's own unit.
  • The definition of one unit. Is a repeat of the same prompt a separate unit? Is each engine a separate unit?
  • Engines included, and which are paid add ons.
  • Markets or locations included, and the price of an extra one.
  • Seats included and the price per extra seat.
  • Overage rate beyond the included volume.
  • Raw answer export: included, restricted to a higher plan, or unavailable.
  • Contract term, notice period, refund terms.
  • Taxes and payment method, which for an Indian buyer changes the invoice meaningfully.

Field four is the one that decides your bill. Under the test spec below, 15 prompts is 135 answer runs. A tool that charges per run and a tool that charges per prompt will quote you numbers that differ by roughly an order of magnitude for exactly the same test, and both will be honest.

Test methodology

The spec, in full, so that you and a vendor can run the same thing and compare.

The prompt set. Exactly 15 prompts, written in the words your buyers use, saved to a file with a version number and today's date. Five of the fifteen should be category questions with no brand in them, five should be comparison questions naming two or three competitors but never you, and five should be problem questions that a buyer asks before they know the category exists. That split matters because the three kinds behave differently and a set made only of category questions flatters everybody.

The blind rule. Your own brand name appears in none of the fifteen. The reason is measured. On 27 July 2026, two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent altogether from the litigation due diligence questions it most wanted to win. The first run was discarded, because the only thing it measured was our own prompt.

The engines. Name them and report them separately. Never publish a blended score across engines.

The repeats. Three runs per prompt per engine, in the same session window, giving 135 answer runs for three engines.

The conditions. Recorded on every row: date, time, country and city, language, signed in state, device, and whether any account personalisation or memory feature was active.

The labels. Four states per run: brand named in the answer text, own domain linked, own page is the top source, and material used with no attribution. The fourth state is real. On 17 September 2026 Claude admitted it had used two specific arguments from clawlaw.in pages and dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement. No tool we know of detects that state, so a human has to read a sample.

The search claim column. Record whatever the system says about whether it searched, and mark it unverified. In three separate batches of our September 2026 audit, Perplexity withdrew its own earlier statement, saying it could not honestly substantiate the claim that it had run a live search for each question.

The tool comparison itself. Give each tool the same 15 prompts and the same engines, then score it on six things you can actually verify: does it let you enter your exact prompts, does it enforce or at least permit the blind rule, does it let you set the repeat count, does it report per engine, does it export the raw answers with timestamps, and how far its counts differ from your hand labels on the same 135 runs.

Results

We publish no comparative scores for the six tools, because we have not run the benchmark against them. Every cell is empty and the spec above is what we are offering instead. When we run it, the results will be published with the prompt set version, the engines, the dates and our hand labels next to each tool's numbers, so the disagreement is visible rather than hidden.

What we can report is dated and comes from our own two programmes.

  • Our own visibility: 0 of 6. Six blind questions about our own category, September 2026. Perplexity audited its own answers and reported it had cited or recommended aiknowsus.com in none of the six.
  • 161 recorded statements that we were not cited, across 24 batches and 72 conversations, September 2026, in the engines' own self audits.
  • Mention order for the tools in this category, counted across the whole September 2026 capture: Semrush most often, then Profound, then Peec, then Otterly, then Scrunch, with the established search tools appearing alongside the specialist ones rather than below them. That is a measure of how often engines name each tool, not of how well any of them works.
  • The engine refused to rank them. In batch after batch it declined to name any competitor as winning most often, saying its own previous answers had not produced a comparable live tested result to support such a claim. September 2026, Perplexity.

The six tools, named, with only what we can defend said about each. AI Knows Us is first because it is ours and because it is built to enforce the spec above, which is a design claim you can check, not a benchmark result.

  • AI Knows Us, our own product. Frozen versioned prompt sets, the blind rule enforced, configurable repeats, per engine reporting, every rate printed with its numerator, denominator and date.
  • Profound, a specialist AI visibility platform, second by mention frequency in our capture. Not benchmarked by us.
  • Otterly, a specialist tool in the same category, fourth by mention frequency. Not benchmarked by us.
  • Semrush, an established search platform with AI answer reporting, first by mention frequency. Not benchmarked by us.
  • Peec, a specialist tool, third by mention frequency. Not benchmarked by us.
  • Scrunch, a specialist tool, fifth by mention frequency. Not benchmarked by us.

Total cost with required add ons

The number the spec above lets us state without inventing anything: executing the disclosed 15 prompt test by hand across three engines costs nothing in software. 135 answer runs, entered by a person into the engines' own interfaces, with screenshots saved to a folder and labels typed into a spreadsheet. The costs are real and they are not licence fees.

  • Person time, which is the whole cost of the hand version. Estimate it yourself from your own first five prompts rather than trusting anybody's figure, including ours.
  • Storage for screenshots and text captures. Negligible at this size.
  • A second labeller for a sample, half an hour, which is what makes your ground truth defensible.

Once a tool is involved, the total cost of the same test is the sum of these seven lines, and you fill them from the pricing worksheet above.

  • Base plan for the billing term you will actually accept.
  • Extra prompt volume, if 135 runs a month exceeds the included allowance.
  • Per engine add ons for any engine not in the base plan.
  • Per market add ons, if you report on more than one location.
  • Extra seats beyond the included number.
  • The export or API tier, if raw answers are gated above your plan. Treat this as required rather than optional, because without exports you cannot audit the number.
  • Taxes and payment charges.

Write that total next to zero, which is what the hand version costs, and the question becomes clear: is the tool buying you scale and time, and is it giving you raw data you could defend? Those are good reasons to pay. A prettier chart is not.

What this benchmark cannot settle

It cannot tell you which tool is most accurate in general, only which one agrees most closely with your hand labels on your 15 prompts in your market on those dates. It cannot detect the use without attribution state, which is documented above with a date. It cannot make two tools' historical numbers comparable, because their past prompt sets and repeat counts differ. It cannot hold still: engines change, so a benchmark is valid for the period it was run in and needs rerunning. And no tool in the list, ours included, can promise you a position in an AI answer.

Reproduction and sources

To reproduce: the 15 prompt set with its five plus five plus five split, the blind rule, three named engines, three repeats, the four label states, the recorded conditions, and the six tool checks. All of it is inside your control and none of it requires our product.

Sources for every figure quoted: the aiknowsus.com audit of September 2026, 24 batches and 72 conversations, captured outside this repository, and the clawlaw.in programme recorded in GEO_BASELINE_RESULTS_2026-07-27.md and the assistant audit files of 17 September 2026. Counts were produced by a script over those files. Disclosure: AI Knows Us is our product and is listed first. Version: 29 September 2026, first publication, to be updated when the tool benchmark is actually run and when any vendor correction is received.

Common questions

Profound or Otterly, which should we pick?

We cannot answer that from evidence, and we will not answer it from impressions. Run the 15 prompt hand audit first, then put the identical prompts into both trials and compare each tool's counts with your own labels on the same runs. The one that matches your reading and exports the raw answers is the one to buy.

Why does the repeat count matter so much to the price?

Because it multiplies the run count. 15 prompts on three engines with three repeats is 135 runs rather than 15, and vendors meter differently. Ask for a quote based on 135 runs a month and see which quotes change.

Can we skip the hand audit and trust the tool?

You can, and then you have no way to tell a tool error from a real change. The hand audit is the only ground truth available to you, and 135 runs is one focused afternoon.

Is a blended visibility score across engines ever useful?

For a headline slide, maybe. For deciding what to do next, no, because the engines pick sources differently and a blend hides which one you are absent from. Report per engine and keep the blend out of the working document.

What if a tool will not let us use our own prompts?

Then it is measuring its own prompt set, not your market, and its numbers cannot be compared with your hand audit or with your own history. For most buyers that is a reason to walk away.

How often should the benchmark be rerun?

Quarterly for the tool comparison, monthly for your own visibility numbers, with the prompt set frozen in both cases. Any change to the prompt set starts a new version and a new baseline, and the old numbers are reported separately rather than joined to the new ones.

What to do first

Write the 15 prompts today, five category, five comparison, five problem, with your brand in none of them. Run them on three engines, three times each, and label the 135 answers yourself. Then start the trials. Walking into a vendor demo with your own ground truth changes the conversation completely, and it costs you an afternoon and no money.

See what AI says about you.

The first scan is free and takes about 20 seconds.

Free. No card. We ask 5 real buyer questions on 2 AI apps.