AI describes my business incorrectly: a 30 query fact accuracy audit with an evidence log
How to build the fact sheet, count errors against a stated total, and why we publish no error fraction of our own yet.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
When an assistant names your business and gets the details wrong, the fix starts with a fact sheet you can prove, not with a complaint. Write down every testable fact about your business with the source that proves each one, ask 30 buyer questions blind, and then score each answer fact by fact, so the result is a fraction with both numbers printed: facts stated incorrectly, out of testable facts assessed, on a date. As of 29 September 2026 we have not completed a fact by fact accuracy audit on either of our own domains, so we publish no error fraction, no percentage and no typical error rate. What we publish instead is the whole protocol, the six error types, the correction order, and the dated errors we did find in our own programme, which are from legal research software sold in India and must not be read as a rate for anything.
What counts as an incorrect business description?
An error is only an error against a fact you can prove, so the definition comes first. We use six error types and they are scored separately, because the fix for each one is different.
- Wrong fact. A stated detail that contradicts your proven fact sheet: a price, a founding year, a location, a coverage claim, a feature that does not exist.
- Outdated fact. Correct at some earlier point and not now. The fix is usually on a page of yours that still says the old thing.
- Missing fact. A material detail absent from an answer where it changes the decision, such as a price or a limitation.
- Wrong category. You are described as a kind of business you are not, which is the error most likely to lose you the enquiry.
- Confused identity. Your details mixed with another company of a similar name, or your product attributed to another vendor.
- Unsupported inference. A statement the assistant appears to have reasoned to, with no source, such as a claim about who your customers are.
Two rules keep the scoring stable. Score only facts that appear on your fact sheet, so that an answer is never marked wrong for something you never established. And score each fact once per answer, not once per mention, or a repeated error inflates the count.
Build a verifiable fact sheet
This is the part that makes the audit possible. A fact without a proof is an opinion, and an opinion cannot be scored.
Record eight fields per fact.
- The fact in one short sentence.
- Its category: identity, location, product, price, coverage, people, compliance, or contact.
- The proof: a public URL, a document, or a registry entry.
- Whether that proof is publicly reachable without a login.
- The date the fact was last verified.
- Who inside the company owns it.
- Whether the fact appears on a page of yours that a machine can read without running scripts.
- Every other public place the fact appears, so contradictions are visible.
Count the sheet and print the count, because it becomes the denominator of the whole audit. Thirty facts is a normal size for a small company. If the sheet has 30 facts and an answer touches 11 of them, that answer is scored out of 11 and the page says so.
Run two checks on the sheet before you test anything. First, fetch each page holding a fact without running scripts and confirm the fact is in what comes back. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts had run, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead of from the company's own site. Second, look for contradictions between your own public pages. The same day, ChatGPT found that the clawlaw.in website and its app store listing carried different plan names and different prices for the same product, and it said so inside its answer.
And remove anything you cannot prove. On 6 August 2026 a headline figure on the clawlaw.in main site could not be true, and when Claude checked it against public numbers it advised a buyer against the product, after which the company's accurate claims stopped counting for that answer. An unprovable claim on your own site does more damage than a missing one.
Test assistant answers
One. Write 30 questions and freeze them. A mix of three kinds: questions that ask what your company is and does, questions a buyer asks about the category, and questions about the specific facts most likely to be wrong, such as price and coverage. Save the file with a version number and a date and do not edit it during the audit.
Two. Keep your brand out of the category questions, and use it only in the identity questions. Identity questions have to name you, because you are asking what the assistant thinks you are. Category questions must not, and mixing the two in one denominator is the most common mistake. Report them as two separate sets with separate counts. The reason is measured: on 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in, the branded run returned the company first on almost every question, and the blind run put it second by breadth and absent from the litigation due diligence questions it most wanted to win. The branded run was discarded.
Three. Three repetitions per question per assistant, same day, all answers kept. Errors are not stable. The same question can produce a correct answer and an incorrect one within minutes, and a single ask will tell you that things are fine when they are not.
Four. Print the denominator arithmetic twice. Once for answers: 30 questions times 3 repetitions times 3 assistants is 270 answers. Once for facts: if those answers touch 640 scoreable fact instances in total and 52 are wrong, the result is 52 of 640 fact instances scored incorrect, on the named dates, with the answer count printed beside it. Both denominators go on the page, because an error count against answers and an error count against facts are different numbers and they get confused constantly.
Five. Log twelve fields per answer. Question text, file version, identity or category set, assistant, model version if shown, timestamp, country and language setting, signed in or out, the full answer text saved, every URL cited exactly as given, the fact instances scored, and the error types found.
Six. Two scorers, then publish the disagreement count. Have a second person score at least 20 answers without seeing the first scores. Most disagreements will be about missing facts and unsupported inferences, which is why both have written definitions above.
Seven. Record any claim about searching as a claim. In the aiknowsus.com audit of September 2026, Perplexity withdrew its own earlier statement when asked how many of the questions it had actually searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. Three separate batches produced that retraction. It matters here because an answer built without a search reflects training data, and correcting your website will not change it quickly.
Results and error taxonomy
Fact by fact accuracy audits completed by us: zero, as of 29 September 2026. So we publish no error fraction, no percentage, no typical rate and no ranking of which error type is most common. Every one of those would be invented, and on a page about being described incorrectly that would be the worst possible failure.
The dated errors we did find, all in legal research software sold in India, and all about public pages. These are individual findings, not a rate.
- Wrong source for a price. 6 August 2026, Claude and ChatGPT: the pricing page served no prices without scripts, so the prices the assistants quoted came from an app store listing.
- Contradiction between two public price lists. 6 August 2026, ChatGPT: different plan names and different prices for the same product on the website and the app store listing, and the assistant said so in its answer.
- An unsupportable headline figure. 6 August 2026, Claude: the figure could not be true, the assistant checked it against public numbers and advised a buyer against the product, and the company's accurate claims stopped counting for that answer.
- Correct figures treated as advocacy. 17 September 2026, Claude: two comparison pages stated competitors' prices with no link, no date and no source, the figures were correct on independent checking, and the pages were still discounted.
- Material used with attribution dropped. 17 September 2026, Claude: two specific arguments drawn from clawlaw.in pages were used and the attribution was dropped in both places, described as a citation lapse rather than a ranking judgement.
The nearest real fraction we hold, with its sector named. 1 of 18: across eighteen blind commercial questions on 18 August 2026, ChatGPT made clawlaw.in the top source on exactly one of them. Legal research software sold in India. That is a citation fraction and not an accuracy fraction, and we are printing it because it is the nearest real fraction we have rather than because it answers this page's question. It is not a figure for any other sector.
What has not been run, dated. As of 29 September 2026 we have not built a published fact sheet with proofs for either of our own domains, have not scored answers fact by fact, and have not run the audit on Gemini or Copilot at all. The first audit we complete will be published with both denominators, the disagreement count and the dates.
Trace each claim to its source
An error you cannot trace is an error you cannot fix. For each incorrect fact, work through the following five steps in order.
- Ask the assistant for its sources for that specific claim, and save the URLs exactly as given. Treat the list as a claim about its own behaviour, not as a complete source set.
- Open each URL and find the sentence that produced the error. Often it is a page of yours that still says the old thing.
- Check whether your own page is readable without scripts. If your correct fact is injected by a script and somebody else's stale fact is in plain text, the stale one wins.
- Check for a second public source of yours that contradicts the first. App store listings, directory profiles, partner pages and old PDFs are the usual culprits.
- Accept that some sources are invisible. On 17 September 2026 Claude confirmed that clawlaw.in pages had ranked in its raw results and shaped what it wrote while the site still appeared only as a name inside a list. A page can be used and not shown, so a source list is a sample and not the full set.
Correction order
Fix in this order, because each step makes the next one work better.
- One. Remove every claim you cannot prove, starting with numbers. This is first because of the 6 August 2026 result: one impossible figure led an assistant to advise against the product, and the accurate claims stopped counting.
- Two. Make the facts readable without scripts, on the pages that hold them. A fact a crawler cannot see does not exist for anything that reads your site without a browser.
- Three. Reconcile your own public pages, one sheet, one owner per row, so that no two of your pages disagree about a price, a plan name or a coverage claim.
- Four. Add a date and a source to every figure, because correct and undated was still discounted on 17 September 2026.
- Five. Publish the facts that answers keep getting wrong as their own page, stated plainly with proofs, so there is something specific to quote.
- Six. Only then ask the platforms to correct anything. Report the incorrect answer through the feedback route each assistant provides, with your evidence and the timestamp, and log what you sent and when.
Doing step six first is the common mistake. A correction request about a fact your own pages still contradict is unlikely to survive.
Retest and publish the change log
Retest with the file unchanged, four weeks after the corrections, and report the new counts against the same two denominators. Keep both months side by side rather than merging them, because the model version and the question set version both need to be visible next to each figure.
Then publish the change log on the page: the date of each correction, what was changed, and the counts before and after. Two honesty rules go with it. A change after a change is not proof of a cause, because the assistant may have changed its own behaviour that month. And where the correction produced no measurable change, publish that too, because the absence is as informative as the improvement and almost nobody publishes it.
What this audit cannot establish
- It cannot produce a general accuracy rate for any assistant. It produces a count against your fact sheet, your questions and your dates.
- It cannot tell you why an answer was wrong. Source lists are partial, and an assistant's account of its own behaviour can be withdrawn, which we have dated.
- It cannot fix an answer built from training data quickly. On 27 July 2026, 0 of 78 blind questions triggered a live search in ChatGPT, so website corrections had nothing to reach that day.
- It cannot score facts you have not proven. The fact sheet is the limit of the audit, by design.
- Our dated errors are one company in legal technology in India. They are individual findings, not a rate, and they are not evidence about your sector.
- Nothing here promises a correction, or a position in an AI answer.
Sources and change log
Sources. The six error types, the eight fact fields, the twelve answer fields and the correction order are ours. Every dated finding comes from the clawlaw.in programme, recorded in Tier_1/GEO_BASELINE_RESULTS_2026-07-27.md, Tier_1/GEO_GAP_ANALYSIS_2026-08-18.md and Tier_1/claude_response_17_09_audit.md, or from the aiknowsus.com audit of September 2026 across 24 batches and 72 conversations. The pricing page, app store contradiction and impossible claim findings concern public pages and can be checked from outside. The capture wide counts cannot, until we publish the captures.
Change log. 29 September 2026, first publication, stating that no fact by fact audit has been completed and that no error fraction is published. The first completed audit will appear here with the fact sheet size, the answer count arithmetic, the fact instance count, the errors by type, the disagreement count between two scorers, and the dates.
Common questions
An assistant says our price is wrong. What do we check first?
Fetch your own pricing page without running scripts and read what comes back. On 6 August 2026 that single check explained why assistants were quoting an app store listing for clawlaw.in prices instead of the company's own pricing page. If your number is not in the response, nothing downstream can be right.
Should we report the error to the assistant?
Yes, after the first five correction steps, with your evidence and the timestamp, and log what you sent. Reporting first, while your own pages still contradict the fact, wastes the request.
How many facts should the sheet hold?
As many as are testable and proven, and print the count. Thirty is a normal size for a small company. What matters is that every fact has a public proof and a verification date, because the sheet is the denominator of everything that follows.
Why not publish the error rate you found in your own audit?
Because we have not run a fact by fact audit yet, and we would rather print that sentence with a date on it than a plausible figure. The nearest real fraction we hold is 1 of 18 from 18 August 2026 in legal research software in India, and it measures citation rather than accuracy, which is exactly why it is labelled that way above.
The answers were correct last month and wrong this month. Is that possible?
Yes, and it is why the repetition count and the monthly retest exist. Answers to identical words differ within a day, and model versions change. Keep every answer, report each month as its own count, and never merge two months into one denominator.
Does an assistant using our content without naming us count as an error?
Not as a factual error, and it should still have a slot in your log. On 17 September 2026 Claude admitted it had used two specific arguments from clawlaw.in pages and dropped the attribution in both places. If your log has no field for that, you will record a zero where something real happened.
What to do first
Write the fact sheet today, with a public proof and a verification date against every fact, and print the count. Then fetch each page holding those facts without running scripts and mark the facts that are missing from what comes back. That list of missing facts is the highest value work available to you this week, and it is the reason most descriptions go wrong in the first place.