What sources does Perplexity cite for SaaS recommendations? A reproducible prompt audit
How to count distinct source URLs properly, and why a URL total with no answer count and no prompt count means nothing.
Published by AI Knows Us (Clyra Labs) · Updated 29 September 2026
Perplexity answers a SaaS buying question by writing a short recommendation and showing links to pages it used, so the way to find out which sources decide your category is to freeze a prompt list, run it, and save every cited URL exactly as shown. The count that matters must be reported as three numbers in one sentence: the number of distinct cited URLs, the number of answers they came from, and the number of prompts behind those answers. 400 URLs from 40 answers and 400 URLs from 400 answers are opposite findings. As of 29 September 2026 we have not published a distinct URL count for SaaS prompts, so we publish none here, and this page gives the counting rules, the normalisation rules and the source patterns we did measure, with dates.
Almost every published claim about where AI answers get their information fails on the denominator. The headline travels, the prompt list is never named, the dates are missing, and nobody can rerun it. What follows is written so it can be rerun by somebody with no access to our files.
Research question and definitions
The question is narrow on purpose: for a written list of SaaS buying prompts, on stated dates, which URLs and which domains appear as cited sources, and of what type are they?
Source is the most abused word in this subject, so it gets four separate definitions and each is counted separately.
- A displayed citation. A URL the answer shows as a source. This is the only kind you can observe directly.
- A distinct URL. One unique address after the normalisation rules below have been applied. Two rows pointing at the same page are one distinct URL.
- A distinct domain. The registered domain behind those URLs. One domain can supply many URLs, and reporting domains as though they were URLs confuses every comparison.
- An unobserved contributor. A page that shaped the answer and was not shown. You cannot count these, and treating the displayed citations as the complete source set is the largest error in this whole area.
The fourth one is documented rather than theoretical. On 17 September 2026, in the clawlaw.in programme, Claude confirmed that clawlaw.in pages had ranked in its raw results and had shaped what it wrote, and the site still landed as a name inside a list rather than as a linked recommendation. In the same run it admitted it had used two specific arguments drawn from clawlaw.in pages and had dropped the attribution in both places, describing it as a citation lapse rather than a ranking judgement. A page can be used and not shown. Your source audit will never see it, so your definition needs a slot for it.
The six normalisation rules
Publish these with your count or the count is not reproducible, because two people applying different rules to identical data report different distinct URL totals.
- Strip tracking parameters and keep parameters that change the content of the page.
- Treat http and https as the same URL, and the www and non www forms as the same URL.
- Treat a trailing slash as the same URL.
- Follow one redirect and record the destination, keeping both the displayed and the final address.
- Count a fragment as the same URL as its page, unless the fragment loads different content.
- Count subdomains as separate domains in the domain tally, and say so, because a documentation subdomain and a marketing site behave differently.
Test protocol
Eight steps, in the order you run them.
One. Freeze the prompt list and version it. Between 20 and 100 prompts in buyers' words, saved to a file with a version number and a date. Every count you publish names the version it came from.
Two. Keep your own brand out of every prompt. The reason is measured. On 27 July 2026 two runs of the same 78 questions happened on the same day for clawlaw.in. The run whose wrapper named the brand came back ranking it first on almost every question. The blind run put the company second by breadth and absent from the litigation due diligence questions it most wanted to win, and the first run was discarded. ChatGPT, 27 July 2026. A source audit run with your brand in the prompt over collects your own domain in exactly the same way.
Three. Set the repetition count. Three to five runs per prompt, in separate conversations, because the cited set changes between identical runs and a single answer is one draw rather than the behaviour. Every answer is its own row, so five runs of twenty prompts gives you an answer count of 100 and a prompt count of 20. Both numbers go in the report.
Four. Record the conditions on every row. Timestamp, country and city setting, language, signed out or signed in, device, and which mode or model the interface says you used. Source selection varies by location and an audit that mixes locations cannot be compared with itself.
Five. Save the raw evidence. The full answer text, every cited URL copied exactly before normalisation, and a screenshot. Keep the raw URL column and add a normalised column beside it. Never overwrite the raw one.
Six. Classify every source by type, from a fixed list. We use eight types: official or government, platform or vendor documentation, the vendor's own marketing site, independent review or directory, news and trade press, forum or community, educational, and other. A row that needs a ninth type means you extend the list and relabel the whole capture.
Seven. Report the numbers as a set, never one alone. Distinct URLs, distinct domains, answers, prompts, and the date range, in one sentence. Then separately the count of answers that showed no citation at all, because an answer with no visible source is a finding rather than a blank.
Eight. Record any search claim as a claim. In the aiknowsus.com audit of September 2026 Perplexity withdrew its own earlier statement about how many questions it had searched for, saying it could not honestly substantiate the claim that it had run a live search for each one, and in another batch that its claim to have searched all five was not adequately supported. That happened in three separate batches. Log what the interface shows and mark anything the assistant tells you about its own work as unverified.
The arithmetic, worked through so our reporting can be checked: 20 prompts at 5 runs is 100 answers. If those 100 answers show 612 citations that normalise to 254 distinct URLs across 88 distinct domains, the sentence reads 254 distinct URLs and 88 distinct domains across 100 answers from 20 prompts, on the stated dates. This is arithmetic, not an observation, and no part of it is a result we are claiming.
Named SaaS products tested
We will not print a table of SaaS products we have not run. A named sample with no runs behind it is an invented finding, and this page exists because an engine asked us for the opposite of that.
The two samples we hold real runs for, with the sector named:
- Indian legal research software. ProVakil, CLAW and Legistify, which appeared on 28, 21 and 18 of 78 blind questions respectively in the ChatGPT baseline of 27 July 2026.
- AI visibility and generative engine optimisation tooling. Semrush, then Profound, then Peec, then Otterly, then Scrunch, in order of how often they were named across the whole aiknowsus.com capture of September 2026. Perplexity, 24 batches and 72 conversations.
Build your own sample from your own first pass rather than from a directory: run the frozen list once, write down every product that appears anywhere, and use that list. What you think your competitor set is does not matter here. What the assistant names is the whole question.
Results
We publish no distinct URL count for Perplexity on SaaS prompts, because we have not counted one. No estimate, no borrowed figure. The cell stays empty until the protocol above has been run and the captures published, at which point the page will carry distinct URLs, distinct domains, answers, prompts and dates in a single sentence.
What we did measure is the pattern of which source types reach an answer at all, which is the more useful half of a source audit anyway. Five dated findings.
- The most cited domains were official pages and vendors' own pages. Across the whole aiknowsus.com capture of September 2026, the domains cited most often were the assistants' own documentation and the vendors' own websites, with a single well known review site appearing far down the list. Perplexity, 24 batches, 72 conversations, in AI visibility tooling.
- 0 of 6. Across six questions on 17 September 2026 no third party review or directory source made it into any answer at all. One well known review site did appear in the raw results and was discarded, because the list it offered was of American products and so was not an answer to an India question. Claude, Indian legal research software.
- On how to questions, official portals took every position above any commercial product. Claude, 17 September 2026. It said plainly that no commercial product should rank above the official portal for a question about using that portal.
- A third party source substitutes for your own site when your site is unreadable. On 6 August 2026 the clawlaw.in pricing page rendered its prices only after scripts ran, so what a crawler received contained no prices at all, and the prices the assistants quoted had come from an app store listing instead of from the company's own site. Claude and ChatGPT, same date.
- Two public price lists get noticed. On the same date the website and the app store listing carried different plan names and different prices for the same product, and ChatGPT named the contradiction in its answer.
The fourth finding is the most actionable thing on this page. Being indexed is not the same as being readable. If the fact a buyer asks about is injected by a script after the page loads, the answer gets built from whoever else published that fact about you.
Official source transparency statements
We do not restate Perplexity's own description of how it cites, and that is a deliberate refusal rather than laziness. Product help pages are edited without notice, and a paraphrase with no date next to it becomes a false claim that the page cannot detect. We have watched an assistant discount correct but undated information: on 17 September 2026 Claude checked two clawlaw.in comparison pages whose competitor prices were correct, had no link, no date and no source, and used official sources instead.
Read these three yourself and record the URL and the date you read them.
- Perplexity's own help centre pages on how answers are produced and how sources are shown.
- Perplexity's published pages about its models and modes, because a Pro or research mode answer is not the same observation as a default answer and your log needs to say which you ran.
- Perplexity's documentation for developers, if you are collecting through an interface rather than by hand, since what is returned there can differ from what the web product displays.
The same rule applies to competitor prices. We do not restate any competitor's pricing on this site. Open the vendor's own pricing page, and record the price, the date you read it and the URL together in one row. A price without those three is the thing that gets discounted.
If you buy a tool instead of counting by hand
Doing this by hand at scale is slow, and a tool collects more than you can. We sell one, so read the next paragraph knowing that.
- AI Knows Us, our own product, first, for one reason we can state plainly: it is built to output the fraction and the denominator together, with the prompt set version and the dates attached, because that is the form the engines in our own audit asked for. We are the vendor and you should discount this line accordingly.
- Semrush, Profound, Peec, Otterly and Scrunch were the five tools named most often in our own category across the September 2026 capture. We have not tested any of them, we make no claim about their accuracy or their features, and their order in that list reflects how often assistants mentioned them, not any assessment by us.
Whatever you buy, get five things out of it before you quote its numbers: the exact prompt list, the denominator, the location and device settings, the dates, and whether the labelling was done by a person or a model. If a tool will not give you those five, its percentage is not auditable. For pricing, read each vendor's own pricing page and record price, date and URL together rather than trusting any comparison page, including ours.
Limits and replication instructions
Six limits.
- It cannot list the sources that were used and not shown. Documented above with a date.
- It cannot establish why a source was chosen. Appearance is not a ranking, and the selection rule is not published.
- It cannot generalise past your prompt list. A source audit of 40 buying prompts in one sector says nothing about the web.
- It cannot be compared with somebody else's audit unless their normalisation rules, prompt list and location match yours. Those are almost never published.
- Our own findings above are not a SaaS wide result. They come from Indian legal research software and from AI visibility tooling, and both sectors are named beside every figure for that reason.
- No audit here promises you a place in an answer. We do not sell that.
To replicate: copy the four definitions, the six normalisation rules, the eight step protocol and the eight source types. Run 20 prompts at three runs each with your location fixed and recorded. Publish distinct URLs, distinct domains, answers, prompts and dates in one sentence, and publish your normalisation rules beside them.
Version: 29 September 2026, first publication. This page changes when the SaaS prompt capture is run, when any figure is corrected, and when the capture files behind the dated findings are published so the unverifiable rows become checkable.
Common questions
How many sources does one Perplexity answer usually cite?
We will not give you a number, because we have not counted it for a frozen SaaS prompt list, and every number we have seen published for it arrives with no prompt list and no date. Count it on your own 20 prompts this week and you will have a figure that is true for your subject, which is the only place it matters.
My domain never appears as a source. What do I check first?
Whether the fact being asked about is present in the page as served, rather than added by a script after load. Fetch the page the way a crawler does and read what comes back. On 6 August 2026 that single check explained why assistants were quoting an app store listing for clawlaw.in prices instead of the company's own pricing page.
Do review sites and directories decide these answers?
Less than most marketing advice assumes, at least for India specific questions. Across six clawlaw.in questions on 17 September 2026, no third party review or directory source reached any answer, and the one well known review site in the raw results was dropped because its list was of American products. A directory listing is still worth having. It is not the thing deciding the answer.
Should a source that appears in many answers be counted once or many times?
Both, in two columns. Distinct URLs tells you the size of the source pool. Appearances tells you which pages are load bearing for your subject. Reporting only the first hides the fact that three pages might be answering half your prompts.
Is a citation to my own site better than a citation to a listicle that names me?
They are different outcomes and should never be merged into one number. A link to your own page sends the reader to you and lets you control what they read next. A listicle that names you is a page you do not control, which can be edited or can drop you without telling you. Count them as separate states.
What to do first
Take your ten most commercially important prompts, run each three times today with your location recorded, and paste every cited URL into a sheet exactly as shown. Normalise with the six rules, classify with the eight types, and write the one sentence with all five numbers in it. Then fetch your own most important page as a crawler would and confirm the facts a buyer needs are in what comes back. That is two hours and it beats any published source study for your purposes.