Short answer
When live search is on, an AI app turns your question into one or more search queries, fetches a small set of results, reads them, and writes an answer from what they say. The pages it relies on become its sources. You cannot see the full ranking logic, but you can influence three things: whether your pages are findable and readable by the AI app's crawler, whether they answer the exact question clearly, and whether the third-party pages that get retrieved mention you accurately.
On this page
Two modes: memory and search
Every answer comes from one of two places, or a mix:
| Mode | Where the facts come from | How fresh | Can you change it quickly? |
|---|---|---|---|
| Memory (training data) | Text the model saw during training | Up to the model's training cutoff | No. Only when a new model is trained |
| Live search (grounding) | Pages fetched for this question | As fresh as the search index | Yes, by changing what those pages say |
Providers document live search for their APIs: Google calls it grounding with Google Search for Gemini, OpenAI and Anthropic each offer a web search tool, and Perplexity's Sonar models are built around search. Consumer apps decide on their own when to search, so a buyer may get either mode.
The retrieval pipeline, step by step
- Query writing. The AI app rewrites the buyer's question into search queries. A question about "billing apps for chemists in Lucknow" may become several shorter searches.
- Search. Those queries run against a search index. Google's own systems use Google Search. Other AI apps use their own index or partners.
- Fetching. The AI app fetches some result pages. This uses a named crawler or fetcher, such as OAI-SearchBot for ChatGPT search, Claude-SearchBot or Claude-User for Claude, and PerplexityBot or Perplexity-User for Perplexity.
- Reading. The text is extracted from the HTML. Text that only appears after heavy JavaScript may be missed.
- Writing. The model writes an answer, pulling names, facts and phrasing from the pages it read.
- Citing. Many AI apps show links to some or all pages used.
Each step is a place where you can drop out. If your page is not in the search results, it is never fetched. If it blocks the crawler, it is never read. If the answer is buried, the model may skip it.
What kinds of pages tend to get used
No AI app publishes a full list of ranking factors. From provider documentation and the GEO research, a few patterns are reasonable to act on:
- Pages that directly answer the question near the top, in plain words.
- Pages with specific, checkable facts: names, prices with dates, steps, tables.
- Pages that cite their own sources, which the KDD 2024 GEO paper found helped visibility in its tests.
- Pages that already do well in search for the query.
- For "best X" questions, list and comparison pages that name several options.
- For local questions, directory and map listings with consistent details.
Why third-party pages matter so much
Ask an AI app "what is the best CRM for a small real estate agency" and look at the sources. You will usually see comparison blogs, review sites, directories and forums, with fewer vendor home pages. That makes sense: an AI app trying to be neutral prefers pages that compare several options.
This means your visibility often depends on pages you do not own. If three of the five sources for a question are list articles and you are on none of them, better content on your own site may not be enough.
AI apps differ
Each AI app uses its own search setup, its own models and its own citation style. The same question can get different names and different sources on each. This is why you should track AI apps separately rather than as one blended number.
| AI app | Search behind it (documented) | Crawler or fetcher names to know |
|---|---|---|
| ChatGPT | OpenAI web search | OAI-SearchBot (search), ChatGPT-User (user actions), GPTBot (training) |
| Gemini and AI Overviews | Google Search | Googlebot for Search; Google-Extended is a separate control for Gemini training and grounding |
| Perplexity | Perplexity's own search | PerplexityBot (search), Perplexity-User (user actions) |
| Claude | Claude web search tool | Claude-SearchBot (search), Claude-User (user actions), ClaudeBot (training) |
What you can do this week
- Check robots.txt and make sure you are not blocking the search crawlers above by accident.
- Pick five buyer questions. For each, write down which page on your site answers it. If none, note it.
- Run each question a few times on two AI apps with search on. Save the source URLs.
- List the third-party domains that appear most often.
- For your own pages, move the direct answer to the top.
- For third-party pages, plan honest outreach or profile updates.
How we help
We record every source domain each AI app used for each check, so the list of pages you need to be on builds itself over time.
Frequently asked questions
Do AI apps use Google results?
Gemini and AI Overviews use Google Search, as Google documents. Other AI apps use their own search systems or partners. Their exact setups are not fully public.
If I block GPTBot, will ChatGPT stop citing me?
OpenAI says GPTBot controls training use, and OAI-SearchBot controls whether your site appears in ChatGPT search answers. They are separate. Blocking OAI-SearchBot is what removes you from search answers.
Why does the same question give different sources each time?
Query rewriting, search results and model sampling can all vary between runs. That is normal, and it is why you measure rates over many checks.
Does my website need to be popular to be cited?
Not necessarily. A clear, specific page that answers a narrow question well can be picked, especially for long, detailed questions.
Can I pay an AI app to be cited?
Organic answers are not sold as sites to be listed on. Some providers are testing ads, which are separate and labelled. GEO is about organic answers.
Sources
Last verified: 17 September 2026. Facts on this page come from the public documents below. Brands such as Tallybook are fictional and used only as examples.
- Gemini API: Grounding with Google Search
- OpenAI API: Web search tool
- Claude API: Web search tool
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google: Common crawlers, including Google-Extended
- Google Search Central: AI features and your website
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024)