Short answer
Technical GEO means making sure AI search crawlers can fetch and read your pages. Allow the search agents (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) in robots.txt, and decide separately about training agents (GPTBot, ClaudeBot, Google-Extended). Serve your main text in the HTML, not only via client-side JavaScript. Keep an up-to-date XML sitemap. Add JSON-LD for Organization, Product, Article, FAQPage and Dataset where it matches the page. An llms.txt file is an optional extra, not a requirement.
On this page
Search agents vs training agents
Most AI companies now run separate agents for search and for model training. Blocking one does not block the other. Here is what each provider documents:
| Agent | Provider | Used for (per provider docs) |
|---|---|---|
| OAI-SearchBot | OpenAI | Surfacing sites in ChatGPT search. Opted-out sites are not shown in ChatGPT search answers |
| GPTBot | OpenAI | Training foundation models |
| ChatGPT-User | OpenAI | User-initiated visits. robots.txt may not apply |
| Claude-SearchBot | Anthropic | Improving search result quality for Claude users |
| Claude-User | Anthropic | Fetching pages when a Claude user asks |
| ClaudeBot | Anthropic | Collecting content for model training |
| PerplexityBot | Perplexity | Surfacing and linking sites in Perplexity search. Not used for training |
| Perplexity-User | Perplexity | User-initiated visits. Generally ignores robots.txt |
| Google-Extended | A control token for Gemini Apps and Vertex AI training and grounding. Does not affect Google Search inclusion or ranking |
A sample robots.txt
This example allows search agents everywhere, allows training agents too, and keeps private areas closed. Change the training lines to Disallow: / if you do not want your content used for training.
# Search and answer agents User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Googlebot Allow: / Disallow: /app/ Disallow: /api/ # Training agents (your choice) User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended Allow: / Disallow: /app/ Disallow: /api/ # Everyone else User-agent: * Allow: / Disallow: /app/ Disallow: /api/ Sitemap: https://www.example.com/sitemap.xml
- Put the file at the root:
https://yourdomain.com/robots.txt. - Remember that a crawler follows the most specific group that names it, so repeat your Disallow rules in each group.
- Check your CDN, firewall or bot-protection settings too. They can block agents that robots.txt allows.
- Test the file with a robots.txt checker, and check your server logs for the agent names.
Serve real HTML
Google's JavaScript SEO guide explains that Google can render JavaScript, but rendering can be delayed and not every crawler renders pages. Many AI fetchers read the raw HTML. If your main content only appears after client-side JavaScript runs, some agents may see an empty page.
- Use server-side rendering or static generation for content pages.
- Check with
curl https://yourdomain.com/pagethat the main text is in the response. - Do not hide key facts behind tabs that load on click.
- Return real 404 status codes for missing pages.
- Keep pages fast and avoid interstitials that block content.
Sitemaps
An XML sitemap lists the URLs you want crawled. Google's documentation describes it as especially useful for large sites, new sites and sites with rich media.
- Include only canonical, indexable URLs.
- Set an accurate lastmod date when a page really changes.
- Reference the sitemap in robots.txt.
- Submit it in Google Search Console and Bing Webmaster Tools.
Structured data with JSON-LD
Structured data describes your content in a standard vocabulary from schema.org. Google recommends JSON-LD and requires the markup to match what is visible on the page. Google also says no special markup is needed to appear in AI features, so treat schema as a clarity aid, not a switch.
| Type | Use it on | Key properties |
|---|---|---|
| Organization | Home or About page | name, url, logo, sameAs (your official profiles), contactPoint |
| Product or SoftwareApplication | Product pages | name, description, offers with price and priceCurrency |
| Article | Guides and blog posts | headline, author, datePublished, dateModified |
| FAQPage | Pages with a visible FAQ | mainEntity with Question and acceptedAnswer |
| Dataset | Original data reports | name, description, creator, license, temporalCoverage |
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Tallybook",
"url": "https://tallybook.example",
"description": "Tallybook is a GST billing and inventory app for small retail shops in India.",
"sameAs": ["https://www.linkedin.com/company/tallybook-example"]
}The example above uses a fictional brand. Use your real entity sentence as the description, and list only profiles you actually control in sameAs.
llms.txt
llms.txt is a proposal, first published by Jeremy Howard in September 2024, for a markdown file at /llms.txt that gives language models a short guide to a site. The format starts with an H1 (the only required part), then an optional short summary in a blockquote, then H2 sections listing links with short notes. An "Optional" section holds links an agent can skip.
# Tallybook > GST billing and inventory app for small retail shops in India. ## Product - [Features](https://tallybook.example/features): what the app does - [Pricing](https://tallybook.example/pricing): plans and GST treatment ## Guides - [Offline billing](https://tallybook.example/learn/offline-billing): how offline mode works ## Optional - [Company history](https://tallybook.example/about)
Technical checklist
- robots.txt allows the search agents you want.
- CDN and firewall do not block those agents.
- Main content is in the server HTML.
- Sitemap is current and referenced in robots.txt.
- Each content page has one canonical URL.
- Organization JSON-LD is on the home or About page.
- Article and FAQPage JSON-LD match visible content.
- Pages you do not want quoted use nosnippet or noindex as needed.
- Optional: an llms.txt file with your key pages.
How we help
Our scan reads your robots.txt and pages the way a simple fetcher would, and flags blocked agents, empty HTML and missing structured data.
Frequently asked questions
If I block GPTBot, am I removed from ChatGPT?
OpenAI describes GPTBot as the training crawler. Its search visibility is tied to OAI-SearchBot. Blocking only GPTBot does not, by OpenAI's description, remove you from ChatGPT search answers.
Does Google-Extended control AI Overviews?
No. Google says Google-Extended does not affect inclusion in Google Search. AI Overviews use normal Search controls such as nosnippet and noindex.
Do I need an llms.txt file?
It is optional. It is a proposal, not a standard that the main AI apps have documented using for answers. Clear, crawlable pages matter more.
Which schema type matters most?
Organization on your home or About page is a sound start, because it states who you are. Then add Article and FAQPage on guides where they match the page.
Do Perplexity-User and ChatGPT-User follow robots.txt?
Both providers say these user-initiated agents may not follow robots.txt, because a person asked for the page. Anthropic says its bots respect robots.txt directives, and that site owners can control access by Claude-User.
Sources
Last verified: 17 September 2026. Facts on this page come from the public documents below. Brands such as Tallybook are fictional and used only as examples.
- OpenAI: Overview of OpenAI crawlers
- Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers
- Google: Common crawlers, including Google-Extended
- Google Search Central: AI features and your website
- RFC 9309: Robots Exclusion Protocol
- Google Search Central: Introduction to robots.txt
- Google Search Central: JavaScript SEO basics
- Google Search Central: Sitemaps overview
- Google Search Central: Introduction to structured data
- Google Search Central: Organization structured data
- schema.org: Organization
- schema.org: Product
- schema.org: Article
- schema.org: FAQPage
- schema.org: Dataset
- llms.txt proposal