Technical

Technical GEO setup: crawlers, llms.txt, schema and rendering

robots.txt rules for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended, plus llms.txt, JSON-LD, server rendering and sitemaps.

Updated 9 min readPublished by AI Knows Us (Clyra Labs), which sells GEO software

Short answer

Technical GEO means making sure AI search crawlers can fetch and read your pages. Allow the search agents (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) in robots.txt, and decide separately about training agents (GPTBot, ClaudeBot, Google-Extended). Serve your main text in the HTML, not only via client-side JavaScript. Keep an up-to-date XML sitemap. Add JSON-LD for Organization, Product, Article, FAQPage and Dataset where it matches the page. An llms.txt file is an optional extra, not a requirement.

On this page
  1. Search agents vs training agents
  2. A sample robots.txt
  3. Serve real HTML
  4. Sitemaps
  5. Structured data with JSON-LD
  6. llms.txt
  7. Technical checklist
  8. How we help
  9. Frequently asked questions
  10. Sources

Search agents vs training agents

Most AI companies now run separate agents for search and for model training. Blocking one does not block the other. Here is what each provider documents:

AgentProviderUsed for (per provider docs)
OAI-SearchBotOpenAISurfacing sites in ChatGPT search. Opted-out sites are not shown in ChatGPT search answers
GPTBotOpenAITraining foundation models
ChatGPT-UserOpenAIUser-initiated visits. robots.txt may not apply
Claude-SearchBotAnthropicImproving search result quality for Claude users
Claude-UserAnthropicFetching pages when a Claude user asks
ClaudeBotAnthropicCollecting content for model training
PerplexityBotPerplexitySurfacing and linking sites in Perplexity search. Not used for training
Perplexity-UserPerplexityUser-initiated visits. Generally ignores robots.txt
Google-ExtendedGoogleA control token for Gemini Apps and Vertex AI training and grounding. Does not affect Google Search inclusion or ranking

A sample robots.txt

This example allows search agents everywhere, allows training agents too, and keeps private areas closed. Change the training lines to Disallow: / if you do not want your content used for training.

# Search and answer agents
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Googlebot
Allow: /
Disallow: /app/
Disallow: /api/

# Training agents (your choice)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Allow: /
Disallow: /app/
Disallow: /api/

# Everyone else
User-agent: *
Allow: /
Disallow: /app/
Disallow: /api/

Sitemap: https://www.example.com/sitemap.xml
  1. Put the file at the root: https://yourdomain.com/robots.txt.
  2. Remember that a crawler follows the most specific group that names it, so repeat your Disallow rules in each group.
  3. Check your CDN, firewall or bot-protection settings too. They can block agents that robots.txt allows.
  4. Test the file with a robots.txt checker, and check your server logs for the agent names.

Serve real HTML

Google's JavaScript SEO guide explains that Google can render JavaScript, but rendering can be delayed and not every crawler renders pages. Many AI fetchers read the raw HTML. If your main content only appears after client-side JavaScript runs, some agents may see an empty page.

  • Use server-side rendering or static generation for content pages.
  • Check with curl https://yourdomain.com/page that the main text is in the response.
  • Do not hide key facts behind tabs that load on click.
  • Return real 404 status codes for missing pages.
  • Keep pages fast and avoid interstitials that block content.

Sitemaps

An XML sitemap lists the URLs you want crawled. Google's documentation describes it as especially useful for large sites, new sites and sites with rich media.

  1. Include only canonical, indexable URLs.
  2. Set an accurate lastmod date when a page really changes.
  3. Reference the sitemap in robots.txt.
  4. Submit it in Google Search Console and Bing Webmaster Tools.

Structured data with JSON-LD

Structured data describes your content in a standard vocabulary from schema.org. Google recommends JSON-LD and requires the markup to match what is visible on the page. Google also says no special markup is needed to appear in AI features, so treat schema as a clarity aid, not a switch.

TypeUse it onKey properties
OrganizationHome or About pagename, url, logo, sameAs (your official profiles), contactPoint
Product or SoftwareApplicationProduct pagesname, description, offers with price and priceCurrency
ArticleGuides and blog postsheadline, author, datePublished, dateModified
FAQPagePages with a visible FAQmainEntity with Question and acceptedAnswer
DatasetOriginal data reportsname, description, creator, license, temporalCoverage
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Tallybook",
  "url": "https://tallybook.example",
  "description": "Tallybook is a GST billing and inventory app for small retail shops in India.",
  "sameAs": ["https://www.linkedin.com/company/tallybook-example"]
}

The example above uses a fictional brand. Use your real entity sentence as the description, and list only profiles you actually control in sameAs.

llms.txt

llms.txt is a proposal, first published by Jeremy Howard in September 2024, for a markdown file at /llms.txt that gives language models a short guide to a site. The format starts with an H1 (the only required part), then an optional short summary in a blockquote, then H2 sections listing links with short notes. An "Optional" section holds links an agent can skip.

# Tallybook

> GST billing and inventory app for small retail shops in India.

## Product
- [Features](https://tallybook.example/features): what the app does
- [Pricing](https://tallybook.example/pricing): plans and GST treatment

## Guides
- [Offline billing](https://tallybook.example/learn/offline-billing): how offline mode works

## Optional
- [Company history](https://tallybook.example/about)

Technical checklist

  1. robots.txt allows the search agents you want.
  2. CDN and firewall do not block those agents.
  3. Main content is in the server HTML.
  4. Sitemap is current and referenced in robots.txt.
  5. Each content page has one canonical URL.
  6. Organization JSON-LD is on the home or About page.
  7. Article and FAQPage JSON-LD match visible content.
  8. Pages you do not want quoted use nosnippet or noindex as needed.
  9. Optional: an llms.txt file with your key pages.

How we help

Our scan reads your robots.txt and pages the way a simple fetcher would, and flags blocked agents, empty HTML and missing structured data.

Frequently asked questions

If I block GPTBot, am I removed from ChatGPT?

OpenAI describes GPTBot as the training crawler. Its search visibility is tied to OAI-SearchBot. Blocking only GPTBot does not, by OpenAI's description, remove you from ChatGPT search answers.

Does Google-Extended control AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search. AI Overviews use normal Search controls such as nosnippet and noindex.

Do I need an llms.txt file?

It is optional. It is a proposal, not a standard that the main AI apps have documented using for answers. Clear, crawlable pages matter more.

Which schema type matters most?

Organization on your home or About page is a sound start, because it states who you are. Then add Article and FAQPage on guides where they match the page.

Do Perplexity-User and ChatGPT-User follow robots.txt?

Both providers say these user-initiated agents may not follow robots.txt, because a person asked for the page. Anthropic says its bots respect robots.txt directives, and that site owners can control access by Claude-User.

Sources

Last verified: 17 September 2026. Facts on this page come from the public documents below. Brands such as Tallybook are fictional and used only as examples.

  1. OpenAI: Overview of OpenAI crawlers
  2. Anthropic: Does Anthropic crawl data from the web, and how can site owners block the crawler?
  3. Perplexity: Perplexity crawlers
  4. Google: Common crawlers, including Google-Extended
  5. Google Search Central: AI features and your website
  6. RFC 9309: Robots Exclusion Protocol
  7. Google Search Central: Introduction to robots.txt
  8. Google Search Central: JavaScript SEO basics
  9. Google Search Central: Sitemaps overview
  10. Google Search Central: Introduction to structured data
  11. Google Search Central: Organization structured data
  12. schema.org: Organization
  13. schema.org: Product
  14. schema.org: Article
  15. schema.org: FAQPage
  16. schema.org: Dataset
  17. llms.txt proposal

Check your own AI visibility.

Enter your website. The first scan is free and asks real buyer questions with live search on.

Free. No card. We ask 5 real buyer questions on 2 AI apps.