Agent-Readiness Grade

Agent-Readiness Grade / Tools / robots.txt for AI crawlers

robots.txt generator for AI crawlers

Pick a rule for each AI crawler token; the robots.txt updates as you choose and shows how the grade's crawler-access check would score it. Every token, purpose and robots.txt statement comes from the vendor's own documentation, read 2026-09-30. Runs in your browser; nothing you choose is sent.

Checked 2026-09-30 against grader version 1.3.0 and each vendor's crawler documentation.

TokenOperator, purpose and robots.txt behaviour (from the vendor)Rule
GPTBot
in the grade
OpenAI · model training. Crawls content that may be used to train OpenAI's generative AI foundation models; disallowing it signals the content should not be used for that training. Honors robots.txt. Independent of OAI-SearchBot. vendor docs · rule page
OAI-SearchBot
in the grade
OpenAI · search. Surfaces websites in ChatGPT search; sites opted out are not shown in ChatGPT search answers, though they can still appear as navigational links. Honors robots.txt; OpenAI says changes take about 24 hours. vendor docs · rule page
ChatGPT-UserOpenAI · user-triggered fetch. Visits a page when a ChatGPT user or a Custom GPT asks; not used for automatic crawling or to decide what appears in search. OpenAI says robots.txt rules may not apply, because a user initiated the fetch. vendor docs · rule page
ClaudeBot
in the grade
Anthropic · model training. Collects web content that could contribute to training Anthropic's models; restricting it signals future content should be excluded from training datasets. Honors robots.txt and the non-standard Crawl-delay; set it on every subdomain. vendor docs · rule page
Claude-SearchBotAnthropic · search. Navigates the web to improve search result quality; disabling it prevents indexing for search optimization. Anthropic says its bots honor robots.txt. vendor docs
Claude-UserAnthropic · user-triggered fetch. Fetches pages when a person asks Claude a question; disabling it prevents retrieval of your content for user queries. Anthropic says its bots honor robots.txt, with no exception for this one. vendor docs · rule page
PerplexityBot
in the grade
Perplexity · search. Surfaces and links websites in Perplexity search results; Perplexity says it is not used to crawl content for AI foundation models. Perplexity recommends allowing it to appear in results; changes take up to 24 hours. vendor docs · rule page
Perplexity-UserPerplexity · user-triggered fetch. Visits a page to answer a user's question in Perplexity; not used for web crawling or training. Perplexity says it generally ignores robots.txt, because a user requested the fetch. vendor docs
Google-Extended
in the grade
Google · control token. Not a crawler: tells Google whether content it crawls may be used to train Gemini models and for grounding. Does not affect Google Search inclusion or ranking. Exists only as a robots.txt token; there is no Google-Extended user agent in requests. vendor docs · rule page
Applebot-ExtendedApple · control token. Opts content out of training Apple's generative foundation models; Applebot still crawls for Spotlight, Siri and Safari search. A robots.txt token; Apple says disallowing it opts out of training. vendor docs
Amazonbot
in the grade
Amazon · products and training. Improves Amazon's products and services and may be used to train Amazon AI models (Amzn-SearchBot is Amazon's separate search crawler). Honors robots.txt, not crawl-delay; about 24 hours to reflect changes, cached copies up to 30 days. vendor docs · rule page
CCBot
in the grade
Common Crawl · open web archive. Builds Common Crawl's free, open repository of web crawl data that anyone can analyze. Honors robots.txt; Common Crawl warns that other crawlers impersonate CCBot. vendor docs · rule page
Bytespider
in the grade
ByteDance · not documented. ByteDance publishes no documentation for Bytespider (its /en/bytespider page returned 404 when checked). Not documented; the lines follow the robots.txt convention for the token. vendor docs · rule page
meta-externalagentMeta · training and indexing. Crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Meta says a disallow blocks it; crawlers may cache robots.txt for up to 24 hours. vendor docs

Updates as you choose when JavaScript is on.

Your robots.txt

Crawler access in the grade: 2/2. None of the 8 tokens it reads (GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, Google-Extended, Amazonbot, CCBot, Bytespider) is blocked.

Put it at the root of every host you want covered (https://example.com/robots.txt, https://www.example.com/robots.txt). Changes take up to about 24 hours to reach OpenAI, Perplexity and Amazon, by their own documentation. A crawler that finds a group for its own token ignores the * group, which is why private paths are repeated in each allowed group.

What this does

Guide for this file: How to allow AI crawlers in robots.txt: GPTBot, ClaudeBot, OAI-SearchBot. Other generators: llms.txt generator · agent-card.json generator · security.txt generator.

Check your site

After publishing, grade the site to see what a crawler sees.

$49 AI Visibility Full Report

The fixes on this site are free. The paid next step is the $49 AI Visibility Full Report (what ChatGPT, Claude and Perplexity say about your brand, with a prioritized fix list) from aivisibility.agentexchange.work. It includes:

Get the Full Report — $49 Stripe checkout; you enter your brand and site right after paying.