# Fix crawler access: allow the AI crawlers you want in robots.txt

> The crawler-access check reads your `/robots.txt` the way a crawler does and reports, for eight AI crawler tokens, whether each may read your site. It is the one check where a deliberate choice (refusing model training) costs points, so this guide shows the lines that pass and what each block means.

Checked 2026-09-30 against Agent-Readiness Grade 1.3.0. HTML version: https://grade.agentexchange.work/fix/robots-txt-ai-crawlers

## What the grade checks

- `GET https://example.com/robots.txt` with a 7-second timeout, parsed by user-agent group for eight tokens: GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, Google-Extended, Amazonbot, CCBot and Bytespider.
- A group that names the token wins over the `User-agent: *` group. `Disallow: /` or `Disallow: /*` without a matching `Allow: /` counts as blocked; narrower path rules count as allowed and are reported as partial.
- **2/2** when none of the eight is blocked, **1/2** when one to three are, **0/2** when four or more are. No robots.txt (a 404) or an HTML page in its place scores 2/2, because everything is allowed by default. A 5xx answer scores 0/2: [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309) tells crawlers to treat an unreachable robots.txt as a full disallow.

2 of the 12 points.

## Why it matters for AI agents and crawlers

Each vendor documents a separate token per purpose, so one robots.txt can welcome AI search while refusing model training. [OpenAI](https://developers.openai.com/api/docs/bots) uses OAI-SearchBot for ChatGPT search and GPTBot for training, and says a site opted out of OAI-SearchBot is not shown in ChatGPT search answers. [Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) separates ClaudeBot (training), Claude-SearchBot (search quality) and Claude-User (pages a person asked for). [Perplexity](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) says PerplexityBot surfaces sites in its results and is not used for foundation-model training.

[Google](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) says Google-Extended is a control token with no crawler of its own: blocking it keeps content out of Gemini training and grounding and does not affect inclusion or ranking in Google Search. [Apple](https://support.apple.com/en-us/119829) documents the same pattern for Applebot-Extended.

Blanket blocks are often accidental. [Cloudflare's managed robots.txt](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/) setting prepends `Disallow: /` groups for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. Six of those are among the eight tokens this check reads, so a site with the setting on scores 0/2 here.

robots.txt is a request, not an access control. Vendors that publish IP lists (OpenAI, Anthropic, Perplexity, Amazon, Common Crawl) let you verify who is really asking, and a change takes time to apply: OpenAI, Perplexity and Amazon each mention about 24 hours.

## How to fix it

### 1. Decide per purpose

Sort the tokens into what you want: AI search and answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus the user-triggered fetchers ChatGPT-User, Claude-User and Perplexity-User), model training (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent; Amazon says Amazonbot may be used to train its models), or everything. The [robots.txt generator](https://grade.agentexchange.work/tools/robots-txt-ai-crawlers) shows how each choice scores before you publish.

### 2. Write explicit groups

Name the tokens you allow so a later edit to `User-agent: *` cannot block them by accident. A group may list several user-agent lines ([RFC 9309](https://www.rfc-editor.org/rfc/rfc9309) section 2.1). If the `*` group has path rules such as `Disallow: /admin/`, repeat them in every named group: a crawler that finds its own group ignores the `*` group.

robots.txt: every AI crawler allowed (crawler access 2/2):

```text
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Amazonbot
User-agent: CCBot
Allow: /
Disallow: /admin/

User-agent: *
Allow: /
Disallow: /admin/

Sitemap: https://example.com/sitemap.xml
```

robots.txt: AI search allowed, training refused (crawler access 0/2: four of the eight graded tokens blocked):

```text
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
Disallow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml
```

### 3. Publish it on your stack

robots.txt must answer at the root of each host as plain text. `www.example.com` and every other subdomain need their own file; Anthropic asks for the opt-out on every subdomain you want covered.

#### Static site

Save the file as `robots.txt` in the web root, next to `index.html`. Web servers send `.txt` files as `text/plain`.

#### WordPress

Without a physical file, WordPress serves a virtual robots.txt. Add groups with the [`robots_txt` filter](https://developer.wordpress.org/reference/hooks/robots_txt/) in a must-use plugin, or upload a physical `robots.txt` to the WordPress root folder (it replaces the virtual one, including the sitemap line WordPress adds).

wp-content/mu-plugins/ai-crawlers.php:

```php
<?php
// Appends AI crawler groups to WordPress's virtual robots.txt.
add_filter( 'robots_txt', function ( $output, $public ) {
    $output .= "\nUser-agent: OAI-SearchBot\nUser-agent: PerplexityBot\nUser-agent: Claude-SearchBot\nAllow: /\n";
    return $output;
}, 10, 2 );
```

#### Next.js

Use the [robots file convention](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots): `app/robots.ts` returns rules, and `userAgent` accepts a string or an array.

app/robots.ts:

```ts
import type { MetadataRoute } from 'next'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      { userAgent: ['OAI-SearchBot', 'ChatGPT-User', 'Claude-SearchBot', 'Claude-User', 'PerplexityBot'], allow: '/' },
      { userAgent: ['GPTBot', 'ClaudeBot', 'Google-Extended', 'Amazonbot', 'CCBot'], allow: '/' },
      { userAgent: '*', allow: '/', disallow: '/admin/' },
    ],
    sitemap: 'https://example.com/sitemap.xml',
  }
}
```

#### Cloudflare

In the dashboard, Security → Settings → Bot traffic → "Set your preference to block training in robots.txt" is the [managed robots.txt](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/). When it is on, Cloudflare prepends Disallow groups for eight AI crawlers to your file: turn it off if you want them allowed, or keep it and accept 0/2 here. Blocking at the edge ([AI Crawl Control](https://developers.cloudflare.com/ai-crawl-control/)) is separate and does not show in robots.txt. A Worker-served site returns the file itself:

Cloudflare Worker:

```js
const ROBOTS = `User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml
`;

export default {
  async fetch(request, env) {
    if (new URL(request.url).pathname === "/robots.txt") {
      return new Response(ROBOTS, { headers: { "content-type": "text/plain; charset=utf-8" } });
    }
    return env.ASSETS.fetch(request); // your static assets binding, if the Worker has one
  },
};
```

## How to verify

Fetch the file the way the grade does and check what each token resolves to.

```sh
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt | grep -i -E '^(user-agent|allow|disallow|sitemap):'
```

`200 text/plain` (or a 404, which allows everything). Then re-grade: the crawler-access evidence lists every token as allowed, disallowed or unmentioned.

Re-grade: https://grade.agentexchange.work/grade?url=example.com&fresh=1

## Questions

### Does blocking GPTBot remove my site from ChatGPT search?

Not according to OpenAI: GPTBot and OAI-SearchBot are independent. GPTBot governs training; OAI-SearchBot governs whether pages are shown in ChatGPT search answers, where opted-out sites can still appear as navigational links.

### Does blocking Google-Extended affect Google Search?

Google says it does not: Google-Extended is not used for inclusion or ranking in Google Search. It controls use of crawled content for Gemini training and grounding, and it has no user agent of its own.

### Do ChatGPT-User, Claude-User and Perplexity-User obey robots.txt?

OpenAI says robots.txt rules may not apply to ChatGPT-User because a person initiated the fetch, and Perplexity says Perplexity-User generally ignores robots.txt for the same reason. Anthropic says its bots, Claude-User included, honor robots.txt. None of the three is scored by this check.

### Why did my grade fall after I turned on Cloudflare's AI bot setting?

The managed robots.txt adds Disallow: / for eight crawlers, six of which (Amazonbot, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot) are among the eight tokens this check reads. Four or more blocked scores 0/2.

### How long until crawlers see a change?

OpenAI, Perplexity and Amazon each state about 24 hours; Meta says robots.txt may be cached for up to 24 hours, and Amazon may use a copy cached for up to 30 days. The grade reads the live file on every run.

## Sources

- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309) (IETF)
- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI)
- [Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic)
- [Google's common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) (Google)
- [Perplexity crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers) (Perplexity)
- [About Applebot](https://support.apple.com/en-us/119829) (Apple)
- [Amazonbot, Amzn-SearchBot and Amzn-User](https://developer.amazon.com/amazonbot) (Amazon)
- [CCBot](https://commoncrawl.org/ccbot) (Common Crawl)
- [Meta web crawlers](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/) (Meta)
- [robots.txt setting (managed robots.txt)](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/) (Cloudflare)
- [AI Crawl Control](https://developers.cloudflare.com/ai-crawl-control/) (Cloudflare)
- [robots_txt filter hook](https://developer.wordpress.org/reference/hooks/robots_txt/) (WordPress Developer Resources)
- [robots.txt file convention](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots) (Next.js)

## Related

- [llms.txt](https://grade.agentexchange.work/fix/llms-txt.md): A Markdown guide to your key pages at /llms.txt, served as text, not as your HTML 404.
- [sitemap.xml](https://grade.agentexchange.work/fix/sitemap-xml.md): An XML sitemap at /sitemap.xml, or a Sitemap: line in robots.txt pointing anywhere.
- [Readable without JavaScript](https://grade.agentexchange.work/fix/readable-without-javascript.md): The home page as a plain fetch sees it: enough visible text, no empty #root, no challenge.
- [robots.txt generator for AI crawlers](https://grade.agentexchange.work/tools/robots-txt-ai-crawlers)
- [All fix guides](https://grade.agentexchange.work/fix)
