Agent-Readiness Grade

Agent-Readiness Grade / Fix guides / robots.txt for AI crawlers

Fix crawler access: allow the AI crawlers you want in robots.txt

The crawler-access check reads your /robots.txt the way a crawler does and reports, for eight AI crawler tokens, whether each may read your site. It is the one check where a deliberate choice (refusing model training) costs points, so this guide shows the lines that pass and what each block means.

Checked 2026-09-30 against grader version 1.3.0 and the 13 sources listed below.

Area: Crawler access (robots.txt). 2 of the 12 points.

What the grade checksWhy it mattersHow to fix itVerifyQuestionsSources

What the grade checks

  • GET https://example.com/robots.txt with a 7-second timeout, parsed by user-agent group for eight tokens: GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot, Google-Extended, Amazonbot, CCBot and Bytespider.
  • A group that names the token wins over the User-agent: * group. Disallow: / or Disallow: /* without a matching Allow: / counts as blocked; narrower path rules count as allowed and are reported as partial.
  • 2/2 when none of the eight is blocked, 1/2 when one to three are, 0/2 when four or more are. No robots.txt (a 404) or an HTML page in its place scores 2/2, because everything is allowed by default. A 5xx answer scores 0/2: RFC 9309 tells crawlers to treat an unreachable robots.txt as a full disallow.

2 of the 12 points. The full scoring rules are on the methodology page.

Why it matters for AI agents and crawlers

Each vendor documents a separate token per purpose, so one robots.txt can welcome AI search while refusing model training. OpenAI uses OAI-SearchBot for ChatGPT search and GPTBot for training, and says a site opted out of OAI-SearchBot is not shown in ChatGPT search answers. Anthropic separates ClaudeBot (training), Claude-SearchBot (search quality) and Claude-User (pages a person asked for). Perplexity says PerplexityBot surfaces sites in its results and is not used for foundation-model training.

Google says Google-Extended is a control token with no crawler of its own: blocking it keeps content out of Gemini training and grounding and does not affect inclusion or ranking in Google Search. Apple documents the same pattern for Applebot-Extended.

Blanket blocks are often accidental. Cloudflare's managed robots.txt setting prepends Disallow: / groups for Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent. Six of those are among the eight tokens this check reads, so a site with the setting on scores 0/2 here.

robots.txt is a request, not an access control. Vendors that publish IP lists (OpenAI, Anthropic, Perplexity, Amazon, Common Crawl) let you verify who is really asking, and a change takes time to apply: OpenAI, Perplexity and Amazon each mention about 24 hours.

How to fix it

  1. Decide per purpose

    Sort the tokens into what you want: AI search and answers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus the user-triggered fetchers ChatGPT-User, Claude-User and Perplexity-User), model training (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, meta-externalagent; Amazon says Amazonbot may be used to train its models), or everything. The robots.txt generator shows how each choice scores before you publish.

  2. Write explicit groups

    Name the tokens you allow so a later edit to User-agent: * cannot block them by accident. A group may list several user-agent lines (RFC 9309 section 2.1). If the * group has path rules such as Disallow: /admin/, repeat them in every named group: a crawler that finds its own group ignores the * group.

    robots.txt: every AI crawler allowed (crawler access 2/2)

    User-agent: GPTBot
    User-agent: OAI-SearchBot
    User-agent: ChatGPT-User
    User-agent: ClaudeBot
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: PerplexityBot
    User-agent: Perplexity-User
    User-agent: Google-Extended
    User-agent: Amazonbot
    User-agent: CCBot
    Allow: /
    Disallow: /admin/
    
    User-agent: *
    Allow: /
    Disallow: /admin/
    
    Sitemap: https://example.com/sitemap.xml

    robots.txt: AI search allowed, training refused (crawler access 0/2: four of the eight graded tokens blocked)

    User-agent: OAI-SearchBot
    User-agent: ChatGPT-User
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: PerplexityBot
    User-agent: Perplexity-User
    Allow: /
    
    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: Applebot-Extended
    User-agent: CCBot
    User-agent: meta-externalagent
    Disallow: /
    
    User-agent: *
    Allow: /
    
    Sitemap: https://example.com/sitemap.xml
  3. Publish it on your stack

    robots.txt must answer at the root of each host as plain text. www.example.com and every other subdomain need their own file; Anthropic asks for the opt-out on every subdomain you want covered.

    Static site

    Save the file as robots.txt in the web root, next to index.html. Web servers send .txt files as text/plain.

    WordPress

    Without a physical file, WordPress serves a virtual robots.txt. Add groups with the robots_txt filter in a must-use plugin, or upload a physical robots.txt to the WordPress root folder (it replaces the virtual one, including the sitemap line WordPress adds).

    wp-content/mu-plugins/ai-crawlers.php

    <?php
    // Appends AI crawler groups to WordPress's virtual robots.txt.
    add_filter( 'robots_txt', function ( $output, $public ) {
        $output .= "\nUser-agent: OAI-SearchBot\nUser-agent: PerplexityBot\nUser-agent: Claude-SearchBot\nAllow: /\n";
        return $output;
    }, 10, 2 );

    Next.js

    Use the robots file convention: app/robots.ts returns rules, and userAgent accepts a string or an array.

    app/robots.ts

    import type { MetadataRoute } from 'next'
    
    export default function robots(): MetadataRoute.Robots {
      return {
        rules: [
          { userAgent: ['OAI-SearchBot', 'ChatGPT-User', 'Claude-SearchBot', 'Claude-User', 'PerplexityBot'], allow: '/' },
          { userAgent: ['GPTBot', 'ClaudeBot', 'Google-Extended', 'Amazonbot', 'CCBot'], allow: '/' },
          { userAgent: '*', allow: '/', disallow: '/admin/' },
        ],
        sitemap: 'https://example.com/sitemap.xml',
      }
    }

    Cloudflare

    In the dashboard, Security → Settings → Bot traffic → "Set your preference to block training in robots.txt" is the managed robots.txt. When it is on, Cloudflare prepends Disallow groups for eight AI crawlers to your file: turn it off if you want them allowed, or keep it and accept 0/2 here. Blocking at the edge (AI Crawl Control) is separate and does not show in robots.txt. A Worker-served site returns the file itself:

    Cloudflare Worker

    const ROBOTS = `User-agent: OAI-SearchBot
    User-agent: PerplexityBot
    Allow: /
    
    User-agent: *
    Allow: /
    
    Sitemap: https://example.com/sitemap.xml
    `;
    
    export default {
      async fetch(request, env) {
        if (new URL(request.url).pathname === "/robots.txt") {
          return new Response(ROBOTS, { headers: { "content-type": "text/plain; charset=utf-8" } });
        }
        return env.ASSETS.fetch(request); // your static assets binding, if the Worker has one
      },
    };

Free generator: robots.txt generator for AI crawlers. Allow or block 14 AI crawler tokens, with each vendor's documentation linked.

How to verify

Fetch the file the way the grade does and check what each token resolves to.

shell

curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt | grep -i -E '^(user-agent|allow|disallow|sitemap):'

200 text/plain (or a 404, which allows everything). Then re-grade: the crawler-access evidence lists every token as allowed, disallowed or unmentioned.

Then re-grade your site: the result lists the evidence for this check.

$49 AI Visibility Full Report

The fixes on this site are free. The paid next step is the $49 AI Visibility Full Report (what ChatGPT, Claude and Perplexity say about your brand, with a prioritized fix list) from aivisibility.agentexchange.work. It includes:

  • 8 real buyer questions tested across ChatGPT-class models
  • Competitor share-of-voice: who AI names, how often, versus you
  • Full GEO site audit with prioritized, specific fixes
  • Agent-Readiness Score: crawler access, llms.txt, schema, discovery manifest
  • Custom 30/60/90-day action plan to get cited by ChatGPT, Perplexity and Google AI Overviews
  • Shareable report, generated in about 60 seconds after checkout

Get the Full Report — $49 Stripe checkout; you enter your brand and site right after paying.

Questions

Does blocking GPTBot remove my site from ChatGPT search?

Not according to OpenAI: GPTBot and OAI-SearchBot are independent. GPTBot governs training; OAI-SearchBot governs whether pages are shown in ChatGPT search answers, where opted-out sites can still appear as navigational links.

Does blocking Google-Extended affect Google Search?

Google says it does not: Google-Extended is not used for inclusion or ranking in Google Search. It controls use of crawled content for Gemini training and grounding, and it has no user agent of its own.

Do ChatGPT-User, Claude-User and Perplexity-User obey robots.txt?

OpenAI says robots.txt rules may not apply to ChatGPT-User because a person initiated the fetch, and Perplexity says Perplexity-User generally ignores robots.txt for the same reason. Anthropic says its bots, Claude-User included, honor robots.txt. None of the three is scored by this check.

Why did my grade fall after I turned on Cloudflare's AI bot setting?

The managed robots.txt adds Disallow: / for eight crawlers, six of which (Amazonbot, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot) are among the eight tokens this check reads. Four or more blocked scores 0/2.

How long until crawlers see a change?

OpenAI, Perplexity and Amazon each state about 24 hours; Meta says robots.txt may be cached for up to 24 hours, and Amazon may use a copy cached for up to 30 days. The grade reads the live file on every run.

Sources

Primary documentation, read 2026-09-30. Vendors change these pages; follow the link before relying on a detail.