Agentic web census · AI access lookup · rank #2 · fetched 2026-09-30
Can AI assistants read cloudflare.com?
cloudflare.com's robots.txt lets ChatGPT, Claude and Perplexity open pages when a user asks. (robots.txt served from www.cloudflare.com)
- Training crawlers blocked: none
- llms.txt: published (plus llms-full.txt)
- A2A agent card: published
- security.txt: published
- Content-Signal:
ai-train=yes, search=yes, ai-input=yes - robots.txt looks like Cloudflare's managed file
| AI agent | Operator and purpose | robots.txt verdict |
|---|---|---|
| User-triggered assistants | ||
| ChatGPT-User (named) | OpenAI (a user asked ChatGPT to open a page) | allowed |
| Claude-User | Anthropic (a user asked Claude to open a page) | allowed |
| Perplexity-User | Perplexity (a user asked Perplexity to open a page) | allowed |
| MistralAI-User | Mistral (user-requested fetch) | allowed |
| DuckAssistBot | DuckDuckGo (DuckAssist answers) | allowed |
| Search crawlers | ||
| OAI-SearchBot | OpenAI (ChatGPT search) | allowed |
| Claude-SearchBot | Anthropic (Claude search) | allowed |
| PerplexityBot (named) | Perplexity (search index) | allowed |
| Googlebot | Google Search | allowed |
| Bingbot | Microsoft Bing (also feeds ChatGPT search) | allowed |
| Training crawlers | ||
| GPTBot (named) | OpenAI (training) | allowed |
| ClaudeBot | Anthropic (training) | allowed |
| Google-Extended (named) | Google (Gemini training and grounding control) | allowed |
| CCBot (named) | Common Crawl (open dataset used to train many models) | allowed |
| Bytespider | ByteDance | allowed |
| Applebot-Extended | Apple (AI training control) | allowed |
| meta-externalagent | Meta (AI crawler) | allowed |
| anthropic-ai (named) | Anthropic (legacy token) | allowed |
| cohere-ai (named) | Cohere | allowed |
| Amazonbot | Amazon | allowed |
"Blocked" means the site's robots.txt disallows the whole site for that agent, by name or through the * group. It is a request: OpenAI says robots.txt may not apply to user-initiated ChatGPT requests, and sites can also block agents at the network edge, which robots.txt does not show.
Show it on your site
updates with each census
<a href="https://grade.agentexchange.work/access/cloudflare.com"><img src="https://grade.agentexchange.work/access/cloudflare.com.svg" alt="AI access: cloudflare.com" height="20"></a>
Is this your site?
Get the full agent-readiness grade Change these rules · robots.txt generator for AI crawlers