Agentic web census · AI access lookup · rank #3813 · fetched 2026-09-30
Can AI assistants read wiki.gg?
wiki.gg's robots.txt lets ChatGPT, Claude and Perplexity open pages when a user asks.
- Training crawlers blocked: GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Applebot-Extended, meta-externalagent, anthropic-ai, cohere-ai, Amazonbot
- llms.txt: published
- A2A agent card: not published
- security.txt: published
- Content-Signal:
search=yes,ai-train=no,use=reference - robots.txt looks like Cloudflare's managed file
| AI agent | Operator and purpose | robots.txt verdict |
|---|---|---|
| User-triggered assistants | ||
| ChatGPT-User | OpenAI (a user asked ChatGPT to open a page) | allowed |
| Claude-User | Anthropic (a user asked Claude to open a page) | allowed |
| Perplexity-User | Perplexity (a user asked Perplexity to open a page) | allowed |
| MistralAI-User | Mistral (user-requested fetch) | allowed |
| DuckAssistBot | DuckDuckGo (DuckAssist answers) | allowed |
| Search crawlers | ||
| OAI-SearchBot | OpenAI (ChatGPT search) | allowed |
| Claude-SearchBot | Anthropic (Claude search) | allowed |
| PerplexityBot | Perplexity (search index) | allowed |
| Googlebot | Google Search | allowed |
| Bingbot | Microsoft Bing (also feeds ChatGPT search) | allowed |
| Training crawlers | ||
| GPTBot (named) | OpenAI (training) | blocked |
| ClaudeBot (named) | Anthropic (training) | blocked |
| Google-Extended (named) | Google (Gemini training and grounding control) | blocked |
| CCBot (named) | Common Crawl (open dataset used to train many models) | blocked |
| Bytespider (named) | ByteDance | blocked |
| Applebot-Extended (named) | Apple (AI training control) | blocked |
| meta-externalagent (named) | Meta (AI crawler) | blocked |
| anthropic-ai (named) | Anthropic (legacy token) | blocked |
| cohere-ai (named) | Cohere | blocked |
| Amazonbot (named) | Amazon | blocked |
"Blocked" means the site's robots.txt disallows the whole site for that agent, by name or through the * group. It is a request: OpenAI says robots.txt may not apply to user-initiated ChatGPT requests, and sites can also block agents at the network edge, which robots.txt does not show.
Show it on your site
updates with each census
<a href="https://grade.agentexchange.work/access/wiki.gg"><img src="https://grade.agentexchange.work/access/wiki.gg.svg" alt="AI access: wiki.gg" height="20"></a>
Is this your site?
Get the full agent-readiness grade Change these rules · robots.txt generator for AI crawlers