Agentic web census · AI access lookup · rank #7062 · fetched 2026-09-30
Can AI assistants read hackernoon.com?
hackernoon.com's robots.txt tells Perplexity-User not to open its pages.
- Training crawlers blocked: GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, meta-externalagent, anthropic-ai, cohere-ai, Amazonbot
- llms.txt: published
- A2A agent card: not published
- security.txt: not published
| AI agent | Operator and purpose | robots.txt verdict |
|---|---|---|
| User-triggered assistants | ||
| ChatGPT-User (named) | OpenAI (a user asked ChatGPT to open a page) | allowed |
| Claude-User (named) | Anthropic (a user asked Claude to open a page) | allowed |
| Perplexity-User | Perplexity (a user asked Perplexity to open a page) | blocked |
| MistralAI-User (named) | Mistral (user-requested fetch) | blocked |
| DuckAssistBot (named) | DuckDuckGo (DuckAssist answers) | allowed |
| Search crawlers | ||
| OAI-SearchBot (named) | OpenAI (ChatGPT search) | allowed |
| Claude-SearchBot (named) | Anthropic (Claude search) | allowed |
| PerplexityBot (named) | Perplexity (search index) | allowed |
| Googlebot (named) | Google Search | allowed |
| Bingbot (named) | Microsoft Bing (also feeds ChatGPT search) | allowed |
| Training crawlers | ||
| GPTBot (named) | OpenAI (training) | blocked |
| ClaudeBot (named) | Anthropic (training) | blocked |
| Google-Extended (named) | Google (Gemini training and grounding control) | blocked |
| CCBot (named) | Common Crawl (open dataset used to train many models) | blocked |
| Bytespider (named) | ByteDance | blocked |
| Applebot-Extended (named) | Apple (AI training control) | allowed |
| meta-externalagent (named) | Meta (AI crawler) | blocked |
| anthropic-ai (named) | Anthropic (legacy token) | blocked |
| cohere-ai (named) | Cohere | blocked |
| Amazonbot (named) | Amazon | blocked |
"Blocked" means the site's robots.txt disallows the whole site for that agent, by name or through the * group. It is a request: OpenAI says robots.txt may not apply to user-initiated ChatGPT requests, and sites can also block agents at the network edge, which robots.txt does not show.
Show it on your site
updates with each census
<a href="https://grade.agentexchange.work/access/hackernoon.com"><img src="https://grade.agentexchange.work/access/hackernoon.com.svg" alt="AI access: hackernoon.com" height="20"></a>
Is this your site?
Get the full agent-readiness grade Change these rules · robots.txt generator for AI crawlers