Agent-Readiness Grade / Reports / ai-crawler-census-2026-09
AI crawlers at the edge of one zone, Aug 29 to Sep 28 2026
Thirty days of edge logs from one Cloudflare zone: 34,052 requests from named AI crawlers, spoofed ChatGPT-User and Claude-User strings probing /.env, real crawlers spending their budget on robots.txt and sitemaps, and new hostnames found within hours through certificate logs.
Published 2026-09-29 by Agent Exchange. Measurement window 2026-08-29 to 2026-09-28; data pulled 2026-09-28.
Method
Source: Cloudflare GraphQL httpRequestsAdaptiveGroups for the zone agentexchange.work, grouped by user-agent string, for the 30 days from 2026-08-29 to 2026-09-28. The top 100 user agents account for 1,468,345 requests of about 1.66 million in total. Requests were classified by their user-agent string, so every count here is a lower bound for spoofing: a client can claim any name. "AI crawlers and answer engines" counts 15,403 requests within the top-100 sample and 34,052 when the crawler names are matched across all user agents. The zone had 29 subdomains during the window. Three new hosts (inference, tools, memory) created on 2026-09-28 were watched separately for their first 48 hours (2026-09-26T22:40Z to 2026-09-28T22:40Z) to see who finds a hostname nobody has linked to.
Table 1. Who sends the requests (top-100 user agents, 30 days)
| Kind | Requests | Share |
|---|---|---|
| Agent-economy monitors and directory crawlers | 511,520 | 34.8% |
| Our own Sentinel probing our own store | 289,532 | 19.7% |
| Scripted clients (python-requests, httpx, aiohttp, node, curl) | 260,732 | 17.8% |
| Browser-like user agents (humans plus headless scanners) | 240,825 | 16.4% |
| Other | 102,286 | 7.0% |
| Blank user agent | 44,053 | 3.0% |
| AI crawlers and answer engines (top-100 sample) | 15,403 | 1.05% |
| Search and SEO crawlers | 3,994 | 0.27% |
Table 2. Named crawlers matched across all user agents (34,052 requests)
Name matching catches every string containing the token, including search and SEO crawlers that were matched alongside the AI ones. The kind column links to the vendor-documented purpose where we have a rule page.
| User agent | Requests | Documented kind |
|---|---|---|
| Amazonbot | 8,318 | training |
| ClaudeBot | 5,568 | training |
| SemrushBot | 2,717 | |
| OAI-SearchBot | 2,197 | search |
| GPTBot | 1,957 | training |
| Applebot | 1,937 | |
| AhrefsBot | 1,411 | |
| Grok | 1,131 | |
| ChatGPT-User | 1,000 | user-triggered fetch |
| CCBot | 950 | open web archive |
| Googlebot | 896 | |
| PerplexityBot | 804 | search |
| bingbot | 681 | |
| xAI | 537 | |
| Bytespider | 521 | not published |
| YouBot | 479 | |
| DuckAssistBot | 436 | |
| Mistral | 412 | |
| cohere-ai | 407 | |
| Google-Extended | 347 | training and grounding control token |
| GoogleOther | 347 | |
| Perplexity-User | 343 | |
| Claude-User | 323 | user-triggered fetch |
| Claude-SearchBot | 288 | |
| meta-externalagent | 45 |
Findings
1. Real crawlers spend their budget on robots.txt and sitemaps
ClaudeBot fetched the store host's three sitemaps 1,029 times and its robots.txt 333 times in the window. OAI-SearchBot fetched robots.txt on 14 subdomains, between 66 and 108 times each. GPTBot's most-fetched content page was try.agentexchange.work/, with 113 requests. With 29 subdomains, most of them without a checkout, the crawl budget is spread across pages that cannot convert a visit into anything.
2. The user-triggered agent strings are spoofed
Requests carrying ChatGPT-User, Claude-User and Perplexity-User had, as their top paths, /.env, /.ssh/id_ed25519, /actuator/... and /gradle.properties, all answered 429 or 404. Only 36 ChatGPT-User hits reached a real page (the root) and 18 Claude-User hits reached the aivisibility host's /mcp. There is no measurable "an assistant fetched our page for a user" traffic in this zone; the strings are worn by vulnerability scanners. OpenAI, Anthropic and Perplexity all publish IP lists for their fetchers (see the rule pages), which is the only way to tell the real ones apart.
3. New hostnames are found within hours, by security scanners and ClaudeBot
Certificate Transparency logs publish every new custom hostname. Within hours of creation, leakix, ForestEngine, TrashHound and ClaudeBot arrived on the new hosts. ClaudeBot crawled both product hosts inside a day with no link pointing at them. The x402-economy monitors (CarbonMonitor, SentinelOracle, mcpbeat, 402explorer) did not find the new hosts; only x402-list-monitor came, because of the one directory submission made.
Table 3. First 48 hours of three new hostnames (created 2026-09-28)
| Host | Requests | Who came | What they fetched |
|---|---|---|---|
| inference.agentexchange.work | 307 | leakix l9scan 131, Windows-Chrome-like 79, x402-list-monitor 12 (the one submission made), curl 10, ClaudeBot 6, ForestEngine 4, TrashHound wildcard resolver 3 | / 75, /.well-known/x402 12, /v1/models 10, scanner paths (/.env, /config.json) |
| tools.agentexchange.work | 207 | leakix 131, browser-like 40, ClaudeBot 6, ForestEngine 4, TrashHound 3 | / 49, /skill/SKILL.md 7, /.well-known/x402 6, /robots.txt 5 |
| memory.agentexchange.work (minutes old) | 25 | our own probes 14, ForestEngine 4, TrashHound 3 | / and 401s |
4. AI crawlers are a rounding error next to the agent-economy monitors
Monitors and directory crawlers built around x402 and MCP sent 511,520 requests (34.8% of the sample). They probe liveness and prices; they do not buy: the store issued 567,251 payment challenges in the window and settled 54 payments (0.01%). Our own Sentinel added 289,532 requests (19.7%) verifying a store that sold nothing. On the aivisibility host, the free MCP tool received 56,968 /mcp requests against 221 human /check calls.
Table 4. Named monitors and scripted clients (requests, 30 days)
| User agent | Requests |
|---|---|
| CarbonMonitor/0.1 healthcheck | 145,114 |
| sentinel-verify/1.0 (ours) | 136,504 |
| node | 132,447 |
| sentinel/1.0 (ours) | 129,454 |
| NicheAgentEarner/0.11 | 63,681 |
| SentinelOracle/0.1 (glimind.com) | 61,110 |
| x402-list-monitor/1.0 | 54,077 |
| mcpbeat/0.1 | 36,631 |
| curl/8.7.1 | 30,207 |
| python-requests (two versions) | 38,839 |
| sentinel-paywall-probe (ours) | 23,574 |
| python-httpx | 21,039 |
| mako-pulse-prober | 19,752 |
| 402explorer/0.1 (discover.paygent.net) | 19,713 |
| forum-labs-trust-prober | 19,401 |
| mpp32-indexer (agentrateindicators.com) | 19,140 |
| JarvisClaw-HealthCheck | 17,314 |
| agent-tools.cloud-crawler | 13,241 |
| aiohttp | 12,535 |
| nohumans.directory-probe | 10,909 |
| AgenstryBot (compare host) | 3,373 |
| rokmcp-collector | 2,322 |
| OvernightMoneyScout (try host) | 1,725 |
Limitations
- Classification is by user-agent string only; no IP or reverse-DNS verification was applied, so the AI-crawler counts include spoofed strings and the monitor counts include anything that named itself honestly.
- Table 1 covers the top 100 user agents (1,468,345 of about 1.66 million requests); the long tail is not classified.
- One zone, one operator, thirty days. The zone sells to agents and runs its own probes, which inflates machine traffic relative to a content site.
- The 48-hour discovery measurement covers three hosts created on one day.
What to do with this if you run a site
- Verify user-triggered fetchers by IP list, not by name: ChatGPT-User, Claude-User.
- Keep robots.txt and sitemaps small and correct; they are the pages real crawlers fetch most. Check yours with a free Agent-Readiness Grade.
- Expect a new hostname to be crawled within a day of its certificate being issued, whether or not you link to it.