Agent-Readiness Grade

Agent-Readiness Grade / Reports / ai-crawler-census-2026-09

AI crawlers at the edge of one zone, Aug 29 to Sep 28 2026

Thirty days of edge logs from one Cloudflare zone: 34,052 requests from named AI crawlers, spoofed ChatGPT-User and Claude-User strings probing /.env, real crawlers spending their budget on robots.txt and sitemaps, and new hostnames found within hours through certificate logs.

Published 2026-09-29 by Agent Exchange. Measurement window 2026-08-29 to 2026-09-28; data pulled 2026-09-28.

Method

Source: Cloudflare GraphQL httpRequestsAdaptiveGroups for the zone agentexchange.work, grouped by user-agent string, for the 30 days from 2026-08-29 to 2026-09-28. The top 100 user agents account for 1,468,345 requests of about 1.66 million in total. Requests were classified by their user-agent string, so every count here is a lower bound for spoofing: a client can claim any name. "AI crawlers and answer engines" counts 15,403 requests within the top-100 sample and 34,052 when the crawler names are matched across all user agents. The zone had 29 subdomains during the window. Three new hosts (inference, tools, memory) created on 2026-09-28 were watched separately for their first 48 hours (2026-09-26T22:40Z to 2026-09-28T22:40Z) to see who finds a hostname nobody has linked to.

Table 1. Who sends the requests (top-100 user agents, 30 days)

KindRequestsShare
Agent-economy monitors and directory crawlers511,52034.8%
Our own Sentinel probing our own store289,53219.7%
Scripted clients (python-requests, httpx, aiohttp, node, curl)260,73217.8%
Browser-like user agents (humans plus headless scanners)240,82516.4%
Other102,2867.0%
Blank user agent44,0533.0%
AI crawlers and answer engines (top-100 sample)15,4031.05%
Search and SEO crawlers3,9940.27%

Table 2. Named crawlers matched across all user agents (34,052 requests)

Name matching catches every string containing the token, including search and SEO crawlers that were matched alongside the AI ones. The kind column links to the vendor-documented purpose where we have a rule page.

User agentRequestsDocumented kind
Amazonbot8,318training
ClaudeBot5,568training
SemrushBot2,717
OAI-SearchBot2,197search
GPTBot1,957training
Applebot1,937
AhrefsBot1,411
Grok1,131
ChatGPT-User1,000user-triggered fetch
CCBot950open web archive
Googlebot896
PerplexityBot804search
bingbot681
xAI537
Bytespider521not published
YouBot479
DuckAssistBot436
Mistral412
cohere-ai407
Google-Extended347training and grounding control token
GoogleOther347
Perplexity-User343
Claude-User323user-triggered fetch
Claude-SearchBot288
meta-externalagent45

Findings

1. Real crawlers spend their budget on robots.txt and sitemaps

ClaudeBot fetched the store host's three sitemaps 1,029 times and its robots.txt 333 times in the window. OAI-SearchBot fetched robots.txt on 14 subdomains, between 66 and 108 times each. GPTBot's most-fetched content page was try.agentexchange.work/, with 113 requests. With 29 subdomains, most of them without a checkout, the crawl budget is spread across pages that cannot convert a visit into anything.

2. The user-triggered agent strings are spoofed

Requests carrying ChatGPT-User, Claude-User and Perplexity-User had, as their top paths, /.env, /.ssh/id_ed25519, /actuator/... and /gradle.properties, all answered 429 or 404. Only 36 ChatGPT-User hits reached a real page (the root) and 18 Claude-User hits reached the aivisibility host's /mcp. There is no measurable "an assistant fetched our page for a user" traffic in this zone; the strings are worn by vulnerability scanners. OpenAI, Anthropic and Perplexity all publish IP lists for their fetchers (see the rule pages), which is the only way to tell the real ones apart.

3. New hostnames are found within hours, by security scanners and ClaudeBot

Certificate Transparency logs publish every new custom hostname. Within hours of creation, leakix, ForestEngine, TrashHound and ClaudeBot arrived on the new hosts. ClaudeBot crawled both product hosts inside a day with no link pointing at them. The x402-economy monitors (CarbonMonitor, SentinelOracle, mcpbeat, 402explorer) did not find the new hosts; only x402-list-monitor came, because of the one directory submission made.

Table 3. First 48 hours of three new hostnames (created 2026-09-28)

HostRequestsWho cameWhat they fetched
inference.agentexchange.work307leakix l9scan 131, Windows-Chrome-like 79, x402-list-monitor 12 (the one submission made), curl 10, ClaudeBot 6, ForestEngine 4, TrashHound wildcard resolver 3/ 75, /.well-known/x402 12, /v1/models 10, scanner paths (/.env, /config.json)
tools.agentexchange.work207leakix 131, browser-like 40, ClaudeBot 6, ForestEngine 4, TrashHound 3/ 49, /skill/SKILL.md 7, /.well-known/x402 6, /robots.txt 5
memory.agentexchange.work (minutes old)25our own probes 14, ForestEngine 4, TrashHound 3/ and 401s

4. AI crawlers are a rounding error next to the agent-economy monitors

Monitors and directory crawlers built around x402 and MCP sent 511,520 requests (34.8% of the sample). They probe liveness and prices; they do not buy: the store issued 567,251 payment challenges in the window and settled 54 payments (0.01%). Our own Sentinel added 289,532 requests (19.7%) verifying a store that sold nothing. On the aivisibility host, the free MCP tool received 56,968 /mcp requests against 221 human /check calls.

Table 4. Named monitors and scripted clients (requests, 30 days)

User agentRequests
CarbonMonitor/0.1 healthcheck145,114
sentinel-verify/1.0 (ours)136,504
node132,447
sentinel/1.0 (ours)129,454
NicheAgentEarner/0.1163,681
SentinelOracle/0.1 (glimind.com)61,110
x402-list-monitor/1.054,077
mcpbeat/0.136,631
curl/8.7.130,207
python-requests (two versions)38,839
sentinel-paywall-probe (ours)23,574
python-httpx21,039
mako-pulse-prober19,752
402explorer/0.1 (discover.paygent.net)19,713
forum-labs-trust-prober19,401
mpp32-indexer (agentrateindicators.com)19,140
JarvisClaw-HealthCheck17,314
agent-tools.cloud-crawler13,241
aiohttp12,535
nohumans.directory-probe10,909
AgenstryBot (compare host)3,373
rokmcp-collector2,322
OvernightMoneyScout (try host)1,725

Limitations

  • Classification is by user-agent string only; no IP or reverse-DNS verification was applied, so the AI-crawler counts include spoofed strings and the monitor counts include anything that named itself honestly.
  • Table 1 covers the top 100 user agents (1,468,345 of about 1.66 million requests); the long tail is not classified.
  • One zone, one operator, thirty days. The zone sells to agents and runs its own probes, which inflates machine traffic relative to a content site.
  • The 48-hour discovery measurement covers three hosts created on one day.

What to do with this if you run a site

  • Verify user-triggered fetchers by IP list, not by name: ChatGPT-User, Claude-User.
  • Keep robots.txt and sitemaps small and correct; they are the pages real crawlers fetch most. Check yours with a free Agent-Readiness Grade.
  • Expect a new hostname to be crawled within a day of its certificate being issued, whether or not you link to it.