Agent-Readiness Grade

Agent-Readiness Grade / Fix guides / sitemap.xml

Fix sitemap: publish /sitemap.xml or declare one in robots.txt

A sitemap is the complete URL inventory of your site. Crawlers, AI crawlers included, use it to find pages no link points to, and it is one of the files they fetch most.

Checked 2026-09-30 against grader version 1.3.0 and the 5 sources listed below.

Area: Machine-readable surfaces. 1 of the 2 points in the machine-readable surfaces area.

What the grade checksWhy it mattersHow to fix itVerifyQuestionsSources

What the grade checks

  • GET https://example.com/sitemap.xml with a 7-second timeout. It counts when the answer is HTTP 200 and XML (a <urlset> or a <sitemapindex>), not an HTML page.
  • A Sitemap: line in robots.txt counts too, wherever the file lives, so a WordPress site whose robots.txt declares /wp-sitemap.xml gets the point.

1 of the 2 points in the machine-readable surfaces area. The full scoring rules are on the methodology page.

Why it matters for AI agents and crawlers

The sitemap protocol limits one file to 50,000 URLs and 50 MB uncompressed; larger sites list several sitemaps in a sitemap index. The protocol also defines the Sitemap: line in robots.txt, which is how crawlers find a sitemap that is not at /sitemap.xml.

In our own 30-day crawler census, sitemaps and robots.txt were what real AI crawlers fetched most: ClaudeBot fetched one host's three sitemaps 1,029 times and its robots.txt 333 times.

Google documents the same format, so one file serves search engines and AI crawlers alike.

How to fix it

  1. Build the file

    List every page you want found, with the date it last changed. Most site generators and CMSs can produce it for you.

    sitemap.xml

    <?xml version="1.0" encoding="UTF-8"?>
    <urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
      <url>
        <loc>https://example.com/</loc>
        <lastmod>2026-09-30</lastmod>
      </url>
      <url>
        <loc>https://example.com/pricing</loc>
        <lastmod>2026-09-12</lastmod>
      </url>
    </urlset>
  2. Publish it and declare it in robots.txt

    Serve it as XML and add one line to robots.txt so crawlers find it wherever it lives.

    Static site

    Generate it at build time (static site generators have sitemap plugins) or write it by hand for a small site, save it as sitemap.xml in the web root, and add the line to robots.txt.

    robots.txt

    Sitemap: https://example.com/sitemap.xml

    WordPress

    Since WordPress 5.5, core serves a sitemap index at /wp-sitemap.xml and references it in the virtual robots.txt (WordPress core announcement); SEO plugins may replace it with their own. Check that robots.txt carries a Sitemap: line; if you uploaded a physical robots.txt, add the line yourself.

    robots.txt

    Sitemap: https://example.com/wp-sitemap.xml

    Next.js

    Use the sitemap file convention and name it in app/robots.ts with sitemap: 'https://example.com/sitemap.xml'.

    app/sitemap.ts

    import type { MetadataRoute } from 'next'
    
    export default function sitemap(): MetadataRoute.Sitemap {
      return [
        { url: 'https://example.com/', lastModified: new Date() },
        { url: 'https://example.com/pricing', lastModified: new Date() },
      ]
    }

    Cloudflare

    Deploy sitemap.xml with your static assets. A Worker that renders pages can build it from the same route list:

    Cloudflare Worker

    const PAGES = ["/", "/pricing", "/docs"];
    
    function sitemapXml(origin) {
      const urls = PAGES.map((p) => `  <url><loc>${origin}${p}</loc></url>`).join("\n");
      return `<?xml version="1.0" encoding="UTF-8"?>\n<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">\n${urls}\n</urlset>\n`;
    }
    
    // inside fetch(): if (url.pathname === "/sitemap.xml")
    //   return new Response(sitemapXml(url.origin), { headers: { "content-type": "application/xml; charset=utf-8" } });

How to verify

Check the file and the robots.txt line.

shell

curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://example.com/sitemap.xml
curl -s https://example.com/robots.txt | grep -i '^sitemap:'

200 with an XML content type, or a Sitemap: line in robots.txt. Either earns the point.

Then re-grade your site: the result lists the evidence for this check.

$49 AI Visibility Full Report

The fixes on this site are free. The paid next step is the $49 AI Visibility Full Report (what ChatGPT, Claude and Perplexity say about your brand, with a prioritized fix list) from aivisibility.agentexchange.work. It includes:

  • 8 real buyer questions tested across ChatGPT-class models
  • Competitor share-of-voice: who AI names, how often, versus you
  • Full GEO site audit with prioritized, specific fixes
  • Agent-Readiness Score: crawler access, llms.txt, schema, discovery manifest
  • Custom 30/60/90-day action plan to get cited by ChatGPT, Perplexity and Google AI Overviews
  • Shareable report, generated in about 60 seconds after checkout

Get the Full Report — $49 Stripe checkout; you enter your brand and site right after paying.

Questions

Does the sitemap have to be at /sitemap.xml?

No. Any location works when robots.txt declares it with a Sitemap: line, and the grade accepts either.

How many URLs can one sitemap hold?

50,000 URLs and 50 MB uncompressed per file. Above that, split it and list the parts in a sitemap index.

Do AI crawlers read sitemaps?

In our 30-day crawler census they did: ClaudeBot fetched one host's three sitemaps 1,029 times, more than any content page.

Sources

Primary documentation, read 2026-09-30. Vendors change these pages; follow the link before relying on a detail.