# Fix sitemap: publish /sitemap.xml or declare one in robots.txt

> A sitemap is the complete URL inventory of your site. Crawlers, AI crawlers included, use it to find pages no link points to, and it is one of the files they fetch most.

Checked 2026-09-30 against Agent-Readiness Grade 1.3.0. HTML version: https://grade.agentexchange.work/fix/sitemap-xml

## What the grade checks

- `GET https://example.com/sitemap.xml` with a 7-second timeout. It counts when the answer is HTTP 200 and XML (a `<urlset>` or a `<sitemapindex>`), not an HTML page.
- A `Sitemap:` line in robots.txt counts too, wherever the file lives, so a WordPress site whose robots.txt declares `/wp-sitemap.xml` gets the point.

1 of the 2 points in the machine-readable surfaces area.

## Why it matters for AI agents and crawlers

The [sitemap protocol](https://www.sitemaps.org/protocol.html) limits one file to 50,000 URLs and 50 MB uncompressed; larger sites list several sitemaps in a sitemap index. The protocol also defines the `Sitemap:` line in robots.txt, which is how crawlers find a sitemap that is not at `/sitemap.xml`.

In our own [30-day crawler census](https://grade.agentexchange.work/reports/ai-crawler-census-2026-09), sitemaps and robots.txt were what real AI crawlers fetched most: ClaudeBot fetched one host's three sitemaps 1,029 times and its robots.txt 333 times.

[Google](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) documents the same format, so one file serves search engines and AI crawlers alike.

## How to fix it

### 1. Build the file

List every page you want found, with the date it last changed. Most site generators and CMSs can produce it for you.

sitemap.xml:

```xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/</loc>
    <lastmod>2026-09-30</lastmod>
  </url>
  <url>
    <loc>https://example.com/pricing</loc>
    <lastmod>2026-09-12</lastmod>
  </url>
</urlset>
```

### 2. Publish it and declare it in robots.txt

Serve it as XML and add one line to robots.txt so crawlers find it wherever it lives.

#### Static site

Generate it at build time (static site generators have sitemap plugins) or write it by hand for a small site, save it as `sitemap.xml` in the web root, and add the line to robots.txt.

robots.txt:

```text
Sitemap: https://example.com/sitemap.xml
```

#### WordPress

Since WordPress 5.5, core serves a sitemap index at `/wp-sitemap.xml` and references it in the virtual robots.txt ([WordPress core announcement](https://make.wordpress.org/core/2020/07/22/new-xml-sitemaps-functionality-in-wordpress-5-5/)); SEO plugins may replace it with their own. Check that robots.txt carries a `Sitemap:` line; if you uploaded a physical robots.txt, add the line yourself.

robots.txt:

```text
Sitemap: https://example.com/wp-sitemap.xml
```

#### Next.js

Use the [sitemap file convention](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/sitemap) and name it in `app/robots.ts` with `sitemap: 'https://example.com/sitemap.xml'`.

app/sitemap.ts:

```ts
import type { MetadataRoute } from 'next'

export default function sitemap(): MetadataRoute.Sitemap {
  return [
    { url: 'https://example.com/', lastModified: new Date() },
    { url: 'https://example.com/pricing', lastModified: new Date() },
  ]
}
```

#### Cloudflare

Deploy `sitemap.xml` with your static assets. A Worker that renders pages can build it from the same route list:

Cloudflare Worker:

```js
const PAGES = ["/", "/pricing", "/docs"];

function sitemapXml(origin) {
  const urls = PAGES.map((p) => `  <url><loc>${origin}${p}</loc></url>`).join("\n");
  return `<?xml version="1.0" encoding="UTF-8"?>\n<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">\n${urls}\n</urlset>\n`;
}

// inside fetch(): if (url.pathname === "/sitemap.xml")
//   return new Response(sitemapXml(url.origin), { headers: { "content-type": "application/xml; charset=utf-8" } });
```

## How to verify

Check the file and the robots.txt line.

```sh
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://example.com/sitemap.xml
curl -s https://example.com/robots.txt | grep -i '^sitemap:'
```

`200` with an XML content type, or a `Sitemap:` line in robots.txt. Either earns the point.

Re-grade: https://grade.agentexchange.work/grade?url=example.com&fresh=1

## Questions

### Does the sitemap have to be at /sitemap.xml?

No. Any location works when robots.txt declares it with a Sitemap: line, and the grade accepts either.

### How many URLs can one sitemap hold?

50,000 URLs and 50 MB uncompressed per file. Above that, split it and list the parts in a sitemap index.

### Do AI crawlers read sitemaps?

In our 30-day crawler census they did: ClaudeBot fetched one host's three sitemaps 1,029 times, more than any content page.

## Sources

- [Sitemaps XML format](https://www.sitemaps.org/protocol.html) (sitemaps.org)
- [Build and submit a sitemap](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) (Google Search Central)
- [New XML sitemaps functionality in WordPress 5.5](https://make.wordpress.org/core/2020/07/22/new-xml-sitemaps-functionality-in-wordpress-5-5/) (Make WordPress Core)
- [sitemap.xml file convention](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/sitemap) (Next.js)
- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309) (IETF)

## Related

- [robots.txt for AI crawlers](https://grade.agentexchange.work/fix/robots-txt-ai-crawlers.md): robots.txt groups for eight AI crawler tokens; a blanket Disallow: / counts as blocked.
- [llms.txt](https://grade.agentexchange.work/fix/llms-txt.md): A Markdown guide to your key pages at /llms.txt, served as text, not as your HTML 404.
- [All fix guides](https://grade.agentexchange.work/fix)
