XML Sitemaps: How to Build, Validate and Submit One That Google Uses
Published · 9 min read
An XML sitemap is a list of the URLs on your site that you want search engines to know about, with an optional last-modified date for each. Google treats it as a discovery aid, not a ranking signal. A good sitemap contains only canonical, indexable pages that return HTTP 200, stays under 50,000 URLs and 50 MB per file, carries a lastmod you can stand behind, is referenced from robots.txt, and is submitted once in Search Console.
What a sitemap is, and what it is not
Google's sitemap overview describes a sitemap as a file where you provide information about the pages, videos and other files on your site, which search engines read to crawl the site more efficiently. The same page states that using a sitemap does not guarantee that the items in it will be crawled or indexed. In our reading, those two sentences settle what a sitemap is for: it tells Google which URLs exist and, if you say so honestly, when they last changed.
What a sitemap does not do:
- It does not raise the ranking of any URL listed in it.
- It does not force Google to index a page that Google considers low value or duplicate.
- It does not override a
noindextag, arobots.txtblock, or a canonical tag that points elsewhere. - It does not replace internal links. Google's documentation says a sitemap helps most when a site is large, new, or has pages that are poorly linked to one another.
A sitemap is cheap insurance. It is not a strategy.
A minimal valid sitemap.xml
The format is defined by the sitemaps.org protocol, which Google, Bing and other engines follow. A file with two URLs looks like this:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-01</lastmod>
</url>
<url>
<loc>https://www.example.com/pricing</loc>
<lastmod>2026-08-14</lastmod>
</url>
</urlset>The rules that matter, taken from Google's build-a-sitemap page:
- The file must be UTF-8 encoded.
- Each
<loc>must be a full, absolute URL including the scheme and host./pricingis not valid;https://www.example.com/pricingis. - Characters such as
&in URLs must be entity-escaped (&). - A sitemap can only list URLs at or below the directory it lives in. A sitemap at
/blog/sitemap.xmlmay not list/pricing. Placing the file at the site root avoids the problem entirely. - Use one host consistently. If the canonical site is
https://www.example.com, do not listhttp://example.com/…URLs.
The limits Google documents
These numbers come from Google's large sitemaps page and the build page linked above. They are hard limits; a file that exceeds them is rejected, not truncated.
| Limit | Value | Notes |
|---|---|---|
| URLs per sitemap file | 50,000 | Counted per file, not per site |
| Size per sitemap file | 50 MB uncompressed | You may gzip the file; the limit applies to the uncompressed content |
| Sitemaps per index file | 50,000 | An index may reference sitemaps on other hosts if you own them in Search Console |
| Size per index file | 50 MB uncompressed | Same rule as sitemap files |
| Nesting | One level | An index may list sitemaps, not other indexes |
A site with 30,000 pages fits in one file. A site with 120,000 pages needs at least three files and an index. Splitting earlier, by content type, is common and makes the Search Console reports easier to read; Google does not require it.
What belongs in a sitemap, and what does not
The test for every URL: would you be happy to see it in Google's index exactly as listed? If not, it does not belong in the sitemap.
| Include | Exclude |
|---|---|
| Canonical URLs (the version the canonical tag points to) | Non-canonical duplicates, including http:// and non-www variants |
| Pages that return HTTP 200 | URLs that redirect (301, 302) or return 404, 410, 5xx |
Indexable pages with no noindex directive | Pages marked noindex or blocked in robots.txt |
| Pages you want users to land on from search | Parameter URLs (?sort=, ?utm_, session IDs), internal search results, filter combinations |
| Paginated pages if each has distinct, useful content | Login, cart, checkout, account and thank-you pages |
The canonical rule deserves emphasis. If /shoes?color=red has a canonical tag pointing to /shoes, only /shoes goes in the sitemap. Google's documentation lists the sitemap as one of the signals it uses to choose a canonical, so listing both versions sends a mixed message. Our canonical tags guide covers how the two signals interact.
lastmod, priority and changefreq: what Google actually reads
The protocol defines three optional fields per URL. Google reads one of them, and only conditionally.
lastmod. Google's build-a-sitemap page says it uses the lastmod value if it is consistently and verifiably accurate, and gives the example of comparing it with the page's last modification. So a sitemap where every lastmod equals the build timestamp is worse than one with no lastmod at all: Google learns the field is unreliable for your site and stops trusting it. Set it from the content's real last-edited date, in W3C datetime format (2026-09-01 or 2026-09-01T10:15:00+00:00). Do not bump it for template edits or sidebar changes.
priority and changefreq. Google's documentation states that it ignores both values. They are harmless, but any effort spent tuning them is wasted. Omit them.
Sitemap index files for large sites
When a site needs more than one sitemap file, or when you want to separate content types, you publish a sitemap index: a file that lists sitemap files instead of page URLs.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemap-pages.xml</loc>
<lastmod>2026-09-01</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemap-blog.xml</loc>
<lastmod>2026-09-12</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemap-products-1.xml</loc>
<lastmod>2026-09-13</lastmod>
</sitemap>
</sitemapindex>The lastmod on an index entry should be the most recent lastmod inside that child sitemap, and the same honesty rule applies. You submit the index URL to Search Console once; Google fetches each child and reports on them separately.
Referencing the sitemap in robots.txt
Search Console submission tells Google. The Sitemap: line in robots.txt tells every crawler that reads the file, including Bing and the AI search crawlers that fetch robots.txt before crawling. Google's robots.txt documentation supports the directive; it takes a full URL, may appear anywhere in the file, and may be repeated.
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xmlPoint it at the index if you have one. Because the reference is a full URL, the sitemap may live on a different host or path than the directory rule above would otherwise allow. While you are in robots.txt, confirm you have not blocked the sitemap path itself, and check what the file says to AI crawlers; our robots.txt guide for AI crawlers walks through that.
Submitting in Search Console and reading the reports
Submission is a one-time step per sitemap URL:
- Open Search Console, choose the property, and go to Indexing → Sitemaps.
- Enter the path relative to the property (
sitemap.xmlorsitemap-index.xml) and click Submit. - Wait. The report updates after Google fetches the file, which can take minutes or days. Resubmitting does nothing; Google re-fetches on its own schedule, and your
lastmodvalues tell it what changed.
Google's Sitemaps report help lists three statuses: Success, Has errors, and Couldn't fetch. The report also shows how many URLs were discovered from the file. That count is discovery, not indexing.
For indexing, open Indexing → Pages and filter to All submitted pages. This is the view that matters. It shows, for the URLs in your sitemap only, which are indexed and which are not, with a reason for each. Google documents every reason in the Page indexing report help. Reasons such as Excluded by noindex tag, Page with redirect or Alternate page with proper canonical tag appearing under submitted pages mean the sitemap contains URLs that fail the include/exclude table above. Fix the sitemap, not the pages.
Troubleshooting
| Symptom | Usual cause | What to do |
|---|---|---|
| Couldn't fetch | Wrong path, sitemap blocked in robots.txt, non-200 response, bot protection challenging Googlebot, or Google has not tried yet | Check the URL with curl -I; confirm HTTP 200 and an XML content type; confirm robots.txt does not disallow the path; wait 24 hours before assuming failure |
| Has errors: invalid XML or unsupported format | Malformed XML, unescaped & in a URL, wrong namespace, HTML error page served at the sitemap URL, byte-order mark or whitespace before the XML declaration | Validate the file with an XML parser; view source rather than the rendered page; confirm the first bytes are <?xml |
| Has errors: URL not allowed | Sitemap lists URLs on another host or above its own path | Move the sitemap to the site root, or reference it from robots.txt |
| Discovered – currently not indexed | Google knows the URL from the sitemap but postponed crawling, often to avoid overloading the site, or because it judged the URL low priority | Check server response times; remove thin, near-duplicate and parameter URLs from the sitemap; add internal links to the pages that matter |
| Crawled – currently not indexed | Google fetched the page and chose not to index it; a quality judgement, not a sitemap problem | Improve or consolidate the page; a sitemap change will not fix this |
| Submitted URL marked noindex / has redirect / not found | The sitemap lists URLs that fail the include/exclude rules | Regenerate the sitemap from the set of canonical, 200, indexable pages only |
In our reading of Google's help text, Discovered is a crawl-priority signal and Crawled is a quality signal. Neither is fixed by resubmitting the sitemap. Both are usually improved by having fewer, better URLs in it.
Where the sitemap fits in a technical check
A sitemap is one item on a longer list. Our technical SEO checklist puts it alongside canonical tags, robots.txt, HTTPS and response codes, all of which make the site legible to crawlers before any content work begins. The free scan on our home page checks whether a sitemap exists and whether robots.txt references one, along with the robots.txt rules for 22 AI crawlers and the canonical tag on the page you enter. It reads one page and one robots.txt; it does not crawl the site or validate the URLs inside the sitemap, so the Search Console reports above remain the source of truth for that. The full list is in the 18 checks explained.
Frequently asked questions
Does an XML sitemap improve rankings?
No. Google's documentation describes a sitemap as a way to help Google discover and crawl URLs more efficiently, and states that including a URL does not guarantee it will be crawled or indexed. A sitemap can get a page found sooner; it does not make that page rank higher. Ranking depends on the page itself and on the signals pointing to it.
How many URLs can one sitemap file contain?
Google's documented limit is 50,000 URLs or 50 MB uncompressed per sitemap file, whichever is reached first. A site with more URLs splits them across several sitemap files and lists those files in a sitemap index, which itself may reference up to 50,000 sitemaps and be up to 50 MB.
Should I include lastmod, changefreq and priority?
Include lastmod only if it reflects the date the page content genuinely last changed. Google says it uses the lastmod value when it is consistently and verifiably accurate, and ignores it when it is not. Google's documentation states that it ignores changefreq and priority, so they can be omitted; they do no harm, but they do nothing.
What does 'Discovered – currently not indexed' mean for sitemap URLs?
Google found the URL, usually through the sitemap, but has not crawled it yet. Google's Page indexing report help says this typically happens because crawling was postponed, often to avoid overloading the site, and that the URL will be revisited. If it persists across many URLs, the usual causes are thin or duplicate pages, slow responses, or too many low-value URLs in the sitemap.