问答
How many URLs can one sitemap hold?
A single sitemap file is capped at 50,000 URLs and 50 MB uncompressed. Past that you split with a sitemap index, and how you split matters.
更新于 2026年4月3日 · 约 4 分钟
The short answer: 50,000 URLs and 50 MB uncompressed per file. Those are the two hard limits in the sitemaps.org protocol, and they apply to every sitemap, including a sitemap index. Exceed either and the file is rejected rather than trimmed.
The two limits
| Limit | Value | Applies to |
|---|---|---|
| URLs per sitemap | 50,000 | Every <urlset> file |
| Uncompressed size | 50 MB (52,428,800 bytes) | Every sitemap file |
| Sitemaps per index | 50,000 | <sitemapindex> files |
| Encoding | UTF-8 | All sitemap files |
You may serve the file gzipped — sitemap.xml.gz is common — but the 50 MB cap is measured on the uncompressed bytes. A gzip file that expands past the limit still fails.
The size limit bites before the URL limit
On a content site, 50 MB is reached well before 50,000 URLs. Each <url> block with a <loc>, a <lastmod> and indentation runs roughly 120–160 bytes, so a technical sitemap with image extensions or long URLs can approach the byte limit at a fraction of the URL count. If you are adding <image:image> or <news:news> entries, watch the size, not just the count.
Splitting with a sitemap index
When you outgrow one file, the fix is a sitemap index: a file that lists other sitemap files. It uses <sitemapindex> instead of <urlset>:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://example.com/sitemaps/products-1.xml</loc>
<lastmod>2026-03-30</lastmod>
</sitemap>
<sitemap>
<loc>https://example.com/sitemaps/blog-1.xml</loc>
<lastmod>2026-04-01</lastmod>
</sitemap>
</sitemapindex>
The index is what you submit in Search Console and what you reference in robots.txt under Sitemap:. The child files are only discovered through it.
Choosing a sensible granularity
You are not required to fill each file to 50,000 URLs. The limit is a ceiling, not a target. Common, defensible splits:
| Split by | Why it usually works |
|---|---|
| Content type (products, posts, docs) | A single generator owns each file; regeneration is cheap |
| Section or locale | Errors stay isolated to one part of the site |
| Rolling groups of a few thousand | Keeps diffs small in version control |
| Static vs dynamic | Pages that never change stop being rewritten on every deploy |
A file of a few thousand URLs regenerates fast, is easy to diff, and, when it fails validation, points you straight at the section that broke. Pushing every URL into one 50,000-entry file makes every one of those jobs harder for no benefit.
What happens when you go over the limit
A file that breaches either limit is not trimmed to fit. The parser rejects the document, so none of its entries are read — a sitemap holding 60,000 URLs delivers zero URLs, not 50,000. That failure is why the size check matters as much as the count: an over-limit file looks correct in a browser and produces nothing in a crawler.
The rejection is usually silent. The server returns a 200 because it is happy to serve the bytes; the failure happens when a crawler tries to parse them. That is the argument for validating the file rather than eyeballing it.
Do more URLs get crawled faster?
No. The number of URLs in a sitemap does not buy crawl priority. Crawl attention follows the site’s overall signals — link equity, freshness and response times — not the length of the list. A 50,000-entry sitemap of slow pages and errors is crawled more slowly than a 5,000-entry sitemap of clean 200s. The URL limit is a ceiling to stay under, not a target to fill.
Common mistakes
Splitting at exactly 50,000. The next page you publish breaks the file. Leave headroom.
Counting URLs but ignoring bytes. A file under 50,000 URLs can still exceed 50 MB once image or news extensions are added. Measure the uncompressed size.
Submitting child sitemaps individually. When an index exists, submit the index. Listing children too doubles the reporting and splits your coverage numbers.
Nesting indexes. An index must reference <urlset> sitemap files, not other indexes. Chaining index → index → urlset is not supported.
Linking the index from robots.txt but submitting a child. The discovery path and the reported path should point at the same file.
Forgetting lastmod on index entries. It is optional, but a correct value helps crawlers fetch only the children that changed.
Where this tool fits
The sitemap validator reads a sitemap or an index, reports the entry count and file size, and flags entries that break the protocol. Run it before you publish a split so you catch an over-limit file while it is still your local copy.
Frequently asked questions
▸ What is the maximum number of URLs in a sitemap?
50,000 URLs per file, per the sitemaps.org protocol. If you have more, split into multiple sitemap files and link them from a sitemap index.
▸ What is the size limit for a sitemap file?
50 MB uncompressed, or 52,428,800 bytes. You may serve the file gzipped, but the limit applies to the uncompressed content.
▸ How many sitemaps can a sitemap index hold?
The same ceiling applies: an index may reference up to 50,000 sitemap files and must itself stay under 50 MB uncompressed.
▸ Should I split at 50,000 URLs exactly?
No. Splitting at 50,000 means every future URL breaks the file. Most sites split by section or content type at a few thousand URLs per file, which keeps diffs small and errors easy to isolate.
▸ Does a bigger sitemap get crawled slower?
Size is not a direct ranking or speed factor. What matters is that the file is reachable, well under the limits, and that its URLs return 200. A bloated file with redirects and errors is a crawl-budget problem regardless of count.