指南
The sitemap that actually gets crawled
A valid sitemap.xml can still be ignored. Declaring it in robots.txt, submitting it in Search Console and using lastmod decide whether it gets crawled.
更新于 2026年4月2日 · 约 3 分钟
A sitemap.xml is a list of URLs you would like a search engine to crawl. It is a hint, not an instruction, and a file can be perfectly valid XML while being quietly ignored. What separates a sitemap that gets crawled from one that does not is rarely the formatting — it is where the file lives, what it points at, and whether a crawler can trust the URLs inside it.
What a sitemap actually does
Three things, and no more:
- It tells crawlers the listed URLs exist, which matters most for pages that are poorly linked internally.
- It can carry a
lastmodhint so a crawler can prioritise what changed. - It gives you a per-file report in Search Console when something breaks.
It does not force indexing, it does not override noindex, and it does not guarantee a crawl schedule.
Reference it in robots.txt
The fastest win is a single line at the end of your root robots.txt:
# 声明 sitemap 位置,必须是绝对 URL
Sitemap: https://example.com/sitemap.xml
The line must sit in the root robots.txt at https://example.com/robots.txt, and the value must be an absolute URL, not /sitemap.xml. If you use a sitemap index, point at the index, not at every child file.
Submit it in Search Console
Robots.txt discovery and a Search Console submission are complementary. Add the sitemap under Sitemaps, then read the report: it tells you whether the file was fetched, when it was last read, and how many URLs were discovered. A crawl error there is immediate feedback that the file is unreachable or malformed.
List only canonical, 200 URLs
This is where most sitemaps leak. The file should contain exactly the set of pages that are canonical, indexable and served with a 200. Anything else dilutes it.
| URL state | Belongs in the sitemap? |
|---|---|
200, canonical, indexable | Yes |
301 / 302 redirect | No — list the destination instead |
404 / 410 | No |
Page carrying noindex | No |
| Duplicate that canonicalises elsewhere | No — list the canonical URL |
| Paginated series | Usually just the hub, unless each page is independently valuable |
If a URL in your sitemap redirects, replace it with the target. If it 404s, remove it. The same rule applies to URLs that are blocked by robots.txt: a sitemap entry for a disallowed URL cannot be crawled, so it is noise.
Get lastmod right
lastmod is the field people abuse most. Include it only when it reflects a real content change, and use a valid W3C datetime:
<url>
<loc>https://example.com/guide/getting-started</loc>
<lastmod>2026-03-28</lastmod>
</url>
A date-only value is fine. What is not fine is stamping the build time on every URL: if lastmod says today for the whole site after a typo fix in the footer, and the page content never changed, crawlers stop trusting the signal and may ignore it entirely. When in doubt, omit lastmod rather than lie.
Common mistakes
Declaring the sitemap as a relative path. Sitemap: /sitemap.xml is not valid; crawlers expect a full URL.
Submitting each child sitemap separately. If you have a sitemap index, submit the index. Submitting children as well creates duplicate reporting and clutter.
Letting the sitemap list URLs the site cannot serve. A sitemap that 404s half its entries teaches Google that the file is unreliable. Keep it to pages that return 200 without a login.
Setting lastmod on every deploy. Automation makes this easy and wrong. Only bump it when the page content genuinely changed.
Forgetting to update the sitemap after removing pages. Stale entries point at gone URLs. Regenerate the sitemap as part of the deploy, or validate it as a release step.
Assuming submission equals indexing. A submitted sitemap only queues URLs for crawling. Indexing still depends on the page itself.
Where this tool fits
The sitemap validator parses your sitemap.xml for format errors and then crawls every listed URL, reporting which entries redirect, error or return something other than a 200. That turns the table above into a concrete checklist: the URLs that should not be there are the ones it flags.
Frequently asked questions
▸ Why is my sitemap not being crawled?
The usual causes are a sitemap that is not referenced in robots.txt or submitted in Search Console, and one that lists redirected, non-canonical or non-200 URLs. Validate the XML, confirm the file loads without a login, then check what the URLs inside it actually resolve to.
▸ Do I need to submit a sitemap if it is already in robots.txt?
You do not have to, but submitting it in Search Console is still worth doing. The robots.txt reference helps discovery; a submitted sitemap gives you status reporting, error counts and last-read dates in the report.
▸ Should a sitemap list every page on the site?
No. It should list the canonical, indexable, 200-status URLs you want indexed. Listing noindex pages, redirects or error pages wastes crawl attention and can make the file less trusted.
▸ Does lastmod help rankings?
Not directly. Google uses lastmod as a hint about whether a URL is worth re-crawling. When the value is inaccurate or changes on every deploy, crawlers learn to ignore it.
▸ Can a sitemap include a noindex page?
It can be parsed, but it should not be there. A noindex URL in a sitemap is a contradiction: the file asks for the page to be indexed while the page refuses. Google typically ignores the entry, and repeated conflicts reduce trust in the whole file.