实操
How to generate a sitemap (and check it before you ship)
Hand-write one, or let your framework emit it. Concrete Astro and Next.js setups, when each is right, and the self-check to run before the file goes live.
更新于 2026年4月7日 · 约 4 分钟
A sitemap is only as good as the page list behind it. The moment the list changes — a new post, a retired product — a hand-written file goes stale. That is the whole argument for generating it from the same source that generates your pages.
Which approach to use
| Approach | Use it when | Watch out for |
|---|---|---|
| Framework integration (Astro, Next) | Pages come from a build or CMS | Remember to set the production site URL |
| Hand-written XML | A small static site, a few dozen pages | Goes stale the first time you forget |
| Generated from a database | Thousands of dynamic URLs | Exceeding the 50,000 / 50 MB limits without an index |
If your site is larger than a handful of static pages, pick a generated sitemap and treat the file as build output, not source.
Astro: @astrojs/sitemap
Install the integration and set site to your canonical origin — without it the integration cannot build absolute URLs:
// astro.config.mjs —— 中文注释:site 必须是绝对 URL
import { defineConfig } from 'astro/config';
import sitemap from '@astrojs/sitemap';
export default defineConfig({
site: 'https://example.com',
integrations: [
sitemap({
// 排除不需要收录的路径
filter: (page) => !page.includes('/drafts/'),
}),
],
});
At build time the integration emits sitemap-index.xml plus child sitemap files. It splits automatically — it has an entryLimit option with a default that keeps a margin under the protocol cap, so you are not expected to manage the split by hand. Run the build and you will find the files in dist/.
Next.js: app/sitemap.ts
The App Router turns a route file into a sitemap. Export a default function that returns MetadataRoute.Sitemap:
// app/sitemap.ts
import type { MetadataRoute } from 'next';
export default function sitemap(): MetadataRoute.Sitemap {
return [
{
url: 'https://example.com',
lastModified: new Date('2026-04-01'),
changeFrequency: 'weekly',
priority: 1,
},
{
url: 'https://example.com/about',
lastModified: new Date('2026-03-20'),
},
];
}
Next serves this at /sitemap.xml. Two cautions. First, changeFrequency and priority are widely ignored by Google, so do not spend effort tuning them. Second, if your URL list is large, use the generateSitemaps helper to emit several files instead of one oversized list. If the list comes from a CMS, fetch it in this function rather than passing a static array, so the sitemap tracks reality.
Hand-written, for the small case
For a static site of a few dozen pages, a plain file is fine as long as you commit to updating it:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
<lastmod>2026-04-01</lastmod>
</url>
</urlset>
Keep every <loc> absolute, and only bump <lastmod> when the page itself changed.
Self-check before you ship
Generation proves the file exists, not that it is correct. Run three checks:
# 1. Did the build actually produce the file, and what does it look like?
curl -s https://example.com/sitemap.xml | head -n 20
# 2. Is it served as XML, not HTML?
curl -sI https://example.com/sitemap.xml \
| grep -iE '^(HTTP|content-type)'
You want a 200 and a content-type of application/xml or text/xml. A text/html response means a framework or SPA fallback served your index page instead. Then run the file through the validator: parse it for format errors, and crawl the URLs to confirm they return 200 and are not redirects.
Where the generated file should live
Framework output normally lands at /sitemap.xml or /sitemap-index.xml at the site root. Keep it there. Crawlers look for /sitemap.xml by convention and expect a root-level file; a copy nested under /assets/ is only found if it is declared explicitly. If your build writes the files somewhere else, confirm the served URL returns 200 before you reference it.
Keeping the file fresh
A sitemap is a build artifact, so regenerate it on every deploy that changes the page list rather than on a weekly cron. A schedule drifts out of sync the moment someone publishes between runs. Tie generation to the build and validate the output as the final pipeline step, so a broken or stale file never reaches production unnoticed. For a CMS-backed list, invalidate the cache on publish instead of regenerating on every crawl.
Common mistakes
Forgetting the site URL. In Astro, a missing site produces relative <loc> values, which are invalid. Set the production origin.
Generating at request time on every hit. If a CMS-backed sitemap regenerates on every crawl, cache it and invalidate on publish instead.
Including drafts and private routes. Generated sitemaps often pick up everything in the router. Add an explicit filter.
Shipping without checking the content type. A sitemap served as HTML fails silently; the crawler reads markup where it expected XML.
Trusting priority and changeFrequency. Google ignores both. Do not architect around them.
Letting the sitemap and the site drift. The file must be regenerated as part of the same deploy that adds or removes pages.
Where this tool fits
The sitemap validator is the last check in the pipeline: point it at the generated file and it reports format errors, entry counts and any listed URL that is not a clean 200.
Frequently asked questions
▸ Should I hand-write or auto-generate a sitemap?
Auto-generate whenever your pages come from a CMS, a database or a build step, because the sitemap must track the page list as it changes. Hand-writing only makes sense for a small, static site of a few dozen pages that rarely changes.
▸ How do I add a sitemap to Astro?
Set the site option in astro.config.mjs, then add the @astrojs/sitemap integration. The integration writes sitemap-index.xml and the child sitemaps into the build output automatically.
▸ How do I create a sitemap in Next.js?
Add an app/sitemap.ts file that exports a default function returning MetadataRoute.Sitemap. Next serves it at /sitemap.xml and rebuilds it on each request or build.
▸ Does every generated sitemap get indexed?
No. Generation only produces the file. You still need to reference it in robots.txt or submit it in Search Console, and the URLs inside must return 200 for the entries to have any effect.
▸ Should the sitemap include pages behind a login?
No. Crawlers cannot authenticate, so those URLs return a redirect to a login page or a 403. Leave private pages out; a sitemap is for publicly reachable pages.