对比
robots.txt vs meta robots vs X-Robots-Tag
Three mechanisms, three scopes. Which blocks crawling, which blocks indexing, how they combine, and the trap that hides a noindex behind a robots.txt rule.
更新于 2026年4月3日 · 约 4 分钟
Three ways to tell a crawler what to do with a URL, and they are constantly confused because all of them can contain the word noindex-adjacent directives. They differ in where they live, what they cover, and — most importantly — whether they stop crawling or stop indexing.
Scope at a glance
| robots.txt | meta robots | X-Robots-Tag | |
|---|---|---|---|
| Where it lives | A file at the host root | A tag in the HTML <head> | An HTTP response header |
| Scope | Whole site, per user-agent, by path | A single page | A response, path or whole site |
| Controls | Crawling | Indexing | Indexing |
| Works on non-HTML | Yes | No | Yes |
| Read only if fetched | — | Yes, requires a crawl | Yes, requires a crawl |
| Typical use | Block a directory | Keep one page out of the index | Site-wide or file-type indexing control |
The dividing line is the middle row. robots.txt controls crawling. meta robots and X-Robots-Tag control indexing. They are not alternatives to each other; they operate on different stages of the pipeline.
Crawl vs index, in one box
| Block crawling (robots.txt) | Block indexing (noindex) | |
|---|---|---|
| Page is fetched? | No | Yes |
| Page can appear in results? | Yes, as a URL-only listing | No |
| Tag is ever read? | Not applicable | Yes |
A disallowed page can still surface as a bare URL with no snippet, because the crawler never fetched it and other pages link to it. A noindexed page is fetched, read, and dropped from the index. If the goal is “remove this from search”, noindex is the directive you want.
The classic trap: disallow plus noindex
The most common mistake is combining the two and expecting the noindex to work.
# robots.txt (disallow) + a meta tag on the same URL
# Disallow: /old-landing/
<!-- old-landing.html -->
<meta name="robots" content="noindex">
The crawler sees Disallow: /old-landing/, never requests the page, and therefore never reads the noindex. The URL can stay in the index indefinitely, per the table above. To actually de-index a disallowed URL:
- Remove the
Disallowso the page can be crawled. - Serve
noindexon the page. - Wait for the crawler to fetch it and drop it from the index.
- Once it is out, you may re-add the
Disallowif you still want to block crawling.
Order matters, and step 3 takes time, so plan de-indexing well before a deadline.
When to reach for which
- Use robots.txt to stop a crawler from spending budget on paths you never want fetched — search results, cart pages, internal APIs.
- Use meta robots for a single page you want excluded from the index while staying crawlable, such as a thank-you page or a staging copy you cannot block site-wide.
- Use X-Robots-Tag when the directive must apply on the response side: site-wide via config, a whole file type, or a non-HTML asset like a PDF or image that cannot hold a meta tag.
# X-Robots-Tag 作用于响应头,可覆盖整个路径
location /downloads/ {
add_header X-Robots-Tag "noindex";
}
Because it lives at the proxy or server layer, X-Robots-Tag is easy to set once and forget. That is also its danger: a broad rule can silently deindex everything beneath it.
How they combine
They stack rather than override, each at its own stage:
robots.txtdecides whether the request happens at all.- If fetched,
X-Robots-Tagandmeta robotsboth feed indexing directives, and the most restrictive result wins when they disagree. - Neither header nor meta overrides a
Disallow— a blocked URL is simply never read.
The practical rule: if a directive must be seen, the page has to be crawlable. Anything that blocks the fetch also hides every indexing tag on it.
Common mistakes
Expecting Disallow to de-index. It prevents crawling, not indexing. Use noindex to remove a page from results.
Putting noindex on a disallowed URL. The tag is never read. Fix the order: allow, noindex, wait, re-block if needed.
Setting noindex and noarchive together by habit. none already means noindex, nofollow; stacking redundant directives makes the intent harder to read, not stronger.
Putting a meta tag on a PDF. It cannot carry one. Use X-Robots-Tag, or the file will never be de-indexed on request.
Assuming a CDN passes headers through unchanged. A proxy can strip or rewrite headers. Verify the response the crawler actually receives.
Blocking crawling to “save” a page. Blocking a URL you want indexed stops the crawl and can starve it of the signals it needs. Block budget waste, not your landing pages.
Where this tool fits
The robots.txt tester resolves whether a given user-agent may crawl a URL, which tells you whether any meta or header directive on that page can ever be read. Run it first: if the URL is disallowed, no noindex you write will take effect.
Frequently asked questions
▸ What is the difference between noindex and disallow?
Disallow, set in robots.txt, blocks crawling: the crawler never fetches the URL. Noindex, set in a meta robots tag or X-Robots-Tag header, blocks indexing: the page is fetched and read, then left out of the index.
▸ Which wins if robots.txt and a meta robots tag disagree?
They act at different stages, so "winning" is the wrong frame. robots.txt decides whether the page is fetched. If it is fetched, noindex decides whether it is indexed. A disallowed page is never crawled, so its meta tag is never read.
▸ Does X-Robots-Tag override meta robots?
They are both index-level directives and are generally combined, with the most restrictive result applied. Using both on the same URL is redundant; pick one and check the live response.
▸ Why is my noindex not working on a disallowed page?
Because the crawler cannot fetch a disallowed URL to read the tag. To apply a noindex, remove the disallow, let the page be crawled and de-indexed, then re-block it if you want.
▸ Can X-Robots-Tag be used on non-HTML files?
Yes, and that is its advantage. A PDF, image or video cannot carry a meta robots tag, but it can be served with an X-Robots-Tag header, so the header is the only in-page way to control indexing for those files.