跳到主要内容

对比

robots.txt vs meta robots vs X-Robots-Tag

Three mechanisms, three scopes. Which blocks crawling, which blocks indexing, how they combine, and the trap that hides a noindex behind a robots.txt rule.

更新于 2026年4月3日 · 约 4 分钟

Three ways to tell a crawler what to do with a URL, and they are constantly confused because all of them can contain the word noindex-adjacent directives. They differ in where they live, what they cover, and — most importantly — whether they stop crawling or stop indexing.

Scope at a glance

robots.txtmeta robotsX-Robots-Tag
Where it livesA file at the host rootA tag in the HTML <head>An HTTP response header
ScopeWhole site, per user-agent, by pathA single pageA response, path or whole site
ControlsCrawlingIndexingIndexing
Works on non-HTMLYesNoYes
Read only if fetched—Yes, requires a crawlYes, requires a crawl
Typical useBlock a directoryKeep one page out of the indexSite-wide or file-type indexing control

The dividing line is the middle row. robots.txt controls crawling. meta robots and X-Robots-Tag control indexing. They are not alternatives to each other; they operate on different stages of the pipeline.

Crawl vs index, in one box

Block crawling (robots.txt)Block indexing (noindex)
Page is fetched?NoYes
Page can appear in results?Yes, as a URL-only listingNo
Tag is ever read?Not applicableYes

A disallowed page can still surface as a bare URL with no snippet, because the crawler never fetched it and other pages link to it. A noindexed page is fetched, read, and dropped from the index. If the goal is “remove this from search”, noindex is the directive you want.

The classic trap: disallow plus noindex

The most common mistake is combining the two and expecting the noindex to work.

# robots.txt (disallow) + a meta tag on the same URL
# Disallow: /old-landing/
<!-- old-landing.html -->
<meta name="robots" content="noindex">

The crawler sees Disallow: /old-landing/, never requests the page, and therefore never reads the noindex. The URL can stay in the index indefinitely, per the table above. To actually de-index a disallowed URL:

  1. Remove the Disallow so the page can be crawled.
  2. Serve noindex on the page.
  3. Wait for the crawler to fetch it and drop it from the index.
  4. Once it is out, you may re-add the Disallow if you still want to block crawling.

Order matters, and step 3 takes time, so plan de-indexing well before a deadline.

When to reach for which

  • Use robots.txt to stop a crawler from spending budget on paths you never want fetched — search results, cart pages, internal APIs.
  • Use meta robots for a single page you want excluded from the index while staying crawlable, such as a thank-you page or a staging copy you cannot block site-wide.
  • Use X-Robots-Tag when the directive must apply on the response side: site-wide via config, a whole file type, or a non-HTML asset like a PDF or image that cannot hold a meta tag.
# X-Robots-Tag 作用于响应头,可覆盖整个路径
location /downloads/ {
    add_header X-Robots-Tag "noindex";
}

Because it lives at the proxy or server layer, X-Robots-Tag is easy to set once and forget. That is also its danger: a broad rule can silently deindex everything beneath it.

How they combine

They stack rather than override, each at its own stage:

  1. robots.txt decides whether the request happens at all.
  2. If fetched, X-Robots-Tag and meta robots both feed indexing directives, and the most restrictive result wins when they disagree.
  3. Neither header nor meta overrides a Disallow — a blocked URL is simply never read.

The practical rule: if a directive must be seen, the page has to be crawlable. Anything that blocks the fetch also hides every indexing tag on it.

Common mistakes

Expecting Disallow to de-index. It prevents crawling, not indexing. Use noindex to remove a page from results.

Putting noindex on a disallowed URL. The tag is never read. Fix the order: allow, noindex, wait, re-block if needed.

Setting noindex and noarchive together by habit. none already means noindex, nofollow; stacking redundant directives makes the intent harder to read, not stronger.

Putting a meta tag on a PDF. It cannot carry one. Use X-Robots-Tag, or the file will never be de-indexed on request.

Assuming a CDN passes headers through unchanged. A proxy can strip or rewrite headers. Verify the response the crawler actually receives.

Blocking crawling to “save” a page. Blocking a URL you want indexed stops the crawl and can starve it of the signals it needs. Block budget waste, not your landing pages.

Where this tool fits

The robots.txt tester resolves whether a given user-agent may crawl a URL, which tells you whether any meta or header directive on that page can ever be read. Run it first: if the URL is disallowed, no noindex you write will take effect.

Frequently asked questions

▸ What is the difference between noindex and disallow?

Disallow, set in robots.txt, blocks crawling: the crawler never fetches the URL. Noindex, set in a meta robots tag or X-Robots-Tag header, blocks indexing: the page is fetched and read, then left out of the index.

▸ Which wins if robots.txt and a meta robots tag disagree?

They act at different stages, so "winning" is the wrong frame. robots.txt decides whether the page is fetched. If it is fetched, noindex decides whether it is indexed. A disallowed page is never crawled, so its meta tag is never read.

▸ Does X-Robots-Tag override meta robots?

They are both index-level directives and are generally combined, with the most restrictive result applied. Using both on the same URL is redundant; pick one and check the live response.

▸ Why is my noindex not working on a disallowed page?

Because the crawler cannot fetch a disallowed URL to read the tag. To apply a noindex, remove the disallow, let the page be crawled and de-indexed, then re-block it if you want.

▸ Can X-Robots-Tag be used on non-HTML files?

Yes, and that is its advantage. A PDF, image or video cannot carry a meta robots tag, but it can be served with an X-Robots-Tag header, so the header is the only in-page way to control indexing for those files.