跳到主要内容

指南

robots.txt explained: syntax, wildcards and precedence

Every directive that matters, how Google groups rules by user-agent, why the longest match wins and Allow breaks ties, and which crawlers ignore the file.

更新于 2026年4月2日 · 约 4 分钟

robots.txt is a plain-text file that lives at the root of your host and tells well-behaved crawlers which paths they may request. It is a request, not a lock, but it is the first place to look when a crawler refuses to visit a page you want indexed — or happily visits one you want gone.

Where it lives and who reads it

The file must be at https://example.com/robots.txt. A copy at /blog/robots.txt does nothing: crawlers only fetch the root file, and one root file applies to the entire host. It is a public file, so do not put anything secret in it — it is often the first thing an attacker reads to find admin paths.

Groups and rules

The file is a list of groups. Each group starts with one or more User-agent lines and is followed by the rules that apply to them, until a blank line ends the group.

# 允许所有爬虫抓取除 /admin 之外的全部内容
User-agent: *
Disallow: /admin/
Allow: /admin/public/

# 单独声明一个爬虫
User-agent: Googlebot
Disallow: /no-google/

Sitemap: https://example.com/sitemap.xml

Rules to remember:

  • An empty Disallow: line means allow everything. Disallow: / means block everything.
  • Comments start with # and can follow a value on the same line.
  • User-agent names are matched case-insensitively, but instructions say to use the canonical casing.
  • The Sitemap: line is not part of a group and can appear anywhere.

Wildcards and end anchors

Two characters do the heavy lifting, and both are extensions rather than part of the original convention:

PatternMatchesExample
*Any sequence of characters, including noneDisallow: /*.pdf$
$The end of the URLDisallow: /page$ matches /page but not /page2
Plain pathA prefix of the pathDisallow: /tmp matches /tmp, /tmp/x

Disallow: /page$ blocks exactly /page and leaves /page/ crawlable, because $ demands that the URL end right there. Without the anchor, Disallow: /page also blocks /page2, /pageant and anything else sharing that prefix. The anchor is the difference between a precise rule and a wide one.

The precedence rule

When a URL matches more than one rule, crawlers pick the most specific match — the longest pattern — and when an Allow and a Disallow of equal length both match, Allow wins.

User-agent: *
Disallow: /shop/
Allow: /shop/sale/

Here /shop/sale/boots matches both rules; the longer Allow path wins, so the page is crawlable. /shop/cart matches only the Disallow, so it is blocked. This is how you open a hole in a broad block: write the exception longer than the rule it escapes.

Groups are resolved the same way. When several groups match one crawler, Google uses the single most specific group — the one whose User-agent line is the closest match — and ignores the others.

Crawl-delay and other non-standard lines

Crawl-delay asks a crawler to wait a number of seconds between requests. It is not in the original convention, and Google ignores it entirely — Googlebot decides its own crawl rate from server response times. Some other crawlers honour the value, so it can help against a scraper that respects it, but it is not a lever for controlling Google. Throttle Googlebot through Search Console’s crawl-rate setting instead.

Who ignores robots.txt

Plenty of clients never read the file:

  • Malicious scrapers and content harvesters treat an open endpoint as open.
  • Some SEO and aggregator tools fetch regardless of the rules.
  • Anything hitting the site from a browser or a script is outside the protocol.

For crawlers that do comply — Googlebot, Bingbot and most mainstream search bots — the file is a genuine instruction. For everyone else it is advice. If a path must not be requested at all, use authentication, a WAF rule or IP-level blocking rather than a Disallow.

Common mistakes

Believing Disallow deindexes. It blocks crawling, not indexing. A disallowed URL can still be indexed as a URL-only result.

Blocking the whole site while testing. Disallow: / on production is a classic staging-config leak. Check the live file, not the one you remember writing.

Rules that never match. A Disallow written against the wrong path, or with a stray space, quietly does nothing. Test the exact URL you care about.

Using Crawl-delay to slow Googlebot. Google ignores it; the block is wasted.

Case mismatch. Paths are case-sensitive on most servers. /Admin/ and /admin/ are different rules.

Layering a WAF on top and forgetting it. A block that never reaches robots.txt still stops the crawl, and you will debug the wrong file.

Where this tool fits

The robots.txt tester evaluates a URL against your rules for a chosen user-agent and tells you whether the result is allowed, showing which matching rule decided it. That is the fastest way to confirm a wildcard or an Allow exception behaves the way you intended.

Frequently asked questions

▸ Where does robots.txt have to live?

At the root of the host, at https://example.com/robots.txt. Crawlers only read the root file; a copy in a subdirectory is ignored. One file applies to the whole host and port.

▸ What is the difference between Allow and Disallow?

Disallow blocks a crawler from a matching path; Allow permits one. They are often used together to carve an exception out of a broader Disallow, because Allow wins when both rules match with equal specificity.

▸ How do wildcards work in robots.txt?

A * matches any sequence of characters, and a trailing $ anchors the match to the end of the URL. Both are supported by Google and Bing but are not part of the original 1994 robots.txt convention.

▸ Does Google support Crawl-delay?

No. Google ignores Crawl-delay and manages crawl rate itself. Some other crawlers honour it, so it can still be useful, but do not rely on it to control Googlebot.

▸ Does a Disallow line remove a page from search results?

No. Disallow controls crawling, not indexing. A disallowed URL can still appear in results as a URL-only listing if other pages link to it, because the crawler never fetches it to read a noindex.