指南
robots.txt explained: syntax, wildcards and precedence
Every directive that matters, how Google groups rules by user-agent, why the longest match wins and Allow breaks ties, and which crawlers ignore the file.
更新于 2026年4月2日 · 约 4 分钟
robots.txt is a plain-text file that lives at the root of your host and tells well-behaved crawlers which paths they may request. It is a request, not a lock, but it is the first place to look when a crawler refuses to visit a page you want indexed — or happily visits one you want gone.
Where it lives and who reads it
The file must be at https://example.com/robots.txt. A copy at /blog/robots.txt does nothing: crawlers only fetch the root file, and one root file applies to the entire host. It is a public file, so do not put anything secret in it — it is often the first thing an attacker reads to find admin paths.
Groups and rules
The file is a list of groups. Each group starts with one or more User-agent lines and is followed by the rules that apply to them, until a blank line ends the group.
# 允许所有爬虫抓取除 /admin 之外的全部内容
User-agent: *
Disallow: /admin/
Allow: /admin/public/
# 单独声明一个爬虫
User-agent: Googlebot
Disallow: /no-google/
Sitemap: https://example.com/sitemap.xml
Rules to remember:
- An empty
Disallow:line means allow everything.Disallow: /means block everything. - Comments start with
#and can follow a value on the same line. - User-agent names are matched case-insensitively, but instructions say to use the canonical casing.
- The
Sitemap:line is not part of a group and can appear anywhere.
Wildcards and end anchors
Two characters do the heavy lifting, and both are extensions rather than part of the original convention:
| Pattern | Matches | Example |
|---|---|---|
* | Any sequence of characters, including none | Disallow: /*.pdf$ |
$ | The end of the URL | Disallow: /page$ matches /page but not /page2 |
| Plain path | A prefix of the path | Disallow: /tmp matches /tmp, /tmp/x |
Disallow: /page$ blocks exactly /page and leaves /page/ crawlable, because $ demands that the URL end right there. Without the anchor, Disallow: /page also blocks /page2, /pageant and anything else sharing that prefix. The anchor is the difference between a precise rule and a wide one.
The precedence rule
When a URL matches more than one rule, crawlers pick the most specific match — the longest pattern — and when an Allow and a Disallow of equal length both match, Allow wins.
User-agent: *
Disallow: /shop/
Allow: /shop/sale/
Here /shop/sale/boots matches both rules; the longer Allow path wins, so the page is crawlable. /shop/cart matches only the Disallow, so it is blocked. This is how you open a hole in a broad block: write the exception longer than the rule it escapes.
Groups are resolved the same way. When several groups match one crawler, Google uses the single most specific group — the one whose User-agent line is the closest match — and ignores the others.
Crawl-delay and other non-standard lines
Crawl-delay asks a crawler to wait a number of seconds between requests. It is not in the original convention, and Google ignores it entirely — Googlebot decides its own crawl rate from server response times. Some other crawlers honour the value, so it can help against a scraper that respects it, but it is not a lever for controlling Google. Throttle Googlebot through Search Console’s crawl-rate setting instead.
Who ignores robots.txt
Plenty of clients never read the file:
- Malicious scrapers and content harvesters treat an open endpoint as open.
- Some SEO and aggregator tools fetch regardless of the rules.
- Anything hitting the site from a browser or a script is outside the protocol.
For crawlers that do comply — Googlebot, Bingbot and most mainstream search bots — the file is a genuine instruction. For everyone else it is advice. If a path must not be requested at all, use authentication, a WAF rule or IP-level blocking rather than a Disallow.
Common mistakes
Believing Disallow deindexes. It blocks crawling, not indexing. A disallowed URL can still be indexed as a URL-only result.
Blocking the whole site while testing. Disallow: / on production is a classic staging-config leak. Check the live file, not the one you remember writing.
Rules that never match. A Disallow written against the wrong path, or with a stray space, quietly does nothing. Test the exact URL you care about.
Using Crawl-delay to slow Googlebot. Google ignores it; the block is wasted.
Case mismatch. Paths are case-sensitive on most servers. /Admin/ and /admin/ are different rules.
Layering a WAF on top and forgetting it. A block that never reaches robots.txt still stops the crawl, and you will debug the wrong file.
Where this tool fits
The robots.txt tester evaluates a URL against your rules for a chosen user-agent and tells you whether the result is allowed, showing which matching rule decided it. That is the fastest way to confirm a wildcard or an Allow exception behaves the way you intended.
Frequently asked questions
▸ Where does robots.txt have to live?
At the root of the host, at https://example.com/robots.txt. Crawlers only read the root file; a copy in a subdirectory is ignored. One file applies to the whole host and port.
▸ What is the difference between Allow and Disallow?
Disallow blocks a crawler from a matching path; Allow permits one. They are often used together to carve an exception out of a broader Disallow, because Allow wins when both rules match with equal specificity.
▸ How do wildcards work in robots.txt?
A * matches any sequence of characters, and a trailing $ anchors the match to the end of the URL. Both are supported by Google and Bing but are not part of the original 1994 robots.txt convention.
▸ Does Google support Crawl-delay?
No. Google ignores Crawl-delay and manages crawl rate itself. Some other crawlers honour it, so it can still be useful, but do not rely on it to control Googlebot.
▸ Does a Disallow line remove a page from search results?
No. Disallow controls crawling, not indexing. A disallowed URL can still appear in results as a URL-only listing if other pages link to it, because the crawler never fetches it to read a noindex.