实操
How to block AI crawlers with robots.txt
The user-agent tokens to disallow — GPTBot, ClaudeBot, CCBot and the rest — plus why robots.txt is only a request, and what to do when a bot ignores it.
更新于 2026年4月10日 · 约 4 分钟
Adding a Disallow for an AI crawler is a two-minute change and, on its own, a weak one. The file is worth writing — compliant crawlers do read it — but knowing exactly what it does and does not do is what keeps you from relying on it as a security control.
The tokens worth knowing
There is no single “AI crawler” user-agent. Each vendor publishes its own tokens, and they serve different purposes — training, live user-triggered fetches, and search are often separate.
| Token | Operator | Roughly covers |
|---|---|---|
GPTBot | OpenAI | Crawling for model training |
ChatGPT-User | OpenAI | Pages fetched on behalf of a user in chat |
OAI-SearchBot | OpenAI | Indexing for ChatGPT search |
ClaudeBot | Anthropic | Claude’s crawler |
anthropic-ai | Anthropic | Older Anthropic crawler token |
CCBot | Common Crawl | The shared corpus many models train on |
PerplexityBot | Perplexity | Perplexity’s index |
Bytespider | ByteDance | ByteDance’s crawler |
Google-Extended | AI training and grounding, separate from Search | |
Applebot-Extended | Apple | Apple’s AI training uses |
This list shifts as vendors add and retire tokens, so verify each one against the operator’s current documentation before you publish. Google-Extended is a control token, not a crawler: it does not stop Googlebot from indexing your site for Search, only from using it for Gemini training and grounding.
The robots.txt to write
One group per token. A shared User-agent: * group will not do, because named groups take precedence over the wildcard:
# 屏蔽主流 AI 训练与抓取爬虫
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: PerplexityBot
Disallow: /
User-agent: Google-Extended
Disallow: /
Sitemap: https://example.com/sitemap.xml
If you want search engines to keep crawling while AI training is blocked, use the AI-specific tokens above and leave Googlebot and Bingbot alone. If you want the middle ground, some teams disallow the training crawlers but allow the user-triggered ones, so a person asking a chatbot about a page still gets an answer. That is a policy call, not a technical one.
Why robots.txt is only a request
robots.txt is a voluntary convention. Nothing enforces it:
- It stops crawlers that choose to obey — the same compliant bots that were unlikely to be hostile.
- It does nothing to a scraper that never reads it. An open, public URL is fetchable by anyone with
curl. - It is not authentication, not a firewall, and not a licence.
It is still worth writing, because it is the clearest, machine-readable statement of your intent, and mainstream vendors do honour it. Just do not expect it to be a lock.
Harder controls
When you need the block to actually hold, enforce it below robots.txt:
# nginx —— 按 User-Agent 直接拒绝,返回 403
map $http_user_agent $block_ai {
default 0;
"~*GPTBot" 1;
"~*ClaudeBot" 1;
"~*CCBot" 1;
"~*Bytespider" 1;
}
server {
if ($block_ai) { return 403; }
# ...
}
Layered with the user-agent check, consider:
- WAF or bot management. Provider-level rules that block by IP range and request behaviour, not just the string a bot claims to be.
- Rate limiting. Caps the damage from a crawler that ignores robots.txt and hammers the site.
- Authentication. The only block that is a real block. A page behind a login cannot be scraped without credentials.
- Terms and licensing. A machine-readable content licence and clear terms give you a legal position; they are not enforcement, but they matter when you pursue one.
Remember the trade-off: blocking AI crawlers can keep your content out of AI answers and summaries as well as out of training. Decide which you want before you ship the rules.
Common mistakes
Relying on User-agent: *. Named groups override the wildcard for those crawlers only if the named group exists. List each AI token explicitly.
Assuming Google-Extended affects Search. It governs AI training and grounding, not your normal search presence. Blocking it will not remove your pages from results.
Treating robots.txt as security. It is a request. If a path must not be served, enforce it at the edge or behind a login.
Blocking the user-triggered bot by accident. Disallowing ChatGPT-User or ClaudeBot can remove your site from human-initiated AI answers, not just training. Pick tokens deliberately.
Writing a token that no longer exists. Vendors rename and retire agents. Check the list periodically.
Ignoring caching. A CDN may serve a cached robots.txt; purge it after a change or crawlers keep seeing the old rules.
Where this tool fits
The robots.txt tester checks whether a specific user-agent — GPTBot, ClaudeBot or a token you type in — is allowed for a given URL under your current rules. Use it to confirm each new group resolves to “disallowed” before you rely on the file.
Frequently asked questions
▸ How do I block GPTBot in robots.txt?
Add a group with User-agent: GPTBot and Disallow: / . To block several AI crawlers, repeat a group per token. Each group needs its own User-agent line followed by the disallow rules.
▸ Which robots.txt tokens block AI crawlers?
Common tokens include GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, CCBot, PerplexityBot, Bytespider and Google-Extended. The list changes, so confirm each token against the vendor's current documentation.
▸ Does robots.txt actually stop AI scraping?
Only for crawlers that choose to comply. robots.txt is a convention with no enforcement; anyone can ignore it. Treat it as a signal of intent, not a security control.
▸ Does blocking Google-Extended remove my site from AI answers?
Google-Extended controls the use of your content for Gemini model training and grounding, not normal Search indexing or AI Overviews. Blocking it does not remove your pages from standard search results.
▸ What is a stronger alternative to robots.txt?
A WAF or bot-management rule that blocks by IP range and behaviour, server-level user-agent rejection, or requiring authentication. These enforce the block instead of relying on the crawler to honour it.