跳到主要内容

实操

How to block AI crawlers with robots.txt

The user-agent tokens to disallow — GPTBot, ClaudeBot, CCBot and the rest — plus why robots.txt is only a request, and what to do when a bot ignores it.

更新于 2026年4月10日 · 约 4 分钟

Adding a Disallow for an AI crawler is a two-minute change and, on its own, a weak one. The file is worth writing — compliant crawlers do read it — but knowing exactly what it does and does not do is what keeps you from relying on it as a security control.

The tokens worth knowing

There is no single “AI crawler” user-agent. Each vendor publishes its own tokens, and they serve different purposes — training, live user-triggered fetches, and search are often separate.

TokenOperatorRoughly covers
GPTBotOpenAICrawling for model training
ChatGPT-UserOpenAIPages fetched on behalf of a user in chat
OAI-SearchBotOpenAIIndexing for ChatGPT search
ClaudeBotAnthropicClaude’s crawler
anthropic-aiAnthropicOlder Anthropic crawler token
CCBotCommon CrawlThe shared corpus many models train on
PerplexityBotPerplexityPerplexity’s index
BytespiderByteDanceByteDance’s crawler
Google-ExtendedGoogleAI training and grounding, separate from Search
Applebot-ExtendedAppleApple’s AI training uses

This list shifts as vendors add and retire tokens, so verify each one against the operator’s current documentation before you publish. Google-Extended is a control token, not a crawler: it does not stop Googlebot from indexing your site for Search, only from using it for Gemini training and grounding.

The robots.txt to write

One group per token. A shared User-agent: * group will not do, because named groups take precedence over the wildcard:

# 屏蔽主流 AI 训练与抓取爬虫
User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Sitemap: https://example.com/sitemap.xml

If you want search engines to keep crawling while AI training is blocked, use the AI-specific tokens above and leave Googlebot and Bingbot alone. If you want the middle ground, some teams disallow the training crawlers but allow the user-triggered ones, so a person asking a chatbot about a page still gets an answer. That is a policy call, not a technical one.

Why robots.txt is only a request

robots.txt is a voluntary convention. Nothing enforces it:

  • It stops crawlers that choose to obey — the same compliant bots that were unlikely to be hostile.
  • It does nothing to a scraper that never reads it. An open, public URL is fetchable by anyone with curl.
  • It is not authentication, not a firewall, and not a licence.

It is still worth writing, because it is the clearest, machine-readable statement of your intent, and mainstream vendors do honour it. Just do not expect it to be a lock.

Harder controls

When you need the block to actually hold, enforce it below robots.txt:

# nginx —— 按 User-Agent 直接拒绝,返回 403
map $http_user_agent $block_ai {
    default 0;
    "~*GPTBot"        1;
    "~*ClaudeBot"     1;
    "~*CCBot"         1;
    "~*Bytespider"    1;
}

server {
    if ($block_ai) { return 403; }
    # ...
}

Layered with the user-agent check, consider:

  • WAF or bot management. Provider-level rules that block by IP range and request behaviour, not just the string a bot claims to be.
  • Rate limiting. Caps the damage from a crawler that ignores robots.txt and hammers the site.
  • Authentication. The only block that is a real block. A page behind a login cannot be scraped without credentials.
  • Terms and licensing. A machine-readable content licence and clear terms give you a legal position; they are not enforcement, but they matter when you pursue one.

Remember the trade-off: blocking AI crawlers can keep your content out of AI answers and summaries as well as out of training. Decide which you want before you ship the rules.

Common mistakes

Relying on User-agent: *. Named groups override the wildcard for those crawlers only if the named group exists. List each AI token explicitly.

Assuming Google-Extended affects Search. It governs AI training and grounding, not your normal search presence. Blocking it will not remove your pages from results.

Treating robots.txt as security. It is a request. If a path must not be served, enforce it at the edge or behind a login.

Blocking the user-triggered bot by accident. Disallowing ChatGPT-User or ClaudeBot can remove your site from human-initiated AI answers, not just training. Pick tokens deliberately.

Writing a token that no longer exists. Vendors rename and retire agents. Check the list periodically.

Ignoring caching. A CDN may serve a cached robots.txt; purge it after a change or crawlers keep seeing the old rules.

Where this tool fits

The robots.txt tester checks whether a specific user-agent — GPTBot, ClaudeBot or a token you type in — is allowed for a given URL under your current rules. Use it to confirm each new group resolves to “disallowed” before you rely on the file.

Frequently asked questions

▸ How do I block GPTBot in robots.txt?

Add a group with User-agent: GPTBot and Disallow: / . To block several AI crawlers, repeat a group per token. Each group needs its own User-agent line followed by the disallow rules.

▸ Which robots.txt tokens block AI crawlers?

Common tokens include GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, CCBot, PerplexityBot, Bytespider and Google-Extended. The list changes, so confirm each token against the vendor's current documentation.

▸ Does robots.txt actually stop AI scraping?

Only for crawlers that choose to comply. robots.txt is a convention with no enforcement; anyone can ignore it. Treat it as a signal of intent, not a security control.

▸ Does blocking Google-Extended remove my site from AI answers?

Google-Extended controls the use of your content for Gemini model training and grounding, not normal Search indexing or AI Overviews. Blocking it does not remove your pages from standard search results.

▸ What is a stronger alternative to robots.txt?

A WAF or bot-management rule that blocks by IP range and behaviour, server-level user-agent rejection, or requiring authentication. These enforce the block instead of relying on the crawler to honour it.