Robots.txt for AI Crawlers: Which Bots to Allow and Why
Allow the crawlers that fetch pages to answer questions: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and Bingbot, plus the user-initiated fetchers ChatGPT-User and Claude-User. Training crawlers, GPTBot and ClaudeBot, are a business choice, and blocking them doesn't affect search answers. Google-Extended is the exception: Google says it covers Gemini training and grounding in the Gemini apps, but not Google Search.
Key takeaways
- Blocking OAI-SearchBot keeps you out of ChatGPT search answers. Blocking GPTBot only opts you out of training.
- Anthropic and Perplexity also separate search from training: allow Claude-SearchBot and PerplexityBot to stay citable.
- Google-Extended covers Gemini training and grounding in Gemini apps. Google says it doesn't affect Google Search, AI Overviews or AI Mode.
- robots.txt is a request, not access control, and user-initiated fetchers such as ChatGPT-User and Perplexity-User may not follow it.
Which AI crawlers exist, and what does each do?
AI companies run several bots with different jobs. Search bots fetch pages to answer questions now and cite them. Training bots collect content for future models. User-initiated fetchers open a page because a person asked.
| User agent | Run by | Job | If you block it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search | You aren't shown in ChatGPT search answers, though you can still appear as a navigational link |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user's action calls for it | OpenAI says robots.txt rules may not apply, and it isn't used to decide whether you appear in Search |
| GPTBot | OpenAI | Collects content that may train its models | Your content shouldn't be used for training. Search is unaffected |
| Claude-SearchBot | Anthropic | Improves the quality of Claude's search results | May reduce your visibility and accuracy in Claude's search results |
| Claude-User | Anthropic | Fetches a page when someone asks Claude a question | May reduce your visibility for user-directed search |
| ClaudeBot | Anthropic | Collects web content that may contribute to training | Signals your future content should be excluded from training |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity's search results | You risk dropping out of Perplexity's search results. Perplexity says the bot isn't used to train foundation models |
| Perplexity-User | Perplexity | Visits a page when a user asks | Perplexity says it generally ignores robots.txt, since a user requested the fetch |
| Googlebot | Crawls for Google Search, including its AI features | You leave Google Search, AI Overviews and AI Mode | |
| Google-Extended | A token for Gemini training and grounding in the Gemini apps and on Vertex AI | No effect on Search. Google's wording covers Gemini training and grounding | |
| Bingbot | Microsoft | Bing's main crawler, which Copilot's answers draw on | You may not be indexed or selected for grounding in Bing and Copilot |
OpenAI says its systems take about 24 hours to adjust to a robots.txt change, and Perplexity says up to 24 hours.
Should you allow the search crawlers?
Yes, if you want to be cited. These are the bots that decide whether you can appear in an AI answer. Both OpenAI and Perplexity recommend allowing their search bots and their published IP ranges, which matters if a CDN or firewall stands in front of your site. Many sites block these bots by accident with a blanket rule written to stop AI training, so check yours with the free AI Crawler Checker.
Should you block the training crawlers?
It's your call, and the search crawlers don't depend on it. OpenAI says its bots' settings are independent, so you can allow OAI-SearchBot while disallowing GPTBot. Anthropic and Perplexity separate search from training in the same way.
Reasons to block: you don't want your content in future models. Reasons to allow: it's the only way your content can end up in future training data, though no vendor says that improves your recommendations, so don't count on it.
Google-Extended needs more care. Google describes it as a token for managing whether content it crawls may be used to train future Gemini models and for grounding in the Gemini apps and Grounding with Google Search on Vertex AI. So "block training only" isn't really on offer for Gemini. If you want Gemini to be able to use your pages when it answers, leave Google-Extended allowed. Google says it has no effect on inclusion in Google Search. To stay out of AI Overviews and AI Mode, use the Search generative AI setting in Search Console. See how to get recommended by Gemini.
What does a good robots.txt look like?
This version keeps you citable and opts out of OpenAI and Anthropic training:
# Search crawlers and user-requested fetchers: allow
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Googlebot
Allow: /
User-agent: Bingbot
Allow: /
# Training crawlers: opt out
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
One trap: a crawler obeys the group that names it and falls back to the User-agent: * group only when none does. If you keep private paths disallowed under *, repeat those paths in every group above, or the named bots will ignore them. That is how the standard (RFC 9309) works. The free AI Robots.txt Generator does it for you.
What can robots.txt not do?
- It isn't access control. The standard says its rules are not a form of access authorization. Put private content behind a login.
- It doesn't bind user-initiated fetchers. OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it.
- IP blocking can backfire. Anthropic warns that blocking its IP addresses can stop its bots reading your robots.txt, so the opt-out may not hold.
- It doesn't remove a URL. OpenAI says that if it learns of a disallowed page's URL elsewhere, it may still show just the link and title in ChatGPT Atlas, and recommends a noindex meta tag, which the crawler must be allowed to read.
- It applies per host. Anthropic says to add rules to every subdomain you want covered.
How do you roll it out?
- Open
yoursite.com/robots.txtand read what's there. Look for aUser-agent: *group withDisallow: /. - Decide on training for each company.
- Generate or edit the file and publish it at the root of every host you run.
- Allow the vendors' published IP ranges at your CDN or firewall.
- Re-check after a day with the AI Crawler Checker, then confirm in your server logs that the search bots arrive.
Frequently asked questions
Does blocking GPTBot remove me from ChatGPT?
No. GPTBot is OpenAI's training crawler. ChatGPT search relies on OAI-SearchBot, and OpenAI says the two settings are independent.
Does blocking Google-Extended remove me from AI Overviews?
No. Google says Google-Extended doesn't impact inclusion in Google Search. It does cover Gemini training and grounding in the Gemini apps, so it can affect whether Gemini uses your pages.
Do AI crawlers obey robots.txt?
OpenAI, Anthropic and Perplexity each document robots.txt as the way to manage their crawlers, and Anthropic says its bots honor it. User-initiated fetchers are the exception: OpenAI says robots.txt may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores it. The standard itself is a request, not an enforcement mechanism.
How long do robots.txt changes take to work?
OpenAI says about 24 hours for its systems to adjust, and Perplexity says up to 24 hours. Allow time before judging a change.
Is there a separate crawler for Copilot?
Not in Bing's published list, which names Bingbot, AdIdxBot, BingPreview, MicrosoftPreview and BingVideoPreview. Bing's guidelines say to allow Bingbot to crawl and render your content, since blocking it can keep pages out of grounding results.
Sources
- OpenAI: Overview of OpenAI crawlers
- Claude Help Center: Does Anthropic crawl data from the web, and how can site owners block the crawler?
- Perplexity: Perplexity crawlers (PerplexityBot and Perplexity-User)
- Google Search Central: Google's common crawlers
- Search Console Help: Search generative AI control
- Bing Webmaster Tools: Which crawlers does Bing use?
- Bing Webmaster Tools: Webmaster Guidelines
- RFC 9309: Robots Exclusion Protocol