Block the training crawlers if you do not want your content used to train models, and leave the search crawlers and user-triggered fetchers alone if you want your pages to appear in AI answers. OpenAI and Anthropic run a separate bot for training, for search and for fetching a page a user asked about, so refusing GPTBot or ClaudeBot costs nothing in their AI search. GPTBot and OAI-SearchBot are two separate decisions.
Google and Apple work differently. One crawler does everything, and a control token (Google-Extended, Applebot-Extended) decides whether the content may be used for training. Disallowing the token does not remove a site from Google Search or from Apple's search features. Disallowing Googlebot or Applebot does. One exception is covered below: Google-Extended also governs grounding in Gemini Apps.
Two things make the difference easy to lose. Since 15 September 2026, Cloudflare's "Block" setting also stops Googlebot, Bingbot and Applebot. And the ready-made ai.robots.txt list on GitHub blocks every search and user-fetch crawler we track.
Three jobs: training, search and user fetch
An AI company can visit your site for three reasons, and the major operators document a separate control for each:
- Training. Collecting pages that may be used to train or fine-tune a model. Refusing it keeps future content out of the training set. It does not remove you from anything a user sees today.
- Search. Building an index that an AI search product answers from. OpenAI says sites that opt out of
OAI-SearchBot"will not be shown in ChatGPT search answers, though can still appear as navigational links." - User fetch. Retrieving a page because a person asked for it, in the moment. Anthropic says that disabling
Claude-User"prevents our system from retrieving your content in response to a user query".
Site owners treat these differently when they are given the choice. Cloudflare reported on 15 September 2026 that fewer than 1% of the sites it serves block search bots, while 17% enable some mechanism to block training. A single "block AI" switch hides that difference.
Which user agent does which job
These are the tokens you write in a User-agent: line, with what their operators say about them. The quotes come from the operators' pages, read on 24 September 2026.
| Operator | Token | Job | What the operator says blocking it does |
|---|---|---|---|
| OpenAI | GPTBot | Training | "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models." |
| OpenAI | OAI-SearchBot | Search | Opted-out sites "will not be shown in ChatGPT search answers". |
| OpenAI | ChatGPT-User | User fetch | "Because these actions are initiated by a user, robots.txt rules may not apply." |
| Anthropic | ClaudeBot | Training | The site's "future materials should be excluded from our AI model training datasets." |
| Anthropic | Claude-SearchBot | Search | Prevents indexing for search, which "may reduce your site's visibility and accuracy in user search results." |
| Anthropic | Claude-User | User fetch | Prevents retrieving your content "in response to a user query". |
| Perplexity | PerplexityBot | Search | Surfaces sites in Perplexity results; "not used to crawl content for AI foundation models." |
| Perplexity | Perplexity-User | User fetch | "This fetcher generally ignores robots.txt rules." |
Google-Extended | Control token | Governs Gemini training and grounding in Gemini Apps; "does not impact a site's inclusion in Google Search". | |
| Apple | Applebot-Extended | Control token | "Applebot-Extended does not crawl webpages." Pages that disallow it "can still be included in search results." |
OpenAI puts the independence in one sentence: "Each setting is independent of the others." Blocking GPTBot does not block OAI-SearchBot, and allowing OAI-SearchBot does not allow training.
Many other operators publish crawlers too, including Meta, Amazon, Mistral and DuckDuckGo. The same rule applies to them: look up the token on the operator's own page and read what it is for before you name it. How to write robots.txt rules for AI crawlers lists every token our scanner recognises, grouped by purpose.
Google and Apple: one crawler, a control token
Googlebot and Applebot are the crawlers. Google-Extended and Applebot-Extended are product tokens that exist only to be named in robots.txt. Apple says it directly: "Applebot-Extended does not crawl webpages." Google describes Google-Extended as "a standalone product token".
Because AI Overviews and AI Mode are Google Search features, blocking Google-Extended does not take a site out of them. Since 31 August 2026 Google has offered a separate Search Console toggle for them worldwide, and says that toggle "will not be used as a ranking signal for search results outside of these generative AI Search features."
Google-Extended covers more than training. Google's page says it also governs "grounding (providing content from the Google Search index to the model at prompt time…)" in Gemini Apps and in Grounding with Google Search on Vertex AI. Disallowing it keeps your content out of grounded answers in those products as well as out of training data. Google offers no separate token for refusing training while staying available for grounding.
A robots.txt that blocks training and keeps AI search
This file refuses the four training controls from the table and leaves everything else open. example.com is a placeholder.
# Keep content out of model training
# (Google-Extended also covers grounding in Gemini Apps)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
# Everyone else, including OAI-SearchBot, Claude-SearchBot,
# PerplexityBot and the user-triggered fetchers
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
Sitemap: https://example.com/sitemap.xml
RFC 9309, the Robots Exclusion Protocol, lets one group start with several User-agent lines, so the four tokens share one Disallow. Product tokens are matched case-insensitively. Content-Signal is explained further down. It is not part of the protocol. RFC 9309 says such records "MUST NOT interfere with the parsing of explicitly defined records", so a crawler that does not know the line still reads the group's rules.
Changes take a while to reach the operators. OpenAI says "it can take ~24 hours from a site's robots.txt update for our systems to adjust", and Perplexity says "up to 24 hours".
A named group replaces the * group
Adding a group for one crawler can open paths you meant to keep closed. Under RFC 9309 a crawler obeys the group that names it, and it falls back to User-agent: * only "if no matching group exists." Groups are not layered. So this file:
User-agent: *
Disallow: /admin/
User-agent: OAI-SearchBot
Allow: /
lets OAI-SearchBot into /admin/, because its own group has no rule for that path, and a URI with no matching rule is allowed. If you name a crawler, repeat every Disallow it should still obey inside its group.
To see how your own file treats each of these crawlers, run it through the robots.txt checker.
Before you copy ai.robots.txt
The ai.robots.txt project on GitHub publishes a ready-made robots.txt that disallows every crawler it lists. Its README is clear about the scope: "This list contains AI-related crawlers of all types, regardless of purpose."
We compared the file with our own crawler registry, which was checked against each operator's documentation on 24 September 2026. The version we read was last changed on 7 September 2026 (commit 2acefa38cc). It has 175 User-agent lines naming 172 distinct tokens. Among them are all 11 search crawlers and all 9 user-triggered fetchers in our registry: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User and the rest. It also names Applebot, whose data Apple says powers "the search technology integrated into many user experiences in Apple’s ecosystem including Spotlight, Siri, and Safari."
The project does what it says. Copying the file is a decision to leave AI search and user-requested fetches, and Apple's search features, along with training. If that is your decision, the file is a convenient way to make it. If you only meant to refuse training, use a list of training tokens instead.
If you use Cloudflare: "Block" changed on 15 September 2026
Cloudflare's AI crawler controls now separate three behaviours: Search, Training and Agent (Cloudflare's term for user-directed fetches and browser agents). On 15 September 2026 it changed what "Block" means. In Cloudflare's words, "Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training."
The new setting for refusing training while staying in search is Disallow AI Training. It publishes the preference in robots.txt through Cloudflare's Bot Preference Sync, keeps what Cloudflare calls "Accountable" mixed-use crawlers allowed for search, and blocks other training crawlers. Bing does not yet read a no-training preference from robots.txt; Cloudflare says Microsoft is targeting early 2027. The same post says the legacy "Block AI Bots" toggle is being deprecated.
Cloudflare says existing settings were migrated automatically. A domain that only ever used the old "Block AI Bots" toggle, set to Block, now has Search on Allow, Training on Disallow AI Training and Agent on "Block on pages with ads". That last value refuses user-triggered fetchers on any page Cloudflare detects serving ads. With these settings on, the robots.txt crawlers receive can differ from the one in your repository, because Cloudflare adds its own lines. Fetch the live file from your domain and read it.
User-triggered fetchers may not read robots.txt at all
Nothing in robots.txt stops a crawler that chooses to ignore it. RFC 9309 says so: "These rules are not a form of access authorization." Most operators also say that their user-triggered fetchers treat it as optional:
- OpenAI, on
ChatGPT-User: "robots.txt rules may not apply." - Perplexity, on
Perplexity-User: it "generally ignores robots.txt rules." - Google, on its user-triggered fetchers: they "generally ignore robots.txt rules."
- Anthropic says its bots,
Claude-Userincluded, "respect 'do not crawl' signals by honoring industry standard directives in robots.txt."
So a Disallow for a user fetcher does not guarantee that the page stays out of every answer. It is still a clear statement of intent, and some operators honour it. If the problem is load rather than use, robots.txt is the wrong tool on its own. Anthropic documents Crawl-delay for ClaudeBot, and rate limits work regardless of what a crawler reads. A firewall challenge, though, also stops the crawlers you wanted to keep.
Saying it in words: Content Signals and AIPREF
robots.txt answers "may you fetch this?". It cannot say "you may index this for search, but not train on it" to a crawler that does both. Two vocabularies try to fill that gap:
- Content Signals is Cloudflare's convention:
Content-Signal: search=yes, ai-input=yes, ai-train=no. It definessearch,ai-input(content used in AI answers, such as grounding) andai-train. Its IETF draft has expired and was never endorsed by the IETF, so the syntax comes from Cloudflare's own deployment and documentation. - AIPREF is the IETF working group's vocabulary,
Content-Usage: train-ai=n. The keys are reversed,train-aiagainst Cloudflare'sai-train, so the two are easy to confuse. It is an active working-group draft (draft-ietf-aipref-vocab-08, 14 September 2026), intended as a Proposed Standard and not yet an RFC.
Both are declarations of preference, and neither blocks anything. They are worth adding because they state your terms in a form a crawler can parse, in the file every crawler already reads.
How the Good for Bots Score treats these choices
The Good for Bots Score follows the same split. Blocking training crawlers never costs points: it is the site owner's call. Blocking a search or user-fetch crawler from the homepage is a penalty, because the site has declared that it does not want to be read for AI answers. We charge only tokens whose operators document them, so an old token from a legacy blocklist costs nothing. Naming AI crawlers explicitly in robots.txt earns points, and so does declaring your training terms with Content Signals or AIPREF.
Our own crawler, GoodForBotsBot, reads robots.txt before anything else. If you refuse it, the report says "opted out" and shows no score.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI says GPTBot and OAI-SearchBot are independent: "Each setting is independent of the others." ChatGPT search uses OAI-SearchBot, and user-requested fetches use ChatGPT-User.
Does blocking Google-Extended remove my site from AI Overviews?
No. Google says Google-Extended "does not impact a site's inclusion in Google Search", and AI Overviews are part of Search. Google's separate Search Console control for AI Overviews and AI Mode has been available worldwide since 31 August 2026. Google-Extended does govern grounding in Gemini Apps and on Vertex AI, as well as Gemini training.
Will ChatGPT-User or Perplexity-User obey a Disallow rule?
Not necessarily. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores robots.txt rules", because a person asked for the page. Anthropic says its bots, Claude-User included, honour robots.txt.
Is the ai.robots.txt list safe to copy?
Only if you want to leave AI search. The version from 7 September 2026 disallows every search and user-fetch crawler in our registry, and Applebot, alongside the training crawlers. Its README says it lists crawlers "of all types, regardless of purpose."
Sources
- RFC 9309: Robots Exclusion Protocol — IETF
- Overview of OpenAI crawlers — OpenAI
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic
- Perplexity crawlers — Perplexity
- Google's common crawlers: Google-Extended — Google
- Google's user-triggered fetchers — Google
- New opportunities, control and insights for website owners — Google
- About Applebot — Apple
- Have it both ways: stay discoverable in search while disallowing AI training — Cloudflare
- Managed robots.txt and Content Signals — Cloudflare
- A Vocabulary For Expressing AI Usage Preferences (draft-ietf-aipref-vocab) — IETF AIPREF Working Group
- ai.robots.txt — GitHub

