---
title: "Which AI crawlers to block in robots.txt, and which to keep"
description: "Blocking GPTBot or ClaudeBot refuses model training and leaves AI search alone. Blocking OAI-SearchBot or Claude-User keeps you out of their answers."
url: "https://goodforbots.com/blog/which-ai-crawlers-to-block"
image: "https://goodforbots.com/images/blog/which-ai-crawlers-to-block.webp"
author: "Mateusz Pawlica"
category: "AI Crawlers"
date: "2026-09-24"
lastFactChecked: "2026-09"
readingMinutes: 10
---

# Which AI crawlers to block in robots.txt, and which to keep

Blocking GPTBot or ClaudeBot refuses model training and leaves AI search alone. Blocking OAI-SearchBot or Claude-User keeps you out of their answers.

Block the training crawlers if you do not want your content used to train models, and leave the search crawlers and user-triggered fetchers alone if you want your pages to appear in AI answers. OpenAI and Anthropic run a separate bot for training, for search and for fetching a page a user asked about, so refusing `GPTBot` or `ClaudeBot` costs nothing in their AI search. `GPTBot` and `OAI-SearchBot` are two separate decisions.

Google and Apple work differently. One crawler does everything, and a control token (`Google-Extended`, `Applebot-Extended`) decides whether the content may be used for training. Disallowing the token does not remove a site from Google Search or from Apple's search features. Disallowing `Googlebot` or `Applebot` does. One exception is covered below: `Google-Extended` also governs grounding in Gemini Apps.

Two things make the difference easy to lose. Since 15 September 2026, Cloudflare's "Block" setting also stops Googlebot, Bingbot and Applebot. And the ready-made ai.robots.txt list on GitHub blocks every search and user-fetch crawler we track.

## Three jobs: training, search and user fetch

An AI company can visit your site for three reasons, and the major operators document a separate control for each:

- **Training.** Collecting pages that may be used to train or fine-tune a model. Refusing it keeps future content out of the training set. It does not remove you from anything a user sees today.
- **Search.** Building an index that an AI search product answers from. OpenAI says sites that opt out of `OAI-SearchBot` "will not be shown in ChatGPT search answers, though can still appear as navigational links."
- **User fetch.** Retrieving a page because a person asked for it, in the moment. Anthropic says that disabling `Claude-User` "prevents our system from retrieving your content in response to a user query".

Site owners treat these differently when they are given the choice. Cloudflare reported on 15 September 2026 that fewer than 1% of the sites it serves block search bots, while 17% enable some mechanism to block training. A single "block AI" switch hides that difference.

## Which user agent does which job

These are the tokens you write in a `User-agent:` line, with what their operators say about them. The quotes come from the operators' pages, read on 24 September 2026.

| Operator   | Token               | Job           | What the operator says blocking it does                                                                         |
| ---------- | ------------------- | ------------- | --------------------------------------------------------------------------------------------------------------- |
| OpenAI     | `GPTBot`            | Training      | "Disallowing GPTBot indicates a site's content should not be used in training generative AI foundation models." |
| OpenAI     | `OAI-SearchBot`     | Search        | Opted-out sites "will not be shown in ChatGPT search answers".                                                  |
| OpenAI     | `ChatGPT-User`      | User fetch    | "Because these actions are initiated by a user, robots.txt rules may not apply."                                |
| Anthropic  | `ClaudeBot`         | Training      | The site's "future materials should be excluded from our AI model training datasets."                           |
| Anthropic  | `Claude-SearchBot`  | Search        | Prevents indexing for search, which "may reduce your site's visibility and accuracy in user search results."    |
| Anthropic  | `Claude-User`       | User fetch    | Prevents retrieving your content "in response to a user query".                                                 |
| Perplexity | `PerplexityBot`     | Search        | Surfaces sites in Perplexity results; "not used to crawl content for AI foundation models."                     |
| Perplexity | `Perplexity-User`   | User fetch    | "This fetcher generally ignores robots.txt rules."                                                              |
| Google     | `Google-Extended`   | Control token | Governs Gemini training and grounding in Gemini Apps; "does not impact a site's inclusion in Google Search".    |
| Apple      | `Applebot-Extended` | Control token | "Applebot-Extended does not crawl webpages." Pages that disallow it "can still be included in search results."  |

OpenAI puts the independence in one sentence: "Each setting is independent of the others." Blocking `GPTBot` does not block `OAI-SearchBot`, and allowing `OAI-SearchBot` does not allow training.

Many other operators publish crawlers too, including Meta, Amazon, Mistral and DuckDuckGo. The same rule applies to them: look up the token on the operator's own page and read what it is for before you name it. [How to write robots.txt rules for AI crawlers](https://goodforbots.com/guides/robots-txt-ai-crawlers) lists every token our scanner recognises, grouped by purpose.

### Google and Apple: one crawler, a control token

`Googlebot` and `Applebot` are the crawlers. `Google-Extended` and `Applebot-Extended` are product tokens that exist only to be named in `robots.txt`. Apple says it directly: "Applebot-Extended does not crawl webpages." Google describes `Google-Extended` as "a standalone product token".

Because AI Overviews and AI Mode are Google Search features, blocking `Google-Extended` does not take a site out of them. Since 31 August 2026 Google has offered a separate Search Console toggle for them worldwide, and says that toggle "will not be used as a ranking signal for search results outside of these generative AI Search features."

`Google-Extended` covers more than training. Google's page says it also governs "grounding (providing content from the Google Search index to the model at prompt time…)" in Gemini Apps and in Grounding with Google Search on Vertex AI. Disallowing it keeps your content out of grounded answers in those products as well as out of training data. Google offers no separate token for refusing training while staying available for grounding.

## A robots.txt that blocks training and keeps AI search

This file refuses the four training controls from the table and leaves everything else open. `example.com` is a placeholder.

`/robots.txt`

```txt
# Keep content out of model training
# (Google-Extended also covers grounding in Gemini Apps)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# Everyone else, including OAI-SearchBot, Claude-SearchBot,
# PerplexityBot and the user-triggered fetchers
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /

Sitemap: https://example.com/sitemap.xml
```

RFC 9309, the Robots Exclusion Protocol, lets one group start with several `User-agent` lines, so the four tokens share one `Disallow`. Product tokens are matched case-insensitively. `Content-Signal` is explained further down. It is not part of the protocol. RFC 9309 says such records "MUST NOT interfere with the parsing of explicitly defined records", so a crawler that does not know the line still reads the group's rules.

Changes take a while to reach the operators. OpenAI says "it can take \~24 hours from a site's robots.txt update for our systems to adjust", and Perplexity says "up to 24 hours".

### A named group replaces the `*` group

Adding a group for one crawler can open paths you meant to keep closed. Under RFC 9309 a crawler obeys the group that names it, and it falls back to `User-agent: *` only "if no matching group exists." Groups are not layered. So this file:

`/robots.txt`

```txt
User-agent: *
Disallow: /admin/

User-agent: OAI-SearchBot
Allow: /
```

lets `OAI-SearchBot` into `/admin/`, because its own group has no rule for that path, and a URI with no matching rule is allowed. If you name a crawler, repeat every `Disallow` it should still obey inside its group.

To see how your own file treats each of these crawlers, run it through the [robots.txt checker](https://goodforbots.com/tools/robots-txt-checker).

## Before you copy ai.robots.txt

The [ai.robots.txt](https://github.com/ai-robots-txt/ai.robots.txt) project on GitHub publishes a ready-made `robots.txt` that disallows every crawler it lists. Its README is clear about the scope: "This list contains AI-related crawlers of all types, regardless of purpose."

We compared the file with our own crawler registry, which was checked against each operator's documentation on 24 September 2026. The version we read was last changed on 7 September 2026 (commit `2acefa38cc`). It has 175 `User-agent` lines naming 172 distinct tokens. Among them are all 11 search crawlers and all 9 user-triggered fetchers in our registry: `OAI-SearchBot`, `ChatGPT-User`, `Claude-SearchBot`, `Claude-User`, `PerplexityBot`, `Perplexity-User` and the rest. It also names `Applebot`, whose data Apple says powers "the search technology integrated into many user experiences in Apple’s ecosystem including Spotlight, Siri, and Safari."

The project does what it says. Copying the file is a decision to leave AI search and user-requested fetches, and Apple's search features, along with training. If that is your decision, the file is a convenient way to make it. If you only meant to refuse training, use a list of training tokens instead.

## If you use Cloudflare: "Block" changed on 15 September 2026

Cloudflare's AI crawler controls now separate three behaviours: Search, Training and Agent (Cloudflare's term for user-directed fetches and browser agents). On 15 September 2026 it changed what "Block" means. In Cloudflare's words, "Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training."

The new setting for refusing training while staying in search is **Disallow AI Training**. It publishes the preference in `robots.txt` through Cloudflare's Bot Preference Sync, keeps what Cloudflare calls "Accountable" mixed-use crawlers allowed for search, and blocks other training crawlers. Bing does not yet read a no-training preference from `robots.txt`; Cloudflare says Microsoft is targeting early 2027. The same post says the legacy "Block AI Bots" toggle is being deprecated.

Cloudflare says existing settings were migrated automatically. A domain that only ever used the old "Block AI Bots" toggle, set to Block, now has Search on Allow, Training on Disallow AI Training and Agent on "Block on pages with ads". That last value refuses user-triggered fetchers on any page Cloudflare detects serving ads. With these settings on, the `robots.txt` crawlers receive can differ from the one in your repository, because Cloudflare adds its own lines. Fetch the live file from your domain and read it.

## User-triggered fetchers may not read robots.txt at all

Nothing in `robots.txt` stops a crawler that chooses to ignore it. RFC 9309 says so: "These rules are not a form of access authorization." Most operators also say that their user-triggered fetchers treat it as optional:

- OpenAI, on `ChatGPT-User`: "robots.txt rules may not apply."
- Perplexity, on `Perplexity-User`: it "generally ignores robots.txt rules."
- Google, on its user-triggered fetchers: they "generally ignore robots.txt rules."
- Anthropic says its bots, `Claude-User` included, "respect 'do not crawl' signals by honoring industry standard directives in robots.txt."

So a `Disallow` for a user fetcher does not guarantee that the page stays out of every answer. It is still a clear statement of intent, and some operators honour it. If the problem is load rather than use, `robots.txt` is the wrong tool on its own. Anthropic documents `Crawl-delay` for `ClaudeBot`, and rate limits work regardless of what a crawler reads. A firewall challenge, though, also stops the crawlers you wanted to keep.

## Saying it in words: Content Signals and AIPREF

`robots.txt` answers "may you fetch this?". It cannot say "you may index this for search, but not train on it" to a crawler that does both. Two vocabularies try to fill that gap:

- **Content Signals** is Cloudflare's convention: `Content-Signal: search=yes, ai-input=yes, ai-train=no`. It defines `search`, `ai-input` (content used in AI answers, such as grounding) and `ai-train`. Its IETF draft has expired and was never endorsed by the IETF, so the syntax comes from Cloudflare's own deployment and documentation.
- **AIPREF** is the IETF working group's vocabulary, `Content-Usage: train-ai=n`. The keys are reversed, `train-ai` against Cloudflare's `ai-train`, so the two are easy to confuse. It is an active working-group draft (`draft-ietf-aipref-vocab-08`, 14 September 2026), intended as a Proposed Standard and not yet an RFC.

Both are declarations of preference, and neither blocks anything. They are worth adding because they state your terms in a form a crawler can parse, in the file every crawler already reads.

## How the Good for Bots Score treats these choices

The [Good for Bots Score](https://goodforbots.com/standards) follows the same split. Blocking training crawlers never costs points: it is the site owner's call. Blocking a search or user-fetch crawler from the homepage is a [penalty](https://goodforbots.com/standards#ai-crawler-access), because the site has declared that it does not want to be read for AI answers. We charge only tokens whose operators document them, so an old token from a legacy blocklist costs nothing. Naming AI crawlers explicitly in `robots.txt` [earns points](https://goodforbots.com/standards#ai-crawler-rules), and so does [declaring your training terms](https://goodforbots.com/standards#content-signals) with Content Signals or AIPREF.

Our own crawler, [GoodForBotsBot](https://goodforbots.com/bot), reads `robots.txt` before anything else. If you refuse it, the report says "opted out" and shows no score.

See which AI crawlers your robots.txt lets in [Check a site](https://goodforbots.com/#scan)

## Frequently asked questions

### Does blocking GPTBot remove my site from ChatGPT search?

No. OpenAI says GPTBot and OAI-SearchBot are independent: "Each setting is independent of the others." ChatGPT search uses OAI-SearchBot, and user-requested fetches use ChatGPT-User.

### Does blocking Google-Extended remove my site from AI Overviews?

No. Google says Google-Extended "does not impact a site's inclusion in Google Search", and AI Overviews are part of Search. Google's separate Search Console control for AI Overviews and AI Mode has been available worldwide since 31 August 2026. Google-Extended does govern grounding in Gemini Apps and on Vertex AI, as well as Gemini training.

### Will ChatGPT-User or Perplexity-User obey a Disallow rule?

Not necessarily. OpenAI says robots.txt rules "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores robots.txt rules", because a person asked for the page. Anthropic says its bots, Claude-User included, honour robots.txt.

### Is the ai.robots.txt list safe to copy?

Only if you want to leave AI search. The version from 7 September 2026 disallows every search and user-fetch crawler in our registry, and Applebot, alongside the training crawlers. Its README says it lists crawlers "of all types, regardless of purpose."

**Sources**

1. [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html) — IETF
2. [Overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots) — OpenAI
3. [Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) — Anthropic
4. [Perplexity crawlers](https://docs.perplexity.ai/guides/bots) — Perplexity
5. [Google's common crawlers: Google-Extended](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) — Google
6. [Google's user-triggered fetchers](https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers) — Google
7. [New opportunities, control and insights for website owners](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/) — Google
8. [About Applebot](https://support.apple.com/en-us/119829) — Apple
9. [Have it both ways: stay discoverable in search while disallowing AI training](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/) — Cloudflare
10. [Managed robots.txt and Content Signals](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/) — Cloudflare
11. [A Vocabulary For Expressing AI Usage Preferences (draft-ietf-aipref-vocab)](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/) — IETF AIPREF Working Group
12. [ai.robots.txt](https://github.com/ai-robots-txt/ai.robots.txt) — GitHub

## About the author

**Mateusz Pawlica**, Founder of Good for Bots. Mateusz has been building digital products since 2014 — games and mobile apps first, then websites, directories and AI automations for businesses. He runs [Pawlica Web & AI](https://pawlicaweb.pl) in Tychy, Poland. Before Good for Bots he built two directories of his own, SaaS Cubes and Mapa Oświatowa, where llms.txt files, Markdown versions of pages and structured data are part of the everyday work rather than a checklist. [LinkedIn](https://pl.linkedin.com/in/mateusz-pawlica-65238816b) · [pawlicaweb.pl](https://pawlicaweb.pl)

## Structured data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "WebPage",
      "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#webpage",
      "url": "https://goodforbots.com/blog/which-ai-crawlers-to-block",
      "name": "Which AI crawlers to block in robots.txt, and which to keep",
      "description": "Blocking GPTBot or ClaudeBot refuses model training and leaves AI search alone. Blocking OAI-SearchBot or Claude-User keeps you out of their answers.",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://goodforbots.com/#website"
      },
      "breadcrumb": {
        "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#breadcrumb"
      },
      "mainEntity": {
        "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#article"
      },
      "datePublished": "2026-09-24",
      "dateModified": "2026-09-24"
    },
    {
      "@type": "BreadcrumbList",
      "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#breadcrumb",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "name": "Home",
          "item": "https://goodforbots.com/"
        },
        {
          "@type": "ListItem",
          "position": 2,
          "name": "Blog",
          "item": "https://goodforbots.com/blog"
        },
        {
          "@type": "ListItem",
          "position": 3,
          "name": "Which AI crawlers to block in robots.txt, and which to keep",
          "item": "https://goodforbots.com/blog/which-ai-crawlers-to-block"
        }
      ]
    },
    {
      "@type": "BlogPosting",
      "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#article",
      "url": "https://goodforbots.com/blog/which-ai-crawlers-to-block",
      "mainEntityOfPage": {
        "@id": "https://goodforbots.com/blog/which-ai-crawlers-to-block#webpage"
      },
      "headline": "Which AI crawlers to block in robots.txt, and which to keep",
      "description": "Blocking GPTBot or ClaudeBot refuses model training and leaves AI search alone. Blocking OAI-SearchBot or Claude-User keeps you out of their answers.",
      "image": {
        "@type": "ImageObject",
        "url": "https://goodforbots.com/images/blog/which-ai-crawlers-to-block.webp",
        "width": {
          "@type": "QuantitativeValue",
          "value": 1280,
          "unitCode": "E37",
          "unitText": "px"
        },
        "height": {
          "@type": "QuantitativeValue",
          "value": 720,
          "unitCode": "E37",
          "unitText": "px"
        },
        "caption": "An ink drawing of a picket fence with two gates: a small robot carrying a stack of papers waits at the latched one, while a second robot steps through the other, swung open."
      },
      "datePublished": "2026-09-24",
      "dateModified": "2026-09-24",
      "author": {
        "@id": "https://goodforbots.com/#person-mateusz-pawlica"
      },
      "publisher": {
        "@id": "https://goodforbots.com/#organization"
      },
      "isPartOf": {
        "@id": "https://goodforbots.com/blog#blog"
      },
      "articleSection": "AI Crawlers",
      "inLanguage": "en",
      "wordCount": 2227,
      "timeRequired": "PT10M"
    },
    {
      "@type": "Person",
      "@id": "https://goodforbots.com/#person-mateusz-pawlica",
      "name": "Mateusz Pawlica",
      "givenName": "Mateusz",
      "familyName": "Pawlica",
      "jobTitle": "Founder of Good for Bots",
      "description": "Web developer who builds directories and AI-powered products, and the founder of Good for Bots.",
      "url": "https://goodforbots.com/about#founder",
      "sameAs": [
        "https://pl.linkedin.com/in/mateusz-pawlica-65238816b",
        "https://pawlicaweb.pl/o-mnie"
      ],
      "knowsAbout": [
        "llms.txt",
        "robots.txt",
        "AI crawlers",
        "Markdown content negotiation",
        "Structured data",
        "Technical SEO",
        "Web directories",
        "Next.js",
        "TypeScript",
        "Generative AI"
      ],
      "worksFor": {
        "@id": "https://goodforbots.com/#organization"
      }
    },
    {
      "@type": "Blog",
      "@id": "https://goodforbots.com/blog#blog",
      "url": "https://goodforbots.com/blog",
      "name": "AI visibility & web standards blog",
      "description": "Web standards for AI crawlers and agents, explained from the specifications: llms.txt, robots.txt, Markdown negotiation and more, with data from our own scans.",
      "inLanguage": "en",
      "publisher": {
        "@id": "https://goodforbots.com/#organization"
      },
      "isPartOf": {
        "@id": "https://goodforbots.com/#website"
      }
    },
    {
      "@type": "Organization",
      "@id": "https://goodforbots.com/#organization",
      "name": "Good for Bots",
      "url": "https://goodforbots.com",
      "logo": "https://goodforbots.com/icon-512.png",
      "founder": {
        "@id": "https://goodforbots.com/#person-mateusz-pawlica"
      },
      "contactPoint": {
        "@type": "ContactPoint",
        "contactType": "customer support",
        "email": "support@goodforbots.com",
        "url": "https://goodforbots.com/contact"
      }
    },
    {
      "@type": "WebSite",
      "@id": "https://goodforbots.com/#website",
      "name": "Good for Bots",
      "url": "https://goodforbots.com",
      "description": "Check if AI crawlers can access and read your website, free. Get a score out of 100 with every finding and fix prompts, then browse reports in the directory.",
      "inLanguage": "en",
      "publisher": {
        "@id": "https://goodforbots.com/#organization"
      },
      "potentialAction": {
        "@type": "SearchAction",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://goodforbots.com/directory?q={q}"
        },
        "query-input": "required maxlength=120 name=q"
      }
    }
  ]
}
```
