---
title: "How to write robots.txt rules for AI crawlers"
description: "Write robots.txt rules that treat AI training, search and user-triggered crawlers separately, with tested examples, the current token list and common fixes."
url: "https://goodforbots.com/guides/robots-txt-ai-crawlers"
date: "2026-09-24"
reviewed: "2026-09-25"
---

# How to write robots.txt rules for AI crawlers

By Good for Bots · Reviewed 2026-09-25 · [Crawler access](https://goodforbots.com/guides.md?category=crawler-access)

Give AI crawlers `User-agent` groups by name and decide per purpose: training, search, or a fetch a person asked for. Refusing a training crawler such as `GPTBot` keeps your content out of that operator's future training data and leaves its AI search alone. Refusing a search crawler or a user-triggered fetcher, such as `OAI-SearchBot` or `Claude-User`, asks that operator to leave your pages out of its answers.

`robots.txt` is a request, not a lock. Crawlers that follow RFC 9309 obey it, some user-triggered fetchers say they may not, and a firewall can refuse a crawler the file lets in. Whether to refuse training is your decision; [which AI crawlers to block, and which to keep](https://goodforbots.com/blog/which-ai-crawlers-to-block) sets out what each choice costs. This guide covers writing the rules and checking that crawlers receive them.

## Start with a complete example

The file below is for **Harbour, a fictional service** that refuses model training and stays open to AI search. `example.com` is a placeholder. Replace the paths and the sitemap address with your own.

```txt
# robots.txt for Harbour, a fictional service. example.com is a placeholder.

# Model training: refused. Google-Extended also covers grounding in Gemini Apps.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /

# Every other crawler, including AI search crawlers and user-triggered
# fetchers such as OAI-SearchBot, Claude-SearchBot and ChatGPT-User.
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Content-Usage: train-ai=n
Disallow: /account/
Disallow: /search

Sitemap: https://example.com/sitemap.xml
```

The first group names five training crawlers and control tokens and refuses them the whole site. The second, `User-agent: *`, covers every crawler not named elsewhere, AI search crawlers and user-triggered fetchers included. It keeps them out of the account area and internal search results, and states the owner's terms in two vocabularies explained below. This file passes our robots.txt, AI crawler rules, crawler access and usage-preference checks at the versions reviewed for this guide.

To allow training instead, keep the first group but give it the same `Disallow` lines as the `*` group, and change `ai-train=no` to `ai-train=yes` and `train-ai=n` to `train-ai=y`. An `Allow: /` alone in that group would open `/account/` to those crawlers, for the reason in the next section. Deleting the group also allows training, but the file then names no AI crawler, and [our AI crawler rules check](https://goodforbots.com/standards#ai-crawler-rules) looks for an explicit position either way.

## How crawlers read robots.txt

[RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html), the Robots Exclusion Protocol, defines how a crawler finds the file and applies it. These parts decide whether rules for AI crawlers do what you meant:

- **One file per origin.** The file lives at `/robots.txt`, all lowercase, at the top of the host. Google spells out the consequence: its rules "apply only to the host, protocol, and port number where the robots.txt file is hosted." A subdomain such as `shop.example.com` needs its own file.
- **A group starts with its `User-agent` lines.** One group can name several crawlers, as the example does. It ends at the next `User-agent` line after its rules; a blank line does not end it.
- **Names match regardless of case.** "Crawlers MUST use case-insensitive matching", so `gptbot` and `GPTBot` name the same crawler. Write the token alone, without a version such as `/1.0`: a product token may contain only letters, underscores and hyphens.
- **One set of rules applies.** A crawler combines every group that names it and ignores the others. It obeys `User-agent: *` only "if no matching group exists", so a named group never inherits the `*` rules.
- **The longest path wins.** RFC 9309 defines the most specific match as "the match that has the most octets." When an `Allow` and a `Disallow` are equivalent, the `Allow` should win. A path no rule matches is allowed, and paths should be compared case-sensitively. `*` matches any run of characters and `$` anchors the end of the path.
- **Other lines do not change the rules.** `Sitemap`, `Crawl-delay`, `Content-Signal` and `Content-Usage` are not part of the protocol, and parsing such records "MUST NOT interfere" with the defined ones.

The HTTP answer counts as much as the text. A 4xx means the file is unavailable, and the crawler "MAY access any resources on the server." A 5xx or a network error means the crawler "MUST assume complete disallow". Crawlers "SHOULD NOT use the cached version for more than 24 hours," and they must parse at least 500 KiB; Google ignores anything after that limit.

The file is public. RFC 9309 warns that "listing paths in the robots.txt file exposes them publicly." A `Disallow` line hides nothing from a person or a crawler that ignores it, so protect private areas with authentication.

## Choose AI crawler tokens by purpose

The tables come from the same maintained knowledge base used by our scanner. They distinguish documented controls, HTTP identities, software tools and unconfirmed names. Each entry links to its evidence and review date. The scoring decision is shown separately: a bot can have several uses, and appearing in this catalogue does not automatically affect the score. Where no robots.txt token is confirmed, do not turn the displayed name into a rule.

### Training crawlers and control tokens

These controls express a training opt-out. Refusing them costs nothing in our score. Check each entry's status: Coherebot is a hypothetical example in Cohere's documentation, not an announced active crawler.

`Google-Extended` and `Applebot-Extended` are control tokens, not crawlers. Apple says "Applebot-Extended does not crawl webpages": `Googlebot` and `Applebot` fetch the pages, and the token tells Google or Apple how it may use them. Google-Extended also covers grounding in Gemini Apps and Grounding with Google Search on Vertex AI, and Google offers no token that refuses training alone. Never disallow `Googlebot` or `Applebot` to refuse AI training: that removes the site from Google Search or from Apple's search features.

| Name and robots.txt control                                                    | Use and status                                                                                                   | Evidence                                                                                                                                                                 |
| ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **GPTBot** · OpenAI. `GPTBot`                                                  | Collects content that may be used to train generative models. Training opt-out stays neutral.                    | [Operator source](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24.                                                                                     |
| **ClaudeBot** · Anthropic. `ClaudeBot`                                         | Collects content that may be used to train generative models. Training opt-out stays neutral.                    | [Operator source](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). Reviewed 2026-09-24. |
| **MistralAI-Training** · Mistral. `MistralAI-Training`                         | Collects content that may be used to train generative models. Training opt-out stays neutral.                    | [Operator source](https://docs.mistral.ai/robots). Reviewed 2026-09-24.                                                                                                  |
| **KimiBot** · Moonshot AI. `KimiBot`                                           | Collects content that may be used to train generative models. Training opt-out stays neutral.                    | [Operator source](https://www.kimi.ai/policies/kimi-crawlers). Reviewed 2026-09-24.                                                                                      |
| **AI2Bot**[^crawler-ai2bot] · Ai2. `AI2Bot`                                    | Collects web content to train open language models. Training opt-out stays neutral.                              | [Operator source](https://allenai.org/crawler). Reviewed 2026-09-24.                                                                                                     |
| **Applebot-Extended**[^crawler-applebot-extended] · Apple. `Applebot-Extended` | Controls training use of content collected by Applebot. Training opt-out stays neutral.                          | [Operator source](https://support.apple.com/en-us/119829). Reviewed 2026-09-24.                                                                                          |
| **Webzio-extended**[^crawler-webzio-extended] · Webz.io. `Webzio-extended`     | Controls whether Webz.io data can be used for AI or ML training. Training opt-out stays neutral.                 | [Operator source](https://webz.io/blog/company/from-omgilibot-to-the-webzbot-duo-a-powerful-leap-for-ethical-and-comprehensive-data-collection/). Reviewed 2026-09-24.   |
| **Coherebot**[^crawler-coherebot] · Cohere. No confirmed robots.txt token      | An example of how a future Cohere training crawler could be refused. Status: example. No crawler access penalty. | [Operator source](https://docs.cohere.com/v2/docs/cohere-web-crawlers). Reviewed 2026-09-24.                                                                             |

### AI search crawlers

These discover sources for AI answers, retrieve them at query time, or control the use of already indexed pages. OpenAI says sites that opt out of `OAI-SearchBot` "will not be shown in ChatGPT search answers, though can still appear as navigational links." Refusing one of these crawlers the homepage is a [penalty in our score](https://goodforbots.com/standards#ai-crawler-access): the file declares that the site does not want to be read for AI answers.

| Name and robots.txt control                                                                         | Use and status                                                                                                    | Evidence                                                                                                                                                                 |
| --------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **OAI-SearchBot** · OpenAI. `OAI-SearchBot`                                                         | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24.                                                                                     |
| **Claude-SearchBot** · Anthropic. `Claude-SearchBot`                                                | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). Reviewed 2026-09-24. |
| **PerplexityBot** · Perplexity. `PerplexityBot`                                                     | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://docs.perplexity.ai/docs/resources/perplexity-crawlers). Reviewed 2026-09-24.                                                                   |
| **MistralAI-Index** · Mistral. `MistralAI-Index`                                                    | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://docs.mistral.ai/robots). Reviewed 2026-09-24.                                                                                                  |
| **Kimi-SearchBot** · Moonshot AI. `Kimi-SearchBot`                                                  | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://www.kimi.ai/policies/kimi-crawlers). Reviewed 2026-09-24.                                                                                      |
| **Amzn-SearchBot**[^crawler-amzn-searchbot] · Amazon. `Amzn-SearchBot`                              | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://developer.amazon.com/amazonbot). Reviewed 2026-09-24.                                                                                          |
| **Meta-WebIndexer** · Meta. `Meta-WebIndexer`                                                       | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers). Reviewed 2026-09-24.                                                   |
| **ExaSearchBot**[^crawler-exasearchbot] · Exa. `ExaSearchBot`                                       | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://crawler.exa.ai/). Reviewed 2026-09-24.                                                                                                         |
| **AIWebIndex**[^crawler-aiwebindex] · Lyrenth. `AIWebIndex`                                         | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://lyrenth.com/crawler-policy). Reviewed 2026-09-24.                                                                                              |
| **DuckAssistBot**[^crawler-duckassistbot] · DuckDuckGo. `DuckAssistBot`                             | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot). Reviewed 2026-09-24.                                                              |
| **ShapBot**[^crawler-shapbot] · Parallel. `ShapBot`                                                 | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty.   | [Operator source](https://docs.parallel.ai/resources/crawler). Reviewed 2026-09-24.                                                                                      |
| **YandexAdditional**[^crawler-yandexadditional] · Yandex. `YandexAdditional`, `YandexAdditionalBot` | Controls use of indexed pages in Search with Yandex AI answers. Refusal counts toward the crawler access penalty. | [Operator source](https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots). Reviewed 2026-09-24.                                                      |

### User-triggered fetchers

These fetch a page when a person asks about it. OpenAI, Perplexity, Meta, Amazon and Google say their user-triggered fetchers may fetch a page robots.txt refuses them; OpenAI's wording for `ChatGPT-User` is "robots.txt rules may not apply." Anthropic says its bots, `Claude-User` included, honour robots.txt. A refusal still states your intent, and our score counts it the same way as a search crawler's.

| Name and robots.txt control                                                   | Use and status                                                                                       | Evidence                                                                                                                                                                 |
| ----------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **ChatGPT-User**[^crawler-chatgpt-user] · OpenAI. `ChatGPT-User`              | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24.                                                                                     |
| **Claude-User**[^crawler-claude-user] · Anthropic. `Claude-User`              | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). Reviewed 2026-09-24. |
| **Perplexity-User**[^crawler-perplexity-user] · Perplexity. `Perplexity-User` | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://docs.perplexity.ai/docs/resources/perplexity-crawlers). Reviewed 2026-09-24.                                                                   |
| **MistralAI-User**[^crawler-mistralai-user] · Mistral. `MistralAI-User`       | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://docs.mistral.ai/robots). Reviewed 2026-09-24.                                                                                                  |
| **Kimi-User**[^crawler-kimi-user] · Moonshot AI. `Kimi-User`                  | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://www.kimi.ai/policies/kimi-crawlers). Reviewed 2026-09-24.                                                                                      |
| **Amzn-User**[^crawler-amzn-user] · Amazon. `Amzn-User`                       | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://developer.amazon.com/amazonbot). Reviewed 2026-09-24.                                                                                          |
| **Diffbot-User**[^crawler-diffbot-user] · Diffbot. `Diffbot-User`             | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | [Operator source](https://www.diffbot.com/docs/crawl/faq/robots-txt). Reviewed 2026-09-24.                                                                               |

### Agents

Agents browse and act for a person. Our score does not charge a refusal. An HTTP identity or product name does not establish a robots.txt control. Google lists `Google-Agent` among its user-triggered fetchers, which "generally ignore robots.txt rules", and OpenAI's crawler page names no token for its agent. To control agents, use your firewall or bot-management rules.

| Name and robots.txt control                                                     | Use and status                                                                         | Evidence                                                                                                                              |
| ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| **Google-Agent**[^crawler-google-agent] · Google. No confirmed robots.txt token | Navigates websites and performs actions requested by users. No crawler access penalty. | [Operator source](https://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers). Reviewed 2026-09-24. |

### Extraction services

These crawl on behalf of customers who buy extraction, indexing or retrieval. Our score does not charge a refusal. Software packages may let callers choose their own identity; a package name is not necessarily a robots.txt token.

| Name and robots.txt control                                                                                | Use and status                                                                                             | Evidence                                                                                                                                                               |
| ---------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **FirecrawlAgent**[^crawler-firecrawlagent] · Firecrawl. `FirecrawlAgent`                                  | Controls Firecrawl site crawling for content extraction. No crawler access penalty.                        | [Operator source](https://www.firecrawl.dev/). Reviewed 2026-09-24.                                                                                                    |
| **Webzio**[^crawler-webzio] · Webz.io. `Webzio`                                                            | Collects and structures content for web data products. No crawler access penalty.                          | [Operator source](https://webz.io/blog/company/from-omgilibot-to-the-webzbot-duo-a-powerful-leap-for-ethical-and-comprehensive-data-collection/). Reviewed 2026-09-24. |
| **Google-CloudVertexBot**[^crawler-google-cloudvertexbot] · Google. `Google-CloudVertexBot`                | Crawls requested by site owners to build Vertex AI Agents. No crawler access penalty.                      | [Operator source](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers). Reviewed 2026-09-24.                                          |
| **Brightbot**[^crawler-brightbot] · Bright Data. No confirmed robots.txt token                             | Collects web data for Bright Data services. No crawler access penalty.                                     | [Operator source](https://brightdata.com/brightbot). Reviewed 2026-09-24.                                                                                              |
| **img2dataset**[^crawler-img2dataset] · img2dataset contributors. No confirmed robots.txt token            | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Operator source](https://github.com/rom1504/img2dataset). Reviewed 2026-09-24.                                                                                        |
| **ApifyWebsiteContentCrawler**[^crawler-apifywebsitecontentcrawler] · Apify. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Operator source](https://apify.com/apify/website-content-crawler/input-schema). Reviewed 2026-09-24.                                                                  |
| **Crawl4AI**[^crawler-crawl4ai] · Crawl4AI contributors. No confirmed robots.txt token                     | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Operator source](https://docs.crawl4ai.com/advanced/advanced-features/). Reviewed 2026-09-24.                                                                         |

### General search

These entries document general search uses. Inclusion in this catalogue does not add an AI crawler access penalty.

| Name and robots.txt control                             | Use and status                                                                             | Evidence                                                                                                                      |
| ------------------------------------------------------- | ------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- |
| **Diffbot**[^crawler-diffbot] · Diffbot. `Diffbot`      | Builds a general search engine and the Diffbot Knowledge Graph. No crawler access penalty. | [Operator source](https://www.diffbot.com/docs/crawl/faq/robots-txt). Reviewed 2026-09-24.                                    |
| **Googlebot**[^crawler-googlebot] · Google. `Googlebot` | Crawls for Google Search and associated search products. No crawler access penalty.        | [Operator source](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers). Reviewed 2026-09-24. |

### Mixed uses

One control can cover several uses. Read its research note before deciding: refusing training may also affect another use. The table shows our scoring decision separately from the operator's description.

| Name and robots.txt control                                                            | Use and status                                                                                                         | Evidence                                                                                                                      |
| -------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| **Meta-ExternalFetcher**[^crawler-meta-externalfetcher] · Meta. `Meta-ExternalFetcher` | Fetches user-requested links and supports agentic product functions. Refusal counts toward the crawler access penalty. | [Operator source](https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers). Reviewed 2026-09-24.        |
| **Google-Extended**[^crawler-google-extended] · Google. `Google-Extended`              | Controls Gemini training and grounding uses of crawled content. Training opt-out stays neutral.                        | [Operator source](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers). Reviewed 2026-09-24. |
| **Meta-ExternalAgent**[^crawler-meta-externalagent] · Meta. `Meta-ExternalAgent`       | Collects content for foundation model training and product indexing. Training opt-out stays neutral.                   | [Operator source](https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers). Reviewed 2026-09-24.        |
| **Amazonbot**[^crawler-amazonbot] · Amazon. `Amazonbot`                                | Collects content for product improvements and possible AI model training. Training opt-out stays neutral.              | [Operator source](https://developer.amazon.com/amazonbot). Reviewed 2026-09-24.                                               |
| **CCBot**[^crawler-ccbot] · Common Crawl. `CCBot`                                      | Builds an open web dataset used for research and model development. Training opt-out stays neutral.                    | [Operator source](https://commoncrawl.org/ccbot). Reviewed 2026-09-24.                                                        |
| **Applebot**[^crawler-applebot] · Apple. `Applebot`                                    | Collects content for Apple search features and possible model training. No crawler access penalty.                     | [Operator source](https://support.apple.com/en-us/119829). Reviewed 2026-09-24.                                               |

### Advertising

Advertising checks serve submitted ads. They are separate from organic discovery and carry no crawler access penalty.

| Name and robots.txt control                                                 | Use and status                                                                            | Evidence                                                                             |
| --------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| **OAI-AdsBot**[^crawler-oai-adsbot] · OpenAI. No confirmed robots.txt token | Validates submitted ad pages and helps determine ad relevance. No crawler access penalty. | [Operator source](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24. |

### Unconfirmed names and unclear purposes

These records preserve the research trail. They are not recommended robots.txt directives and do not affect our score. An unavailable source or undocumented token does not prove that a product was retired.

| Name and robots.txt control                                                                                                         | Use and status                                                                                             | Evidence                                                                                                                                                                                                                                                                  |
| ----------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **TerraCotta**[^crawler-terracotta] · Ceramic AI. `TerraCotta`                                                                      | A documented crawler whose precise content use is not established. No crawler access penalty.              | [Operator source](https://github.com/CeramicTeam/CeramicTerracotta). Reviewed 2026-09-24.                                                                                                                                                                                 |
| **anthropic-ai**[^crawler-anthropic-ai] · Anthropic (unconfirmed token). No confirmed robots.txt token                              | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24. |
| **Bytespider**[^crawler-bytespider] · Unconfirmed. No confirmed robots.txt token                                                    | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **FacebookBot**[^crawler-facebookbot] · Meta (unconfirmed token). No confirmed robots.txt token                                     | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                   |
| **cohere-training-data-crawler**[^crawler-cohere-training-data-crawler] · Cohere (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://docs.cohere.com/v2/docs/cohere-web-crawlers). Reviewed 2026-09-24.                                                                                                                                                                         |
| **Ai2Bot-Dolma**[^crawler-ai2bot-dolma] · Ai2 (unconfirmed token). No confirmed robots.txt token                                    | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://allenai.org/crawler) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                     |
| **DeepSeekBot**[^crawler-deepseekbot] · Unconfirmed. No confirmed robots.txt token                                                  | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **QwenBot**[^crawler-qwenbot] · Unconfirmed. No confirmed robots.txt token                                                          | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **PanguBot**[^crawler-pangubot] · Unconfirmed. No confirmed robots.txt token                                                        | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **Timpibot**[^crawler-timpibot] · Unconfirmed. No confirmed robots.txt token                                                        | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://timpi.io/crawler) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                        |
| **omgili**[^crawler-omgili] · Webz.io. No confirmed robots.txt token                                                                | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Operator source](https://webz.io/blog/company/from-omgilibot-to-the-webzbot-duo-a-powerful-leap-for-ethical-and-comprehensive-data-collection/). Reviewed 2026-09-24.                                                                                                    |
| **omgilibot**[^crawler-omgilibot] · Webz.io. No confirmed robots.txt token                                                          | Retained for research; no active scanner token is established. Status: retired. No crawler access penalty. | [Operator source](https://webz.io/blog/company/from-omgilibot-to-the-webzbot-duo-a-powerful-leap-for-ethical-and-comprehensive-data-collection/). Reviewed 2026-09-24.                                                                                                    |
| **YouBot**[^crawler-youbot] · You.com (unconfirmed token). No confirmed robots.txt token                                            | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://about.you.com/youbot/) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                   |
| **Operator**[^crawler-operator] · OpenAI (unconfirmed token). No confirmed robots.txt token                                         | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24.                                                                                                                                                                                 |
| **chatgpt-agent**[^crawler-chatgpt-agent] · OpenAI (unconfirmed token). No confirmed robots.txt token                               | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://developers.openai.com/api/docs/bots). Reviewed 2026-09-24.                                                                                                                                                                                 |
| **GoogleAgent-Mariner**[^crawler-googleagent-mariner] · Google (unconfirmed token). No confirmed robots.txt token                   | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers). Reviewed 2026-09-24.                                                                                                                                |
| **NovaAct**[^crawler-novaact] · Unconfirmed. No confirmed robots.txt token                                                          | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **Devin**[^crawler-devin] · Unconfirmed. No confirmed robots.txt token                                                              | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **Manus-User**[^crawler-manus-user] · Unconfirmed. No confirmed robots.txt token                                                    | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **TwinAgent**[^crawler-twinagent] · Unconfirmed. No confirmed robots.txt token                                                      | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **ApifyBot**[^crawler-apifybot] · Unconfirmed. No confirmed robots.txt token                                                        | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **Crawlspace**[^crawler-crawlspace] · Unconfirmed. No confirmed robots.txt token                                                    | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **HenkBot**[^crawler-henkbot] · Unconfirmed. No confirmed robots.txt token                                                          | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **Claude-Web**[^crawler-claude-web] · Anthropic (unconfirmed token). No confirmed robots.txt token                                  | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) · [Additional source](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24. |
| **iAskBot**[^crawler-iaskbot] · Unconfirmed. No confirmed robots.txt token                                                          | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |
| **iaskspider**[^crawler-iaskspider] · Unconfirmed. No confirmed robots.txt token                                                    | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | [Unconfirmed identity](https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.json). Reviewed 2026-09-24.                                                                                                                                                        |

## State your terms with Content Signals or Content-Usage

`Allow` and `Disallow` say whether a crawler may fetch a path. They cannot tell a crawler that does several jobs which of them you accept. Two vocabularies say that in `robots.txt`:

- **Content Signals** is Cloudflare's convention: `Content-Signal: search=yes, ai-input=yes, ai-train=no`. [contentsignals.org](https://contentsignals.org/) defines `search` as building a search index and returning links and short excerpts, `ai-input` as putting content into an AI model, for example for grounding, and `ai-train` as "Training or fine-tuning AI models." Values are `yes` and `no`, and Cloudflare says "the absence of a signal conveys no meaning." The IETF draft it came from has expired.
- **AIPREF** is the IETF working group's vocabulary: `Content-Usage: train-ai=n`. Its categories are `train-ai`, `ai-use` and `search`, with the values `y` and `n`. Keys are case-sensitive, and `train-ai` reverses Cloudflare's `ai-train`. [draft-ietf-aipref-vocab](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/) is an active working-group draft intended as a Proposed Standard. Version 08, of 14 September 2026, added `ai-use` and notes that its contents "DO NOT REFLECT CONSENSUS of the Working Group."

Both are declarations. Neither blocks a request, and a crawler that does not implement them ignores them. Write them inside a group, and remember that a crawler reads only one set of rules: a crawler with its own named group does not see the line under `User-agent: *`. The AIPREF attachment draft's example shows this, and contentsignals.org's example for specific crawlers repeats the line in each group. [Our usage-preference check](https://goodforbots.com/standards#content-signals) accepts either vocabulary.

## Publish the file where crawlers look

Serve the file at `https://your-domain.example/robots.txt` with HTTP 200, as UTF-8 `text/plain`. Where it comes from depends on your stack:

- **A static file** in the web root, or `app/robots.txt` in a Next.js app. Next.js can also generate the file from `app/robots.ts`; since version 16.3 a rule's `other` field adds lines such as `Content-Signal`, "scoped to the rule's `User-Agent` block."
- **WordPress** generates `/robots.txt` when no file exists at that path, and plugins change its output through the `robots_txt` filter.
- **Shopify** generates the file, and the `robots.txt.liquid` theme template customises it.
- **Webflow** and **Wix** have a robots.txt editor in their SEO settings. Squarespace users cannot edit the file. Its "Block known artificial intelligence crawlers" setting adds a fixed list that includes `DuckAssistBot` and `YouBot`, so check the published file against the controls you intended to set. DuckAssistBot retrieves sources for AI-assisted search answers.
- **Cloudflare** can put its own lines in front of yours. Its Bot Preference Sync, introduced in August 2026 and on by default for new customers, prepends them "to the existing material".

Then read what crawlers receive, not what is in your repository. Replace the example host and run:

```sh
curl --include https://your-domain.example/robots.txt
```

Check that the status is 200, the `Content-Type` is `text/plain` and the body starts with your rules, not with `<!doctype html>`. A redirect is acceptable if it ends at the file quickly: RFC 9309 asks crawlers to follow "at least five consecutive redirects." Compare the body with your source file. A line added by a CDN or a plugin is a rule crawlers obey all the same.

The [robots.txt checker](https://goodforbots.com/tools/robots-txt-checker) fetches the file from our side and shows how it treats each AI crawler.

Changes take time to arrive. RFC 9309 allows a cached copy for up to 24 hours. OpenAI, for search results, and Amazon say about 24 hours, Perplexity and Meta say up to 24 hours, and DuckDuckGo says a change "will take effect after 72 hours".

## Make sure the firewall agrees with the file

`robots.txt` only asks. A firewall, a bot-management product or a CDN setting decides what a crawler receives, and it can refuse a crawler the file allows:

- **Cloudflare.** Since 15 September 2026, "Block" in AI Crawl Control applies to mixed-use crawlers, "including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training." To refuse training only, Cloudflare offers "Disallow AI Training", which publishes the preference in `robots.txt` and keeps what it calls accountable mixed-use crawlers allowed for search.
- **Vercel.** The AI bots managed ruleset is off by default. Its Deny action "blocks all traffic identified as coming from AI bots", and the ruleset covers bots crawling for training, search and user-generated fetches.
- **AWS WAF.** Bot Control's `CategoryAI` rule blocks AI bots "regardless of whether the bots are verified or unverified."

A crawler that receives a challenge page or a 403 gets no content, whatever `robots.txt` says. [Our bot challenge check](https://goodforbots.com/standards#waf-challenge) looks for that in front of our own crawler, and [our crawler page](https://goodforbots.com/bot#cloudflare) explains what to look for in your security events.

## Slow crawlers down instead of refusing them

If the problem is load rather than use, `Crawl-delay` asks a crawler to wait between requests. It is not part of RFC 9309, and support varies. Google, Apple and Amazon say their crawlers do not follow it. Bing and Common Crawl say theirs do, and Anthropic says it supports "the non-standard Crawl-delay extension". A rate limit at your server works whether a crawler reads the file or not, and HTTP 429 with `Retry-After` tells a well-behaved client when to come back. [Our crawler](https://goodforbots.com/bot#controls) honours `Crawl-delay`.

## Fix common mistakes

### A blocklist that also refuses AI search

```txt
# Meant to refuse AI training. It also refuses AI search and user fetches.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Disallow: /

User-agent: *
Disallow: /account/
```

The comment says what the author wanted; the tokens say something else. `OAI-SearchBot`, `Claude-SearchBot` and `PerplexityBot` are search crawlers, and `ChatGPT-User`, `Claude-User` and `Perplexity-User` are user-triggered fetchers. Our crawler access check fails this file for all six. Keep `GPTBot` and `ClaudeBot` in the group and delete the other lines, so those crawlers fall back to `User-agent: *`.

Ready-made lists often mix purposes. The [ai.robots.txt](https://github.com/ai-robots-txt/ai.robots.txt) project says its list "contains AI-related crawlers of all types, regardless of purpose." Check each token against the tables above before you copy a list.

### A named group that drops your other rules

```txt
User-agent: *
Disallow: /account/

User-agent: OAI-SearchBot
Allow: /
```

The author meant to welcome `OAI-SearchBot` and also let it into `/account/`. The crawler finds a group that names it, so it never reads `User-agent: *`, and its own group has no rule for `/account/`. Repeat every `Disallow` a named crawler should still obey inside its group. Our crawler access check tests only the homepage path, so it passes this file; only you can see that a private path is now open.

### A file that names no AI crawler

```txt
User-agent: *
Disallow: /account/

Sitemap: https://example.com/sitemap.xml
```

This file is valid. Under RFC 9309 it lets every AI crawler in, except to `/account/`, and it says nothing about training, search or user fetches. [Our AI crawler rules check](https://goodforbots.com/standards#ai-crawler-rules) fails it because it takes no explicit position on AI crawlers, and [our usage-preference check](https://goodforbots.com/standards#content-signals) fails it because it declares no terms. The complete example above answers both. If you are content to let every crawler in, name the ones you have considered in a group with the same `Disallow` lines as `*`.

### An HTML page at /robots.txt

```html
<!doctype html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Harbour</title>
    <script type="module" src="/assets/index.js"></script>
  </head>
  <body>
    <div id="root"></div>
  </body>
</html>
```

Single-page apps and catch-all routes often answer every path with the application shell and HTTP 200. A crawler finds no rules in it, and our robots.txt check fails it as an HTML page. Serve a real text file at the path, ahead of the application's fallback route.

### A server error at /robots.txt

A 5xx answer is worse than no file. RFC 9309 tells crawlers to "assume complete disallow" when the file is unreachable because of a server or network error, so an outage or a broken route refuses every compliant crawler. Google documents its own variant: it stops crawling for 12 hours, then uses the last good copy for up to 30 days. If you have no rules to publish, a 404 lets crawlers in; a file with your rules is better.

## Keep the rules current

Operators add, rename and retire crawlers. When a new token shows up in your server logs, read what its operator says it does before you name it, and check whether it trains, searches or fetches for a person. After any deploy that touches routing, a CDN rule or a platform setting, fetch the live file again.

[Scan your site](https://goodforbots.com/#scan) to see the `robots.txt` our crawler receives and how each check below reads it. The report shows declared rules and what our own crawler met; it cannot see how your firewall answers other crawlers.

## Sources and review

Reviewed on 24 September 2026 against these sources:

- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html), IETF Proposed Standard
- [Google: How Google interprets the robots.txt specification](https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec)
- [OpenAI: Overview of OpenAI crawlers](https://developers.openai.com/api/docs/bots)
- [Anthropic: crawler controls](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
- [Perplexity: crawlers](https://docs.perplexity.ai/docs/resources/perplexity-crawlers)
- [Google's common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers)
- [Google's user-triggered fetchers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-user-triggered-fetchers)
- [Apple: About Applebot](https://support.apple.com/en-us/119829)
- [Meta: web crawlers](https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers)
- [Amazon: Amazonbot](https://developer.amazon.com/amazonbot), [Common Crawl: FAQ](https://commoncrawl.org/faq) and [DuckDuckGo: DuckAssistBot](https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot)
- [Content Signals](https://contentsignals.org/) and [Cloudflare: The Content Signals Policy](https://blog.cloudflare.com/content-signals-policy/)
- [draft-ietf-aipref-vocab](https://datatracker.ietf.org/doc/draft-ietf-aipref-vocab/) and [draft-ietf-aipref-attach](https://datatracker.ietf.org/doc/draft-ietf-aipref-attach/), IETF working-group drafts
- [Cloudflare: stay discoverable in search while disallowing AI training](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/) and [Bot Preference Sync](https://blog.cloudflare.com/bot-preference-sync/)
- [Vercel: bot management](https://vercel.com/docs/bot-management) and [AWS WAF Bot Control rule group](https://docs.aws.amazon.com/waf/latest/developerguide/aws-managed-rule-groups-bot.html)
- [Next.js: robots.txt](https://nextjs.org/docs/app/api-reference/file-conventions/metadata/robots), [WordPress: robots\_txt filter](https://developer.wordpress.org/reference/hooks/robots_txt/), [Shopify: robots.txt.liquid](https://shopify.dev/docs/storefronts/themes/architecture/templates/robots-txt-liquid), [Webflow](https://help.webflow.com/hc/en-us/articles/41954080897683-Set-robots-txt-rules), [Wix](https://support.wix.com/en/article/editing-your-sites-robotstxt-file) and [Squarespace](https://support.squarespace.com/hc/en-us/articles/360022347072-Request-that-AI-models-exclude-your-site)
- [Bing: To crawl or not to crawl, that is BingBot's question](https://blogs.bing.com/webmaster/2012/5/To-crawl-or-not-to-crawl,-that-is-BingBot-s-questi/)

The crawler tables come from our scanner's registry. The examples are ours, and each is tested against the checks listed below at the versions reviewed for this guide.

## Crawler research notes

[^crawler-amzn-searchbot]: Amazon documents fallback to other search bots when this token is unnamed. Our access
    check uses RFC 9309 wildcard fallback; it does not emulate that vendor extension.

[^crawler-exasearchbot]: Exa documents fallback to major search engine rules before the wildcard group. Our
    access check uses RFC 9309 wildcard fallback. Exabot is a different historical crawler,
    not an alias.

[^crawler-aiwebindex]: The current policy says both scheduled crawling and AIWebIndex-Agent on-demand index
    fills follow the AIWebIndex robots token. The HTTP names do not establish two
    independent access routes. The policy also discusses corpus licensing and separate
    training reservations; it says Lyrenth does not train foundation models itself.

[^crawler-duckassistbot]: DuckDuckGo describes real-time crawling for AI-assisted search answers, with no model
    training. It is classified as ai-search, rather than user-fetch.

[^crawler-shapbot]: Parallel documents discovery and indexing for its web APIs. The old registry placed this
    bot under extraction; the current catalogue records its search role.

[^crawler-yandexadditional]: Yandex lists YandexAdditional and YandexAdditionalBot together as controls over already
    indexed content. They are one catalogue entry and one penalty route. Either explicitly
    refused alias records a refusal; conflicting aliases are conservatively treated as
    refused. This is our declared-policy reading, not an implementation of Yandex-specific
    fallback rules.

[^crawler-ai2bot]: Ai2 documents AI2Bot, including the digit. Matching retains the scanner's existing
    product-token normalization; the catalogue compiler rejects collisions after that
    normalization.

[^crawler-applebot-extended]: This token does not fetch pages. Refusing it does not remove pages from Apple search results.

[^crawler-webzio-extended]: The operator describes validation and tagging of data collected by Webzio, not a second full-content crawler.

[^crawler-coherebot]: Cohere says it is not currently using crawlers for foundation-model training and lists
    no active bot. Coherebot appears in a hypothetical robots.txt example, not an active
    crawler announcement.

[^crawler-diffbot]: Diffbot explicitly distinguishes general search crawling from Diffbot-User requests and
    says it does not crawl to train foundation models. A general-search entry does not
    automatically become an AI access penalty.

[^crawler-googlebot]: Included for reference. This catalogue does not introduce another Googlebot access
    penalty; existing indexing checks own their separate rules.

[^crawler-chatgpt-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-claude-user]: Anthropic says its bots honour robots.txt, including this user-triggered fetcher.

[^crawler-perplexity-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-mistralai-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-kimi-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-amzn-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-diffbot-user]: User-triggered retrieval is separate from automatic crawling. A robots.txt rule
    expresses intent; it does not prove this fetcher will refuse a user request.

[^crawler-google-agent]: Google documents an HTTP identity among user-triggered fetchers, which generally ignore
    robots.txt. The page does not publish a dedicated robots.txt control for this agent; its
    HTTP name is not promoted to a scanner token.

[^crawler-firecrawlagent]: Firecrawl documents this robots.txt directive for its crawl endpoint. Callers can pass
    custom headers; the catalogue does not claim every request has this HTTP identity.

[^crawler-webzio]: Webz.io introduced Webzio and Webzio-extended as replacements for its older Omgilibot
    system. Webzio-extended separately controls training eligibility.

[^crawler-google-cloudvertexbot]: Google says this control affects owner-requested Vertex AI crawling, not Google Search.
    Its documented fallback includes Googlebot; the scanner does not add a visibility
    penalty for this service.

[^crawler-brightbot]: The operator publishes an HTTP identity and collectors.txt controls. That does not
    establish a robots.txt token contract, so this entry is informational.

[^crawler-img2dataset]: An image dataset downloader, not one centrally operated crawler. Its project name is not
    a documented robots.txt token.

[^crawler-apifywebsitecontentcrawler]: The actor documentation says it uses no specific user agent identifier. Respecting
    robots.txt is configurable and defaults to false. The actor name must not become a
    scanner token.

[^crawler-crawl4ai]: A configurable crawling library run by its users. The reviewed documentation does not
    establish the project name as one shared robots.txt identity.

[^crawler-meta-externalfetcher]: Meta says this fetcher may bypass robots.txt on user-requested fetches. Its user-fetch
    route remains in the access policy; its additional agent role does not create another
    charge.

[^crawler-google-extended]: No separate HTTP User-Agent. Refusal does not affect inclusion in Google Search. The
    scanner keeps blocking neutral because this control also refuses training;
    classification is independent of that scoring decision.

[^crawler-meta-externalagent]: Meta documents training and direct indexing for product improvement. This is not a
    training-only claim. Blocking remains neutral; Meta-WebIndexer is the separately
    documented AI search route.

[^crawler-amazonbot]: Amazon says collected content may train its AI models. Amzn-SearchBot and Amzn-User are
    separate controls. Blocking Amazonbot remains neutral.

[^crawler-ccbot]: Common Crawl publishes crawl data for reuse, rather than operating a consumer answer
    engine. Its dataset has uses beyond training. Blocking remains neutral under the
    existing training opt-out policy.

[^crawler-applebot]: Applebot-Extended separately controls foundation-model training. Applebot is included
    for reference without adding a new AI crawler access penalty.

[^crawler-oai-adsbot]: Only pages submitted as ads are visited. The operator excludes foundation-model
    training. This is an advertising identity, not an organic search access route.

[^crawler-terracotta]: The operator repository documents the token and robots compliance, but the reviewed page
    does not establish AI search or training use. The old search classification is not
    retained as a fact.

[^crawler-anthropic-ai]: The current Anthropic reference documents ClaudeBot, Claude-SearchBot and Claude-User,
    but not this spelling. Absence is not proof of retirement; retained as an unconfirmed
    historical name.

[^crawler-bytespider]: The previous registry attributed this name to ByteDance. The community list is a
    discovery lead, not confirmation of the operator, purpose or a working robots.txt token.
    No primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-facebookbot]: The current Meta crawler reference does not document this old spelling. Do not infer
    retirement or equate it with Meta-ExternalAgent.

[^crawler-cohere-training-data-crawler]: This name came from community lists. Cohere currently lists no active training crawler;
    Coherebot is only a hypothetical example. This spelling is not confirmed by its
    reference.

[^crawler-ai2bot-dolma]: The reviewed Ai2 crawler notice publishes AI2Bot. It does not establish this additional
    spelling as a current independent token.

[^crawler-deepseekbot]: The previous registry attributed this name to DeepSeek. The community list is a
    discovery lead, not confirmation of the operator, purpose or a working robots.txt token.
    No primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-qwenbot]: The previous registry attributed this name to Alibaba. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-pangubot]: The previous registry attributed this name to Huawei. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-timpibot]: The attributed operator crawler page returned 404. Community classifications disagree
    about training versus search. Neither purpose nor a current control contract is
    established here.

[^crawler-omgili]: Webz.io describes the move from Omgilibot to Webzio and Webzio-extended. Retained for
    historical blocklists; use the current documented controls. The separate omgili spelling
    has no current token contract in this source.

[^crawler-omgilibot]: Webz.io describes the move from Omgilibot to Webzio and Webzio-extended. Retained for
    historical blocklists; use the current documented controls. The separate omgili spelling
    has no current token contract in this source.

[^crawler-youbot]: The previously cited operator URL returned 404 during this review; the you.com variant
    redirected to sign-in. No replacement primary reference was established. This does not
    prove retirement; the name remains in research, without a penalty.

[^crawler-operator]: The OpenAI crawler reference does not establish this robots.txt token. An agent product
    name must not be converted into a robots token. Retained as an unconfirmed registry
    spelling, not as a claim that the product is inactive.

[^crawler-chatgpt-agent]: The OpenAI crawler reference does not establish this robots.txt token. An agent product
    name must not be converted into a robots token. Retained as an unconfirmed registry
    spelling, not as a claim that the product is inactive.

[^crawler-googleagent-mariner]: The current Google user-triggered fetcher list documents Google-Agent, not this token.
    No alias relationship is inferred.

[^crawler-novaact]: The previous registry attributed this name to Amazon. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-devin]: The previous registry attributed this name to Cognition. The community list is a
    discovery lead, not confirmation of the operator, purpose or a working robots.txt token.
    No primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-manus-user]: The previous registry attributed this name to Manus. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-twinagent]: The previous registry attributed this name to Twin. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-apifybot]: The previous registry attributed this name to Apify. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-crawlspace]: The previous registry attributed this name to Crawlspace. The community list is a
    discovery lead, not confirmation of the operator, purpose or a working robots.txt token.
    No primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-henkbot]: The community list attributes HenkBot to Valyu, while the previous registry said
    unknown. No primary source was established, so the operator remains unconfirmed.

[^crawler-claude-web]: The current Anthropic reference documents ClaudeBot, Claude-SearchBot and Claude-User,
    but not this spelling. Absence is not proof of retirement; retained as an unconfirmed
    historical name.

[^crawler-iaskbot]: The previous registry attributed this name to iAsk. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

[^crawler-iaskspider]: The previous registry attributed this name to iAsk. The community list is a discovery
    lead, not confirmation of the operator, purpose or a working robots.txt token. No
    primary source confirming this exact control was established in this review; it is
    excluded from scanner recognition.

## Current scanner criteria

These criteria come from our current standards catalogue. They describe what Good for Bots checks, including usefulness rules of our own that the format does not require, and what our detection cannot see. Each report keeps the methodology of the scan that produced it.

### robots.txt

active check · Reviewed 2026-09-24

We parse /robots.txt and check delivery and directives, as RFC 9309 defines them: lines may end in CR, LF or CRLF, and a Crawl-delay or Sitemap line does not end a user-agent group. Whether we may scan the site is decided separately, before scoring. A parseable refusal of our bot's access to the homepage ends the scan without a score, even when served under an unusual HTTP status. A rule that refuses us only a particular file, such as /llms.txt or a sitemap, leaves the scan running: we do not request that file, and the check that needs it treats it as unavailable to crawlers. A robots.txt answered with a bare 401 or 403 is treated as no rules for access, as RFC 9309 allows, but it does not pass this check.

**Results:** Pass: the file is delivered at HTTP 200 and has usable directives without parse errors. Warn: it declares nothing or has parsing problems, such as a line without a colon, a rule above the first User-agent line, or an empty or unparseable Sitemap URL. Fail: it is missing, empty, unreadable, an HTML page or delivered under a status other than 200.

**Limitations:** Our requirement for HTTP 200, and our partial credit for a file without directives or with lines we cannot parse, are scoring choices; RFC 9309 asks crawlers only to use the rules they can parse. We read a bounded file sample. User-agent names are compared by their letters, hyphens and underscores, so a digit in a name is ignored. This check does not judge whether the declared rules allow other AI crawlers; separate checks do that.

[Full methodology and sources](https://goodforbots.com/standards#robots-txt)

### Explicit AI crawler rules

active check · Reviewed 2026-09-24

We match user-agent tokens against our crawler registry, combine matching groups and record their purposes and access rules. A wildcard group alone does not count as an explicit AI crawler policy, and neither do Content-Signal or Content-Usage lines: they state preferences to every crawler without naming one, and the content signals check scores them.

**Results:** Pass: at least two crawler purposes are represented, or at least four distinct known crawlers are named. Warn: some are named but neither threshold is met. Fail: no known AI crawlers are named, including when the file states usage preferences for every crawler.

**Limitations:** The thresholds are our criteria, not RFC requirements. Recognition depends on the crawler registry, which we check against the operators' own documentation, and can miss newly introduced names. Earning these points does not mean search crawlers are allowed: that is a separate penalty check.

[Full methodology and sources](https://goodforbots.com/standards#ai-crawler-rules)

### Search crawler access

active check · Reviewed 2026-09-25

We apply robots.txt rules to the homepage path for the search and user-fetch crawler tokens their operators document. Explicit groups, wildcard fallback and rule precedence determine access; training, agent and extraction tokens are not charged. Only a robots.txt served normally, with HTTP 200, is read for this check.

**Results:** Pass: no applicable rules block the selected crawlers. Fail: one or more are blocked, with the exact deduction recorded in the report. N/A: the robots probe was not run. A missing robots file does not itself mean access is blocked.

**Limitations:** This evaluates declared access to the root path, not live requests impersonating other bots. Firewall behaviour and deeper paths can differ. OpenAI, Perplexity, Meta, Amazon and Google say their user-triggered fetchers may fetch a page robots.txt refuses them when a person asks for it, so we charge the declared refusal, not proven invisibility. Recognition depends on our maintained crawler registry, and we charge only tokens their operators document.

[Full methodology and sources](https://goodforbots.com/standards#ai-crawler-access)

### AI usage preferences

active check · Reviewed 2026-09-24

We read every Content-Signal and Content-Usage line in robots.txt, including rules scoped to a path, and the Content-Usage header on the final successful homepage response. Findings identify the header’s response URL. Both carriers share one check and count once. Either vocabulary can satisfy the check; recognised positive and negative preferences count equally. Only yes or no, or y or n for Content-Usage, states a preference; any other value states nothing. Content-Usage is read as the IETF draft specifies: parameters after a value are ignored, and a declaration that does not parse, for example one with uppercase keys, states nothing.

**Results:** Pass: at least one recognised category and value. Warn: a usage declaration exists but its categories or values are not recognised. Fail: no declaration is found.

**Limitations:** HTTP coverage is limited to the final homepage response; its header describes that response, not the whole site. We do not resolve effective permissions across carriers or decide whether a downstream service honours a preference. A rule scoped to a path counts as though it covered the whole site. Drafts and conventions are not all published standards. The check recognises declarations; it is not a legal interpretation of permission.

[Full methodology and sources](https://goodforbots.com/standards#content-signals)

[Check your site](https://goodforbots.com/#scan) · [Try the robots.txt checker](https://goodforbots.com/tools/robots-txt-checker)

## More guides

- [How to serve content AI crawlers can read without JavaScript](https://goodforbots.com/guides/server-rendered-html): Page content, reviewed Sep 25, 2026. Which AI crawlers run JavaScript, how to fix empty app shells in Next.js, Nuxt, SvelteKit or Vite, and how to test the raw HTML response yourself.
- [How to serve Markdown to AI agents with content negotiation](https://goodforbots.com/guides/markdown-negotiation): Page content, reviewed Oct 3, 2026. Return Markdown from the same URL when an AI agent sends Accept: text/markdown, with tested responses, Vary and CDN caching, and nginx and Next.js setups.
- [How to write a useful llms.txt](https://goodforbots.com/guides/llms-txt): Discovery files, reviewed Sep 24, 2026. Write an llms.txt that tells AI agents what your site covers: a complete tested example, the format explained, publishing checks and fixes for common mistakes.

[All guides](https://goodforbots.com/guides)

## Structured data

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "WebPage",
      "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#webpage",
      "url": "https://goodforbots.com/guides/robots-txt-ai-crawlers",
      "name": "How to write robots.txt rules for AI crawlers",
      "description": "Write robots.txt rules that treat AI training, search and user-triggered crawlers separately, with tested examples, the current token list and common fixes.",
      "inLanguage": "en",
      "isPartOf": {
        "@id": "https://goodforbots.com/#website"
      },
      "breadcrumb": {
        "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#breadcrumb"
      },
      "mainEntity": {
        "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#article"
      },
      "datePublished": "2026-09-24",
      "dateModified": "2026-09-25"
    },
    {
      "@type": "BreadcrumbList",
      "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#breadcrumb",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "name": "Home",
          "item": "https://goodforbots.com/"
        },
        {
          "@type": "ListItem",
          "position": 2,
          "name": "Guides",
          "item": "https://goodforbots.com/guides"
        },
        {
          "@type": "ListItem",
          "position": 3,
          "name": "How to write robots.txt rules for AI crawlers",
          "item": "https://goodforbots.com/guides/robots-txt-ai-crawlers"
        }
      ]
    },
    {
      "@type": "TechArticle",
      "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#article",
      "url": "https://goodforbots.com/guides/robots-txt-ai-crawlers",
      "mainEntityOfPage": {
        "@id": "https://goodforbots.com/guides/robots-txt-ai-crawlers#webpage"
      },
      "headline": "How to write robots.txt rules for AI crawlers",
      "description": "Write robots.txt rules that treat AI training, search and user-triggered crawlers separately, with tested examples, the current token list and common fixes.",
      "inLanguage": "en",
      "datePublished": "2026-09-24",
      "dateModified": "2026-09-25",
      "author": {
        "@id": "https://goodforbots.com/#organization"
      },
      "publisher": {
        "@id": "https://goodforbots.com/#organization"
      },
      "image": "https://goodforbots.com/og/guides/robots-txt-ai-crawlers.png?v=1wcpd3b7ysio2"
    },
    {
      "@type": "Organization",
      "@id": "https://goodforbots.com/#organization",
      "name": "Good for Bots",
      "url": "https://goodforbots.com",
      "logo": "https://goodforbots.com/icon-512.png",
      "founder": {
        "@id": "https://goodforbots.com/#person-mateusz-pawlica"
      },
      "contactPoint": {
        "@type": "ContactPoint",
        "contactType": "customer support",
        "email": "support@goodforbots.com",
        "url": "https://goodforbots.com/contact"
      }
    },
    {
      "@type": "WebSite",
      "@id": "https://goodforbots.com/#website",
      "name": "Good for Bots",
      "url": "https://goodforbots.com",
      "description": "Check if AI crawlers can access and read your website, free. Get a score out of 100 with every finding and fix prompts, then browse reports in the directory.",
      "inLanguage": "en",
      "publisher": {
        "@id": "https://goodforbots.com/#organization"
      },
      "potentialAction": {
        "@type": "SearchAction",
        "target": {
          "@type": "EntryPoint",
          "urlTemplate": "https://goodforbots.com/directory?q={q}"
        },
        "query-input": "required maxlength=120 name=q"
      }
    },
    {
      "@type": "Person",
      "@id": "https://goodforbots.com/#person-mateusz-pawlica",
      "name": "Mateusz Pawlica",
      "givenName": "Mateusz",
      "familyName": "Pawlica",
      "jobTitle": "Founder of Good for Bots",
      "description": "Web developer who builds directories and AI-powered products, and the founder of Good for Bots.",
      "url": "https://goodforbots.com/about#founder",
      "sameAs": [
        "https://pl.linkedin.com/in/mateusz-pawlica-65238816b",
        "https://pawlicaweb.pl/o-mnie"
      ],
      "knowsAbout": [
        "llms.txt",
        "robots.txt",
        "AI crawlers",
        "Markdown content negotiation",
        "Structured data",
        "Technical SEO",
        "Web directories",
        "Next.js",
        "TypeScript",
        "Generative AI"
      ],
      "worksFor": {
        "@id": "https://goodforbots.com/#organization"
      }
    }
  ]
}
```
