Give AI crawlers User-agent groups by name and decide per purpose: training, search, or a fetch a person asked for. Refusing a training crawler such as GPTBot keeps your content out of that operator's future training data and leaves its AI search alone. Refusing a search crawler or a user-triggered fetcher, such as OAI-SearchBot or Claude-User, asks that operator to leave your pages out of its answers.
robots.txt is a request, not a lock. Crawlers that follow RFC 9309 obey it, some user-triggered fetchers say they may not, and a firewall can refuse a crawler the file lets in. Whether to refuse training is your decision; which AI crawlers to block, and which to keep sets out what each choice costs. This guide covers writing the rules and checking that crawlers receive them.
Start with a complete example
The file below is for Harbour, a fictional service that refuses model training and stays open to AI search. example.com is a placeholder. Replace the paths and the sitemap address with your own.
# robots.txt for Harbour, a fictional service. example.com is a placeholder.
# Model training: refused. Google-Extended also covers grounding in Gemini Apps.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /
# Every other crawler, including AI search crawlers and user-triggered
# fetchers such as OAI-SearchBot, Claude-SearchBot and ChatGPT-User.
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Content-Usage: train-ai=n
Disallow: /account/
Disallow: /search
Sitemap: https://example.com/sitemap.xml
The first group names five training crawlers and control tokens and refuses them the whole site. The second, User-agent: *, covers every crawler not named elsewhere, AI search crawlers and user-triggered fetchers included. It keeps them out of the account area and internal search results, and states the owner's terms in two vocabularies explained below. This file passes our robots.txt, AI crawler rules, crawler access and usage-preference checks at the versions reviewed for this guide.
To allow training instead, keep the first group but give it the same Disallow lines as the * group, and change ai-train=no to ai-train=yes and train-ai=n to train-ai=y. An Allow: / alone in that group would open /account/ to those crawlers, for the reason in the next section. Deleting the group also allows training, but the file then names no AI crawler, and our AI crawler rules check looks for an explicit position either way.
How crawlers read robots.txt
RFC 9309, the Robots Exclusion Protocol, defines how a crawler finds the file and applies it. These parts decide whether rules for AI crawlers do what you meant:
- One file per origin. The file lives at
/robots.txt, all lowercase, at the top of the host. Google spells out the consequence: its rules "apply only to the host, protocol, and port number where the robots.txt file is hosted." A subdomain such asshop.example.comneeds its own file. - A group starts with its
User-agentlines. One group can name several crawlers, as the example does. It ends at the nextUser-agentline after its rules; a blank line does not end it. - Names match regardless of case. "Crawlers MUST use case-insensitive matching", so
gptbotandGPTBotname the same crawler. Write the token alone, without a version such as/1.0: a product token may contain only letters, underscores and hyphens. - One set of rules applies. A crawler combines every group that names it and ignores the others. It obeys
User-agent: *only "if no matching group exists", so a named group never inherits the*rules. - The longest path wins. RFC 9309 defines the most specific match as "the match that has the most octets." When an
Allowand aDisalloware equivalent, theAllowshould win. A path no rule matches is allowed, and paths should be compared case-sensitively.*matches any run of characters and$anchors the end of the path. - Other lines do not change the rules.
Sitemap,Crawl-delay,Content-SignalandContent-Usageare not part of the protocol, and parsing such records "MUST NOT interfere" with the defined ones.
The HTTP answer counts as much as the text. A 4xx means the file is unavailable, and the crawler "MAY access any resources on the server." A 5xx or a network error means the crawler "MUST assume complete disallow". Crawlers "SHOULD NOT use the cached version for more than 24 hours," and they must parse at least 500 KiB; Google ignores anything after that limit.
The file is public. RFC 9309 warns that "listing paths in the robots.txt file exposes them publicly." A Disallow line hides nothing from a person or a crawler that ignores it, so protect private areas with authentication.
Choose AI crawler tokens by purpose
The tables come from the same maintained knowledge base used by our scanner. They distinguish documented controls, HTTP identities, software tools and unconfirmed names. Each entry links to its evidence and review date. The scoring decision is shown separately: a bot can have several uses, and appearing in this catalogue does not automatically affect the score. Where no robots.txt token is confirmed, do not turn the displayed name into a rule.
Training crawlers and control tokens
These controls express a training opt-out. Refusing them costs nothing in our score. Check each entry's status: Coherebot is a hypothetical example in Cohere's documentation, not an announced active crawler.
Google-Extended and Applebot-Extended are control tokens, not crawlers. Apple says "Applebot-Extended does not crawl webpages": Googlebot and Applebot fetch the pages, and the token tells Google or Apple how it may use them. Google-Extended also covers grounding in Gemini Apps and Grounding with Google Search on Vertex AI, and Google offers no token that refuses training alone. Never disallow Googlebot or Applebot to refuse AI training: that removes the site from Google Search or from Apple's search features.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
GPTBot · OpenAI. GPTBot | Collects content that may be used to train generative models. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
ClaudeBot · Anthropic. ClaudeBot | Collects content that may be used to train generative models. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
MistralAI-Training · Mistral. MistralAI-Training | Collects content that may be used to train generative models. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
KimiBot · Moonshot AI. KimiBot | Collects content that may be used to train generative models. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
AI2Bot1 · Ai2. AI2Bot | Collects web content to train open language models. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
Applebot-Extended2 · Apple. Applebot-Extended | Controls training use of content collected by Applebot. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
Webzio-extended3 · Webz.io. Webzio-extended | Controls whether Webz.io data can be used for AI or ML training. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
| Coherebot4 · Cohere. No confirmed robots.txt token | An example of how a future Cohere training crawler could be refused. Status: example. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
AI search crawlers
These discover sources for AI answers, retrieve them at query time, or control the use of already indexed pages. OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links." Refusing one of these crawlers the homepage is a penalty in our score: the file declares that the site does not want to be read for AI answers.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
OAI-SearchBot · OpenAI. OAI-SearchBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Claude-SearchBot · Anthropic. Claude-SearchBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
PerplexityBot · Perplexity. PerplexityBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
MistralAI-Index · Mistral. MistralAI-Index | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Kimi-SearchBot · Moonshot AI. Kimi-SearchBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Amzn-SearchBot5 · Amazon. Amzn-SearchBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Meta-WebIndexer · Meta. Meta-WebIndexer | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
ExaSearchBot6 · Exa. ExaSearchBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
AIWebIndex7 · Lyrenth. AIWebIndex | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
DuckAssistBot8 · DuckDuckGo. DuckAssistBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
ShapBot9 · Parallel. ShapBot | Discovers or retrieves pages for AI search and cited answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
YandexAdditional10 · Yandex. YandexAdditional, YandexAdditionalBot | Controls use of indexed pages in Search with Yandex AI answers. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
User-triggered fetchers
These fetch a page when a person asks about it. OpenAI, Perplexity, Meta, Amazon and Google say their user-triggered fetchers may fetch a page robots.txt refuses them; OpenAI's wording for ChatGPT-User is "robots.txt rules may not apply." Anthropic says its bots, Claude-User included, honour robots.txt. A refusal still states your intent, and our score counts it the same way as a search crawler's.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
ChatGPT-User11 · OpenAI. ChatGPT-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Claude-User12 · Anthropic. Claude-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Perplexity-User13 · Perplexity. Perplexity-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
MistralAI-User14 · Mistral. MistralAI-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Kimi-User15 · Moonshot AI. Kimi-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Amzn-User16 · Amazon. Amzn-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Diffbot-User17 · Diffbot. Diffbot-User | Retrieves pages in response to a person's request. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Agents
Agents browse and act for a person. Our score does not charge a refusal. An HTTP identity or product name does not establish a robots.txt control. Google lists Google-Agent among its user-triggered fetchers, which "generally ignore robots.txt rules", and OpenAI's crawler page names no token for its agent. To control agents, use your firewall or bot-management rules.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
| Google-Agent18 · Google. No confirmed robots.txt token | Navigates websites and performs actions requested by users. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Extraction services
These crawl on behalf of customers who buy extraction, indexing or retrieval. Our score does not charge a refusal. Software packages may let callers choose their own identity; a package name is not necessarily a robots.txt token.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
FirecrawlAgent19 · Firecrawl. FirecrawlAgent | Controls Firecrawl site crawling for content extraction. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Webzio20 · Webz.io. Webzio | Collects and structures content for web data products. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Google-CloudVertexBot21 · Google. Google-CloudVertexBot | Crawls requested by site owners to build Vertex AI Agents. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| Brightbot22 · Bright Data. No confirmed robots.txt token | Collects web data for Bright Data services. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| img2dataset23 · img2dataset contributors. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| ApifyWebsiteContentCrawler24 · Apify. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| Crawl4AI25 · Crawl4AI contributors. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
General search
These entries document general search uses. Inclusion in this catalogue does not add an AI crawler access penalty.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
Diffbot26 · Diffbot. Diffbot | Builds a general search engine and the Diffbot Knowledge Graph. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Googlebot27 · Google. Googlebot | Crawls for Google Search and associated search products. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Mixed uses
One control can cover several uses. Read its research note before deciding: refusing training may also affect another use. The table shows our scoring decision separately from the operator's description.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
Meta-ExternalFetcher28 · Meta. Meta-ExternalFetcher | Fetches user-requested links and supports agentic product functions. Refusal counts toward the crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Google-Extended29 · Google. Google-Extended | Controls Gemini training and grounding uses of crawled content. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
Meta-ExternalAgent30 · Meta. Meta-ExternalAgent | Collects content for foundation model training and product indexing. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
Amazonbot31 · Amazon. Amazonbot | Collects content for product improvements and possible AI model training. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
CCBot32 · Common Crawl. CCBot | Builds an open web dataset used for research and model development. Training opt-out stays neutral. | Operator source. Reviewed 2026-09-24. |
Applebot33 · Apple. Applebot | Collects content for Apple search features and possible model training. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Advertising
Advertising checks serve submitted ads. They are separate from organic discovery and carry no crawler access penalty.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
| OAI-AdsBot34 · OpenAI. No confirmed robots.txt token | Validates submitted ad pages and helps determine ad relevance. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
Unconfirmed names and unclear purposes
These records preserve the research trail. They are not recommended robots.txt directives and do not affect our score. An unavailable source or undocumented token does not prove that a product was retired.
| Name and robots.txt control | Use and status | Evidence |
|---|---|---|
TerraCotta35 · Ceramic AI. TerraCotta | A documented crawler whose precise content use is not established. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| anthropic-ai36 · Anthropic (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| Bytespider37 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| FacebookBot38 · Meta (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| cohere-training-data-crawler39 · Cohere (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Ai2Bot-Dolma40 · Ai2 (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| DeepSeekBot41 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| QwenBot42 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| PanguBot43 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Timpibot44 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| omgili45 · Webz.io. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| omgilibot46 · Webz.io. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: retired. No crawler access penalty. | Operator source. Reviewed 2026-09-24. |
| YouBot47 · You.com (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| Operator48 · OpenAI (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| chatgpt-agent49 · OpenAI (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| GoogleAgent-Mariner50 · Google (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| NovaAct51 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Devin52 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Manus-User53 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| TwinAgent54 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| ApifyBot55 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Crawlspace56 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| HenkBot57 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| Claude-Web58 · Anthropic (unconfirmed token). No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity · Additional source. Reviewed 2026-09-24. |
| iAskBot59 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
| iaskspider60 · Unconfirmed. No confirmed robots.txt token | Retained for research; no active scanner token is established. Status: unknown. No crawler access penalty. | Unconfirmed identity. Reviewed 2026-09-24. |
State your terms with Content Signals or Content-Usage
Allow and Disallow say whether a crawler may fetch a path. They cannot tell a crawler that does several jobs which of them you accept. Two vocabularies say that in robots.txt:
- Content Signals is Cloudflare's convention:
Content-Signal: search=yes, ai-input=yes, ai-train=no. contentsignals.org definessearchas building a search index and returning links and short excerpts,ai-inputas putting content into an AI model, for example for grounding, andai-trainas "Training or fine-tuning AI models." Values areyesandno, and Cloudflare says "the absence of a signal conveys no meaning." The IETF draft it came from has expired. - AIPREF is the IETF working group's vocabulary:
Content-Usage: train-ai=n. Its categories aretrain-ai,ai-useandsearch, with the valuesyandn. Keys are case-sensitive, andtrain-aireverses Cloudflare'sai-train. draft-ietf-aipref-vocab is an active working-group draft intended as a Proposed Standard. Version 08, of 14 September 2026, addedai-useand notes that its contents "DO NOT REFLECT CONSENSUS of the Working Group."
Both are declarations. Neither blocks a request, and a crawler that does not implement them ignores them. Write them inside a group, and remember that a crawler reads only one set of rules: a crawler with its own named group does not see the line under User-agent: *. The AIPREF attachment draft's example shows this, and contentsignals.org's example for specific crawlers repeats the line in each group. Our usage-preference check accepts either vocabulary.
Publish the file where crawlers look
Serve the file at https://your-domain.example/robots.txt with HTTP 200, as UTF-8 text/plain. Where it comes from depends on your stack:
- A static file in the web root, or
app/robots.txtin a Next.js app. Next.js can also generate the file fromapp/robots.ts; since version 16.3 a rule'sotherfield adds lines such asContent-Signal, "scoped to the rule'sUser-Agentblock." - WordPress generates
/robots.txtwhen no file exists at that path, and plugins change its output through therobots_txtfilter. - Shopify generates the file, and the
robots.txt.liquidtheme template customises it. - Webflow and Wix have a robots.txt editor in their SEO settings. Squarespace users cannot edit the file. Its "Block known artificial intelligence crawlers" setting adds a fixed list that includes
DuckAssistBotandYouBot, so check the published file against the controls you intended to set. DuckAssistBot retrieves sources for AI-assisted search answers. - Cloudflare can put its own lines in front of yours. Its Bot Preference Sync, introduced in August 2026 and on by default for new customers, prepends them "to the existing material".
Then read what crawlers receive, not what is in your repository. Replace the example host and run:
curl --include https://your-domain.example/robots.txt
Check that the status is 200, the Content-Type is text/plain and the body starts with your rules, not with <!doctype html>. A redirect is acceptable if it ends at the file quickly: RFC 9309 asks crawlers to follow "at least five consecutive redirects." Compare the body with your source file. A line added by a CDN or a plugin is a rule crawlers obey all the same.
The robots.txt checker fetches the file from our side and shows how it treats each AI crawler.
Changes take time to arrive. RFC 9309 allows a cached copy for up to 24 hours. OpenAI, for search results, and Amazon say about 24 hours, Perplexity and Meta say up to 24 hours, and DuckDuckGo says a change "will take effect after 72 hours".
Make sure the firewall agrees with the file
robots.txt only asks. A firewall, a bot-management product or a CDN setting decides what a crawler receives, and it can refuse a crawler the file allows:
- Cloudflare. Since 15 September 2026, "Block" in AI Crawl Control applies to mixed-use crawlers, "including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training." To refuse training only, Cloudflare offers "Disallow AI Training", which publishes the preference in
robots.txtand keeps what it calls accountable mixed-use crawlers allowed for search. - Vercel. The AI bots managed ruleset is off by default. Its Deny action "blocks all traffic identified as coming from AI bots", and the ruleset covers bots crawling for training, search and user-generated fetches.
- AWS WAF. Bot Control's
CategoryAIrule blocks AI bots "regardless of whether the bots are verified or unverified."
A crawler that receives a challenge page or a 403 gets no content, whatever robots.txt says. Our bot challenge check looks for that in front of our own crawler, and our crawler page explains what to look for in your security events.
Slow crawlers down instead of refusing them
If the problem is load rather than use, Crawl-delay asks a crawler to wait between requests. It is not part of RFC 9309, and support varies. Google, Apple and Amazon say their crawlers do not follow it. Bing and Common Crawl say theirs do, and Anthropic says it supports "the non-standard Crawl-delay extension". A rate limit at your server works whether a crawler reads the file or not, and HTTP 429 with Retry-After tells a well-behaved client when to come back. Our crawler honours Crawl-delay.
Fix common mistakes
A blocklist that also refuses AI search
# Meant to refuse AI training. It also refuses AI search and user fetches.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Disallow: /
User-agent: *
Disallow: /account/
The comment says what the author wanted; the tokens say something else. OAI-SearchBot, Claude-SearchBot and PerplexityBot are search crawlers, and ChatGPT-User, Claude-User and Perplexity-User are user-triggered fetchers. Our crawler access check fails this file for all six. Keep GPTBot and ClaudeBot in the group and delete the other lines, so those crawlers fall back to User-agent: *.
Ready-made lists often mix purposes. The ai.robots.txt project says its list "contains AI-related crawlers of all types, regardless of purpose." Check each token against the tables above before you copy a list.
A named group that drops your other rules
User-agent: *
Disallow: /account/
User-agent: OAI-SearchBot
Allow: /
The author meant to welcome OAI-SearchBot and also let it into /account/. The crawler finds a group that names it, so it never reads User-agent: *, and its own group has no rule for /account/. Repeat every Disallow a named crawler should still obey inside its group. Our crawler access check tests only the homepage path, so it passes this file; only you can see that a private path is now open.
A file that names no AI crawler
User-agent: *
Disallow: /account/
Sitemap: https://example.com/sitemap.xml
This file is valid. Under RFC 9309 it lets every AI crawler in, except to /account/, and it says nothing about training, search or user fetches. Our AI crawler rules check fails it because it takes no explicit position on AI crawlers, and our usage-preference check fails it because it declares no terms. The complete example above answers both. If you are content to let every crawler in, name the ones you have considered in a group with the same Disallow lines as *.
An HTML page at /robots.txt
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Harbour</title>
<script type="module" src="/assets/index.js"></script>
</head>
<body>
<div id="root"></div>
</body>
</html>
Single-page apps and catch-all routes often answer every path with the application shell and HTTP 200. A crawler finds no rules in it, and our robots.txt check fails it as an HTML page. Serve a real text file at the path, ahead of the application's fallback route.
A server error at /robots.txt
A 5xx answer is worse than no file. RFC 9309 tells crawlers to "assume complete disallow" when the file is unreachable because of a server or network error, so an outage or a broken route refuses every compliant crawler. Google documents its own variant: it stops crawling for 12 hours, then uses the last good copy for up to 30 days. If you have no rules to publish, a 404 lets crawlers in; a file with your rules is better.
Keep the rules current
Operators add, rename and retire crawlers. When a new token shows up in your server logs, read what its operator says it does before you name it, and check whether it trains, searches or fetches for a person. After any deploy that touches routing, a CDN rule or a platform setting, fetch the live file again.
Scan your site to see the robots.txt our crawler receives and how each check below reads it. The report shows declared rules and what our own crawler met; it cannot see how your firewall answers other crawlers.
Sources and review
Reviewed on 24 September 2026 against these sources:
- RFC 9309: Robots Exclusion Protocol, IETF Proposed Standard
- Google: How Google interprets the robots.txt specification
- OpenAI: Overview of OpenAI crawlers
- Anthropic: crawler controls
- Perplexity: crawlers
- Google's common crawlers
- Google's user-triggered fetchers
- Apple: About Applebot
- Meta: web crawlers
- Amazon: Amazonbot, Common Crawl: FAQ and DuckDuckGo: DuckAssistBot
- Content Signals and Cloudflare: The Content Signals Policy
- draft-ietf-aipref-vocab and draft-ietf-aipref-attach, IETF working-group drafts
- Cloudflare: stay discoverable in search while disallowing AI training and Bot Preference Sync
- Vercel: bot management and AWS WAF Bot Control rule group
- Next.js: robots.txt, WordPress: robots_txt filter, Shopify: robots.txt.liquid, Webflow, Wix and Squarespace
- Bing: To crawl or not to crawl, that is BingBot's question
The crawler tables come from our scanner's registry. The examples are ours, and each is tested against the checks listed below at the versions reviewed for this guide.
Crawler research notes
Footnotes
-
Ai2 documents AI2Bot, including the digit. Matching retains the scanner's existing product-token normalization; the catalogue compiler rejects collisions after that normalization. ↩
-
This token does not fetch pages. Refusing it does not remove pages from Apple search results. ↩
-
The operator describes validation and tagging of data collected by Webzio, not a second full-content crawler. ↩
-
Cohere says it is not currently using crawlers for foundation-model training and lists no active bot. Coherebot appears in a hypothetical robots.txt example, not an active crawler announcement. ↩
-
Amazon documents fallback to other search bots when this token is unnamed. Our access check uses RFC 9309 wildcard fallback; it does not emulate that vendor extension. ↩
-
Exa documents fallback to major search engine rules before the wildcard group. Our access check uses RFC 9309 wildcard fallback. Exabot is a different historical crawler, not an alias. ↩
-
The current policy says both scheduled crawling and AIWebIndex-Agent on-demand index fills follow the AIWebIndex robots token. The HTTP names do not establish two independent access routes. The policy also discusses corpus licensing and separate training reservations; it says Lyrenth does not train foundation models itself. ↩
-
DuckDuckGo describes real-time crawling for AI-assisted search answers, with no model training. It is classified as ai-search, rather than user-fetch. ↩
-
Parallel documents discovery and indexing for its web APIs. The old registry placed this bot under extraction; the current catalogue records its search role. ↩
-
Yandex lists YandexAdditional and YandexAdditionalBot together as controls over already indexed content. They are one catalogue entry and one penalty route. Either explicitly refused alias records a refusal; conflicting aliases are conservatively treated as refused. This is our declared-policy reading, not an implementation of Yandex-specific fallback rules. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
Anthropic says its bots honour robots.txt, including this user-triggered fetcher. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
User-triggered retrieval is separate from automatic crawling. A robots.txt rule expresses intent; it does not prove this fetcher will refuse a user request. ↩
-
Google documents an HTTP identity among user-triggered fetchers, which generally ignore robots.txt. The page does not publish a dedicated robots.txt control for this agent; its HTTP name is not promoted to a scanner token. ↩
-
Firecrawl documents this robots.txt directive for its crawl endpoint. Callers can pass custom headers; the catalogue does not claim every request has this HTTP identity. ↩
-
Webz.io introduced Webzio and Webzio-extended as replacements for its older Omgilibot system. Webzio-extended separately controls training eligibility. ↩
-
Google says this control affects owner-requested Vertex AI crawling, not Google Search. Its documented fallback includes Googlebot; the scanner does not add a visibility penalty for this service. ↩
-
The operator publishes an HTTP identity and collectors.txt controls. That does not establish a robots.txt token contract, so this entry is informational. ↩
-
An image dataset downloader, not one centrally operated crawler. Its project name is not a documented robots.txt token. ↩
-
The actor documentation says it uses no specific user agent identifier. Respecting robots.txt is configurable and defaults to false. The actor name must not become a scanner token. ↩
-
A configurable crawling library run by its users. The reviewed documentation does not establish the project name as one shared robots.txt identity. ↩
-
Diffbot explicitly distinguishes general search crawling from Diffbot-User requests and says it does not crawl to train foundation models. A general-search entry does not automatically become an AI access penalty. ↩
-
Included for reference. This catalogue does not introduce another Googlebot access penalty; existing indexing checks own their separate rules. ↩
-
Meta says this fetcher may bypass robots.txt on user-requested fetches. Its user-fetch route remains in the access policy; its additional agent role does not create another charge. ↩
-
No separate HTTP User-Agent. Refusal does not affect inclusion in Google Search. The scanner keeps blocking neutral because this control also refuses training; classification is independent of that scoring decision. ↩
-
Meta documents training and direct indexing for product improvement. This is not a training-only claim. Blocking remains neutral; Meta-WebIndexer is the separately documented AI search route. ↩
-
Amazon says collected content may train its AI models. Amzn-SearchBot and Amzn-User are separate controls. Blocking Amazonbot remains neutral. ↩
-
Common Crawl publishes crawl data for reuse, rather than operating a consumer answer engine. Its dataset has uses beyond training. Blocking remains neutral under the existing training opt-out policy. ↩
-
Applebot-Extended separately controls foundation-model training. Applebot is included for reference without adding a new AI crawler access penalty. ↩
-
Only pages submitted as ads are visited. The operator excludes foundation-model training. This is an advertising identity, not an organic search access route. ↩
-
The operator repository documents the token and robots compliance, but the reviewed page does not establish AI search or training use. The old search classification is not retained as a fact. ↩
-
The current Anthropic reference documents ClaudeBot, Claude-SearchBot and Claude-User, but not this spelling. Absence is not proof of retirement; retained as an unconfirmed historical name. ↩
-
The previous registry attributed this name to ByteDance. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The current Meta crawler reference does not document this old spelling. Do not infer retirement or equate it with Meta-ExternalAgent. ↩
-
This name came from community lists. Cohere currently lists no active training crawler; Coherebot is only a hypothetical example. This spelling is not confirmed by its reference. ↩
-
The reviewed Ai2 crawler notice publishes AI2Bot. It does not establish this additional spelling as a current independent token. ↩
-
The previous registry attributed this name to DeepSeek. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Alibaba. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Huawei. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The attributed operator crawler page returned 404. Community classifications disagree about training versus search. Neither purpose nor a current control contract is established here. ↩
-
Webz.io describes the move from Omgilibot to Webzio and Webzio-extended. Retained for historical blocklists; use the current documented controls. The separate omgili spelling has no current token contract in this source. ↩
-
Webz.io describes the move from Omgilibot to Webzio and Webzio-extended. Retained for historical blocklists; use the current documented controls. The separate omgili spelling has no current token contract in this source. ↩
-
The previously cited operator URL returned 404 during this review; the you.com variant redirected to sign-in. No replacement primary reference was established. This does not prove retirement; the name remains in research, without a penalty. ↩
-
The OpenAI crawler reference does not establish this robots.txt token. An agent product name must not be converted into a robots token. Retained as an unconfirmed registry spelling, not as a claim that the product is inactive. ↩
-
The OpenAI crawler reference does not establish this robots.txt token. An agent product name must not be converted into a robots token. Retained as an unconfirmed registry spelling, not as a claim that the product is inactive. ↩
-
The current Google user-triggered fetcher list documents Google-Agent, not this token. No alias relationship is inferred. ↩
-
The previous registry attributed this name to Amazon. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Cognition. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Manus. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Twin. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Apify. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to Crawlspace. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The community list attributes HenkBot to Valyu, while the previous registry said unknown. No primary source was established, so the operator remains unconfirmed. ↩
-
The current Anthropic reference documents ClaudeBot, Claude-SearchBot and Claude-User, but not this spelling. Absence is not proof of retirement; retained as an unconfirmed historical name. ↩
-
The previous registry attributed this name to iAsk. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
-
The previous registry attributed this name to iAsk. The community list is a discovery lead, not confirmation of the operator, purpose or a working robots.txt token. No primary source confirming this exact control was established in this review; it is excluded from scanner recognition. ↩
Current scanner criteria
These criteria come from our current standards catalogue. They describe what Good for Bots checks, including usefulness rules of our own that the format does not require, and what our detection cannot see. Each report keeps the methodology of the scan that produced it.
robots.txt
active check · Reviewed
We parse /robots.txt and check delivery and directives, as RFC 9309 defines them: lines may end in CR, LF or CRLF, and a Crawl-delay or Sitemap line does not end a user-agent group. Whether we may scan the site is decided separately, before scoring. A parseable refusal of our bot's access to the homepage ends the scan without a score, even when served under an unusual HTTP status. A rule that refuses us only a particular file, such as /llms.txt or a sitemap, leaves the scan running: we do not request that file, and the check that needs it treats it as unavailable to crawlers. A robots.txt answered with a bare 401 or 403 is treated as no rules for access, as RFC 9309 allows, but it does not pass this check.
Results: Pass: the file is delivered at HTTP 200 and has usable directives without parse errors. Warn: it declares nothing or has parsing problems, such as a line without a colon, a rule above the first User-agent line, or an empty or unparseable Sitemap URL. Fail: it is missing, empty, unreadable, an HTML page or delivered under a status other than 200.
Limitations: Our requirement for HTTP 200, and our partial credit for a file without directives or with lines we cannot parse, are scoring choices; RFC 9309 asks crawlers only to use the rules they can parse. We read a bounded file sample. User-agent names are compared by their letters, hyphens and underscores, so a digit in a name is ignored. This check does not judge whether the declared rules allow other AI crawlers; separate checks do that.
Full methodology and sources →Explicit AI crawler rules
active check · Reviewed
We match user-agent tokens against our crawler registry, combine matching groups and record their purposes and access rules. A wildcard group alone does not count as an explicit AI crawler policy, and neither do Content-Signal or Content-Usage lines: they state preferences to every crawler without naming one, and the content signals check scores them.
Results: Pass: at least two crawler purposes are represented, or at least four distinct known crawlers are named. Warn: some are named but neither threshold is met. Fail: no known AI crawlers are named, including when the file states usage preferences for every crawler.
Limitations: The thresholds are our criteria, not RFC requirements. Recognition depends on the crawler registry, which we check against the operators' own documentation, and can miss newly introduced names. Earning these points does not mean search crawlers are allowed: that is a separate penalty check.
Full methodology and sources →Search crawler access
active check · Reviewed
We apply robots.txt rules to the homepage path for the search and user-fetch crawler tokens their operators document. Explicit groups, wildcard fallback and rule precedence determine access; training, agent and extraction tokens are not charged. Only a robots.txt served normally, with HTTP 200, is read for this check.
Results: Pass: no applicable rules block the selected crawlers. Fail: one or more are blocked, with the exact deduction recorded in the report. N/A: the robots probe was not run. A missing robots file does not itself mean access is blocked.
Limitations: This evaluates declared access to the root path, not live requests impersonating other bots. Firewall behaviour and deeper paths can differ. OpenAI, Perplexity, Meta, Amazon and Google say their user-triggered fetchers may fetch a page robots.txt refuses them when a person asks for it, so we charge the declared refusal, not proven invisibility. Recognition depends on our maintained crawler registry, and we charge only tokens their operators document.
Full methodology and sources →AI usage preferences
active check · Reviewed
We read every Content-Signal and Content-Usage line in robots.txt, including rules scoped to a path, and the Content-Usage header on the final successful homepage response. Findings identify the header’s response URL. Both carriers share one check and count once. Either vocabulary can satisfy the check; recognised positive and negative preferences count equally. Only yes or no, or y or n for Content-Usage, states a preference; any other value states nothing. Content-Usage is read as the IETF draft specifies: parameters after a value are ignored, and a declaration that does not parse, for example one with uppercase keys, states nothing.
Results: Pass: at least one recognised category and value. Warn: a usage declaration exists but its categories or values are not recognised. Fail: no declaration is found.
Limitations: HTTP coverage is limited to the final homepage response; its header describes that response, not the whole site. We do not resolve effective permissions across carriers or decide whether a downstream service honours a preference. A rule scoped to a path counts as though it covered the whole site. Drafts and conventions are not all published standards. The check recognises declarations; it is not a legal interpretation of permission.
Full methodology and sources →More guides
Page content
How to serve content AI crawlers can read without JavaScript
Which AI crawlers run JavaScript, how to fix empty app shells in Next.js, Nuxt, SvelteKit or Vite, and how to test the raw HTML response yourself.
Reviewed
Page content
How to serve Markdown to AI agents with content negotiation
Return Markdown from the same URL when an AI agent sends Accept: text/markdown, with tested responses, Vary and CDN caching, and nginx and Next.js setups.
Reviewed
Discovery files
How to write a useful llms.txt
Write an llms.txt that tells AI agents what your site covers: a complete tested example, the format explained, publishing checks and fixes for common mistakes.
Reviewed