Free tool
robots.txt checker
See which AI crawlers your robots.txt allows or blocks, grouped as AI search, user fetches and training, and whether it states AI usage preferences. Free.
What the robots.txt checker checks
It runs these checks from the full scan, under the same rules, and the descriptions below come from our standards catalogue. Unlike the full scan, it never retries through a browser, so a file behind bot protection can fail here and still pass in a report.
robots.txt
We parse /robots.txt and check delivery and directives, as RFC 9309 defines them: lines may end in CR, LF or CRLF, and a Crawl-delay or Sitemap line does not end a user-agent group. Whether we may scan the site is decided separately, before scoring. A parseable refusal of our bot's access to the homepage ends the scan without a score, even when served under an unusual HTTP status. A rule that refuses us only a particular file, such as /llms.txt or a sitemap, leaves the scan running: we do not request that file, and the check that needs it treats it as unavailable to crawlers. A robots.txt answered with a bare 401 or 403 is treated as no rules for access, as RFC 9309 allows, but it does not pass this check.
Results: Pass: the file is delivered at HTTP 200 and has usable directives without parse errors. Warn: it declares nothing or has parsing problems, such as a line without a colon, a rule above the first User-agent line, or an empty or unparseable Sitemap URL. Fail: it is missing, empty, unreadable, an HTML page or delivered under a status other than 200.
Full methodology and sources →Explicit AI crawler rules
We match user-agent tokens against our crawler registry, combine matching groups and record their purposes and access rules. A wildcard group alone does not count as an explicit AI crawler policy, and neither do Content-Signal or Content-Usage lines: they state preferences to every crawler without naming one, and the content signals check scores them.
Results: Pass: at least two crawler purposes are represented, or at least four distinct known crawlers are named. Warn: some are named but neither threshold is met. Fail: no known AI crawlers are named, including when the file states usage preferences for every crawler.
Full methodology and sources →Search crawler access
We apply robots.txt rules to the homepage path for the search and user-fetch crawler tokens their operators document. Explicit groups, wildcard fallback and rule precedence determine access; training, agent and extraction tokens are not charged. Only a robots.txt served normally, with HTTP 200, is read for this check.
Results: Pass: no applicable rules block the selected crawlers. Fail: one or more are blocked, with the exact deduction recorded in the report. N/A: the robots probe was not run. A missing robots file does not itself mean access is blocked.
Full methodology and sources →AI usage preferences
We read every Content-Signal and Content-Usage line in robots.txt, including rules scoped to a path, and the Content-Usage header on the final successful homepage response. Findings identify the header’s response URL. Both carriers share one check and count once. Either vocabulary can satisfy the check; recognised positive and negative preferences count equally. Only yes or no, or y or n for Content-Usage, states a preference; any other value states nothing. Content-Usage is read as the IETF draft specifies: parameters after a value are ignored, and a declaration that does not parse, for example one with uppercase keys, states nothing.
Results: Pass: at least one recognised category and value. Warn: a usage declaration exists but its categories or values are not recognised. Fail: no declaration is found.
Full methodology and sources →Fix what it finds
- How to write robots.txt rules for AI crawlers
Write robots.txt rules that treat AI training, search and user-triggered crawlers separately, with tested examples, the current token list and common fixes.